MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
Quanhao Li, Zhen Xing, Rui Wang, Hui Zhang, Qi Dai, Zuxuan Wu
Introduction
With the rapid development of diffusion models, video generation has witnessed significant progress in recent years. Earlier video generation approaches, such as AnimateDiff and SVD , primarily rely on the UNet structure, which results in videos with limited length and quality. Sora has demonstrated the powerful capabilities of the DiT architecture in text-to-video (T2V) generation. Following this breakthrough, subsequent models leveraging the DiT architecture have achieved higher quality outputs and extended video duration.
While DiT-based models excel at producing high-quality and longer videos, many text-to-video approaches lack precise control over attributes like object movement and camera motion . Fine-grained trajectory-controllable video generation emerges as a solution, which is especially critical for generating controllable videos in real-world scenarios.
Previous trajectory-controllable video generation methods can be categorized based on the type of control signals they use. These include points-based control , optical flow-based control , bounding boxes-based control , masks-based control , and 3D trajectories-based control . However, these methods exhibit several limitations. First, the trajectory control conditions used by these methods are unitary. Each method only support a single type of control signal. However, sparse trajectories (e.g., points and optical flow) result in imprecise control over object shape and size, while dense trajectories (e.g., masks and 3D trajectories) are challenging for users to provide. Second, a publicly available large-scale dataset for trajectory-controllable video generation is lacking. Existing VOS (Video Object Segmentation) datasets suffer from short video lengths , small scales , or few foreground objects . Third, a unified benchmark for evaluating different methods is absent. Besides, previous work focuses solely on video quality and trajectory accuracy while neglecting the impact of the number of moving objects. We argue that controlling fewer or more objects presents different challenges, making it essential to include this factor in the evaluation metrics.
To address these issues, we propose MagicMotion, a controllable video generation framework with dense-to-sparse trajectory guidance. To inject trajectory conditions into the generation process, we utilize an architecture similar to ControlNet called Trajectory ControlNet to encode the trajectory information, which is later added to the original DiT model through a zero-initialized convolution layer. MagicMotion supports three types of trajectory conditions: masks, boxes, and sparse boxes using a progressive training strategy. Experiments show that the model can leverage the knowledge learned in the previous stage to achieve better performance than training from scratch. Additionally, we propose a novel latent segment loss that helps the video generation model better understand the fine-grained shape of objects with minimal computation.
We also construct MagicData, a high-quality public dataset comprising 51K video samples, each annotated with a
In conclusion, the main contributions of our work are summarized as follows:
We present MagicMotion, a trajectory-controllable image-to-video generation model that supports three types of control signals: masks, boxes, and sparse boxes.
We introduce a data curation and filtering mechanism, and construct MagicData, the first public dataset for trajectory-controlled video generation.
We propose MagicBench, a comprehensive benchmark for evaluating trajectory-controllable video generation models for both video quality and trajectory control accuracy across different numbers of controlled objects.
Related Works
Video Diffusion Models Diffusion models have made great progress in image generation , which has led to the rapid development of video generation . VDM is the first to apply diffusion models to video generation. Early works like AnimateDiff and SimDA attempt to insert temporal layers into pretrained T2I model for video generation. Subsequently, VideoCrafter and SVD use large-scale and high-quality data for training, acheiving better performance. However, these methods have trouble on generating long videos with high quality, mainly due to the inherent limitations of the UNet architecture. The emergence of Sora is a significant success, demonstrating the potential of DiT models to generate high-quality videos with tens of seconds. Recent video generation methods are mainly based on DiT architecture, and have achieved great success in the open-source community. However, these methods rely solely on text or image guidance for video generation, lacking precise control over object or camera trajectory, which is crucial for high-quality video generation.
Trajectory Controllable Video Generation Trajectory Controllable Video Generation has recently garnered significant attention for its ability to precisely control object and camera trajectories during video synthesis. Previous methods integrate optical flow maps into video generation through a trajectory encoder. Recent works suggest using point maps as a form of guidance. MotionCtrl processes point maps with a Gaussian filter and employs trainable encoders to encode object trajectories. Trackgo represents objects using a few key points and injects this information via an encoder and a custom-designed adapter structure. Other works employ bounding box to control object trajectories. Boximator employs a trainable self-attention layer to fuse box and visual tokens inspired by GLIGEN . Some training-free methods purpose to modify attention layers or initial noised video latents to inject box signals. Additionally, certain methods explore the potential of 3D trajectories to achieve more sophisticated motion control. LeViTor employs keypoint trajectory maps enriched with depth information, while others construct custom 3D trajectories to represent object movements. However, sparse trajectories lead to imprecise control on objects shape and size, while dense trajectories are difficult for users to provide. In contrast, MagicMotion can control both dense and sparse trajectories, providing users with more flexible control over video generation.
Method
Our work mainly focuses on trajectory-controllable video generation. Given an input image , and several trajectory maps , the model can generate a video in line with the provided trajectories, where T denotes the length of generated video. In the following sections, we first provide a detailed explanation of our model architecture in Section 3.2. Next, we outline our progressive training procedure in Section 3.3. In Section 3.4, we introduce the Latent Segmentation Loss and demonstrate how it enhances the model capabilities on fine-grained object shape. We then describe our dataset curation and filtering pipeline in Section 3.5. Finally, we present an in-depth overview of MagicBench in Section 3.6.
2 Model Architecture
We utilize CogVideoX-5B-I2V as our base image to video model. CogVideoX is built upon a DiT (Diffusion Transformer) architecture, incorporating 3D-Full Attention to generate high-quality videos. As shown in Fig. 2, the model takes an input image and a corresponding video and encodes them into latent representations using a pretrained 3D VAE . Later, is zero-padded to frames and concatenated with a noised version of and then fed into the Diffusion Transformer, where a series of Transformer blocks iteratively denoise it over a predefined number of steps. Finally, the denoised latent is decoded by a 3D VAE decoder to get the output video .
Trajectory ControlNet To ensure that the generated video follows the motion patterns given by the input trajectory maps , we adopt a design similar to ControlNet to inject trajectory condition. As shown in Fig. 2, we employ the 3D VAE encoder to encode the trajectory maps into , which is then concatenated with the encoded video and serves as input to Trajectory ControlNet. Specifically, Trajectory ControlNet is constructed with a trainable copy of all pre-trained DiT blocks to encode the user-provided trajectory information. The output of each Trajectory ControlNet block is then processed through a zero-initialized convolution layer and added to the corresponding DiT block in the base model to provide trajectory guidance.
3 Dense-to-Sparse Training Procedure
Dense trajectory conditions, such as segmentation masks, offer more precise control than sparse conditions like bounding boxes but are less user-friendly. To address this, MagicMotion employs a progressive training procedure, where each stage initializes its model with the weights from the previous stage. This enables three types of trajectory control ranging from dense to sparse. We found that this progressive training strategy helps the model achieve better performance compared to training from scratch with sparse conditions.
Specifically, we adopt the following trajectory conditions across stages: stage1 uses segmentation masks, stage2 uses bounding boxes, and stage3 uses sparse bounding boxes, where fewer than 10 frames have box annotations. Additionally, we always set the first frame of the trajectory condition as a segmentation mask to specify the foreground objects that should be moving.
Our model uses velocity prediction following . Let be the initial video latents, be the gaussian noise, be the noised video latents, and be the model output. The diffusion loss can be written as:
4 Latent Segmentation Loss
Bounding box-based trajectory is able to control an object’s position and size but lacks fine-grained shape perception. To address this, we propose Latent Segmentation Loss, which introduces segmentation mask information during model training and enhances the model’s ability to perceive fine-grained object shapes.
Previous works have leveraged diffusion generation models for perception tasks, demonstrating that the features extracted by diffusion models contain rich semantic information. However, these models generally operate in the pixel space, which leads to extensive computational time and substantial GPU memory.
To incorporate dense trajectory information while keeping computational costs within a reasonable range, we propose utilizing a lightweight segmentation head to predict segmentation masks directly in the latent space, eliminating the need for decoding operations.
Specifically, our segmentation head takes a list of diffusion features from each DiT block, and outputs a latent segmentation mask . We use a light-weight architecture inspired by Panoptic FPN . Each diffusion feature first passes through a convolution layer to extract visual features . The resulting features are then concatenated and processed by another convolution layer followed by an upsampling layer to generate the final latent segmentation mask.
We compute the latent segment loss as the Euclidean distance between and the ground truth mask trajectory latents , which can be written as,
In practice, is only used in stage2 and stage3, providing the model with dense conditions information when trained with sparse conditions. In detail, we set the weight of as 0.5, and the original diffusion loss as 1. In total, our final loss function can be written as,
where is set to 0 in stage1, and 0.5 in stage2 and stage3.
5 Data Pipeline
Trajectory controllable video generation requires a video dataset with trajectory annotations. However, existing large-scale video datasets only provide text annotations and lack trajectory data. Moreover, almost all previous works use privately curated datasets, which are not publicly available.
We present a comprehensive and general data pipeline for generating high-quality video data with both dense (mask) and sparse (bounding box) annotations. As shown in Fig. 3, the pipeline consists of two main stages: the Curation Pipeline and the Filtering Pipeline. The Curation Pipeline is responsible for constructing trajectory information from a video-text dataset, while the Filtering Pipeline ensures that unsuitable videos are removed before training.
Curation Pipeline. We begin our dataset curation process with Pexels , a large-scale video-text dataset containing 396K video clips with text annotations. It encompasses videos featuring diverse subjects, various scenes, and a wide range of movements. We utilize Llama3.1 to extract the foreground moving objects from the textual annotations of each video. As shown in Fig. 3, we input the video’s caption into the language model and prompt it to identify the main foreground objects mentioned in the sentence. If the model determines that the sentence does not contain any foreground objects, it simply returns “empty” and such videos are filtered out. Next, we utilize Grounded-SAM2 , a grounded segmentation model that takes a video along with its main objects as input and generates segmentation masks for each primary object. Each object is consistently annotated with a unique color. Finally, bounding boxes are extracted from each segmentation mask using the coordinates of the top-left and bottom-right corners to draw the corresponding boxes. The color of the bounding box for each object remains consistent with its segmentation mask.
Filtering Pipeline. Many videos contain only static scenes, which are not beneficial for training trajectory-controlled video generation models. To address this, we use optical flow scores to filter out videos with little motion and dynamics. Specifically, we utilize UniMatch to extract optical flow maps between frames and compute the mean absolute value of these flow maps as the optical flow score, representing the video’s motion intensity. However, videos with background movement but static foreground can still have high motion scores. To address this, we further use UniMatch to extract optical flow scores for foreground objects based on segmentation masks and bounding boxes. Videos with low foreground optical flow scores are filtered out, ensuring MagicData includes only videos with moving foreground objects.
The trajectory annotations generated by the curation pipeline require further refinement. As shown in Fig. 3, some videos contain too many foreground objects annotations, or their sizes may be too large or too small. To address this, we regulate these factors within a reasonable range and filter out videos that fall outside the acceptable range.
Specifically, based on extensive manual evaluation, we empirically set the optical flow score threshold to 2.0, limit the number of foreground object annotations from 1 to 3, and constrain the annotated area ratio to a range of 0.008 to 0.83. The whole data curation and filtering pipeline yields us with MagicData, a high quality dataset for trajectory controllable video generation containing 51K videos with both dense and sparse trajectory annotations.
6 MagicBench
Previous works on trajectory-controlled video generation have primarily been validated on DAVIS (which has a relatively small dataset size), VIPSeg (where the annotated frames per video are insufficient), or privately constructed test sets. Thus there is an urgent need for a large-scale, publicly available benchmark to enable fair comparisons across different models in this field. To merge this gap, we use the data pipeline mentioned in Sec. 3.5 to construct MagicBench, a large-scale open benchmark consisting of 600 videos with corresponding trajectory annotations. MagicBench evaluates not only video quality and trajectory accuracy but also takes the number of controlled objects as a key evaluation factor. Specifically, it is categorized into 6 groups based on the number of controlled objects, ranging from 1 to 5 objects and more than 5 objects, with each category containing 100 high-quality videos.
Metrics. For evaluation metrics, we adopt FVD to assess video quality and FID to evaluate image quality, following . To quantify motion control accuracy, we use Mask_IoU and Box_IoU, which measure the accuracy of masks and bounding boxes, respectively. Specifically, given a generated video , we use the groundtruth masks of the first frame as input to SAM2 to predict the masks of the foreground objects in . For each foreground object, we compute the Intersection over Union (IoU) between and groundtruth masks in each frame, then average these values to obtain Mask_IoU. Similarly, we compute the IoU between the predicted and groundtruth bounding boxes for each foreground object in each frame and take the average as Box_IoU.
Experiment
We employ CogVideoX 5B as our base image-to-video model, which is trained to generate a 49-frame video at a resolution of . Each stage of MagicMotion was trained on MagicData for one epoch. The training process consists of three stages. stage1 trains Trajectory ControlNet from scratch. In stage2, Trajectory ControlNet is further refined using the weights from stage1, while Segment Head is trained from scratch. Finally, in stage3, both Trajectory ControlNet and Segment Head continue training initialized with the weights from stage2. All training experiments were conducted on 4 NVIDIA A100-80G GPUs. We employed AdamW as the optimizer, training with a learning rate of and a batch size of 1 on each GPU. During inference, we set steps to 50, the guidance scale to 6, and the weight of Trajectory ControlNet to 1.0 by default.
Datasets.
During training, we use MagicData as our training set. MagicData is annotated with dense to sparse trajectory information using the data pipeline described in Section 3.5. It comprises a total of 51,000
2 Comparison with Other Approaches
For thorough and fair comparisons, we compare our methods against 7 public trajectory controllable I2V methods . Quantitative comparison and qualitative comparison results are shown below.
To compare MagicMotion with previous works, we use the first 49 frames of each video from DAVIS and MagicBench as the ground truth video. Since some methods do not support video generation up to 49 frames in length, we uniformly sample N frames from these 49 frames for evaluation, where N represents the video length that each method support. We leverage the mask and box annotations from these selected frames as trajectory inputs for mask or box-based methods. The center point of each frame’s mask is extracted as input for point or flow-based methods .
As shown in Table 1, our method outperforms all previous approaches across all metrics both on MagicBench and DAVIS, demonstrating its ability to produce higher-quality videos and more precise trajectory control. Additionally, we evaluate each method’s performance on MagicBench based on the number of controlled objects. As shown in Fig. 4, our method achieves the best results across all object number categories, further demonstrating the superiority of our approach.
Qualitative comparison.
Qualitative comparison results are shown in Fig. 5, with the input image, prompt and trajectory provided. As shown in Fig. 5, Tora accurately controls the motion trajectory but struggles to maintain the shape of the objects. While DragAnything , ImageConductor , and MotionI2V have difficulty preserving the consistency of the original subject, resulting in substantial deformation in subsequent frames. Meanwhile, DragNUWA , LeviTor , and SG-I2V frequently produce artifacts and inconsistencies in fine details. In contrast, MagicMotion allows moving objects to follow the specified trajectory smoothly while preserving high video quality.
3 Ablation Studies
In this section, we present ablation studies to validate the effectiveness of our MagicData dataset. Additionally, we demonstrate how our progressive training procedure and latent segment loss enhance the model’s understanding of precise object shapes under sparse control conditions, thereby improving trajectory control accuracy.
Ablations on Dataset. To verify the effectiveness of MagicData, we constructed an ablation dataset by combining two public VOS datasets, MeViS and MOSE . For a fair comparison, we trained MagicMotion stage2 for one epoch using either MagicData or the ablation dataset as the training set, both initialized with the same stage1 weights. We then evaluated the models on both MagicBench and DAVIS.
As shown in Table 2, the model trained on MagicData outperforms the one trained on the ablation dataset across all metrics. Qualitative comparisons are shown in Fig. 6. In this case, we aim to gradually move the boy in the lower right corner to the center of the image. However, not using MagicData results to an unexpected child appears next to the boy. In contrast, the model trained with MagicData performs well, moving the boy along the specified trajectory while maintaining video quality.
Ablations on Progressive Training Procedure. Progressive Training Procedure allows the model to leverage the weights learned in the previous stage, incorporating dense trajectory control information when trained with sparse trajectory conditions. To validate the effectiveness of this approach, we train the model from scratch for one epoch using bounding boxes as trajectory conditions. We then compare its performance with MagicMotion stage2.
As shown in Table 3, excluding Progressive Training Procedure weakens the model’s ability to perceive object shapes, ultimately reducing the accuracy of trajectory control. Qualitative comparisons in Fig. 7 further illustrate these effects, where the model trained without Progressive Training Procedure turns the woman’s head entirely into hair.
Ablations on Latent Segment Loss. Latent Segment Loss makes the model predict dense segmentation masks while training with sparse trajectories, enhancing its ability to perceive fine-grained object shapes under sparse conditions. To evaluate the effectiveness of this technique, we train the model from stage1 for one epoch using bounding boxes as trajectory conditions and compare its performance with MagicMotion stage2. Table 3 shows that the absence of Latent Segment Loss reduces the model’s ability on object shapes, leading to less precise trajectory control. Qualitative comparisons in Fig. 8 further highlight this effect. Without Latent Segment Loss, the woman’s arm in the generated video appears incomplete.
Conclusion
In this paper, we proposed MagicMotion, a trajectory-controlled image-to-video generation method that uses a ControlNet-like architecture to integrate trajectory information into the diffusion transformer. We employed a progressive training strategy, allowing MagicMotion to support three levels of trajectory control: dense masks, bounding boxes and sparse boxes. We also utilized Latent Segment Loss to enhance the ability of the model to perceive fine-grained object shapes when only provided with sparse trajectory conditions. Additionally, we presented MagicData, a high-quality annotated dataset for trajectory-controlled video generation, created through a robust data pipeline. Finally, we introduced MagicBench, a large-scale benchmark for evaluating trajectory-controlled video generation. MagicBench not only assessed video quality and trajectory accuracy but also took the number of controlled objects into account. Extensive experiments on both MagicBench and DAVIS demonstrated the superiority of MagicMotion compared to previous works.
References
Appendix A Additional Experiments
We conducted additional experiments using MagicMotion under various task settings, including camera motion control and video editing. We also generate videos by applying different motion trajectories to a single input image.
As shown in Fig. 9, MagicMotion enables precise control over camera motion, allowing for operations such as rotation, zoom, and pan. In the first row of Fig. 9, we enclose oranges within bounding boxes and apply rotation to the boxes. This results in a video with a simulated camera rotation effect. In the second and third rows, we adjust the size of the foreground object’s bounding box to control its perceived distance from the camera, effectively achieving zoom-in and zoom-out effects. In the last two rows, we shift the bounding box to the left and downward, creating the effect of the camera moving in the opposite direction.
Video Editing.
As shown in Fig. 10, MagicMotion can be applied to video editing to generate high-quality videos. Specifically, we first use FLUX to edit the first frame of the original video, which serves as the input for MagicMotion. Then, we extract the segment mask of the original video and use it as trajectory guidance for the MagicMotion Stage1. Using this approach, we transform a black swan into a diamond swan, make the camel walk in a majestic palace, and turn a hiking backpacker into an astronaut.
Same input image with different trajectories
Extensive experiments have demonstrated that MagicMotion enables objects to move along specified trajectories, generating high-quality videos. To further showcase the capabilities of MagicMotion, we use stage2 of MagicMotion to animate objects from the same input image along different motion trajectories. As shown in Fig. 11, MagicMotion successfully animates two bears, two fish, and the moon, each following their designated paths.
Appendix B Latent Segment Masks
In this section, we provide a more detailed demonstration of Latent Segmentation Masks. Specifically, we use MagicMotion Stage 3 to predict the latent segmentation masks for each frame based on sparse bounding box conditions. As shown in Fig. 12, MagicMotion accurately predicts the Latent Segmentation Masks throughout dynamic scenes, such as a man gradually standing up to face a robot and a boy’s head slowly sinking into the water. This holds true for frames where only the bounding box trajectory is available and even for frames where no trajectory information is provided at all.
Appendix C Additional Comparison results
As shown in Table 4, we provide a comparison of the backbones used by each method, along with the supported video generation length and resolution.
Quantitative Comparisons on different object number
Due to space constraints, we only included radar charts in the main text to compare the performance of different methods in controlling varying numbers of objects on MagicBench. Here, we provide the specific quantitative results. As shown in Table 5, Table 6, and Table 7, MagicMotion consistently outperforms other methods across all metrics by a significant margin, especially when the number of moving objects is large. This demonstrates that other methods exhibit poorer performance when controlling a larger number of objects.
More qualitative comparison results.
In this section, we provide additional qualitative comparison results with previous works. As shown in Fig.13, Fig.14, Fig.15, Fig.16, and Fig. 17, MagicMotion accurately controls object trajectories and generates high-quality videos, while other methods exhibit significant defects. For fully rendered videos, we refer the reader to “Supplementary video.mp4” in supplementary material.
Appendix D Additional Ablation Results.
Here, we provide additional qualitative comparison results from the ablation study. As shown in Fig. 18, not using MagicData for training results in the generation of a woman with an extra hand. Not using the Progressive Training Procedure results in significant defects, such as a dancing woman showing severe issues when turning, with a second face appearing where her hair should be. Additionally, without the Latent Segment Loss, the woman’s lipstick is distorted into a rectangular shape.
Appendix E More Details on MagicData
Here, we provide some detailed statistical information about MagicData. On average, each video in MagicData contains 346 frames, with a typical height of 999 pixels and a width of 1503 pixels. For a more comprehensive understanding of the distribution and variability across the dataset, please refer to Fig. 20, which visualizes the detailed distribution of video frame counts, heights, and widths. During training, these videos are resized to 48 frames and converted to a 720p resolution.
Appendix F More Details on MagicBench
For evaluation purposes, all videos in MagicBench are sampled to 49 frames and resized to a resolution of 720p. MagicBench is categorized into 6 classes based on the number of annotated foreground objects. Below, we provide one video example for each category, offering a more intuitive understanding of MagicBench.