CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video Generation
Qinghe Wang, Yawen Luo, Xiaoyu Shi, Xu Jia, Huchuan Lu, Tianfan Xue, Xintao Wang, Pengfei Wan, Di Zhang, Kun Gai
Introduction
With the advances of diffusion models and large-scale pretraining paradigms, text-to-video generation (T2V) has experienced rapid development. The T2V models empower both artists and novices to create impressive videos merely by providing textual prompts. Subsequently, controllable video generation has gained increasing attention, driven by the growing demand for fine-grained control over the video creation process.
Ideally, a controllable T2V framework should grant users comparable controllability as a professional film director: allowing precise placement of objects within a scene, flexible manipulation of both objects and camera, and intuitive layout control over each rendered frame. However, existing approaches fall short of achieving this vision. Early works typically extend image-based ControlNet to the video domain, guiding generation using condition maps (e.g. depth , semantic , optical flow , or canny edge maps ). Yet, these methods generally rely on pre-existing videos to obtain condition maps, since it is non-trivial to create such condition maps from scratch. Moreover, in terms of controllability, MotionCtrl and Direct-A-Video offer preliminary control over both object and camera movements, but these methods only allow 2D object control which is very different from how filmmakers or video creators plan a shooting in 3D space. More recent works have moved towards integrating 3D-aware signals into video generation . However, these methods are only designed for image-to-video generation or rely on synthetic data from Unreal Engine.
To bridge the aforementioned gap, we present CineMaster, a framework designed for highly controllable text-to-video generation as shown in Fig. 1. Specifically, CineMaster operates in two stages, as presented in Fig. 2. The first stage is an interactive workflow that allows users to specify the requirements (conditions) of a generated video, similar to how filmmakers design a capturing plan. It allows users to describe the primary objects in a scene using a set of 3D bounding boxes with semantic labels. These bounding boxes, along with the camera, can be repositioned across keyframes, allowing users to orchestrate complex motion dynamics. After each modification, CineMaster provides a preview of the rendered frames for iterative refinement until the desired rendered effects are achieved. In the second stage, we finetune a text-to-video diffusion model to generate a video conditioned on the control signals provided in the first stage. Crucially, in addition to camera trajectories and user-provided class labels, we propose to utilize the rendered depth maps of all frames as augmented visual cues. These depth maps explicitly contain the desired 3D layout of each frame, serving as strong guidance for the diffusion model to generate the user-intended video content.
One challenge for this design is the lack of videos with ground-truth 3D bounding box and camera trajectory annotations. To solve this limitation, we further propose an automatic data labeling pipeline, as illustrated in Sec 3.3. Using this pipeline, we build the largest video datasets with both the ground-truth 3D bounding box and 3D camera trajectory annotations.
Finally, to evaluate the controllability of our proposed framework, we conduct extensive experiments, comparing it with existing SOTA methods, and performing ablative studies to validate the effectiveness of our core modules.
Related Work
Controllable Video Generation via Planar Condition Maps. The pioneer works ControlNet and T2I-Adapter introduce the paradigm of conditioning generation on planar maps in the image generation field. Subsequently, many works extend this paradigm to the video domain using different condition maps, e.g. depth maps , human pose maps , semantic maps and optical flow maps . However, they generally assume the existence of such condition maps, but it is indeed non-trivial to create precise condition maps from scratch, especially for novices. Therefore, we carefully design an interactive workflow to help users obtain 3D-aware condition maps in an intuitive way. We also get inspiration from LooseControl , to use 3D bounding box as an appropriate abstract representation of objects in the scene.
Object Motion Control. Previous methods primarily focus on motion control in 2D space. MotionCtrl , DragNUWA , and Tora represent object motion trajectories as sequences of spatial positions, encoding coordinates into dense control maps. Additionally, 2D bounding boxes have been adopted as control signals to enable flexible motion generation in Direct-A-Video and Boximator . Motion control through sketches has also been explored in VideoComposer . While these methods have demonstrated capabilities in object motion control, their control signals limit the controllability only in 2D space. 3DTrajMaster is the first to use 6D pose sequences of objects to control object motion in 3D space.
Camera Motion Control. Camera pose serves as a crucial control signal in video generation, determining which portions of the scene are captured and presented in the final output. MotionCtrl pioneers the integration of camera poses as control signals for camera movement manipulation. Building upon this foundation, CameraCtrl introduces the use of Plücker embeddings of camera poses to enhance motion controllability. These methods are all trained on an indoor dataset RealEstate10K for learning camera motion which limits the ability to generalize to in-the-wild scenes. Direct-A-Video uses data augmentations on static videos to simulate basic camera movements (only pan and zoom movements) and employs Fourier embedder and temporal cross-attention layers to inject camera poses. It cannot generalize to more complex camera movements such as “Anti-clockwise”. Therefore, the development of this field is constrained by the scarcity of large-scale in-the-wild datasets with camera pose annotations.
Joint Motion Control. Based on the preliminary explorations of joint motion control , some concurrent works further advance this field. Motion Prompting leverages 2D point tracking results to represent object motion and camera motion. Perception-as-Control and DaS further capture 3D point tracking results by SpatialTracker to extend the joint control into 3D space. The former denotes objects of reference image as unit spheres with different colors. The latter directly uses the point tracking video as the motion condition. MotionCanvas also measures camera motion by point tracking and renders 2D instance box map as object global motion representation. However, these methods are all designed for the image-to-video generation task which only animates an initial image and cannot plan a shooting in 3D space from scratch. SynFMC uses Unreal Engine to render both 6D pose of objects and camera to construct video datasets for joint motion control, but the limited diversity and the domain gap of UE data restrict the model’s generalizability. To overcome the scarcity of in-the-wild datasets with both 3D object motion and camera pose annotations, we carefully establish an automated data annotation pipeline to extract 3D bounding boxes and camera trajectories from large-scale video data for learning joint motion control. In addition, we provide an interactive workflow to allow users to intuitively manipulate objects and camera in 3D scene.
Method
Controllable text-to-video generation (T2V) targets at providing more conditional guidance beyond textual prompts, thereby enabling fine-grained control over the video generation process. We present CineMaster, which aspires to give users a level of controllability comparable to professional film directors: allowing precise object placement within a scene, flexible manipulation of both objects and camera, and intuitive layout control over each rendered frame. Our proposed CineMaster operates in two stages. In the first stage (Sec 3.1), we propose an interactive workflow for constructing 3D-aware control signals. In the second stage (Sec 3.2), these control signals serve as conditions for T2V models to synthesize the desired video content. Moreover, due to the scarcity of large-scale datasets with 3D bounding box and camera trajectory annotations, we carefully develop an automated data labeling pipeline (Sec 3.3).
The first stage centers on constructing 3D-aware control signals through a user-friendly workflow. Inspired by LooseControl , we employ 3D bounding boxes as the principal form of object representation. Users can freely adjust the size and position of these bounding boxes within the 3D scene. By repositioning bounding boxes and camera across keyframes, users gain intuitive control over object and camera trajectories, effectively dictating the motion dynamics. Another key component of our workflow is the preview mechanism, which lets users examine rendered frames after each modification.
This workflow closely mirrors real-world filmmaking: directors typically arrange actor and camera movements in multiple takes, reviewing footage on monitors to refine the final shots. Once satisfactory rendering effects are achieved, we export camera trajectories and per-frame projected depth maps for use in the subsequent stage. The primary advantage of this system is its 3D-native and intuitive nature. We implement the interactive system using the open-source engine Blender, where users select keyframes for object and camera placement. The system then automatically interpolates trajectories for intermediate frames, providing a seamless and efficient workflow for complex scene setup.
2 Stage 2: Conditional Video Generation
We condition a base text-to-video model on the control signals derived from the first stage. Crucially, beyond using the camera trajectory and object labels as inputs, we also introduce projected depth maps of each frame as augmented visual condition. These depth maps explicitly encode the desired 3D layout, providing strong guidance for the diffusion model to generate accurate video content. To effectively integrate these additional inputs into the T2V model, we design two key components: a semantic layout injection module and camera adapter, as illustrated in Fig. 3.
Base Model. Our model is developed upon a pretrained text-to-video foundation model, which consists of a 3D Variational Auto-Encoder (VAE) , T5 encoder and a transformer-based latent diffusion model . Each basic transformer block is instantiated as a sequence of 2D spatial self-attention, 3D spatial-temporal self-attention, text cross attention and feed-forward network (FFN). The text prompts are encoded as by T5 encoder to guide the generation model. We define a straight forward path between clean data and noised data at timestep with Rectified Flow :
where . The denoising process is defined as a mapping from to by an ordinary differential equation (ODE):
where the velocity is parameterized by the weights of the denoising network. The training process is supervised by Conditional Flow Matching to regress velocity:
While the proposed Semantic Layout ControlNet enables precise control over the 3D position of each entity in generated videos, relying exclusively on 3D bounding boxes might introduce ambiguity between object and camera movements. For instance, if the bounding box of a balloon shifts upward, it could indicate that the ballon is rising, the camera is moving downward, or both. To resolve this ambiguity, we additionally inject explicit camera poses into the generation process, allowing the model to distinguish object motion from camera trajectories more reliably.
3 Dataset Labeling Pipeline
There is a lack of large-scale video datasets with 3D bounding box and camera pose annotations. To train the second-stage network, we design an automated data labeling pipeline as shown in Fig. 4. It takes an in-the-wild video as input and extracts the required class labels, camera trajectories and projected depth maps.
Class Labels. To obtain the class labels for objects present in the scene, we perform instance segmentation for each entity in the video. To achieve open-set instance segmentation, we combine Grounding DINO with SAM 2 , where Grounding DINO produces D bounding boxes guided by entity descriptions. To enhance foreground entity detection, we utilize the multi-modal large model Qwen2 to generate entity descriptions as guidance for Grounding DINO. This process yields D bounding boxes and class labels for each entity in the first frame, which then guides SAM 2 for video segmentation. To address potential issues with overlapping boxes and incorrect class labels from Grounding DINO, we implement crucial post-processing steps: a box IOU filter and feature similarity verification between the regions within D boxes and their assigned class labels.
Camera Trajectories. We employ the SOTA camera pose estimation model MonST3R to obtain camera trajectories throughout the video sequence.
Projected Depth Maps. We employ DepthAnything V2 to generate metric depth maps for the entire video sequence, which are essential for the subsequent inverse projection process. The third step involves inverse projection to obtain D boxes for each entity. We operate under the assumption that each entity maintains a constant volume in the D scene. To address cases where entities may appear partially in certain frames, we identify the optimal frame index for each entity, typically when the entity is most completely visible, to ensure accurate inverse projection and adequate volume representation. For each entity, we combine the instance segmentation mask with the corresponding metric depth map at this optimal frame to generate a D point cloud, from which we derive the minimal-volume D bounding box.
Following the establishment of maximum D boxes for all entities, the final step involves computing temporal-spatial transformations of these boxes within the D scene for all frames. We conduct 3D point tracking by SpatialTracker starting from the optimal frame of th object to the rest frames , and the average inter-frame displacements of all feature points from each object are regarded as the spatial movement of the th object’s 3D box where denotes th frame of th object. Then we can compute the 3D boxes of all objects in all frames. To represent 3D boxes as explicit control signals, we further project the constructed 3D boxes into image space and render depth maps.
Experiments
Training Paradigm. We design a dedicated training strategy consisting of three stages, i.e., 1) training a DiT-based ControlNet on dense depth maps, 2) adapting ControlNet to 3D box datasets, and 3) jointly training Semantic Layout ControlNet and Camera Adapter. Specifically, following LooseControl , we first train our DiT-based ControlNet on 167K videos crawled from the Internet with dense depth maps labeled by DepthAnything V2 . Subsequently, we use our data annotation pipeline to construct a 3D box dataset with 156K videos and 118K images for training Semantic Layout ControlNet. The images are collected from COCO and Object365 which provide more categories and precise instance segmentation annotations. With the metric depth maps measured by DepthAnything V2 , we could obtain the 3D boxes of image datasets by calculating the 3D box with minimal volume for each object in the point cloud. By the image-video joint training, we integrate the spatial layout and semantic information into the Semantic Layout ControlNet which could guide the base model to generate box-aligned videos with specified class labels. We further annotate 99.6K videos out of the 156k videos using our proposed pipeline to obtain camera poses. We also utilize RealEstate10K dataset which features larger camera motion, resulting in 10.4K data samples. We sample data between our dataset and Real-Estate10K with 3:1 ratio for training to enhance the learning of larger camera motion. Based on the merged video dataset with both 3D box and camera pose annotations, we train the Semantic Layout ControlNet and Camera Adapter jointly and master the joint controllability of object motion and camera motion for flexible customized video generation.
Implementation Details. We train CineMaster based on our internal text-to-video generation model with parameters for research purposes. Following NaViT , we pad the videos to the same shape for each batch managed by attention masks during training. Each training video segment contains 77 frames (i.e., 5 seconds) sampled with 15 frames per second (fps). We use Adam optimizer and train on 24 NVIDIA A800 GPUs, with a batch size of 4 and a learning rate of . The three stages of the training process consist of 12,000, 7,000, 6000 steps respectively. During inference, we set the scale of classifier-free guidance as 12.5 and the DDIM steps as 50. We make a trade-off between object motion and camera motion by injecting semantic layout information and camera poses with 25 and 15 steps respectively.
Baselines. We compare CineMaster with existing SOTA methods MotionCtrl and Direct-a-Video which could also control object motion and camera motion simultaneously. To align with different input requirements, we convert our 3D box condition into object trajectories for MotionCtrl and 2D bounding box sequences for Direct-A-Video. In addition, we empirically align the coordinate system and scale of the input camera poses for comparison.
Evaluation Metrics. 1) Object-box alignment: We use Grounding DINO to detect 2D object boxes in generated videos to measure the mean Intersection over Union (mIoU) and measure the trajectory deviation (Traj-D) by calculating the difference of the center points against ground truths. In addition, to evaluate the depth control accuracy of the generated objects, we calculate the average depth of the object regions in each generated frame using SAM 2 and DepthAnything V2 , and measure the depth deviation (Depth-D) by the Root Mean Squared Error (RMSE) with the depth values of the given 3D boxes. 2) Video quality: We employ Fréchet Video Distance (FID) , Fréchet Inception Distance (FID) and CLIP Similarity (CLIP-T) to evaluate the appearance of generated results.
2 Comparison with Other Methods
Qualitative Comparison. As shown in Fig. 5, we show three different feature comparisons: moving object static camera, static object moving camera and moving object moving camera. In the first setting, MotionCtrl moves the camera rightward to align the object trajectory, but fails to either maintain the camera stationary or make the object move. It shows that there is still the camera motion and object motion coupling issue in MotionCtrl. In the third setting, since MotionCtrl is unable to associate multiple trajectories with their respective objects, it generates the “McLaren” that appears to follow the trajectory of the “person” and fails to generate the “person”. Direct-A-Video presents low-quality textures for generating “bus” and “rock” which demonstrates that its control disturbs the generation quality. It exhibits weaker camera movement and box alignment, and produces unexpected shot changes and more artifacts in the third setting. In comparison, the proposed CineMaster performs the best control performance for the control of object motion and camera motion in three settings.
Quantitative Comparison. In addition, we further report the quantitative comparison in Table 1. Since MotionCtrl only uses point trajectory sequences to specify the positions of generated objects and ignores the spatial size, we do not calculate its mIoU. It does not explicitly associate multiple trajectories with their respective objects and suffers from the camera motion and object motion coupling issue. Therefore, it obtains unsatisfactory Traj-D. Direct-A-Video only trains for learning camera movement, and controls the object motion by spatial cross-attention modulation with associated object words and box trajectories to guide the spatial-temporal placement of objects only for inference. The training-free object motion control biases the vanilla inference distribution, resulting in the degradation of generation quality. In addition, no joint training of object motion and camera motion control leads to a gap between training and inference, so it obtains weak mIoU and Traj-D. In contrast, we construct the video dataset with both 3D box and camera pose for joint training, and the proposed CineMaster could harmoniously control the object motion and camera motion simultaneously. Therefore, CineMaster outperforms previous SOTA methods on all metrics. In particular, we achieve significantly higher mIoU and Traj-D, indicating that our framework can generate videos that better follow the user’s spatial design.
3 Ablation Study
We experiment with different training paradigms to validate the effectiveness of the delicate designs in our workflow:
“w/o stage 1”: without training the DiT-based ControlNet on dense depth maps, the training process starts directly from the second training stage.
“w/o semantic”: without semantic injector, this setting does not specify the class labels of 3D boxes.
“Isolated S,C”: Semantic Layout ControlNet and Camera Adapter are trained separately and used together for inference.
“Fix S train C”: this setting first trains Semantic Layout ControlNet to convergence, then freezes its weights and trains the Camera Adapter.
“Joint Train” (our final version): Semantic Layout ControlNet and Camera Adapter are trained simultaneously.
As shown in Table 2, the setting of “w/o stage 1” lacks the fine-grained perception for depth control signal, so it obtains mediocre Depth-D. The positions of the generation objects can only be specified via text prompt in the setting of “w/o semantic”, resulting in the poor mIoU, Traj-D and CLIP-T. The setting of “Isolated S,C” faces a discrepancy between training and inference phases, as the Semantic Layout ControlNet and Camera Adapter are trained separately without cross-module communication, resulting in degraded generation quality supported by lower FVD and FID. Although the setting of “Fix S train C” trains the camera adapter based on the frozen Semantic Layout ControlNet, it still fails to eliminate the coupling of camera motion and object motion already learned in the frozen Semantic Layout ControlNet, leading to suboptimal FVD and FID. Benefiting from the constructed video dataset with 3D box and camera pose labels, we experimentally observe that the setting of “Joint Train” could harmoniously integrate the control for camera motion and object motion and perform the best results on all metrics.
Limitations and Conclusions
Ideally, 3D bounding box can naturally and precisely control the orientation of objects in space. For instance, when we rotate the 3D box of a human, it should produce a video sequence of a human turning around. However, the community currently lacks accurate open-set object pose estimation models. Therefore, we leave this promising functionality as future work. In conclusion, our project stems from the goal of granting users the creation controllability as professional film directors. To this end, we propose CineMaster for highly controllable text-to-video generation. Specifically, we first design a 3D-native workflow that allows users to manipulate objects and camera in an intuitive manner. Then we train a conditional text-to-video diffusion model to synthesize the user-intended videos. We emphasize the importance of adopting projected depth maps as strong visual control signals. Extensive experiments demonstrate that CineMaster achieves controllable and 3D-aware cinematic video generation.