Towards Language-Driven Video Inpainting via Multimodal Large Language Models

Jianzong Wu, Xiangtai Li, Chenyang Si, Shangchen Zhou, Jingkang Yang, Jiangning Zhang, Yining Li, Kai Chen, Yunhai Tong, Ziwei Liu, Chen Change Loy

Introduction

Video inpainting, a technique for restoring missing or corrupted segments in video frames, finds extensive application in areas such as video completion , video restoration , and object removal . Despite continuous advancements in enhancing image quality and temporal coherence of inpainting results , current methods predominantly depend on manually annotated binary masks to identify restoration areas. This manual process is time-consuming and impractical for long videos. While automatic labeling tools, such as segmentation and object tracking models , offer some relief, they often necessitate manual refinement due to imperfect labeling.

Perhaps a more natural way to perform video inpainting is through natural language, as shown in Fig. 1. The task would become much easier if we could leverage natural language descriptions to specify the inpainting areas, like “woman on the left,” thereby preventing the need for pixel-level manual annotations. Importantly, the language-driven setting can benefit from the flexibility of natural language. For example, with richer sentences, one can easily refer to multiple or abstract objects, which is much more effective than labeling masks. Extending from this notion, one could divide the task into two subtasks, namely “Referring Video Inpainting” and “Interactive Video Inpainting.” The former takes simple referring expressions as inputs, and the latter considers more complex conversation-like interactions to accomplish the inpainting task.

To establish a baseline model for the proposed tasks, it is essential to have an appropriate dataset for both training and evaluation. Currently, no publicly available dataset comprises the triplets of original videos, removal expressions, and inpainted videos. In response to this gap, we build a new dataset named the Remove Objects from Videos by Instructions (ROVI) dataset. Specifically, we employ referring object segmentation datasets, which are pre-annotated with object masks and descriptive expressions. These datasets are further augmented with corresponding inpainted videos generated using a state-of-the-art video inpainting model. We find existing referring expressions for interactive video inpainting tasks too simplistic. To address this limitation, we employ Multimodal Large Language Models (MLLMs) to create conversation-like dialogues. These dialogues are designed to simulate real-world scenarios, encompassing user requests and corresponding machine responses. This approach enriches the dataset, making it more representative of the complexity and variability found in practical video inpainting applications.

In addition to the dataset, we introduce the first end-to-end baseline model, Language-Driven Video Inpainting (LGVI), for the proposed tasks. Our model is built upon diffusion-based generative models. In particular, we inflate the text-to-image (T2I) model to become a text-to-video (T2V) architecture by extending the parameters with an additional temporal dimension. We propose an efficient visual conditioning approach that minimally increases the number of parameters. To further enhance our model’s capabilities for the interactive task, we extend the LGVI framework to LGVI-I (Interactive). This extension incorporates an MLLM specifically designed to process and understand user requests phrased in a conversation-like format. The LGVI-I model is trained in an end-to-end manner. This interactive architecture enables the system to interpret complex instructions accurately. As a result, it can produce appropriate inpainting results and relevant responses within the conversational context, thus paving the way for more intuitive and flexible user interactions with interactive video inpainting systems.

In summary, our key contributions are as follows:

We introduce a novel language-driven video inpainting task, significantly reducing reliance on human-labeled masks in video inpainting applications. This task includes two distinct sub-tasks: referring video inpainting and interactive video inpainting.

We propose a dataset to facilitate training and evaluation for the proposed tasks. This dataset is the first of its kind, containing triplets of original videos, removal expressions, and inpainted videos, offering a unique resource for research in this domain.

We present a diffusion-based architecture, LGVI, as a baseline model for the proposed task. We show how one could leverage MLLMs to improve language guidance for interactive video inpainting. To our knowledge, it is the first model to perform end-to-end language-driven video inpainting.

Related Work

Video inpainting. Video inpainting is a technique aimed at restoring or filling missing or corrupted parts in a video plausibly. While related to image inpainting methods , video inpainting techniques extend the problem to the more complex domain of moving visuals. This technique can be applied for various applications, such as object removal, visual restoration, and completion. With the advent of deep learning, visual inpainting networks usually employ convolutional neural networks (CNNs) and generative adversarial networks (GANs) . Recent works also apply vision Transformers to enhance the global interaction among visual features . State-of-the-art methods show strong abilities in restoring missing parts and removing objects. Most of these works require the input of a binary mask to define the restoring area . However, the generation of object-like masks, particularly for lengthy videos, poses a significant and labor-intensive challenge,

Language-driven visual editing. Diffusion-based text-to-image generation models show excellent abilities in generating images and videos following text guidance. Recent studies also achieve image editing , image segmentation and video editing with DMs. Among these, Prompt2Prompt manipulates the cross-attention maps within the model to enable various editing operations such as object modification, addition, and style transfer. InstructPix2Pix leverages this approach to create an image editing dataset. Similarly, Tune-A-Video proposes a training-free architecture to edit videos by language references. However, these works are intended for general visual editing. When applied to more specific challenges, such as language-driven video inpainting, they tend to yield suboptimal results. Figure 2 shows two examples where these models produce inferior results when instructed to remove objects. A few works have explored the image inpainting task using DMs. Repaint still takes the image and mask as input and lets the DM restore the original image. SmartBrush takes mask and text as input to guide a region-controlled generation on the masked area, which aims to generate new concepts rather than remove the object. Recently, Inst-Inpaint proposes a method to perform object removal on images based on the language descriptions. Despite its innovative approach, Inst-Inpaint’s training samples are constrained by a lack of interactive expressions and video resources, which limits its practical effectiveness in complex scenarios.

Multi-Modal Large Language Models. Large language models (LLMs) have demonstrated exceptional performance across a variety of text-based tasks and applications . Recent works extend the capabilities of LLMs to include image processing and computer vision. A notable example is LLaVA , which translates image tokens into a language feature space, thereby transforming the fine-tuned model into an MLLM. This adaptation enables MLLMs to interpret and understand visual content. Subsequent research uses MLLMs in diverse applications, including visual reasoning, object detection, and segmentation . To the best of our knowledge, this study is the first to integrate MLLMs into the domain of language-driven video inpainting.

ROVI Dataset

Table 1 summarizes the differences between ROVI and related datasets. In image and video inpainting research, prevalent training and evaluation datasets mainly include vision-centric ones like Places , CelebA , YouTube-VOS , and DAVIS , without human annotations. These datasets typically employ random masking in training to simulate missing areas for inpainting. However, for object removal tasks, specifically labeled masks are essential. While YouTube-VOS provides object masks, it lacks corresponding inpainting ground truths. The GQA-Inpaint dataset , although rich in object expressions and inpainting results, is limited to image data and does not accommodate video or interactive contexts. Our ROVI dataset addresses these limitations with comprehensive annotations covering a wide array of regions, including object masks, referring expressions, inpainting results, and conversation-like dialogues. Unlike Places and CelebA, which focus on specific image categories like buildings or faces, ROVI encompasses a broader spectrum of general scenes, making it more adaptable for diverse inpainting applications.

2 Dataset Statistics

Figure 3 presents a comprehensive statistical analysis of the ROVI dataset. The dataset encompasses 2,967 videos from A2D-Sentences and 2,683 videos from Refer-YouTube-VOS, divided into train and test splits, as shown in Fig. 3(a). Figure 3(b) illustrates the diversity of referring expressions with word clouds. Figure 3(c) shows several examples of our dataset’s interactive requests and responses, showing the diversity and complexity of the dialogues. Figure 3(d) details the relative sizes of segmentation masks (mask area divided by image area). We drop objects with a relative size larger than 0.25 because the inpainting results for large objects usually have worse qualities. Figure 3(e) analyzes the distribution of object motion. Figure 3(f) delivers a histogram of sentence lengths within the dataset.

3 Dataset Construction Pipeline

Video data selection. As depicted in Fig. 4, we have chosen referring video object segmentation datasets for the source of video data. Referring video object segmentation aims to segment an object referred to by a given language description. These datasets have pre-annotated object masks and descriptive expressions, making them well-suited for the proposed task. Specifically, we select Refer-YouTube-VOS and A2D-Sentences as our data sources.

Annotation pipeline. We use a video inpainting model to generate the inpainting ground truth. Specifically, we choose E2FGVI , a state-of-the-art video inpainting model, to produce the inpainting results. This model, trained on video data, guarantees temporal consistency in the inpainting results.

To further ensure the ground truth is of high quality, we incorporate a human selection process on the hyperparameter of the inpainting method. Specifically, the input mask can be expanded with different pixel sizes, denoted as dd. The bigger the dd is, the larger the input mask is developed so that it may cover the whole object. The best dd value varies through objects, causing an unstable performance in the inpainted videos if set to a fixed value. Therefore, throughout the data generation process, we experiment with various hyperparameters to generate multiple results for each object and involve human annotators to select the best result. More details are provided in Appendix A.

For interactive annotations, we need to collect expressions through chat-style conversations. Unlike the straightforward “remove” sentences, these interactive requests should be implicit, necessitating the model to discern the user’s underlying intent. Rather than relying solely on human annotators to articulate these requests, we explore a more automated approach: we employ LLMs and MLLMs to simulate a human user and generate potential requests and responses. We propose a multi-step approach with details illustrated in Fig. 4. By employing this dual-faceted annotation pipeline, the ROVI dataset is enabled to handle complex user requests.

Methodology

In this section, we introduce our Language-Driven Video Inpainting framework (LGVI) and the MLLM-enhanced LGVI-I (Interactive) architecture.

where Wq\mathbf{W}_{q}, Wk\mathbf{W}_{k}, and Wv\mathbf{W}_{v} are learnable matrices to project the inputs to query, key, and value. The computational complexity of the Temporal Attention module is O(CT2)\mathcal{O}(CT^{2}), while spatial self-attention has a complexity of O(CD2)\mathcal{O}(CD^{2}). Considering T≪DT\ll D. The Temporal Attention module is a time-efficient tool to ensure the consistency of output video sequences.

Diffusion models learn to gradually remove noises in a noised video. During training, the target video Y\mathbf{Y} is added with noises to be the start point of the noised video. Besides the noised target video input, LGVI also takes the original video X\mathbf{X} as a control signal input. Concretely, we encode the original video X\mathbf{X} to the latent space and concatenate its feature with the noised target video in the channel dimension. Note that the noise is added only to the target video latent, and during testing, the noised target video is a randomly generated noise.

where Ldiff\mathcal{L}_{diff} and Lmask\mathcal{L}_{mask} are the diffusion model training objective and mask loss, respectively; c\mathbf{c} is the language guidance features from the referring expressions; M^\hat{\mathbf{M}} is the mask prediction and M\mathbf{M} is the ground truth mask. The parameters λ1\lambda_{1} and λ2\lambda_{2} are loss weights to balance training.

2 LGVI-I with MLLM

In the interactive video inpainting task, models are expected to extract valuable information from complex conversations. To overcome this problem, we propose incorporating MLLMs to extend the LGVI from work to LGVI-I (Interactive). MLLMs have demonstrated strong capabilities in visual comprehension and multimodal reasoning, making them well-suited for our proposed interactive video inpainting task. As shown in Fig. 5, the MLLM takes both the image frame and the chat-style user request as inputs, generating the language response and a pair of special indicators: <<PROMPT>> and <</PROMPT>>. The hidden states of the last layer between these two indicators are then passed through an MM head, implemented as a two-layer linear layer with activation functions. The transformed features are fed to the cross-attention module to guide the U-Net inpainting process. Mathematically, given the input video X\mathbf{X} and user request ss, the computation pipeline of the MLLM can be summarized as follows:

where hp\mathbf{h}_{p} is the MLLM-enhanced language condition, Llm\mathcal{L}_{lm} is language modeling loss, implemented as the Cross-Entropy Loss, w\mathbf{w} is the ground truth sentence, and λ3\lambda_{3} is the loss weight for language modeling loss. By integrating an MLLM into the LGVI framework, the system achieves a higher level of user interactivity. This enables users to guide the visual inpainting process with interactive language instructions, thus establishing a more user-friendly and accessible approach for the interactive video inpainting task.

Experiments

Datasets and metrics. We use the ROVI dataset test set for both the referring video inpainting and interactive video inpainting tasks. The test set contains 478 videos and 758 objects, each equipped with one referring expression and one interactive expression. During the training of our models, we incorporate a referring image inpainting dataset, GQA-Inpaint , to enrich the data vocabulary. We follow video inpainting works to use PSNR and SSIM to assess the statistical similarity between predicted results and ground truth. Additionally, we use VIFD to measure the perceptual similarities. To assess the temporal consistency and smoothness of the generated videos, we also apply the EwarpE_{warp} metric .

Baselines. For the baselines, we select three language-driven image editing methods: InstructPix2Pix , Inst-Inpaint , and MagicBrush . It is worth noting that InstructPix2Pix and MagicBrush are pre-trained on extensive image editing datasets. Inst-Inpaint is proposed to perform referring image inpainting on images. We also compare with Inpaint Anything, a multi-stage method for one-click video inpainting. It uses Segment Anything and OSTrack to produce segmentation masks based on user click, followed by inpainting the masked area using inpainting models . We implement Inpaint Anything*, which facilities Inpaint Anything with GroundingDINO , enabling it to process language inputs.

Implementation details. We initialize the U-Net weights from MagicBrush . The newly introduced modules are trained from scratch. During training, we sample video and image data at a ratio of 3:13:1. For the MLLM, we adopt LLaVA-7B . The learning rates are 3e-5, 1e-4, and 1e-4 for U-Net, mask decoder, and tuned parameters in MLLM, respectively. The loss weights are set to λ1=1\lambda_{1}=1, λ2=0.001\lambda_{2}=0.001, λ3=0.1\lambda_{3}=0.1. These weights differ significantly due to the different types of losses they represent. The input and output video sizes are set to 512 ×\times 320, and the video length is 24. For LGVI, we train 50 epochs on the ROVI dataset with a batch size of 32 for videos and 768 for images. For LGVI-I, we load the LGVI checkpoint and fine-tune it for 50 epochs under the same batch size. All experiments are carried out on 8 NVIDIA A100 GPUs.

2 Referring Video Inpainting

Quatitative results. We report quantitative results on the referring video inpainting task. Compared with baseline models, our model is the first one-stage language-driven video inpainting model. As shown in Tab. 2, our model outperforms MagicBrush in all metrics and achieves on-par results with Inpaint Anything* , even if Inpaint Anything* uses a mask-based inpainting model . The results demonstrate the effectiveness of our model.

Qualitative results. Figure 6 shows qualitative results. We compare with MagicBrush , a robust generalized language-driven image editing model. In the first example, where the language refers to the cat on the right, the MagicBrush model removes all the cats in the scene, while our model successfully inpaints the right cat. In the second example, the referring expression becomes more complex. MagicBrush struggles to identify the object requiring inpainting and removes the wrong object (the balls) in the last frame. In contrast, our model generates a plausible output, demonstrating its superior performance in handling complex language-driven inpainting tasks. Furthermore, in Fig. 7, we compare with Inpaint Anything* on sentences referring to multiple objects or nonexistent objects. Inpaint Anything is driven by a simple combination of referring segmentation and video inpainting models. Thus, it is fixed to produce one mask for each sentence. When referring to multiple or nonexistent objects, it outputs inaccurate results, while our model produces the correct output. This demonstrates the robustness of the language-driven setting.

3 Interactive Video Inpainting

Quatitative results. We report the interactive video inpainting task results in Tab. 3. As shown in the top 5 rows, when the models are trained using referring expressions, their performance drops correspondingly in this task. This is intuitive because interactive expressions are much longer and more implicit. For the MLLM-Enhanced Two-Stage Models, we combine the baseline models with an MLLM in a zero-shot manner. The interactive inputs are transferred into shorter referring expressions by simply prompting the MLLM. These models exhibit improved performances. Our LGVI-I model achieves the highest performance, demonstrating the effectiveness of the proposed architecture.

Qualitative results. Figure 8 presents examples of the interactive video inpainting task. The user requests pose a significant challenge and complexity for the baseline models to comprehend. In particular, Inpaint Anything* predicts incorrect masks, leading to inaccurate results. Similarly, other diffusion-based models struggle to interpret the users’ intentions accurately, resulting in less satisfactory outcomes. In contrast, our LGVI-I model, which harnesses the power of MLLM, consistently delivers the best performance in these challenging scenarios. This underscores the superiority of our proposed approach.

4 Ablation Study

We conduct three ablations involving mask supervision, fine-tuning the entire U-Net, and joint training with images.

The benefit of mask supervision. The models without mask supervision rely solely on the inpainting ground for guidance, lacking an explicit signal to direct the inpainting area. As shown in Fig. 9, the running man remains present in all frames. Notably, although we utilize mask annotation in the ROVI dataset for training, the LGVI framework does not need mask input during inference.

The benefit of fine-tuning the whole U-Net. We fine-tune the whole U-Net instead of only adjusting the parameters of new modules. As shown in Fig. 9, limiting the fine-tuning to only the new modules hinders training convergence, and the model struggles to output expected results. This demonstrates the necessity of tuning the whole model.

The benefit of joint training with images. As shown in Fig. 9, the model without joint training produces outputs with heavier artifacts than ours. This is because enlarging the dataset brings more diversity of objects and scenes to the model. It demonstrates the effectiveness of joint training.

Conclusion, Challenges, and Outlook

In this paper, we propose a novel language-driven video inpainting task that uses language to guide inpainting areas. For training and evaluation, we collect a video dataset, namely ROVI. Comprehensive statistics demonstrate the uniqueness and diversity of our dataset, especially the chat-style interactive conversations generated by powerful LLMs and MLLMs. We further propose a diffusion-based baseline model, LGVI. Quantitative and qualitative experimental results show the effectiveness and robustness of our model.

Challenges. (1) Handling Ambiguity in Language Descriptions. Language-driven video inpainting relies heavily on the accuracy and clarity of language inputs. Ambiguities or vagueness in language descriptions can lead to inaccuracies in inpainting results. Developing models that can intelligently handle or clarify ambiguous language inputs is a significant challenge. (2) Real-Time Processing. Video inpainting in a real-time setting, especially with complex language-driven inputs, is computationally demanding. Diffusion-based models also experience the slow inference problem due to the Markov denoising process. Improving the speed and efficiency of these models without compromising accuracy is a crucial challenge. (3) Scalability and Generalization. Another challenge is ensuring that the model generalizes well across various languages and video types. Models might perform well on the dataset they were trained on but struggle with new, unseen data.

Future work. A promising direction is to resolve ambiguities in language inputs, possibly by using contextual clues from the video or previous language inputs. In addition, researching methods to optimize these models for real-time video inpainting, could be valuable for live broadcasting or interactive media. Another important future work is incorporating interactive user feedback mechanisms that allow the system to learn from corrections or preferences indicated by users, thereby improving the accuracy and relevance of the inpainting results over time. See supplementary for more discussions.

Potential social impacts. This technology could potentially boost creative fields such as film-making, advertising, and content creation. It allows for more seamless editing and creative storytelling, enabling creators to modify and enhance their visual narratives easily. It also comes with negative impacts. For example, it can be used in creating misleading or false media and ethical or moral issues.

Appendix

Overview. Our supplementary includes the following sections:

Appendix A. Details for our dataset annotation process.

Appendix B. Implementation details of the baseline models.

Appendix C. Quantitative ablation study results.

Appendix D. More qualitative results of different models.

Appendix E. The foundations of Latent Diffusion Models and correlation with our model.

Appendix F. Discussions of limitations and challenges.

Video Demo. We also include the video introduction of our work, which shows the visualization demo.

Appendix A Dataset Annotation Details

When generating inpainting results using a video inpainting network , the input mask can be trickily expanded with different pixel sizes, denoted as dd. The bigger dd is, the larger the input mask is developed so that it may cover the whole object. The best dd value varies through objects, causing an unstable performance in the inpainted videos if set to a fixed value. Therefore, throughout the generation process, we experiment with various hyperparameters to generate multiple results for each object and involve human annotators to select the best result. In particular, we generate six samples with d∈d\in for each object. Human annotators are expected to choose the best-looking result in these examples. The object does not enter the ROVI dataset if all examples are evaluated as unqualified. This human labeling process guarantees the high-quality ground truth of the ROVI dataset. Fig. 10 shows an illustration of the human annotation interface.

Additionally, we find there are several misannotations in the Refer-YouTube-VOS dataset. In some videos, the object label matches wrongly with object masks and expressions. For example, a person may correspond to “a dog walking by a person” according to the object labels. So, human annotators also make necessary revisions on the expressions in case of mis-annotations.

A.2 Interactive Annotation Details

In the interactive annotation pipeline, all the generating processes are completed by prompting MLLMs and LLMs without fine-tuning. Firstly, we let an MLLM model generate a detailed description of a given frame. Then, we give the descriptions to an LLM and let it generate a possible user request. Finally, the MLLM generates AI responses according to the request and frame content. The pipeline is shown in Figs. 11, 12 and 13.

Appendix B Baseline Details

For the Inst-Inpaint baseline, we fine-tune the released checkpoint on the ROVI dataset with the same hyperparameters of LGVI. Specifically, we train 50 epochs on 8 × 80GB NVIDIA A100 GPUs with video and image batch sizes of 32 and 768. We resize the input image to 512 × 320, keeping the same as LGVI. For InstructPix2Pix , MagicBrush , and Inpaint Anything* , we directly use their released checkpoints due to they are pre-trained on large scale image editing/inpainting datasets.

Appendix C More Ablation Studies

We show the quantitative ablation results in Tab. 4. The observation is consistent with the ablations in the main paper. The performance without joint training with images drops slightly, except for a slight increase of 0.011 in the VFID metric. This can be attributed to the enhanced diversity of visual sources provided by additional image data. while it brings a compromise between the quality of the results and temporal consistency. The results also demonstrate the necessity of the U-Net inflation and mask supervision modifications of LGVI. The absence of these modifications leads to a noticeable reduction in performance. The most significant performance degradation is observed when the U-Net is not fully fine-tuned, but only the newly added parameters are trained. This decline can be attributed to the intrinsic differences between the inpainting task and the pre-trained image generation task. In the latter, the language input guides the model on what to create, but it does not specify what needs to be removed.

Appendix D More Qualitative Results

Fig. 14 compares the referring video inpainting task. It demonstrates the effectiveness of the proposed LGVI model. Fig. 15 shows the qualitative results of the interactive video inpainting task. Our LGVI-I model outputs both inpainting results and comprehensive text responses.

Appendix E Basics of Diffusion Models

Denoising Diffusion Probabilistic Models (DDPMs). The core of DDPMs involves iteratively adding noise to the data until it becomes a sample from a simple Gaussian distribution. The reverse process, which generates data from the noise, is learned by the model. The forward process, also known as the “noising” process, is typically modeled as a Markov chain that gradually adds Gaussian noise to the data over a sequence of time steps TT:

where x0x_{0} is a sample from the data distribution q(x0)q(x_{0}), xtx_{t} represents the data at time step tt, and αt\alpha_{t} is a variance schedule that determines the amount of noise to add at each step. ϵ\epsilon is the noise sampled from a standard Gaussian distribution N(0,I)\mathcal{N}(0,\mathbf{I}). The reverse process, which is the generative process, aims to learn the distribution of the original data by reversing the noising process. This involves learning a parameterized function θ\theta that models the reverse conditional probability pθ(xt−1∣xt)p_{\theta}(x_{t-1}|x_{t}). The reverse process is described by:

where μθ(xt,t)\mu_{\theta}(x_{t},t) and Σθ(xt,t)\Sigma_{\theta}(x_{t},t) are learned functions that predict the mean and covariance of the distribution for xt−1x_{t-1}, given xtx_{t} and time step tt. The learning of θ\theta is typically done via a variational approach, minimizing a loss function that is a modified version of the Evidence Lower BOund (ELBO). This loss function ensures that the learned reverse process closely approximates the true distribution of the data, which can be simplified as:

The denoising process can incorporate extra guidance, where the model is trained to generate samples conditioned on a set of labels or attributes cc. Typical guidances are language and images . The loss function can be updated as follows:

where cc is guidance to control the generation result. We extend the condition input with a video input X\mathbf{X} to control the inpainting results.

Latent Diffusion Models (LDMs). Latent Diffusion Model (LDM) is a type of generative model that operates on a latent space rather than directly on the data space. The primary idea is to first encode high-dimensional data, like images, into a lower-dimensional latent representation z=E(x)z=\mathcal{E}(x) and then apply the diffusion process within this latent space. A decoder reconstructs the latent back to the pixel distribution x=D(z)x=\mathcal{D}(z).

LGVI Architecture. The core of our model is to produce inpainting results Y^\hat{\mathbf{Y}} driven by language guidance cc and vision input X\mathbf{X}. This core concept is versatile and can be integrated into a variety of existing architectures, including those based on diffusion or transformer paradigms, as long as the model can fuse language and vision inputs. We choose the LDM architecture because of its flexibility in inflating to video modality and improved sample quality due to the reduced dimensionality of the problem.

Appendix F Future Work Discussions

Fig. 16 shows two LGVI and LGVI-I failure cases. The models still face the core challenge of implicit language or description. In the first example, the referring expression is comparatively long and hard to understand, leading to poor performance. In the second example for the interactive task, even if the MLLM outputs a reasonable response and correctly predicts the position of the removed person, the diffusion model does not provide the right output. That is because the current diffusion-based baseline model is a preliminary modification of a pre-trained text-to-image model, which lacks the understanding of precise locations. The failure cases demonstrate that despite the comparatively stronger performance of previous methods, our proposed model is still a baseline in the language-driven video inpainting field. Future work is expected to develop more advanced methods to overcome the challenges.

References