Text2LIVE: Text-Driven Layered Image and Video Editing
Omer Bar-Tal, Dolev Ofri-Amar, Rafail Fridman, Yoni Kasten, Tali Dekel
Introduction
Computational methods for manipulating the appearance and style of objects in natural images and videos have seen tremendous progress, facilitating a variety of editing effects to be achieved by novice users. Nevertheless, research in this area has been mostly focused in the Style-Transfer setting where the target appearance is given by a reference image (or domain of images), and the original image is edited in a global manner . Controlling the localization of the edits typically involves additional input guidance such as segmentation masks. Thus, appearance transfer has been mostly restricted to global artistic stylization or to specific image domains or styles (e.g., faces, day-to-night, summer-to-winter). In this work, we seek to eliminate these requirements and enable more flexible and creative semantic appearance manipulation of real-world images and videos.
Inspired by the unprecedented power of recent Vision-Language models, we use simple text prompts to express the target edit. This allows the user to easily and intuitively specify the target appearance and the object/region to be edited. Specifically, our method enables local, semantic editing that satisfies a given target text prompt (e.g., Fig. 1 and Fig. 2). For example, given the cake image in Fig. 1(b), and the target text: “oreo cake”, our method automatically locates the cake region and synthesizes realistic, high-quality texture that combines naturally with the original image – the cream filling and the cookie crumbs “paint” the full cake and the sliced piece in a semantically-aware manner. As seen, these properties hold across a variety of different edits.
Our framework leverages the representation learned by a Contrastive Language-Image Pretraining (CLIP) model, which has been pre-trained on 400 million text-image examples . The richness of the enormous visual and textual space spanned by CLIP has been demonstrated by various recent image editing methods (e.g., ). However, the task of editing existing objects in arbitrary, real-world images remains challenging. Most existing methods combine a pre-trained generator (e.g., a GAN or a Diffusion model) in conjunction with CLIP. With GANs, the domain of images is restricted and requires to invert the input image to the GAN’s latent space –-a challenging task by itself . Diffusion models overcome these barriers but face an inherent trade-off between satisfying the target edit and maintaining high-fidelity to the original content . Furthermore, it is not straightforward to extend these methods to videos. In this work, we take a different route and propose to learn a generator from a single input–image or video and text prompts.
If no external generative prior is used, how can we steer the generation towards meaningful, high-quality edits? We achieve this via the following two key components: (i) we propose a novel text-guided layered editing, i.e., rather than directly generating the edited image, we represent the edit via an RGBA layer (color and opacity) that is composited over the input. This allows us to guide the content and localization of the generated edit via a novel objective function, including text-driven losses applied directly to the edit layer. For example, as seen in Fig. 2, we use text prompts to express not only the final edited image but also a target effect (e.g., fire) represented by the edit layer. (ii) We train our generator on an internal dataset of diverse image-text training examples by applying various augmentations to the input image and text. We demonstrate that our internal learning approach serves as a strong regularization, enabling high quality generation of complex textures and semi-transparent effects.
We further take our framework to the realm of text-guided video editing. Real-world videos often consist of complex object and camera motion, which provide abundant information about the scene. Nevertheless, achieving consistent video editing is difficult and cannot be accomplished naïvely. We thus propose to decompose the video into a set of 2D atlases using . Each atlas can be treated as a unified 2D image representing either a foreground object or the background throughout the video. This representation significantly simplifies the task of video editing: edits applied to a single 2D atlas are automatically mapped back to the entire video in a consistent manner. We demonstrate how to extend our framework to perform edits in the atlas space while harnessing the rich information readily available in videos.
In summary, we present the following contributions:
An end-to-end text-guided framework for performing localized, semantic edits of existing objects in real-world images.
A novel layered editing approach and objective function that automatically guides the content and localization of the generated edit.
We demonstrate the effectiveness of internal learning for training a generator on a single input in a zero-shot manner.
An extension to video which harnesses the richness of information across time, and can perform consistent text-guided editing.
We demonstrate various edits, ranging from changing objects’ texture to generating complex semi-transparent effects, all achieved fully automatically across a wide-range of objects and scenes.
Related Work
Text-guided image manipulation and synthesis. There has been remarkable progress since the use of conditional GANs in both text-guided image generation , and editing . ManiGAN proposed a text-conditioned GAN for editing an object’s appearance while preserving the image content. However, such multi-modal GAN-based methods are restricted to specific image domains and limited in the expressiveness of the text (e.g., trained on COCO ). DALL-E addresses this by learning a joint image-text distribution over a massive dataset. While achieving remarkable text-to-image generation, DALL-E is not designed for editing existing images. GLIDE takes this approach further, supporting both text-to-image generation and inpainting.
Instead of directly training a text-to-image generator, a recent surge of methods leverage a pre-trained generator, and use a pre-trained CLIP to guide the generation process by text . StyleCLIP and StyleGAN-NADA use a pre-trained StyleGAN2 for image manipulation, by either controlling the GAN’s latent code , or by fine-tuning the StyleGAN’s output domain . However, editing a real input image using these methods requires first tackling the GAN-inversion challenge . Furthermore, these methods can edit images from a few specific domains, and edit images in a global fashion. In contrast, we consider a different problem setting – localized edits that can be applied to real-world images spanning a variety of object and scene categories.
A recent exploratory and artistic trend in the online AI community has demonstrated impressive text-guided image generation. CLIP is used to guide the generation process of a pre-trained generator, e.g., VQ-GAN , or diffusion models . takes this approach a step forward by optimizing the diffusion process itself. However, since the generation is globally controlled by the diffusion process, this method is not designed to support localized edits that are applied only to selected objects.
To enable region-based editing, user-provided masks are used to control the diffusion process for image inpainting . In contrast, our goal is not to generate new objects but rather to manipulate the appearance of existing ones, while preserving the original content. Furthermore, our method is fully automatic and performs the edits directly from the text, without user edit masks.
Several works take a test-time optimization approach and leverage CLIP without using a pre-trained generator. For example, CLIPDraw renders a drawing that matches a target text by directly optimizing a set of vector strokes. To prevent adversarial solutions, various augmentations are applied to the output image, all of which are required to align with the target text in CLIP embedding space. CLIPStyler takes a similar approach for global stylization. Our goal is to perform localized edits, which are applied only to specific objects. Furthermore, CLIPStyler optimizes a CNN that observes only the source image. In contrast, our generator is trained on an internal dataset, extracted from the input image and text. We draw inspiration from previous works that show the effectiveness of internal learning in the context of generation .
Other works use CLIP to synthesize or edit a single 3D representation (NeRF or mesh). The unified 3D representation is optimized through a differentiable renderer: CLIP loss is applied across different 2D rendered viewpoints. Inspired by this approach, we use a similar concept to edit videos. In our case, the “renderer” is a layered neural atlas representation of the video .
Consistent Video Editing. Existing approaches for consistent video editing can be roughly divided into: (i) propagation-based methods, which use keyframes or optical flow to propagate edits through the video, and (ii) video layering-based methods, in which a layered representation of the video is estimated and then edited . For example, Lu et al. estimate omnimattes – RGBA layers that contain a target subject along with their associated scene effects. Omnimattes facilitate a variety of video effects (e.g., object removal or retiming). However, since the layers are computed independently for each frame, it cannot support consistent propagation of edits across time. Kasten et al. address this challenge by decomposing the video into unified 2D atlas layers (foreground and background). Edits applied to the 2D atlases are automatically mapped back to the video, thus achieving temporal consistency with minimal effort. In our work, we treat a pre-trained neural layered atlas model as a video renderer and leverage it for the task of text-guided video editing.
Text-Guided Layered Image and Video Editing
We focus on semantic, localized edits expressed by simple text prompts. Such edits include changing objects’ texture or semantically augmenting the scene with complex semi-transparent effects (e.g., smoke, fire). To this end, we harness the potential of learning a generator from a single input image or video while leveraging a pre-trained CLIP model, which is kept fixed and used to establish our losses . Our task is ill-posed – numerous possible edits can satisfy the target text according to CLIP, some of which include noisy or undesired solutions . Thus, controlling edits’ localization and preserving the original content are both pivotal components for achieving high-quality editing results. We tackle these challenges through the following key components:
Layered editing. Our generator outputs an RGBA layer that is composited over the input image. This allows us to control the content and spatial extent of the edit via dedicated losses applied directly to the edit layer.
Explicit content preservation and localization losses. We devise new losses using the internal spatial features in CLIP space to preserve the original content, and to guide the localization of the edits.
Internal generative prior. We construct an internal dataset of examples by applying augmentations to the input image/video and text. These augmented examples are used to train our generator, whose task is to perform text-guided editing on a larger and more diverse set of examples.
As illustrated in Fig. 3, our framework consists of a generator that takes as input a source image and synthesizes an edit layer, , which consists of a color image and an opacity map . The final edited image is given by compositing the edit layer over :
Our main goal is to generate such that the final composite would comply with a target text prompt . In addition, generating an RGBA layer allows us to use text to further guide the generated content and its localization. To this end, we consider a couple of auxiliary text prompts: which expresses the target edit layer, when composited over a green background, and which specifies a region-of-interest in the source image, and is used to initialize the localization of the edit. For example, in the Bear edit in Fig. 2, “fire out of the bear’s mouth”, “fire over a green screen”, and “mouth”. We next describe in detail how these are used in our objective function.
Objective function. Our novel objective function incorporates three main loss terms, all defined in CLIP’s feature space: (i) , which is the driving loss and encourages to conform with , (ii) , which serves as a direct supervision on the edit layer, and (iii) , a structure preservation loss w.r.t. . Additionally, a regularization term is used for controlling the extent of the edit by encouraging sparse alpha matte . Formally,
where , , and control the relative weights between the terms, and are fixed throughout all our experiments (see Appendix 0.A.3).
Composition loss. reflects our primary objective of generating an image that matches the target text prompt and is given by a combination of a cosine distance loss and a directional loss :
where is the cosine distance between the CLIP embeddings for and . Here, , denote CLIP’s image and text encoders, respectively. The second term controls the direction of edit in CLIP space and is given by: .
Similar to most CLIP-based editing methods, we first augment each image to get several different views and calculate the CLIP losses w.r.t. each of them separately, as in . This holds for all our CLIP-based losses. See Appendix 0.A.2 for details.
Screen loss. The term serves as a direct text supervision on the generated edit layer . We draw inspiration from chroma keying –a well-known technique by which a solid background (often green) is replaced by an image in a post-process. Chroma keying is extensively used in image and video post-production, and there is high prevalence of online images depicting various visual elements over a green background. We thus composite the edit layer over a green background and encourage it to match the text-template over a green screen”, (Fig. 3):
where .
A nice property of this loss is that it allows intuitive supervision on a desired effect. For example, when generating semi-transparent effects, e.g., Bear in Fig. 2, we can use this loss to focus on the fire regardless of the image content by using “fire over a green screen”. Unless specified otherwise, we plug in to our screen text template in all our experiments. Similar to the composition loss, we first apply augmentations on the images before feeding to CLIP.
The term is defined as the Frobenius norm distance between the self-similarity matrices of , and :
Sparsity regularization. To control the spatial extent of the edit, we encourage the output opacity map to be sparse. We follow and define the sparsity loss term as a combination of - and -approximation regularization terms:
where is a smooth approximation that penalizes non zero elements. We fix in all our experiments.
Bootstrapping. To achieve accurate localized effects without user-provided edit mask, we apply a text-driven relevancy loss to initialize our opacity map. Specifically, we use Chefer et al. to automatically estimate a relevancy map can only work with images, so we resize both and to before applying the loss of (8) which roughly highlights the image regions that are most relevant to a given text . We use the relevancy map to initialize by minimizing:
Note that the relevancy maps are noisy, and only provide a rough estimation for the region of interest (Fig. 8(c)). Thus, we anneal this loss during training (see implementation details in Appendix 0.A.3). By training on diverse internal examples along with the rest of our losses, our framework dramatically refines this rough initialization, and produces accurate and clean opacity (Fig. 8(d)).
Training data. Our generator is trained from scratch for each input using an internal dataset of diverse image-text training examples that are derived from the input (Fig. 3 left). Specifically, each training example is generated by randomly applying a set of augmentations to and to . The image augmentations include global crops, color jittering, and flip, while text augmentations are randomly sampled from a predefined text template (e.g., “a photo of”); see Appendix 0.A.2 for details. The vast space of all combinations between these augmentations provides us with a rich and diverse dataset for training. The task is now to learn one mapping function for the entire dataset, which poses a strong regularization on the task. Specifically, for each individual example, has to generate a plausible edit layer from such that the composited image is well described by . We demonstrate the effectiveness of our internal learning approach compared to the test-time optimization approach in Sec. 4.
2 Text to Video Edit Layer
A natural question is whether our image framework can be applied to videos. The key additional challenge is achieving a temporally consistent result. Naïvely applying our image framework on each frame independently yields unsatisfactory jittery results (see Sec. 4). To enforce temporal consistency, we utilize the Neural Layered Atlases (NLA) method , as illustrated in Fig. 4(a). We next provide a brief review of NLA and discuss in detail how our extension to videos.
Importantly, NLA enables consistent video editing: the continuous atlas (foreground or background) is first discretized to a fixed resolution image (e.g., 10001000 px). The user can directly edit the discretized atlas using image editing tools (e.g., Photoshop). The atlas edit is then mapped back to the video, and blended with the original frames, using the predicted UV mappings and foreground opacity. In this work, we are interested in generating atlas edits in a fully automatic manner, solely guided by text.
Text to Atlas Edit Layer. Our video framework leverages NLA as a “video renderer”, as illustrated in Fig. 4. Specifically, given a pre-trained and fixed NLA model for a video, our goal is to generate a 2D atlas edit layer, either for the background or foreground, such that when mapped back to the video, each of the rendered frames would comply with the target text.
Training. A straightforward approach for training is to treat as an image and plug it into our image framework (Sec. 3.1). This approach will result in a temporally consistent result, yet it has two main drawbacks: (i) the atlas often non-uniformly distorts the original structures (see Fig. 4), which may lead to low-quality edits , (ii) solely using the atlas, while ignoring the video frames, disregards the abundant, diverse information available in the video such as different viewpoints, or non-rigid object deformations, which can serve as “natural augmentations” to our generator. We overcome these drawbacks by mapping the atlas edit back to the video and applying our losses on the resulting edited frames. Similar to the image case, we use the same objective function (Eq. 2), and construct an internal dataset directly from the atlas for training.
Results
We tested our method across various real-world, high-resolution images and videos. The image set contains 35 images collected from the web, spanning various object categories, including animals, food, landscapes and others. The video set contains seven videos from DAVIS dataset . We applied our method using various target edits, ranging from text prompts that describe the texture/materials of specific objects, to edits that express complex scene effects such as smoke, fire, or clouds. Sample examples for the inputs along with our results can be seen in Fig. 1, Fig. 2, and Fig. 5 for images, and Fig. 6 for videos. The full set of examples and results are included in the Supplementary Materials (SM). As can be seen, in all examples, our method successfully generates photorealistic textures that are “painted” over the target objects in a semantically aware manner. For example, in red velvet edit (first row in Fig. 5), the frosting is naturally placed on the top. In car-turn example (Fig. 6), the neon lights nicely follow the car’s framing. In all examples, the edits are accurately localized, even under partial occlusions, multiple objects (last row and third row of Fig. 5) and complex scene composition (the dog in Fig. 2). Our method successfully augments the input scene with complex semi-transparent effects without changing irrelevant content in the image (see Fig. 1).
2 Comparison to Prior Work
To the best of our knowledge, there is no existing method tailored for solving our task: text-driven semantic, localized editing of existing objects in real-world images and videos. We illustrate the key differences between our method and several prominent text-driven image editing methods. We consider those that can be applied to a similar setting to ours: editing real-world images that are not restricted to specific domains. Inpainting methods: Blended-Diffusion and GLIDE , both require user-provided editing mask. CLIPStyler, which performs image stylization, and Diffusion+CLIP , and VQ-GAN+CLIP : two baselines that combine CLIP with either a pre-trained VQ-GAN or a Diffusion model. In the SM, we also include additional qualitative comparison to the StyleGAN text-guided editing methods .
Fig. 7 shows representative results, and the rest are included in the SM. As can be seen, none of these methods are designed for our task. The inpainting methods (b-c), even when supplied with tight edit masks, generate new content in the masked region rather than changing the texture of the existing one. CLIPStyler modifies the image in a global artistic manner, rather than performing local semantic editing (e.g., the background in both examples is entirely changed, regardless of the image content). For the baselines (d-f), Diffusion+CLIP can often synthesize high-quality images, but with either low-fidelity to the target text (e), or with low-fidelity to the input image content (see many examples in SM). VQ-GAN+CLIP fails to maintain fidelity to the input image and produces non-realistic images (f). Our method automatically locates the cake region and generates high-quality texture that naturally combines with the original content.
3 Quantitative evaluation
Comparison to image baselines. We conduct an extensive human perceptual evaluation on Amazon Mechanical Turk (AMT). We adopt the Two-alternative Forced Choice (2AFC) protocol suggested in . Participants are shown a reference image and a target editing prompt, along with two alternatives: our result and another baseline result. We consider from the above baselines those not requiring user-masks. The participants are asked: “Which image better shows objects in the reference image edited according to the text”. We perform the survey using a total of 82 image-text combinations. We collected 12,450 user judgments w.r.t. prominent text-guided image editing methods. Table 1 reports the percentage of votes in our favor. As seen, our method outperforms all baselines by a large margin, including those using a strong generative prior.
Comparison to video baselines. We quantify the effectiveness of our key design choices for the video-editing by comparing our video method against: (i) Atlas Baseline: feeding the discretized 2D Atlas to our single-image method (Sec. 3.1), and using the same inference pipeline illustrated in Fig. 4 to map the edited atlas back to frames. (ii) Frames Baseline: treating all video frames as part of a single internal dataset, used to train our generator; at inference, we apply the trained generator independently to each frame.
We conduct a human perceptual evaluation in which we provide participants a target editing prompt and two video alternatives: our result and a baseline. The participants are asked “Choose the video that has better quality and better represents the text”. We collected 2,400 user judgments over 19 video-text combinations and report the percentage of votes in favor of the complete model in table 1. We first note that the Frames baseline produces temporally inconsistent edits. As expected, the Atlas baseline produces temporally consistent results. However, it struggles to generate high-quality textures and often produces blurry results. These observations support our hypotheses mentioned in Sec. 3.2. We refer the reader to the SM for visual comparisons.
4 Ablation Study
Fig. 8(top) illustrates the effect of our relevancy-based bootstrapping (Sec. 3.1). As seen, this component allows us to achieve accurate object mattes, which significantly improves the rough, inaccurate relevancy maps.
We ablate the different loss terms in our objective by qualitatively comparing our results when training with our full objective (Eq. 2) and with a specific loss removed. The results are shown in Fig. 8. As can be seen, without (w/o sparsity), the output matte does not accurately capture the mango, resulting in a global color shift around it. Without (w/o structure), the model outputs an image with the desired appearance but fails to preserve the mango shape fully. Without (w/o screen), the segmentation of the object is noisy (color bleeding from the mango), and the overall quality of the texture is degraded (see SM for additional illustration). Lastly, we consider a test-time optimization baseline by not using our internal dataset but rather inputting to the same input at each training step. As seen, this baseline results in lower-quality edits.
5 Limitations
We noticed that for some edits, CLIP exhibits a very strong bias towards a specific solution. For example, as seen in Fig. 9, given an image of a cake, the text “birthday cake” is strongly associated with candles. Our method is not designed to significantly deviate from the input image layout and to create new objects, and generates unrealistic candles. Nevertheless, in many cases the desired edit can be achieved by using more specific text. For example, the text “moon” guides the generation towards a crescent. By using the text “a bright full moon” we can steer the generation towards a full moon (Fig. 9 left). Finally, as acknowledged by prior works (e.g., ), we also noticed that slightly different text prompts describing similar concepts may lead to slightly different flavors of edits.
On the video side, our method assumes that the pre-trained NLA model accurately represents the original video. Thus, we are restricted to examples where NLA works well, as artifacts in the atlas representation can propagate to our edited video. An exciting avenue of future research may include fine-tuning the NLA representation jointly with our model.
Conclusion
We considered a new problem setting in the context of zero-shot text-guided editing: semantic, localized editing of existing objects within real-world images and videos. Addressing this task requires careful control of several aspects of the editing: the edit localization, the preservation of the original content, and visual quality. We proposed to generate text-driven edit layers that allow us to tackle these challenges, without using a pre-trained generator in the loop. We further demonstrated how to adopt our image framework, with only minimal changes, to perform consistent text-guided video editing. We believe that the key principles exhibited in the paper hold promise for leveraging large-scale multi-modal networks in tandem with an internal learning approach.
Acknowledgments
We thank Kfir Aberman, Lior Yariv, Shai Bagon, and Narek Tumanayan for their insightful comments. We thank Narek Tumanayan for his help with the baselines comparison. This project received funding from the Israeli Science Foundation (grant 2303/20).
References
Appendix 0.A Implementation Details
We provide implementation details for our architecture and training regime.
We base our generator network on the U-Net architecture , with a 7-layer encoder and a symmetrical decoder. All layers comprise Convolutional layers, followed by BatchNorm, and LeakyReLU activation. The intermediate channels dimensions is 128. In each level of the encoder, we add an additional Convolutional layer and concatenate the output features to the corresponding level of the decoder. Lastly, we add a Convolutional layer followed by Sigmoid activation to get the final RGB output.
A.2 Internal Dataset (Sec. 3.1)
We apply data augmentations to the source image and target text , to create multiple internal examples . Specifically, at each training step, we apply a random set of image augmentations to , and augment using a pre-defined set of text templates, as follows:
Random spatial crops: 0.85 and 0.95 of the image size in our image and video frameworks respectively.
Random scaling: aspect ratio preserved scaling, of both spatial dimensions by a random factor, sampled uniformly from the range .
Random horizontal-flipping is applied with probability p=0.5.
Random color jittering: we jitter the global brightness, contrast, saturation and hue of the image.
A.2.2 Text augmentations and the target text prompt T𝑇T
We compose with a random text template, sampled from of a pre-defined list of 14 templates. We designed our text-templates that does not change the semantics of the prompt, yet provide variability in the resulting CLIP embedding e.g.:
At each step, one of the above templates is chosen at random and the target text prompt is plugged in to it and forms our augmented text. By default, our framework uses a single text prompt , but can also support multiple input text prompts describing the same edit, which effectively serve as additional text augmentations (e.g., “crochet swan”, and “knitted swan” can both be used to describe the same edit).
A.3 Training Details
We implement our framework in PyTorch (code will be made available). As described in Sec. 3, we leverage a pre-trained CLIP model to establish our losses. We use the ViT-B/32 pretrained model (12 layers, 32x32 patches), downloaded from the official implementation at GitHub. We optimize our full objective (Eq. 2, Sec. 3.1), with relative weights: , (3 for videos), , ( for videos) and . For bootstrapping, we set the relative weight to be 10, and for the image framework we anneal it linearly throughout the training. We use the MADGRAD optimizer with an initial learning rate of , weight decay of 0.01 and momentum 0.9. We decay the learning rate with an exponential learning rate scheduler with ( for videos), limiting the learning rate to be no less than . Each batch contains (see Sec. 3.1), the augmented source image and target text respectively. Every 75 iterations, we add , to the batch (i.e., do not apply augmentations). The output of is then resized down to [px] maintaining aspect ratio and augmented (e.g., geometrical augmentations) before extracting CLIP features for establishing the losses. We enable feeding to CLIP arbitrary resolution images (i.e., non-square images) by interpolating the position embeddings (to match the size of spatial tokens of a the given image) using bicubic interpolation, similarly to .
Training on an input image of size takes minutes to train on a single GPU (NVIDIA RTX 6000) for a total of 1000 iterations. Training on one video layer (foreground/background) of 70 frames with resolution takes 60 minutes on a single GPU (NVIDIA RTX 8000) for a total of 3000 iterations.
A.4 Video Framework
We further elaborate on the framework’s details described in Sec. 3.2 of the paper.
Atlas Pre-processing. Our framework works on a discretized atlas, which we obtain by rendering the atlas to a resolution of 20002000 px. This is done as in , by querying the pre-trained atlas network in uniformly sampled UV locations. The neural atlas representation is defined within the continuous space, yet the video content may not occupy the entire space. To focus only on the used atlas regions, we crop the atlas prior to training, by mapping all video locations to the atlas and taking their bounding box. Note that for foreground atlas, we map only the foreground pixels in each frame, i.e., pixels for which the foreground opacity is above 0.95; the foreground/background opacity is estimated by the pre-trained neural atlas representation.
Training. As discussed in Sec. 3.2 in the paper, our generator is trained on atlas crops, yet our losses are applied to the resulting edited frames. In each iteration, we crop the atlas by first sampling a video segment of 3 frames and mapping it to the atlas. Formally, we sample a random frame and a random spatial crop size where its top left coordinate is at . As a result we get a set of cropped (spatially and temporally) video locations:
where is the offset between frames.
We augment the atlas crop as well as the target text , as described in Sec. 0.A.2 herein to generate an internal training dataset. To apply our losses, we map back the atlas edit layer to the original video segment and process the edited frames the same way as in the image framework: resizing, applying CLIP augmentations, and applying the final loss function of Eq. 2 in Sec. 3.1 in the paper. To enrich the data, we also include one of the sampled frame crops as a direct input to and apply the losses directly on the output (as in the image case). Similarly to the image framework, every 75 iterations we additionally pass the pair , where is the entire atlas (without augmentations, and without mapping back to frames). For the background atlas, we first downscale it by three due to memory limitations.
Inference. As described in Sec. 3.2, at inference time, the entire atlas is fed into results in . The edit is mapped and combined with the original frames using the process that is described in (Sec. 3.4, Eq. (15),(16)). Note that our generator operates on a single atlas. To produce foreground and background edits, we train two separate generators for each atlas.