ReCo: Region-Controlled Text-to-Image Generation

Zhengyuan Yang, Jianfeng Wang, Zhe Gan, Linjie Li, Kevin Lin, Chenfei Wu, Nan Duan, Zicheng Liu, Ce Liu, Michael Zeng, Lijuan Wang

Introduction

Text-to-image (T2I) generation aims to generate faithful images based on an input text query that describes the image content. By scaling up the training data and model size, large T2I models have recently shown remarkable capabilities in generating high-fidelity images. However, the text-only query allows limited controllability, e.g., precisely specifying the content in a specific region. The naive way of using position-related text words, such as “top left” and “bottom right,” often results in ambiguous and verbose input queries, as shown in Figure 2 (a). Even worse, when the text query becomes long and complicated, or describes an unusual scene, T2I models might overlook certain details and rather follow the visual or linguistic training prior. These two factors together make region control difficult. To get the desired image, users usually need to try a large number of paraphrased queries and pick an image that best fits the desired scene. The process known as “prompt engineering” is time-consuming and often fails to produce the desired image.

The desired region-controlled T2I generation is closely related to the layout-to-image generation . As shown in Figure 2 (b), layout-to-image models take all object bounding boxes with labels from a close set of object vocabulary as inputs. Despite showing promise in region control, they can hardly understand free-form text inputs, nor the region-level combination of open-ended text descriptions and spatial positions. The two input conditions of text and box provide complementary referring capabilities. Instead of separately modeling them as in text-to-image and layout-to-image generations, we study “region-controlled T2I generation” that seamlessly combines these two input conditions. As shown in Figure 2 (c), the new input interface allows users to provide open-ended descriptions for arbitrary image regions, such as precisely placing a “brown glazed chocolate donut” in a specific area.

To this end, we propose ReCo (Region-Controlled T2I) that extends pre-trained T2I models to understand spatial coordinate inputs. The core idea is to introduce an extra set of input position tokens to indicate the spatial positions. The image width/height is quantized uniformly into NbinsN_{\text{bins}} bins. Then, any float-valued coordinate can be approximated and tokenized by the nearest bin. With an extra embedding matrix (EpE_{p}), the position token can be mapped onto the same space as the text token. Instead of designing a text-only query with positional words “in the top red donut” as in Figure 2 (a), ReCo takes region-controlled text inputs “<x1>,<y1>,<x2>,<y2>\tiny{<}x_{\text{1}}\tiny{>},\tiny{<}y_{\text{1}}\tiny{>},\tiny{<}x_{\text{2}}\tiny{>},\tiny{<}y_{\text{2}}\tiny{>} red donut,” where <x>\tiny{<}x\tiny{>},<y>\tiny{<}y\tiny{>} are the position tokens followed by the corresponding free-form text description. We then fine-tune a pre-trained T2I model with EpE_{p} to generate the image from the extended input query. To best preserve the pre-trained T2I capability, ReCo training is designed to be similar to the T2I pre-training, i.e., introducing minimal extra model parameters (EpE_{p}), jointly encoding position and text tokens with the text encoder, and prefixing the image description before the extended regional descriptions in the input query.

Figure 1 visualizes ReCo’s use cases and capabilities. As shown in Figure 1 (a), we can easily control the view (front/side) and type (single-/double-deck) of the “bus” by tweaking position tokens in region-controlled text inputs. Position tokens also allow the user to provide free-form regional descriptions, such as “an orange cat wearing a red hat” at a specific location. Furthermore, we empirically observe that position tokens are less likely to get overlooked or misunderstood than text words. As shown in Figure 1 (b), ReCo has better control over object count, spatial relationship, and size properties, especially when the query is long and complicated, or describes a scene that is less common in real life. In contrast, T2I models may struggle with generating scenes with correct object counts (“ten”), relationships (“boat below traffic light”), relative sizes (“chair larger than airplane”), and camera views (“zoomed out”).

To evaluate the region control, we design a comprehensive experiment benchmark based on a pre-trained regional object classifier and an object detector. The object classifier is applied on the generated image regions, while the detector is applied on the whole image. A higher accuracy means a better alignment between the generated object layout and the region positions in user queries. On the COCO dataset , ReCo shows a better object classification accuracy (42.02%→62.42%42.02\%\rightarrow 62.42\%) and detector averaged precision (2.3→32.02.3\rightarrow 32.0), compared with the T2I model with carefully designed positional words. For image generation quality, ReCo improves the FID from 8.828.82 to 7.367.36, and SceneFID from 15.5415.54 to 6.516.51. Furthermore, human evaluations on PaintSkill show +19.28%+19.28\% and +17.21%+17.21\% accuracy gain in more correctly generating the query-described object count and spatial relationship, indicating ReCo’s capability in helping T2I models to generate challenging scenes.

Our contributions are summarized as follows.

We propose ReCo that extends pre-trained T2I models to understand coordinate inputs. Thanks to the introduced position tokens in the region-controlled input query, users can easily specify free-form regional descriptions in arbitrary image regions.

We instantiate ReCo based on Stable Diffusion. Extensive experiments show that ReCo strictly follows the regional instructions from the input query, and also generates higher-fidelity images.

We design a comprehensive evaluation benchmark to validate ReCo’s region-controlled T2I generation capability. ReCo significantly improves both the region control accuracy and the image generation quality over a wide range of datasets and designed prompts.

Related Work

Text-to-image generation. Text-to-image (T2I) generation aims to generate a high-fidelity image based on an open-ended image description. Early studies adopt conditional GANs for T2I generation. Recent studies have made tremendous advances by scaling up both the data and model size, based on either auto-regressive or diffusion-based models . We build our study on top of the successful large-scale pre-trained T2I models, and explore how to better control the T2I generation by extending a pre-trained T2I model to understand position tokens.

Layout-to-image generation. Layout-to-image studies aim to generate an image from a complete layout, i.e., all bounding boxes and the paired object labels. Early studies adopt GAN-based approaches by properly injecting the encoded layout as the input condition. Recent studies successfully apply the layout query as the input condition to the auto-regressive framework and diffusion models . Our study is related to the layout-to-image generation as both directions require the model to understand coordinate inputs. The major difference is that our design synergetically combines text and box to help T2I generation. Therefore, ReCo can take open-ended regional descriptions and benefit from large-scale T2I pre-training.

Unifying open-ended text and localization conditions. Previous studies have explored unifying open-ended text descriptions with localization referring (box, mask, mouse trace) as the input generation condition. One modeling approach is to separately encode the image description in T2I and the layout condition in layout-to-image, and trains a model to jointly condition on both input types. TRECS takes mouse traces in the localized narratives dataset to better ground open-ended text descriptions with a localized position. Other than taking layout as user-generated inputs, previous studies have also explored predicting layout from text to ease the T2I generation of complex scenes. Unlike the motivation of training another conditional generation model parallel to T2I and layout-to-image, we explore how to effectively extend pre-trained T2I models to understand region queries, leading to significantly better controllability and generation quality than training from scratch. In short, we position ReCo as an improvement for T2I by providing a more flexible input interface and alleviating controllability issues, e.g., being difficult to override data prior when generating unusual scenes, and overlooking words in complex queries.

ReCo Model

Region-Controlled T2I Generation (ReCo) extends T2I models with the ability to understand coordinate inputs. The core idea is to design a unified input token vocabulary containing both text words and position tokens to allow accurate and open-ended regional control. By seamlessly mixing text and position tokens in the input query, ReCo obtains the best from the two worlds of text-to-image and layout-to-image, i.e., the abilities of free-form description and precise position control. In this section, we present our ReCo implementation based on the open-sourced Stable Diffusion (SD) . We start with the SD preliminaries in Section 3.1 and introduce the core ReCo design in Section 3.2.

We take Stable Diffusion as an example to introduce the T2I model that ReCo is built upon. Stable Diffusion is developed upon the Latent Diffusion Model , and consists of an auto-encoder, a U-Net for noise estimation, and a CLIP ViT-L/14 text encoder. For the auto-encoder, the encoder E\mathcal{E} with a down-sampling factor of 88 encodes the image xx into a latent representation z=E(x)z=\mathcal{E}(x) that the diffusion process operates on, and the decoder D\mathcal{D} reconstructs the image x^=D(z)\hat{x}=\mathcal{D}(z) from the latent zz. U-Net is conditioned on denoising timestep tt and text condition τθ(y(T))\tau_{\theta}(y(T)), where y(T)y(T) is the input text query with text tokens TT and τθ\tau_{\theta} is the CLIP ViT-L/14 text encoder that projects a sequence of tokenized texts into the sequence embedding.

The core motivation of ReCo is to explore more effective and interaction-friendly conditioning signals yy, while best preserving the pre-trained T2I capability. Specifically, ReCo extends text tokens with an extra vocabulary specialized for spatial coordinate referring, i.e., position tokens PP, which can be seamlessly used together with text words TT in a single input query yy. ReCo aims to show the benefit of synergetically combining text and position conditions for region-controlled T2I generation.

2 Region-Controlled T2I Generation

ReCo fine-tuning. ReCo extends the text-only query y(T)y(T) with text tokens TT into ReCo input query y(P,T)y(P,T) that combines the text word TT and position token PP. We fine-tune the Stable Diffusion with the same latent diffusion modeling objective , following the notations in Section 3.1:

where ϵθ\epsilon_{\theta} and τθ\tau_{\theta} are the fine-tuned network modules. All model parameters except position token embedding EpE_{p} are initiated from the pre-trained Stable Diffusion model. Both the image description and several regional descriptions are required for ReCo model fine-tuning. For the training data, we run a state-of-the-part captioning model on the cropped image regions (following the annotated bounding boxes) to get the regional descriptions. During fine-tuning, we resize the image with the short edge to 512512 and randomly crop a square region as the input image xx. We will release the generated data and fine-tuned model for reproduction.

We empirically observe that ReCo can well understand the introduced position tokens and precisely place objects at arbitrary specified regions. Furthermore, we find that position tokens can also help ReCo better model long input sequences that contain multiple detailed attribute descriptions, leading to fewer detailed descriptions being neglected or incorrectly generated than the text-only query. By introducing position tokens with a minimal change to the pre-trained T2I model, ReCo obtains the desired region controllability while best preserving the appealing T2I capability.

Experiments

Datasets. We evaluate the model on the COCO , PaintSkill , and LVIS datasets. For input queries, we take image descriptions and boxes from the datasets , and generate regional descriptions with the same captioning model on cropped regions. For COCO , we follow the established setting in the T2I generation that reports the results on a subset of 30,000 captions sampled from the COCO 2014 validation set. We fine-tune stable diffusion with image-text pairs from the COCO 2014 training set. PaintSkill evaluates models’ capabilities on following arbitrarily positioned boxes and generating images with the correct object type/count/relationship. We conduct the T2I inference with validation set prompts, which contain 1,050/2,520/3,528 queries for object recognition, counting, and spatial relationship skills, respectively. LVIS tests if the model understands open-vocabulary regional descriptions, with the object categories unseen in the COCO fine-tuning data. We report the results on the 4,809 LVIS validation images from the COCO 2017 validation set . We do not further fine-tune the model when experimenting on PaintSkill and LVIS to test the generalization capability in out-of-domain data.

Evaluation metrics. We evaluate ReCo with metrics focused on region control accuracy and image generation quality. For region control accuracy, we use Object Classification Accuracy and DETR detector Average Precision (AP) . Object accuracy trains a classifier with ground-truth (GT) image crops to classify the cropped regions on generated images. DETR detector AP detects objects on generated images and compares the results with input object queries. Thus, higher accuracy and AP can indicate a better layout alignment. For image generation quality, we use the Fréchet Inception Distance (FID) to evaluate the image quality. We take SceneFID as an indicator for region-level visual quality, which computes FID on the regions cropped based on input object boxes. We compute FID and SceneFID with the Clean-FID repo against center-cropped COCO images. We further conduct human evaluations on PaintSkill, due to the lack of GT images and effective automatic evaluation metrics.

Implementation details. We fine-tune ReCo from the Stable Diffusion v1.4 checkpoint. We introduce N=1000N=1000 position tokens and increase the max length of the text encoder to 616. The batch size is 2048. We use AdamW optimizer with a constant learning rate of 1e−41e^{-4} to train the model for 20,000 steps, equivalent to around 100 epochs on COCO 2014 training set. The inference is conducted with 50 PLMS steps . We select a classifier-free guidance scale that gives the best region control performance, i.e., 4.0 for ReCo and 7.5 for original Stable Diffusion, detailed in Section 4.3. We do not use CLIP image re-ranking.

2 Region-Controlled T2I Generation Results

COCO. Table 1 reports the region-controlled T2I generation results on COCO. The first row “real images” provides an oracle reference number on applicable metrics. The top part of the table shows the results obtained with the pre-trained Stable Diffusion (SD) model without fine-tuning on COCO, i.e., the zero-shot setting. As shown in the left three columns, we experiment with adding “region description” and “region position” information to the input query in addition to “image description.” Since T2I models can not understand coordinates, we carefully design positional text descriptions, indicated by “text” in the “region position” column. Specifically, we describe a region with one of the three size words (small, medium, large), three possible region aspect ratios (long, square, tall), and nine possible locations (top left, top, …\ldots, bottom right). The bottom part compares the main ReCo model with other variants fine-tuned with the corresponding input queries. The middle three rows report the results on region control accuracy. For AP and AP50{}_{\text{50}}, we use a DETR ResNet-50 object detector trained on COCO to get the detection results on images generated based on the input texts and boxes from the COCO 2017 val5k set . The “object accuracy” column reports the region classification accuracy . The trained ResNet-101 region classifier yields a 71.41%71.41\% oracle 80-class accuracy on real images. The right two columns report the image generation quality metrics, i.e., SceneFID and FID, which evaluate the region and image visual qualities.

One advantage of ReCo is its strong region control capability. As shown in the bottom row, ReCo achieves an AP of 32.032.0, which is close to the real image oracle of 36.836.8. Despite the careful engineering of positional text words, ReCoPosition Word{}_{\text{Position Word}} only achieves an AP of 2.32.3. Similarly, for object region classification, 62.42%62.42\% of the cropped regions on ReCo-generated images can be correctly classified, compared with 42.02%42.02\% of ReCoPosition Word{}_{\text{Position Word}}. ReCo also improves the generated image quality, both at the region and image level. At the region level, ReCo achieves a SceneFID of 6.516.51, indicating strong capabilities in both generating high-fidelity objects and precisely placing them in the queried position. At the image level, ReCo improves the FID from 10.4410.44 to 7.367.36 with the region-controlled text input that provides a localized and more detailed image description. We present additional FID comparisons to state-of-the-art conditional image generation methods in Table 5 (c).

We show representative qualitative results in Figure 4. (a) ReCo can more reliably generate images that involve counting or complex object relationships, e.g., “five birds” and “sitting on a bench.” (b) ReCo can more easily generate images with unique camera views by controlling the relative position and size of object boxes, e.g., “a top-down view of a cat” that T2I models struggle with. (c) Separating detailed regional descriptions with position tokens also helps ReCo better understand long queries and reduce attribute leakage, e.g., the color of the clock and person’s shirt.

PaintSkill. Table 2 shows the skill correctness and region control accuracy evaluations on PaintSkill . Skill correctness evaluates if the generated images contain the query-described object type/count/relationship, i.e., the “object,” “count,” and “spatial” subsets. We use human judges to obtain the skill correctness accuracy. For region control, we use object classification accuracy to evaluate if the model follows those arbitrarily shaped and located object queries. We reuse the COCO region classifier introduced in Table 1.

Based on the human evaluation for “skill correctness,” 87.38%87.38\% and 82.08%82.08\% of ReCo-generated images have the correct object count and spatial relationship (“count” and “spatial”), which is +19.28%+19.28\% and +17.21%+17.21\% more accurate than ReCoPosition Word{}_{\text{Position Word}}, and +26.98%+26.98\% and +32.97%+32.97\% higher than the T2I model with image description only. The skill correctness improvements suggest that region-control text inputs could be an effective interface to help T2I models more reliably generate user-specified scenes. The object accuracy evaluation makes the criteria more strict by requiring the model to follow the exact input region positions, in addition to skills. “ReCo” achieves a strong region control accuracy of 63.40%63.40\% and 67.30%67.30\% on count and skill subsets, surpassing “ReCoPosition Word{}_{\text{Position Word}}” by +38.05%+38.05\% and +44.48%+44.48\%.

PaintSkill contains input queries with randomly assigned object types, locations, and shapes. Because of the minimal constraints, many queries describe challenging scenes that appear less frequently in real life. We observe that ReCo not only precisely follows position queries, but also fits objects and their surroundings naturally, indicating an understanding of object properties. In Figure 5 (a), the three buses with different aspect ratios each have their unique viewing angle and direction, such that the object “bus” fits tightly with the given region. More interestingly, the directions of each bus go nicely with the road, making the image look real to humans. Figure 5 (b) shows challenging cases that require drawing two less commonly co-occurred objects into the same image. ReCo correctly fits “bed” and “fire hydrant,” “boat” and “bus” into the given region. More impressively, ReCo can create a scene that makes the generated image look plausible, e.g., “looking through a window with a bed indoors,” with the commonsense knowledge that “bed” is usually indoor while “fire hydrant” is usually outdoor. The randomly assigned region categories can also lead to objects with unusual relative sizes, e.g., the bag that is larger than the airplane in Figure 5 (c). ReCo shows an understanding of image perspectives by placing smaller objects such as “backpack” and “dog” near the camera position.

LVIS. Table 3 reports the T2I generation results with out-of-vocabulary regional entities. We observe that ReCo can understand open-vocabulary regional descriptions, by transferring the open-vocab capability learned from large-scale T2I pre-training to regional descriptions. ReCo achieves the best SceneFID and object classification accuracy over the 1,203 LVIS classes of 10.0810.08 and 23.42%23.42\%. The results show that the ReCo position tokens can be used with open-vocabulary regional descriptions, despite being trained on COCO with 80 object types. Figure 6 shows examples of generating objects that are not annotated in COCO, e.g., “curtain” and “loveseat” in (a), “ferris wheel” and “clock tower” in (b), “sausage” and “tomato” in (c), “salts” in (d).

Qualitative results. We next qualitatively show ReCo’s other capabilities with manually designed input queries. Figure 1 (a) shows examples of arbitrary object manipulation and regional description control. As shown in the “bus” example, ReCo will automatically adjust the object viewing (from side to front) and type (from single- to double-deck) to reasonably fit the region constraint, indicating the knowledge about object “bus.” ReCo can also understand the free-form regional text and generate “cats” in the specified region with different attributes, e.g., “wearing a red hat,” “pink,” “sleeping,” etc. Figure 7 (a) shows an example of generating images with different object counts. ReCo’s region control provides a strong tool for generating the exact object count, optionally with extra regional texts describing each object. Figure 7 (b) shows how we can use the box size to control the camera view, e.g., the precise control of the exact zoom-in ratio. Figure 7 (c) presents additional examples of images with unusual object relationships.

3 Analysis

Regional descriptions. Alternative to the open-ended free-form texts, regional descriptions can be object indexes from a constrained category set, as the setup in layout-to-image generation . Table 4 compares ReCo with ReCoOD Label{}_{\text{OD Label}} on COCO and LVIS . The leftmost “accuracy” column on COCO shows the major advantage of ReCoOD Label{}_{\text{OD Label}}, i.e., when fine-tuned and tested with the same regional object vocabulary, ReCoOD Label{}_{\text{OD Label}} is +7.28%+7.28\% higher in region control accuracy, compared with ReCo. However, the closed-vocabulary OD labels bring two disadvantages. First, the position tokens in ReCoOD Label{}_{\text{OD Label}} tend to only work with the seen vocabulary, i.e., the 80 COCO categories. When evaluated on other datasets such as LVIS or open-world use cases, the region control performance drops significantly, as shown in the “accuracy” column on LVIS. Second, ReCoOD Label{}_{\text{OD Label}} only works well with constrained object labels, which fail to provide detailed regional descriptions, such as attributes and object relationships. Therefore, ReCoOD Label{}_{\text{OD Label}} helps less in generating high-fidelity images, with FID 1.721.72 and 5.335.33 worse than ReCo on COCO and LVIS. Given the aforementioned limitations, we use the open-ended free-form regional descriptions in ReCo.

Guidance scale and T2I SOTA comparison. Table 5 (a,b) examines how different classifier-free guidance scales influence region control accuracy and image generation quality on the COCO 2014 validation subset . We empirically observe that scale of 1.51.5 yields the best image quality, and a slightly larger scale of 4.04.0 provides the best region control performance. Table 5 (c) compares ReCo with the state-of-the-art T2I methods in the fine-tuned setting. We reduce the guidance scale from the 4.04.0 in Table 1 to 1.51.5 for a fair comparison. We do not use any image-text contrastive models for results re-ranking. ReCo achieves an FID of 5.185.18, compared with 6.986.98 when we fine-tune Stable Diffusion with COCO T2I data without regional description. ReCo also outperforms the real image retrieval baseline and most prior studies .

Limitations. Our method has several limitations. First, ReCo might generate lower-quality images when the input query becomes too challenging, e.g., the unusual giant “dog” in Figure 7 (c). Second, for evaluation purposes, we train ReCo on the COCO training set. Despite preserving the open-vocabulary capability shown on LVIS, the generated image style does bias towards COCO. This limitation can potentially be alleviated by conducting the same ReCo fine-tuning on a small subset of pre-training data used by the same T2I model . We show this ReCo variant in the supplementary material. Finally, ReCo builds upon large-scale pre-trained T2I models such as Stable Diffusion and shares similar possible generation biases.

Conclusion

We have presented ReCo that extends a pre-trained T2I model for region-controlled T2I generation. Our introduced position token allows the precise specification of open-ended regional descriptions on arbitrary image regions, leading to an effective new interface of region-controlled text input. We show that ReCo can help T2I generation in challenging cases, e.g., when the input query is complicated with detailed regional attributes or describes an unusual scene. Experiments validate ReCo’s effectiveness on both region control accuracy and image generation quality.

ReCo: Region-Controlled Text-to-Image Generation (Supplementary Material)

Appendix A ReCo with LAION data

In the main paper, we focus on the ReCo model trained on COCO (ReCoCOCO{}_{\text{COCO}}) to standardize the evaluation process. In this section, we present ReCoLAION{}_{\text{LAION}} that conducts the same ReCo fine-tuning on a small subset of the LAION dataset used by the pre-trained SD model . Figure 8 shows selected ReCoLAION{}_{\text{LAION}}-generated image samples.

Training setup. Instead of using the 414K image-text pairs (83K images) from the COCO 2014 training set, we randomly sample 100K images from the LAION-Aesthetics datasetWe use the first 100K samples with an aesthetics score of 6 or higher following the index in https://huggingface.co/datasets/ChristophSchuhmann/improved_aesthetics_6plus.. We take the Detic object detector to generate the object region predictions. We use a confidence threshold of 0.50.5 and filter out small boxes with a size smaller than 0.03×W×H0.03\times W\times H. Following the setting for ReCoCOCO{}_{\text{COCO}}, we feed all cropped regions to the pre-trained GIT captioning model for regional descriptions. We fine-tune ReCo for 10,000 steps with the same training and inference settings introduced in the main paper.

Qualitative results. Figure 9 shows qualitative results on LVIS . Both ReCoCOCO{}_{\text{COCO}} and ReCoLAION{}_{\text{LAION}} show strong region-controlled T2I generation capabilities. Compared with ReCoCOCO{}_{\text{COCO}}, ReCoLAION{}_{\text{LAION}}-generated images have better image aesthetic scores, thanks to the high-aesthetic fine-tuning data from LAION .

Figure 10 shows qualitative results on LAION-Aesthetics. We run T2I inference on 3K samples indexed after the first 100K samples used for ReCo fine-tuning. ReCoLAION{}_{\text{LAION}} can preserve the pre-trained SD’s capabilities of understanding celebrities, art styles, and open-vocabulary descriptions, and meanwhile extend SD with the appealing new ability of region-controlled T2I generation.

Quantitative results. Table 6 compares ReCoLAION{}_{\text{LAION}} with ReCoCOCO{}_{\text{COCO}} on LVIS . The “COCO Image” column indicates if the COCO image style is seen during ReCo fine-tuning. Automatic metrics show that ReCoCOCO{}_{\text{COCO}} achieves better region control accuracy and image FID. For region control, COCO ground-truth boxes provide a cleaner region specification than Detic-predicted boxes, thus benefiting the controlling accuracy. For the FID evaluation, ReCoCOCO{}_{\text{COCO}} has seen COCO images during ReCo training, leading to better FID scores. Qualitatively, ReCoLAION{}_{\text{LAION}}-generated images show comparable, if not better visual qualities than ReCoCOCO{}_{\text{COCO}}. Overall, both ReCo model variants significantly outperform the original SD model in both region control accuracy and image generation quality.

Appendix B Position Token Cross-Attention

To help interpret how the introduced position tokens operate, Figure 11 visualizes the cross-attention maps between the visual latent zz and token embedding τθ(y(P,T))\tau_{\theta}(y(P,T)). We show the averaged attention maps across all diffusion steps and U-Net blocks. We empirically observe that the four position tokens for each region help the model to gradually localize the specified area by attending to the corner or edge positions of the box region. The position tokens help text tokens to localize and focus on the detailed regional descriptions, e.g., the “green light” in the “traffic light.”

We would like to thank Lin Liang and Faisal Ahmed for their help on human evaluation and data preparation, and Jaemin Cho and Xiaowei Hu for their helpful discussions. We would like to thank Yumao Lu and Xuedong Huang for their support.

References