Going Beyond Nouns With Vision & Language Models Using Synthetic Data
Paola Cascante-Bonilla, Khaled Shehada, James Seale Smith, Sivan Doveh, Donghyun Kim, Rameswar Panda, Gül Varol, Aude Oliva, Vicente Ordonez, Rogerio Feris, Leonid Karlinsky
Introduction
There have been impressive advances in the performance of zero-shot recognition through the use of large-scale pre-trained Vision & Language (VL) models . However, these VL models still face some important challenges in understanding Visual Language Concepts (VLC) beyond object nouns (e.g., recognizing attributes, relations, states) and in terms of compositional reasoning capabilities (i.e.., understanding subtle changes in meaning due to small changes in word order). Recently, several benchmark tests have been devised to demonstrate the extent to which these models lack these capabilities Please also see supplementary material for the expanded set of results of including results for all the most recent open-sourced VL models, all exhibiting poor VLC understanding performance.. As noted in several recent works , this behavior of VL models is likely due to the contrastive pre-training prevalent for all of them and likely inducing ‘bag-of-objects’ kind of representations (for both images and text alike). Indeed, for (even large) random batches of paired image-text samples, the collection of objects (nouns) in the image (or text) is likely to uniquely determine the image (or text) in the batch, making contrastive batch losses focus on the objects (nouns) while regarding other details (attributes, relations, states, word order, etc.) as unnecessary. Intuitively, this impairs VLC understanding and compositional reasoning of the resulting model.
Given the above, a natural question to ask is what is the most effective way to ‘fix’ the VL models to improve their VLC understanding and compositional reasoning performance? An approach proposed in concurrent works , advocates for the use of text augmentation, using language tools to teach a model the importance of non-noun words by manipulating them (e.g., replacing them with incorrect alternatives) and adding the resulting texts to the same batch. Although effective, such augmentation techniques are only easy on the text side and are much harder and prohibitively expensive on the image side. Indeed, finding, collecting, or generating real image samples sharing the objects but differing in their composition, attributes, relations, or states is very difficult. Although significant progress has been achieved with text-based editing , these methods are relatively slow (leveraging diffusion) and not sufficiently stable to allow effective use for augmentation in training pipelines. In this work, therefore, we propose an orthogonal route – VL data synthesis for fixing VL models by targeted demonstration. Specifically, we propose enhancing the VLC and compositionality aspects of the generated visual and text data, in turn using this data for finetuning VL models teaching them to pay closer attention to these aspects. Moreover, besides being largely free and infinitely scalable, synthetic data has an additional advantage – it can also be free from privacy concerns always accompanying real data.
Besides the inherent challenges of realistic data simulation, building synthetic data that can be effectively used to improve VLC and compositionality aspects of VL models pre-trained on massive real data poses additional technical challenges. Unlike the majority of prior work focusing on synthetic visual data generation, we need not only to generate images, but also the text that describes compositional items in a scene. We generate synthetic videos that leverage realistic physical 3D simulation including diverse 3D environments and different 3D objects, human motions, and actions assets , added interaction with objects, and different camera viewpoints. Every frame of these videos is accompanied by rich metadata, allowing using language grammar for generating detailed descriptive captions of any instantaneous scene in each video. These captions, in turn, allow collecting diverse image-text pairs samples contrasting which one to another highlights to the model the importance of the compositional items in the text captions (e.g. different viewpoints or different frames in the same video share objects but may strongly differ in the VLC and other compositional items). While motion assets were used by previous works to generate synthetic data , the visual data was not accompanied by textual captions and was not designed with the need to highlight compositionality in mind. We contribute Synthetic Visual Concepts (SyViC) – a large (million-scale) generated synthetic VL dataset with rich textual captions, easily extensible through our data synthesis code together with all the already generated million-scale synthetic data used in this paper (Figure 1).
In addition to the data synthesis pipeline, we also offer a strategy for effectively leveraging the generated synthetic data, while avoiding forgetting real data alignment and losing the strong a-priori zero-shot capabilities of the model. We propose and extensively ablate a combination of domain adaptation by stylization , parameter efficient fine-tuning , long captions handling, and model averaging methods to reduce forgetting, as well as examine the effect of different aspects of data synthesis and finetuning choices on the gains in VLC and compositionality understanding.
Our contributions can be summarized as follows: (i) we contribute SyViC – a million-scale synthetic dataset with rich textual captions, intended for improving VLC understanding and compositional reasoning in VL models, as well as the methodology and the generation codebase We release our code together with all million-scale synthetic data used in this paper here: https://github.com/uvavision/SyViC for its synthesis and potentially extensibility; (ii) an effective general VL model finetuning strategy enabling effective leveraging of SyViC data for enhancing the aforementioned aspects of strong pre-trained VL models without sacrificing their zero-shot capabilities; (iii) experimental results and extensive ablation study showing significant (over 10% in some cases) improvement in VLC understanding and compositional reasoning respectively, measured on all the recent VL-Checklist, ARO, and Winoground benchmarks and validated on the most popular CLIP model and its derivatives (e.g. the most recent CyCLIP ).
For supplemental materials, readers are referred to the associated arXiv document at [arXiv:2303.17590].
Related Work
Large-scale Vision&Language (VL) Models: Large-scale pre-trained VL models such as CLIP or ALIGN show remarkable success in many zero-shot recognition tasks such as image classification or detection . Despite the continued advancements made in this direction , recent studies (e.g., ) show that existing VL models exhibit limited comprehension of structured vision language concepts (VLC). Yuksekgonul et al. argue that contrastive learning for image-retrieval learns shortcuts and does not learn compositional information. To address this limitation, some approaches investigate how to augment the text captions or images in contrastive learning to enhance the ability of VLC . Smith et al. learn VLC concepts with additional supervised datasets in a continual learning setup. In contrast, we use 3D graphic engines to generate realistic synthetic videos with different compositions and generate corresponding text captions, which allows a VL model to learn compositionality and non-object words such as attributes, actions, relations, etc.
Learning from Synthetic Data. There has been a lot of work on learning from synthetic data in image classification , semantic segmentation , human pose estimation , action recognition , etc. Synthetic data is easy to generate and particularly useful for providing dense annotation such as semantic segmentation and depth estimation since these are prohibitively expensive to annotate manually. Some of the work relies on graphics engines to generate realistic data. Mishra et al. propose a method to learn how to generate task-adaptive synthetic data with the 3D simulation engine. For human-related problems, parametric body models (e.g., SMPL ) can be leveraged, along with motion assets , to generate synthetic human videos for low-level body analysis tasks or action recognition . Similar to our work, seek to associate semantic labels to synthetic images, but different from symbolic action categories, our focus is to assign rich textual descriptions to our generated images.
Since synthetic data suffers from a domain gap such as textures, visual styles, or colors from real images, domain adaptation, and generalization have been proposed to address this issue. Adversarial learning can be used to generate real-like images or feature alignment between synthetic and real data. Additionally, stylization methods are proposed as a style augmentation to make a model robust to diverse styles. In contrast, we manually randomize the visual content including different 3D objects, materials, and color attributes in graphics engines. Then we generate realistic synthetic videos from different domains with corresponding text captions. The generated data can be served as a hard negative augmentation and enhance the ability of VLC.
Method
We first present our synthetic data generation pipeline (Sec. 3.1), then describe how we leverage it for significant gains in VLC understanding and compositional reasoning capabilities of strong pre-trained VL models (Sec. 3.2). Our entire approach is illustrated in detail in Fig. 2.
In this section, we outline the components and the pipeline of our approach used to generate the proposed Synthetic Visual Concepts (SyViC) synthetic VL dataset for improving VLC understanding and compositional reasoning of VL models. Our contributed dataset includes 767,738 image-text pairs, 598K sampled from 1,680 diverse synthetic videos, and the remaining 169K generated as individual static synthetic scenes. Example samples from SyViC are provided in Supplementary.
3D physics-based simulation platform: ThreeDWorld (TDW) , which is built on top of Unity3D, is a multi-modal simulation platform that enables realistic physical interactions. TDW contains objects, 585 unique materials subdivided in metal, cardboard, wood, ceramic, and glass, and over indoor and outdoor scenes (3D environments). For generating synthetic VL data, we start with placing random objects in a scene following the workflow proposed by . We also use their camera positions and configurations to place objects visible inside good empty room perspectives. We group the available 3D object models by assigning dimension-related labels to each object and use the ImageNet category labels available for each object model as its associated text for later caption synthesis.
Camera Viewpoints: To further augment the set of plausible object placements and relations, we simultaneously place to cameras around a specific point of an empty room, and randomly place objects in the scene, allowing us to render images from different views of the same scene further strengthening the compositional aspects of the data as discussed in the introduction (Sec. 1). For each scene (frame), TDW cameras are able to capture RGB images, the corresponding object instance and category semantic segmentation masks, and a depth map. We use these, as well as a range of sensor and physics data representing the state of the world returned by TDW’s API, to enable dense annotations and supervision for each scene (frame) as part of our metadata generation process. We collect all of this information in our metadata and use it to estimate the position of the objects in the scene instead of relying on the 3D coordinates of each object and the camera position.
Digital humans: As we focus on compositionality aspects of images and text pairs, having people in our images is important. However, people models (especially animatable ones) are usually not present in common collections of 3D assets. Existing large-scale synthetic datasets often focus on realistically placing objects in a scene, but typically humans and animals are not included. We first inspected what libraries were available for realistic human synthesis. PeopleSansPeople , a library with 28 human 3D models and 39 unique actions, allows only random human placement, not allowing for humanoid customization or integration of human-object interactions. We leverage TDW support for Skinned Multi-Person Linear Model (SMPL) humanoids. SMPL is a parametric body model that enables a realistic representation of the shape and pose of arbitrary (non-clothed) 3D human bodies with diverse genders and shapes. SMPL models can be easily animated using motion capture data. The pose of the SMPL model is defined by a set of joint angles that determine the position of the corresponding body parts in the 3D space. The extended SMPL-X additionally allows controlling hand articulation and face expressions. Given the available library asset in TDW that enables placing these SMPLs in a scene, we create a stand-alone module to automatically incorporate arbitrary custom animations and unique human textures from the SURREAL and Multi-Garment datasets for clothing the synthetic human models for further enhancing the diversity and compositional features of our data.
Human motion synthesis and handling interactions: In Unity, SMPLs have skeletons and can be driven by motion-capture animations but they don’t have mass or colliders, this means that they are not physics assets since they can walk through other objects without interacting with them. To solve this issue, we add colliders to each body part of the SMPL model and create three asset templates (i.e., male, female, neutral) that contain all mesh configurations. Figure 2 shows some examples of our rigged asset templates with colliders. All other existing 3D models in TDW have colliders and track collisions at runtime through TDW physics-based simulation. Therefore, our collider-enhanced SMPLs are pulled downward by gravity, as well as simulate interaction by reacting naturally to collisions with other objects in the scene during motion simulation. For human actions, we first synthesize a diverse set of human actions from random language descriptions using TEACH , a Transformer-based model that generates a continuous sequence of SMPL motions from a sequence of action text labels. TEACH was trained on AMASS , a large-scale motion-capture (mocap) collection, and BABEL , a dataset that provides textual descriptions for AMASS, including per-frame unique actions annotations. Second, we extend our set of SMPL motions by directly sampling unique human motions from BABEL and AMASS. We export the corresponding mocaps to FBX files, and extract the animations in Unity, enabling them as asset bundles for use with TDW. FBX is a common format that facilitates data exchange between different 3D simulation platforms.
Domain randomization: One of the key qualities of the generated synthetic data, is its ability to highlight the importance of VLCs and compositional properties of the scene (e.g., objects attributes, relations, and more) in the contrastive learning objectives guiding the VL model finetuning. As opposed to methods based on text augmentation that can only enhance those on the text part of the VL image-text pairs, in SyViC construction we can easily manipulate also the visual content. We randomly place to 3D object models in the scene, randomizing their material and color attributes ( choices for each placed object). We randomly place to human avatars in the scene, randomizing their gender and clothing. We set camera poses as explained above, keeping scene ID shared for all the cameras of the same scene, and randomly sample viewpoints of each scene. We use human motion assets (as explained above) and randomly sample a motion sequence for each human avatar out of imported or generated mocaps. Finally, we sample an average of frames from each resulting synthetic video sequence generating image-text pairs describing a scene with the same objects and attributes, but in different arrangements and different corresponding captions thus enhancing the importance of the compositional aspects of the scene in the contrastive loss. We explore the importance of different simulation aspects in our ablation studies in Sec. 4.4.
Metadata-driven caption text synthesis: In addition to RGB frames, we obtain a large collection of rich metadata from the simulation platform, containing information on objects, humanoids, and the scene setting. For each frame, the metadata includes: (i) The world coordinates of each object and humanoid in the scene, including the camera position and viewing direction. (ii) The physical attributes of each object and humanoid in the scene (object physical attributes include color, size, and material; human attributes include the per-frame action label that changes over time and clothing description). (iii) Rendered depth images, instance segmentation masks, and category segmentation masks. Using the metadata, we compute the positional relations between each pair of objects and/or humans by comparing the pixels covered by their segmentation masks as well as their camera coordinates. Then, we use a simple grammar that deterministically maps positional relationships, object attributes, human attributes and action descriptions, and scene descriptions to a well-formed caption. More details on the grammar are provided in Supplementary.
2 Finetuning large-scale pre-trained VL models using synthetic data
In this section, we propose a methodology for effectively leveraging the SyViC synthetic VL data produced as explained in Sec. 3.1. We will use the following notation. Let be the text & image pair admitted by a VL model. The model (e.g., CLIP , CyCLIP ) components are denoted as: (i) image encoder ; (ii) text encoder . In this notation, the text-to-image similarity score is computed as:
where is the cosine similarity (inner product of normalized vectors). We next describe in detail the components of our finetuning strategy. Their merit and tradeoffs are thoroughly investigated in Sec. 4.4, arriving at the conclusion that parameter efficient finetuning + domain adaptive stylization + proposed caption splitting technique are the most effective combination. We also confirm in Sec. 4.4, that model averaging can provide expected trade-offs between VLC understanding and compositional reasoning gains and maintaining zero-shot performance.
Avoiding forgetting through parameter efficient fine-tuning: Inspired by , we use LoRA for VL fine-tuning with reduced forgetting of base model performance. We apply LoRA to adapt the encoders (, ) of a pre-trained VL model by parameterizing the adapted weights corresponding to the original model weights for each layer as:
where for of size , and are rank- matrices of sizes and respectively. These low-rank residual adapters can be applied efficiently during training and collapsed at inference time resulting in zero cost in terms of inference speeds or parameter counts . During finetuning all the base model parameters are frozen and only the LoRA adapters are being learned. Keeping rank low, the number of extra parameters added by all the LoRA adapters is low, consequently leading to significantly reduced forgetting in terms of largely maintaining the zero-shot performance of the original VL model.
Further reducing forgetting via model averaging: introduced an elegant technique to mitigate forgetting in finetuned models. All the parameters of the source model (before finetune) and the final model (after finetune) are averaged between the two models (typically with weight). We evaluate the effect of this on SyViC finetuned models in our ablation Sec. 4.4.
Domain adaption using style transfer: In addition, to mitigate the domain gap introduced by the use of synthetic data, we experiment with two style transfer techniques that align the content and feature statistics of the input frames with randomly-selected real-life images. A pre-trained Adaptive Instance Normalization (AdaIN) enabled encoder-decoder model was used to align the channel-wise statistics of each synthetic frame with a randomly-sampled image from the Human Motion Database (HMDB51) dataset thus generating a stylized synthetic image. We use AdaIN with an interpolation factor . In addition, in order to preserve the color information in the synthetic frames, we first match the color distribution of the sampled style image to that of the synthetic frame . We additionally experimented with MixStyle (using ImageNet as a source of real style images) as an extension of the DA stylization pipeline without observing significant gains over AdaIN.
Handling arbitrary caption length with caption splitting: The captions generated for SyViC are comprehensive: they contain descriptions of every object and/or humanoid visible in the frame as well as the pairwise positional relationship between objects. Intuitively, including these more elaborate (dense) descriptions in our captions gives a clear advantage in terms of promoting VLC understanding and compositionality following the finetuning of a VL model on SyViC. Hence, captions need to be sufficiently long texts that cannot be fully processed by common VL models (e.g. CLIP) text encoders () during training, as those are caped by relatively short max sequence context length (e.g. 77 for CLIP). Therefore, inspired by CLIP multi-caption strategy for inference , during training, we handle arbitrary caption lengths by splitting a given caption into sub-captions that can each be encoded separately and averaging the text features obtained from each sub-caption. In particular, the features of a caption of arbitrary length text is:
where is a sub-caption comprised of one or more sentences that fit into the text encoder max context size.
Losses: We employ the original models (e.g. CLIP and CyCLIP ) contrastive and other losses when training on SyViC with the aforementioned architectural and training protocol changes as explained above.
Experiments
For CLIP, we use the original OpenAI CLIP implementation and checkpoints. We modify their codebase to include LoRA adapters (Sec. 3.2), and use rank 16 in all our experiments. For CyCLIP, we adapt the implementation used in Code and checkpoints kindly shared by the authors.. For both CLIP and CyCLIP, we use a 5e-7 initial learning rate for finetuning and follow a cosine annealing learning rate schedule using an Adam optimizer. For all experiments, we use ViT/32-B as the model architecture and fine-tune it for six epochs on one A100 GPU with a total batch size of image-caption pairs. In addition to the original CLIP data augmentation transforms, we apply heavy random augmentation policies including manipulations in image inversion, contrast, sharpness, equalization, posterization, colorization, brightness, and solarization.
2 Datasets
To test the effectiveness of our proposed SyViC synthetic dataset and the accompanying finetuning approach for improving VL models’ VLC understanding and compositional reasoning capabilities we have evaluated on 3 benchmarks (Winoground , VL-Checklist , and ARO ) consisted of 7 datasets total.
VL-Checklist – is a large-scale dataset comprised of: Visual Genome , SWiG , VAW , and HAKE . Each image of these datasets is associated with two captions, a positive and a negative. The positive caption corresponds to the image and is taken from the source dataset. The negative caption is made from the positive caption by changing one word, so the resulting sentence no longer corresponds to the image. Depending on the word that was changed, VL-Checklist evaluates 7 types of VLC that can be divided into two main groups: (1) Attributes – color, material, size, state, and action, and (2) Relations – spatial or action relation between two objects and/or humans. In the following, we report average results for each of the main (Rel. and Attr.) groups on the combined VL-Checklist dataset. We also detail the individual improvements on all 7 VLC types in Fig. 3 (left).
Winoground – is a small dataset that evaluates the ability of VL models for compositional reasoning, specifically understanding the meaning of the sentence after changing the order of its words. The dataset has 400 samples, each comprised of two images and two texts. The texts have the same words in a different order, each text corresponding to one image in the sample. The Winoground metrics include (a) image score - percent of samples where the model picks the correct text for each image; (b) text score - percent of samples where the model picks the correct image for each text; (c) group score - percent of samples where both text and image score conditions are satisfied jointly. Recently, has analyzed Winoground for the source of its difficulty and found that only of its 400 samples are a valid subset. Other samples are not compositional, ambiguous, related to invisible details, have highly uncommon images or text, or require complex reasoning beyond compositionality. We report results on both the full Winoground and the ‘clean’ 171 images subset from .
ARO – or the Attribution, Relation, and Order benchmark, is a large dataset designed to evaluate the ability of VL models to understand four different types of skills. It consists of Visual Genome Attribution and Visual Genome Relation, which leverages the Visual Genome dataset along with the GQA annotations to test the understanding of properties and relational understanding of objects in complex natural scenes. VG-Relation includes distinct relations with test cases, and VG-Attribution includes unique attribute pairs with test cases. It also leverages the COCO and Flickr30k datasets to evaluate the model sensitivity to select the right caption after applying four different shuffling perturbations (e.g., exchanging nouns and adjectives, or by shuffling trigrams). These tests are performed on the and the images from the respective COCO and Flickr30k test splits.
3 Results
The main results of finetuning CLIP , CyCLIP – one of CLIP’s most recent improvements are summarized in Tables 1 and 2. All models were finetuned using our proposed approach and SyViC synthetic data to obtain their syn-model variants. Each model is compared to its respective source model pre-trained on large-scale real data before finetuning on SyViC. As we can observe, our SyViC synthetic data and the proposed finetuning recipe on this data demonstrate significant improvements over their source baselines. E.g. for CLIP obtaining , , and average absolute improvement in Winoground group score (most difficult average metric), VL-Checklist and ARO respectively. In addition, we illustrate the individual VLC metrics improvements obtained for CLIP in VL-Checklist and ARO benchmarks in Fig. 3 showing up to and respective absolute improvements. This underlines the effectiveness and promise of our method and SyViC synthetic data towards improving VLC understanding and compositional reasoning in VL models. Importantly, as we can see from Table 1, these strong gains come at a very small (under 1%) cost in the zero-shot performance of the respective VL models measured using the standard Elevater benchmark using 21 diverse zero-shot tasks.
4 Ablations
We extensively ablate our SyViC synthetic data and the proposed VL models finetuning approach on this data according to the following points. We use the most popular CLIP model finetuned on our SyViC synthetic dataset evaluated on the largest of the benchmarks - the VL-Checklist to perform our ablations.
SyViC - objects, humans, object attribute randomization – we evaluate the major components that comprise our SyViC synthetic data, namely the importance of the synthetic data to contain humans performing various motions and actions, the importance of having objects with randomized attributes (Sec. 3.1), and the final result of having all types of data combined. The results of this ablation are summarized in Tab. 3. As expected, humans alone cannot teach the model the needed skills only improving relations VLC by a small margin. Additionally, having only objects with randomized attributes improves attribute VLC, yet only improves relations by which is also expected, as many of the relations involve human actions. The best result is observed on the combined dataset with all the components.
SyViC - human clothing – we evaluate the diversity of human clothing comparing 3 levels of diversity: (i) none - using only a uniform color for human models; (ii) basic - using less diverse texture maps from SURREAL ; and (iii) most diverse - using texture maps from Multi-Garment , enriched with clothing colors, human age, and hair color annotations (manually done by us for the textures) which increase captions’ expressivity. Results are presented in Table 4. As expected, the most diverse human textures deliver the best result underlining the importance of this factor. Surprisingly, better human textures improve VL-Checklist Relations metric performance, likely due to the significantly better realism of the Multi-Garment textures.
SyViC - types of object attributes to randomize – Table 5 examines how randomizing different object attributes affects performance. Specifically, we evaluate the randomization of size, material, and color. Interestingly, we find that the best performance is achieved without color randomization. We suspect it is due to unnatural color-object combinations that arise under such randomization, which teach the model wrong beliefs on real objects’ color distributions and go against true object-color associations existing in the VL model following pre-training on the original VL data.
SyViC - types of captioning – we have investigated several variants of ways to obtain textual captions from SyViC metadata (Sec. 3.1). Results are summarized the Supplementary. We compared our proposed metadata grammar-based approach to two cascade methods that paraphrase the captions resulting from the grammar using zero-shot LLM inference (in-context learning). The paraphrased caption is then appended to the original grammar-based caption and consumed through our caption-splitting module (as standalone, open LLM-based paraphrasing is not very high quality). As can be seen, currently paraphrasing has minimal effect, but we posit it will become an important tool as stronger LLMs will become openly available.
SyViC - importance of physics and number of humans in the scene – we also looked into to which extent reliable physical simulation (made available in SyViC through TDW and Unity capabilities) and human-human positional and other relations are important for the observed VL model improvements. In Table 6 we evaluate the effects of removing the physics (resulting in humans or objects floating in space) or removing the multi-human scenes (thus preventing all human-human relations from appearing in the data). As expected, both reliable physics simulation and human-human relations (interactions of sorts) are mostly important to the gains in the Relations metric.
SyViC - number of models and number of samples – for lack of space, this is explored in the Supplementary.
SyViC - finetuning recipe components – in Table 7 we extensively evaluate the different components of the proposed finetuning approach on SyViC that leads to significant improvements on VL-Checklist, ARO, and Winoground VLC and compositional reasoning metrics. We start with vanilla CLIP finetuning on SyViC (row #1), already showing some improvement in VL-Checklist relations metrics, and on the ARO benchmark, yet losing to base CLIP on VL-Checklist attributes metrics. Adding our caption splitting module (Sec. 3.2) allows handling long (arbitrary size) texts outputted by our metadata-driven grammar and consequently utilizes all the caption information re-gaining the attributes performance (row #2). Adding parameter-efficient finetuning (LoRA, Sec. 3.2) regularizes finetuning by forcing smaller (low-rank, low-parameters) updates of the large-scale pre-trained CLIP model, consequently somewhat handling the expected domain gap between the synthetic data of SyViC and the downstream evaluation (real data) tasks. Notably, LoRA does not add any additional parameters to the model, all LoRA adapters are collapsed into the model weights after finetuning. Consequently, we observed significant improvements from adding LoRA in all metrics (row #3) with only minor degradation (0.7%) in ZS evaluation. With adding domain stylization (Sec. 3.2) we observe the best results in all VLC and compositional reasoning metrics improving ARO by 2.8% and keeping (even slightly improving) the advantages on VL-Checklist. Next, we investigate the variations of our best approach (LoRA + domain stylization + caption splitting). First, we investigate a strategy inspired by the LiT approach (row #5). Freezing the visual encoder as expected provides a (small) boost in ZS performance, but the reduced plasticity of the model comes at the price of observing smaller (only 3.1%) improvements on ARO and almost no improvements on the VL-Checklist. This leads us to conclude, that freezing the visual encoder is not a good strategy for SyViC finetuning. Next, we check the model averaging strategy (Sec. 3.2) inspired by (row #6). This does a better job of mitigating ZS forgetting, while at the same time keeping more of the gains on VL-Checklist and ARO. We conclude that model averaging is a good strategy to complement SyViC finetuning, allowing a soft trade-off between mitigating ZS forgetting and VLC and compositionality metrics gains. Finally, we again explore the importance of caption splitting for the best finetuning configuration of SyViC (row #7) and re-confirm its significance as performance drops without it.
Summary & Conclusions
Large vision and language models have dictated the status quo in computer vision and multimodal perception, achieving state-of-the-art results in a number of challenging benchmarks. However, existing models struggle with compositional reasoning and understanding concepts beyond object nouns, such as attributes and relationships. Our work has investigated, for the first time, whether synthetic data can be leveraged to mitigate these shortcomings. We proposed a data generation pipeline, used to create a million-scale dataset of synthetic images and accompanying captions, and an effective fine-tuning strategy with comprehensive analysis to enhance the compositional and concept understanding capabilities of multimodal models, without compromising their zero-shot classification performance.
Limitations. While we have achieved quite promising results in three different benchmarks, our work has limitations. As an example, our graphics simulator has a simplified model of lighting, sensor noise, and reflectance functions compared to the real world, which may impact robustness to color constancy. We believe more advanced domain adaptation and rendering techniques are likely needed to further improve our results. We also think a more detailed study of the scaling laws for synthetic data is a great research direction to fully unlock the potential of our work.
Acknowledgements. This material is based upon work supported by the Defense Advanced Research Projects Agency (DARPA) under Contract No. FA8750-19-C-1001. Any opinions, findings and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the Defense Advanced Research Projects Agency (DARPA). This project was also partially supported by NSF Award IIS-2221943, ANR CorVis ANR-21-CE23-0003-01, and the MIT-IBM Watson AI Lab. We thank Jeremy Schwartz and Esther Alter for their helpful discussions and assistance with the TDW package.
References
Appendix
In this supplementary material, we share our code and provide additional insights and experimental results that were not included in the main paper due to space constraints. In Section A we describe the implementation code for generating SyViC and the code for the proposed finetuning approach on SyViC. In Section B, we analyze the performance of recently open-sourced VL models on VL-Checklist and show they have low performance, demonstrating the need for our improvements. In Section C, we provide additional results on Winoground using CyCLIP (excluded from the main text for space purposes). Section D demonstrates how we can improve BLIP using SyViC. In Section E, we combine our contributions with those of concurrent work of and demonstrate that our approach for improving VLC using synthetic data is orthogonal / complementary to the text-augmentation based methods. Section F provides more dataset details, describing how the metadata from each synthesized scene is used to generate a caption for each image in SyViC. In Section G, we explore combinations of our metadata-driven grammar-based caption generation with paraphrasing using openly available large language models. Section H provides ”SyViC - number of models and number of samples” ablation excluded from the main paper due to lack of space. Finally, in Section I we provide some randomly sampled examples from SyViC.
Appendix A Code
Our code for both SyViC data synthesis and the proposed finetuning approach is included in our project page: https://synthetic-vic.github.io/
Appendix B Expanding VL-Checklist [70] analysis to most recent VL models
As promised in the footnote in the introduction we have evaluated the very recently released open-source VL models, namely: METER (CVPR 22) , X-VLM (ICML 22) , and VLMO (NeurIPS 22) on the most extensive VLC understanding benchmark of VL-Checklist observing average performance of 56.8%, 58.9%, and 54.6% respectively. As noted in the introduction, this relatively low VLC understanding performance (below CLIP ) of the newest (open) VL models illustrates once again the very much needed improvement in this aspect. Consequently, it also underlines the importance of SyViC and the proposed finetuning approach for administering some of this improvement and highlighting the future potential of our approach and synthetic data in general for the VL modeling. We additionally explore the very recent BLIP model and how it could be improved using SyViC and our approach in Section D.
Appendix C Winoground Results for CyCLIP [18]
As promised in the main paper (lines 555-556), we include the Winoground results of CyCLIP not included in the main paper for lack of space. The results are included in Table 8 and, compared to the CyCLIP baseline, demonstrate stable improvements of up to group score for syn-CyCLIP finetuned on SyViC using our proposed approach.
Appendix D Improving BLIP [31] using SyViC
BLIP is a recently released VL model achieving better out-of-the-box performance on VL-Checklist and Winoground compared to CLIP . In Table 9 we show how using our proposed SyViC dataset and the finetuning approach applied to BLIP, significant additional performance gains ( on VL-Checklist and up to on Winoground group score) can be achieved. BLIP is designed and optimized for VL understanding and generation , and has a relatively low zero-shot out-of-the-box performance compared to CLIP (e.g., we observed an over drop in zero-shot comparing baseline BLIP to CLIP and similar drop for syn-BLIP compared to CLIP). In more detail, we employ the retrieval flow of BLIP starting from ViT/B and CapFilt-L base model and use it as the BLIP baseline in Table 9. We finetune BLIP on SyViC following the complete proposed recipe detailed in Section 3.2 of the main paper. We add rank-16 LoRa adapters to both BLIP encoders and the decoder (cross-attention layers in the text encoder). We fine-tune for two epochs with a learning rate of 5e-6 using an Adam optimizer with a weight decay factor of .
Appendix E Exploring a combination with text augmentation methods
As discussed in lines 108-112 in the Introduction of the main paper, concurrent works propose an orthogonal approach of improving VLC understanding performance via using text augmentation while training on additional real VL paired image+text data. These works use language tools to teach a model the importance of non-noun words by manipulating them (replacing words with incorrect alternatives in the text captions of real image+text pairs) and adding the resulting texts to the same batch. In order to show that our proposed approach of improving the VLC understanding performance of VL models using targeted demonstration on both text and image side via generating synthetic data (our SyViC dataset) is truly orthogonal and complementary to the text augmentation methods, we have conducted the following experiment whose results are summarized in Table 10. Specifically, we used code, kindly shared to us by the authors, to combine our SyViC finetuning (as described in Section 3.2 of the main paper) with the LAION experiment of using their text augmentation method both for the real data captions as well as for our SyViC synthetic data captions. More specifically, we finetune (using the method described in Section 3.2 in the main paper, also including negative text augmentations and their additional losses for the negatives as detailed in ) on combined batches containing both LAION text+image pairs and SyViC text+image pairs. The base model is CLIP both for and syn-. As we can see in 10, syn- significantly (up to on Relations and on average) improves the base performance on VL-Checklist (trained on the same LAION data without SyViC) and is roughly matching performance on the ARO and zero-shot evaluations.
Appendix F Metadata-driven caption text synthesis, more details
This section describes how the metadata from each synthesized scene is used to generate a caption for each image in SyViC. We outline a rule-based mechanism to deterministically generate dense captions given:
List of the objects present in the scene, each with its corresponding name and world coordinates.
List of humanoids present in the scene, each with its world coordinates, clothing identifier, and a textual description of the action it performs.
Segmented image that has a label for each pixel corresponding to the object or humanoid it belongs to.
Scene identifier that maps to a textual description of the scene.
We annotate a list of 115 clothing textures from the Multi-Garment and SURREAL . Clothing annotations include a list of textual descriptions such as the colors of the clothes, the hair/beard style, as well as any features that stand out such as logos, tattoos, and accessories. Additionally, we use the original scene descriptions provided by ThreeDWorld’s scene library.
To generate the description of the objects in the image, we use the (3D) world positions of the objects to create positional relations between them. For objects that are horizontally aligned, we generate a description of which object is to the left or right of the other by comparing their corresponding pixels in the segmentation image. Furthermore, we generate a description of which object is in front of the other by translating the world coordinates into camera coordinates and comparing their z-coordinates. Unique identifiers for names are used to as placeholders to obtain those relationships, and object names are filled in once all positional relations are established, using indefinite articles when necessary. This process is applied to every pair of objects present in the scene.
Furthermore, we generate descriptions of humans while referring to them using ordinal numbers. In particular, for each human present in the scene, we retrieve its action description and place it in a sentence (e.g. ”The {first} person {walks forward}”. Additionally, we retrieve the list of textual descriptions associated with the human’s clothing, if exist. We consider each text as a separate sentence.
Finally, we compile a list of sentences containing a caption prefix, an enumeration of the objects, the pairwise positional relations between objects, a scene description, and action and clothing descriptions for each human. We concatenate the sentences together to get a full dense caption of the image. A simplified pseudo-code for generating a caption is shown below:
Evidently, dense captions tend to be way too descriptive and hence noisy to be used fully in VL training. Therefore, we add a sampling option where statements are sampled with certain probabilities following their weights. For example, instead of mentioning all pairwise positional relations, this option allows sampling a number of sentences from the positional relations category.
Appendix G LLM-Based Caption Paraphrasing
We additionally experiment with using the rule-based system to guide the use of large language models for caption generation / paraphrasing. Specifically, we adapt the deterministically-generated caption (as detailed in Section F) into a prompt for instruction-based text completion by replacing the prefix ”This scene contains” (in the synthesized captions) to ”Please describe a scene containing” and adding a suffix for text completion: ”In this scene, we can see”. We use the adapted texts as prompts for language models and generate text completions. We limit the generated texts to 150 tokens and use caption split averaging, as described in Section 3.2 in the main paper. We experiment with the Bloomz 7.1B and Flan-T5 XXL through Huggingface .
Table 11 shows the performance of syn-CLIP trained using different caption generation mechanisms. We do not observe any significant performance gains when using the captions generated by openly avaialble language models that we tried over the rule-based system. This is indeed expected, as the captions generated by current open language models tend to repeat much of the content in the prompt, often correcting verb tenses or adding appropriate punctuation marks, which don’t contribute to the semantic richness of the caption.
However, we remark that additional work on using LLMs for caption generation could investigate more powerful language models, or the use of visual grounding for caption generation as an additional information source, to yield better paraphrasing / captions.
Appendix H Exploration into Synthetic Data Diversity
As promised in Ablations Section 4.4 of the main paper (lines 740-741 in ”SyViC - number of models and number of samples”) we include the effect of the number of synthetic samples and the number of object models used for SyViC generation analysis in Figure 4. These ablations were not included in the main paper due to lack of space. As we can see the performance is improving consistently, both with adding more synthetic images (Fig. 4a) and with adding more 3D models used for synthesis (Fig. 4b).
Appendix I Some Qualitative Examples with Synthetic Humans
In this section, we first showcase qualitative improvement in the compositional capabilities of CLIP after finetuning on SyViC using our proposed approach via GradCAM in Figure 5. Next, we show textured SMPL samples in Figure 6.
Finally, we showcase some visual examples from SyViC along with the dense captions we generate describing human actions and detailed human-object interactions and relative position descriptions in the following pages.