ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
Peng Wang, Shijie Wang, Junyang Lin, Shuai Bai, Xiaohuan Zhou, Jingren Zhou, Xinggang Wang, Chang Zhou
Introduction
Representation models have received considerable attention in computer vision , speech processing , natural language processing , etc. Learning from large amounts of data, representation models demonstrate strong generalization ability in a wide range of downstream tasks. Furthermore, the explosive growth of large-scale language models (LLMs) has sparked an escalating appetite for representation models. Until recently, representation models have shown their bedrock role to unleash LLMs to understand, perceive, and interact with other modalities (e.g., vision) .
Due to the distinct characteristics of different modalities, previous research mainly focuses on building uni-modal representation models with individual architectures and pretraining tasks. Despite achieving excellent results, uni-modal representation models face difficulties in effectively utilizing multi-modal data such as image-text pairs and audio-text pairs, which makes them challenging to extend to multi-modal tasks. With the development of unified architectures and efficient pretraining tasks , recent works have achieved promising results in vision-language learning and audio-language learning . Nevertheless, there is still rare research on developing general models that can be applied to vision, audio, and language modalities. utilize the Multiway Transformer to process both image and text modalities with a unified masked prediction task for pretraining. The masked prediction task requires a pretrained CLIP model to discretize image data, which limits the scalability to other modalities such as audio. proposes a general pretraining method that can be applied to vision, audio, and language modalities without the need for third-party models (e.g., CLIP), but it doesn’t extend the method to multi-modal data.
In this paper, we explore a scalable way to build a general representation model toward unlimited modalities. We advocate that a general representation model should meet the following conditions: 1. The model architecture must be flexible enough to accommodate various modalities and support multi-modal interaction. 2. Pretraining tasks should not only extract information within each modality but also ensure alignment across modalities. 3. Pretraining tasks should be general and straightforward, allowing them to be applied to different modalities.
Driven by these motivations, we propose ONE-PEACE, a model with 4B parameters that can seamlessly align and integrate representations across vision, audio, and language modalities. The architecture of ONE-PEACE consists of multiple modality adapters and a modality fusion encoder. Each modality is equipped with an adapter for converting the raw inputs into feature sequences. The modality fusion encoder operates on feature sequences with Transformer architecture. Each Transformer block contains a shared self-attention layer and multiple modality Feed Forward Networks (FFNs). The self-attention layer enables interaction between the multi-modal features through the attention mechanism, while the modality FFNs facilitate information extraction within modalities. With the clear division of labor in this architecture, extending new modalities only requires the injection of adapters and FFNs.
To pretrain ONE-PEACE, we design two modality-agnostic pretraining tasks. The first one is cross-modal contrastive learning, it contains both vision-language contrastive learning and audio-language contrastive learning, which effectively align the semantic spaces of vision, audio, and language modalities. The second one is intra-modal denoising contrastive learning, it can be viewed as a combination of masked prediction and contrastive learning , where we perform contrastive loss between the fine-grained masked features and visible features, such as image patches, language tokens, or audio waveform features. These tasks collaborate to enhance the model’s fine-tuning performance while also maintaining cross-modal retrieval capability. Furthermore, they are universal for all modalities, which obviates the need for modality-specific designs. With the scaling-friendly model architecture and pretraining tasks, ONE-PEACE has the potential to expand to unlimited modalities.
We conduct comprehensive experiments on different tasks across various modalities, including vision, audio, vision-language, and audio-language tasks. Without using any vision or language pretrained model for initialization, ONE-PEACE achieves leading results in both uni-modal and multi-modal tasks, including image classification ( accuracy on ImageNet w/o privately labeled data), semantic segmentation ( mIoU on ADE20K), audio-text retrieval (outperforming previous SOTAs on AudioCaps and Clotho by a large margin), audio classification ( zero-shot accuracy on ESC-50, accuracy on FSD50K, accuracy on VGGSound w/o visual information), audio question answering ( accuracy on AVQA w/o visual information), image-text retrieval ( I2T R@1 on MSCOCO and I2T R@1 on Flickr30K w/o intermediate finetuning and ranking), and visual grounding (// scores on RefCOCO/+/g test sets).
Related Work
Recent years have witnessed the rapid development of vision-language pretraining. Early approaches relied heavily on region features extracted by object detectors, which is resource&time-consuming. With the increasing popularity of Vision Transformer , numerous works use Transformer to jointly learn vision-language data and demonstrate superior performance in downstream tasks . To facilitate alignment between vision and language modalities, researchers propose various efficient pretraining tasks. Among them, contrastive learning is one of the most representative methods that has been widely adopted in a lot of works . There also emerge some works explore the unified frameworks to handle vision-language tasks . adopts an encoder-decoder model to transform all vision-language tasks into generation tasks. uses contrastive learning and text generation as the pretraining objectives, thus can be applied to image-text retrieval and vision-language generation tasks. employs the Multiway Transformer to process vision-language data, and discretizes images into image tokens through CLIP for joint learning with text tokens.
Audio-Language Pretraining.
There is currently a significant amount of research being conducted in audio-language pretraining. One category of these works focuses on speech-text joint pretraining. For instance, some studies propose to train a unified encoder for speech and text, which utilizes a large amount of unlabeled speech and text with masked prediction tasks and paired speech-text data to learn alignment . There are also some works proposed to jointly pretrain speech and text under the encoder-decoder framework , which can be well applied to generation tasks, such as speech recognition and synthesis. Another category introduces cross-modal contrastive learning to audio-language pretraining . uses CNN14 and BERT to extract audio and text information respectively, and conducts contrastive learning on environmental sound data. further introduces more environmental sound data and trained with HTSAT and RoBERTa . It achieves state-of-the-art results in audio downstream tasks such as audio classification and audio-text retrieval.
Vision-Audio-Language Pretraining.
Recently researchers have begun exploring joint learning of vision, audio, and language modalities. employs a unified Transformer model to jointly learn video, text, and audio through cross-modal contrastive learning. utilize external models (e.g., VQ-VAE and AV-HuBERT ) to discretize the video and audio data, and train the models with masked prediction objectives. proposes a general self-supervised learning method that does not rely on external models. It successfully applies the method to vision, language, and audio modalities, but has not extended to multi-modal data.
Compared to previous works, ONE-PEACE has a flexible architecture that is compatible with multiple modalities. Furthermore, the pretraining tasks of ONE-PEACE are universally applicable without external models Therefore, ONE-PEACE can be easily extended to various modalities.
Method
The model architecture of ONE-PEACE consists of three modality adapters and a modality fusion encoder. The overall architecture is shown in Figure 1.
We design modality adapters to convert different raw signals into unified features. Note that these adapters do not interact with each other, which affords us the flexibility to choose appropriate networks for them, such as Transformers , CNNs , RNNs , etc. We design three lightweight modality adapters for ONE-PEACE:
Vision Adapter (V-Adapter). Given an image, we use a hierarchical MLP (hMLP) stem to patchify the image by gradually increasing the patch size to . There is no interaction between different patches. Then the image patches are flattened into a sequence and prepended with a vision class embedding. By adding the absolute positional embeddings into the image embeddings, the image representation is , where denotes the total number of image patches.
Audio Adapter (A-Adapter). Given an audio, we set the sample rate to 16kHz and normalize the raw audio waveform to zero mean and unit variance. Then the normalized waveform is processed by a convolutional feature extractor to get the audio embeddings. Instead of using the absolute positional embeddings, we use a convolution layer to extract relative position information and add it to the audio embeddings . With a prepended audio class embedding, we obtain the audio representation , where denotes the length of the audio representation.
Language Adapter (L-Adapter). Given a text, we first apply byte-pair encoding (BPE) to transform it to a subword sequence. Two special tokens and are inserted at the beginning and end of the sentence to indicate its start and end. Then an embedding layer is used to embed the subword sequence to the text embeddings. After summing the text embeddings with absolute positional embeddings, we obtain the text representation , where denotes the text sequence length.
Modality Fusion Encoder.
Following previous works , the modality fusion encoder is based on the Transformer architecture . We set up a shared self-attention layer and three modality feed-forward networks (FFNs) in each Transformer block. The shared self-attention layer enables the interaction between different modalities through the attention mechanism. The three modality FFNs (V-FFN, A-FFN, and L-FFN) can further extract information within their respective modalities. To stabilize training and enhance model performance, we make the following improvements:
Sub-LayerNorm. We incorporate Sub-LayerNorm into each Transformer block to enhance training stability. Specifically, We insert layer normalization before the input projection and output projection of each self-attention layer and FFN layer. In our preliminary experiments, we find that Sub-LayerNorm can achieve better performance compared to the Pre-LayerNorm .
GeGLU Activation Function. To further improve performance, we replace the activation function in FFN with GeGLU activation function. The intermediate dimension of FFN is set to times the embedding dimension, which is consistent with the practice of PaLM .
Relative Position Bias (RPB). For positional information, we introduce 1D relative position bias for text and audio, and 2D relative position bias for image . At the pretraining stage, the relative position bias of different self-attention layers is shared. At the fine-tuning stage, we decouple the relative position bias of each self-attention layer and let them inherit the weights of the pretrained relative bias.
LayerScale. We use LayerScale to dynamically adjust the output of each residual block. Specifically, before adding to the residual, we multiply the output of each layer (e.g., self-attention layer and FFN) by a learnable diagonal matrix, whose values will be initialized to . In our preliminary experiments, LayerScale is beneficial for stabilizing training and improving performance.
This ”sharing-separated” architecture enables ONE-PEACE to disassemble into different branches that handle tasks for various modalities. For example, the vision adapter, self-attention layer, and vision FFNs can be combined into the vision branch (V-Branch) to process vision tasks. Similarly, we named other branches as audio branch (A-Branch), language branch (L-Branch), vision-audio branch (VA-Branch), vision-language branch (VL-Branch), audio-language branch (AL-Branch), and vision-audio-language branch (VAL-Branch).
2 Tasks
The pretraining tasks of ONE-PEACE include cross-modal contrastive learning and intra-modal denoising contrastive learning. Cross-modal contrastive learning endows the model with cross-modal retrieval capability, while intra-modal denoising contrastive learning enables the model to achieve superior fine-tuning performance in downstream tasks. An illustration of the pretraining tasks is shown in Figure 2.
Cross-modal contrastive learning is a widely-used pretraining task that effectively aligns the semantic spaces of different modalities. The key idea of this method is to maximize the similarity of related sample pairs across different modalities while minimizing the similarity of unrelated sample pairs. Given a sample pair of arbitrary modalities (e.g., image-text pair or audio-text pair), we extract their features using the corresponding branches of ONE-PEACE. The outputs of the special tokens (e.g., vision class token or language class token) are regarded as global representations. Followed by a linear projection and normalization, we obtain the final representations and . The loss function is shown below:
where is the batch size, are indexes within the batch, and is a learnable temperature parameter (initialized to ). Following previous works , the cross-modal contrastive loss is computed by gathering negative features from all GPU devices. We apply cross-modal contrastive learning to image-text pairs and audio-text pairs, denoted by and respectively.
Intra-Modal Denoising Contrastive Learning.
Cross-modal contrastive learning mainly focuses on aligning features between different modalities. However, it lacks emphasis on the learning of fine-grained details within modalities, leading to suboptimal performance in downstream tasks . To address this issue, we further introduce intra-modal denoising contrastive learning to train ONE-PEACE Intra-modal denoising contrastive learning is similar to , but extends to more modalities.. Intra-modal denoising contrastive learning can be viewed as a combination of masked prediction and contrastive learning, where we perform contrastive loss between the fine-grained masked features and visible features, such as image patches, text tokens, or audio waveform features.
Given a sample of arbitrary modalities, we first encode it into an embedding sequence through the corresponding modality adapter. Then, we randomly mask some units (e.g., text tokens or image patches) within the sequence. Following , we only input the unmasked units to the modality fusion encoder to reduce computation costs and save memory. The encoded unmasked features are concatenated with the learnable mask tokens and fed to a lightweight Transformer decoder, which generates the masked features. We also use the ONE-PEACE model to encode the raw input sample into target features without masking. Finally, we perform the contrastive loss between the masked features and target features, the loss function is shown below:
Where is the representation of the masked unit, is the representation of the target unit, is the stop gradient operation. is the number of masked units within a sample, is the number of whole units within a sample. is a constant temperature value, we set it to . This loss not only encourages the masked units close to the positive units but also gets away from the negative units. As a result, each unit acquires a unique semantic meaning, which makes the model better transfer to downstream tasks .
We apply intra-modal denoising contrastive learning to types of data: image, audio, text, image-text pairs, and audio-text pairs. For image, we randomly mask patches, and the loss function used for this type of data is denoted by . For audio, we sample of all time-steps to be starting indices and mask the subsequent time-steps, and the loss function is denoted by . For text, we randomly mask tokens of the text sequence, and the loss function is denoted by . For image-text pairs, we randomly mask patches of the image and tokens of the text. The unmasked patches and tokens are concatenated together and encoded as masked features. The original image patches and text tokens are also concatenated together and encoded as target features. We then perform contrastive loss on the image patches and text tokens respectively, the average of these two losses is denoted by . For audio-text pairs, we randomly mask time-steps of the audio waveform and tokens of the text. The loss is similar to the above one, we denote it by .
3 Training
The overall pretraining process of ONE-PEACE is divided into two stages: vision-language pretraining and audio-language pretraining. At the vision-language pretraining stage, the model trains on image-text pairs and only updates parameters that are relevant to vision and language modalities. For each image-text pair, we not only utilize them to calculate and , but also separately using the image and text to calculate and respectively. The loss function at this stage is shown below:
At the audio-language pretraining stage, the model trains on audio-text pairs, and we only update A-Adapter, A-FFNs, and other audio-related parameters. The remaining parameters including self-attention layers are totally frozen. Despite not training on image-audio pairs, the semantic space between vision and audio is still aligned by using language as the anchor. The loss function at the audio-language pretraining stage is shown below:
Pretraining Details
The pretraining datasets of ONE-PEACE are divided into two parts: image-text pairs and audio-text pairs. For image-text pairs, we use LAION-2B , a dataset obtained by web crawling. For audio-text pairs, we collect a large amount of open-source environmental sound datasets. To ensure reproducibility, all pretraining datasets are publicly available. We provide more details about the pretraining datasets in Appendix A.1.
Pretraining Settings.
ONE-PEACE is a giant-size model with B parameters. We list the detailed hyper-parameters in Table 1. During pretraining, we introduce a lightweight Transformer decoder to recover the masked units from the visible units. The decoder is similar to the modality-fusion encoder, each block of it also consists of a shared self-attention layer and three modality FFNs. It has layers with hidden size, intermediate size, and attention heads. The model weights of ONE-PEACE are randomly initialized at the beginning, except for the audio feature extractor of A-adapter, for which we use the weights of WavLM’s feature extractor for initialization. We find that incorporating WavLM’s feature extractor significantly improves the model performance. More details about the pretraining settings are provided in Appendix A.2.
Training Acceleration.
We introduce several model acceleration and memory optimization techniques to accelerate the training. Firstly, we use the memory-efficient attention technique implemented in the xformers libraryhttps://github.com/facebookresearch/xformers to improve training speed. Secondly, we use the gradient checkpointing technique to save memory, which allows us to train the model with a larger batch size. Furthermore, we replace the layer normalization with Fused LayerNorm implemented in the Flash Attention libraryhttps://github.com/HazyResearch/flash-attention, and leverage nvFuserhttps://pytorch.org/blog/introducing-nvfuser-a-deep-learning-compiler-for-pytorch to fuse the operations of dropout, LayerScale, stochastic depth, and residual summing, which can bring additional speed improvements. To improve the training speed and prevent gradient overflow issues, we adopt Bfloat16 precision to train ONE-PEACE.
Experiments
We transfer ONE-PEACE to various mainstream vision benchmarks, including image classification, semantic segmentation, object detection, instance segmentation and video action recognition. We provide the implementation details in Appendix B.1.
In our experiments, we assess the image classification transfer performance of ONE-PEACE using the ImageNet-1K dataset, encompassing 1.28 million training images and 50,000 validation images distributed across 1,000 distinct categories. We also use intermediate fine-tuning on ImageNet-21k . As demonstrated in Table 2, ONE-PEACE obtains 89.8 top-1 accuracy on ImageNet with less token length . Note that FD-SwinV2-G, BEiT-3, and EVA all rely on the assistance of an external CLIP model for pretraining, while ONE-PEACE is trained from scratch without the help of external models. Even so, ONE-PEACE is able to achieve better results, which demonstrates its strong transferability.
Semantic Segmentation.
We experiment on ADE20k using ViT-Adapter for task adaptation and Mask2Former as the segmentation head. Following common practice, we first fine-tune the segmentation head on coco-stuff then fine-tune on ADE20k. As demonstrated in Table 3, ONE-PEACE establishes a new state-of-the-art, achieving a mean Intersection over Union (mIoU) of 63.0. This result indicates that ONE-PEACE exhibits exceptional transferring performance in the domain of dense prediction tasks.
Object Detection and Instance Segmentation.
We perform fine-tuning experiments on the COCO 2017 dataset. For the backbone, we employ the ONE-PEACE backbone and use the ViTDet with Cascade Mask-RCNN architecture, which incorporates a straightforward feature pyramid and window attention for addressing object detection and instance segmentation tasks. The model is fine-tuned on the COCO dataset. Soft-NMS is used during the inference stage. As illustrated in Table 4, the instance-level transfer capabilities of ONE-PEACE exhibit a performance that is on par with the current state-of-the-art methods.
Video Action Recognition.
We benchmark ONE-PEACE on Kinetics 400 dataset for video action recognition. Following AIM , we keep the whole model frozen and add several MLP adapters in each transformer layer. We use I3D head as the classification layer. As demonstrated in Table 5, without fine-tuning the full encoder, ONE-PEACE could achieve 88.1 top-1 accuracy, even outperforming CoCa which is pre-trained on privately collected data, and ViT-22B with 14x more parameters.
2 Results on Audio(-Language) Tasks
We evaluate ONE-PEACE on various audio and audio-language tasks, including audio-text retrieval, audio classification, and audio question answering (AQA). The implementation details are provided in Appendix B.2.
Table 6 presents the performance of ONE-PEACE and baseline models in the audio-text retrieval task. As a general representation model, ONE-PEACE achieves SOTA results on both AudioCaps and Clotho datasets, outperforming the previous audio representation model by a large margin. On AudioCaps, ONE-PEACE achieves 21.1% improvement on R@1 in text-to-audio retrieval and 11.4% improvement on R@1 in audio-to-text retrieval. On Clotho, ONE-PEACE achieves 23.1% improvement on R@1 in text-to-audio retrieval and 5.4% on R@1 in audio-to-text retrieval.
Audio Classification & Audio Question Answering.
Table 7 present the results of ONE-PEACE and baseline models in the audio classification and audio question answering (AQA) tasks. On ESC-50, ONE-PEACE achieves zero-shot accuracy, outperforming LAION-CLAP by . On FSD50K, ONE-PEACE significantly outperforms the previous SOTA by . For the VGGSound dataset, which consists of both visual and audio information, we only utilized the audio information and disregarded the visual information. With this setting, ONE-PEACE achieves score, surpassing the previous SOTA by . In the audio question answering task, ONE-PEACE outperforms the previous SOTA by . These results demonstrate the superior ability of ONE-PEACE on audio-related tasks.
3 Results on Vision-Language Tasks
We conduct experiments on various vision-language tasks, including image-text retrieval, visual grounding, visual question answering, and visual reasoning. the implementation details are provided in Appendix B.3.
Table 8 presents the performance of ONE-PEACE and baseline models on the image-text retrieval task. Under the fine-tuning setting, ONE-PEACE achieves the best performance in both MSCOCO and Flickr30K test sets. This indicates that after combining both cross-modal contrastive learning and intra-modal denoising contrastive learning, ONE-PEACE can effectively transfer to downstream retrieval task. Under the zero-shot setting, ONE-PEACE can achieve better or competitive performance compared to previous dual-encoder models like CLIP and Florence. Notice that the results of ONE-PEACE are inferior to CoCa, which might be because ONE-PEACE only trained on billion image-text pairs while CoCa trained on up to billion image-text pairs.
Visual Grounding.
To evaluate the capability of visual grounding, we conduct experiments on RefCOCO, RefCOCO+, and RefCOCOg datasets . Table 9 presents the results of ONE-PEACE and baseline models. It is worth noting that previous SOTA OFA use additional visual grounding datasets for training (i.e., Visual Genome ). Without introducing additional visual grounding datasets, ONE-PEACE still achieves new SOTA results on the datasets. We also compared the visual grounding ability of ONE-PEACE and OFA on an out-of-domain Pokémon picture.We use the Hugging Face spaces demo of OFA: https://huggingface.co/spaces/OFA-Sys/OFA-Visual_Grounding As shown in Figure 3, given a specific description of a Pokémon, both ONE-PEACE and OFA can obtain the correct result. However, when we directly provide the name of the Pokémon, OFA fails to obtain the correct result while ONE-PEACE can give a correct answer.
Vision-Language Understanding.
Table 10 presents the results of ONE-PEACE and baselines on two popular multimodal understanding tasks: visual question answering (VQA ) and visual reasoning (NLVR-2 ). For the VQA task, ONE-PEACE achieves a score of on the test-dev set and on the test-std set, outperforming previous strong baselines like CoCa and BLIP-2. For the NLVR2 task, ONE-PEACE surpasses CoCa with gains of 1.7 and 1.3 on the dev set and test-P set respectively. Notice that our results on both tasks are lower than BEiT-3. This may be attributed to two reasons: Firstly, BEiT-3 is pretrained on in-domain datasets such as MSCOCO and Visual Genome , which usually results in better downstream finetuning effects. Secondly, BEiT-3 incorporates pure text data for pretraining, which improves its language understanding ability and consequently enhances its multimodal understanding ability. In addition, OFA and BLIP-2 have shown that combined with language pretrained models can improve performance on multimodal understanding tasks. Therefore, we will explore the combination of ONE-PEACE and language pretrained models in the future.
4 Ablation Study
For the following ablation experiments, we utilize VIT-B/16 as the model backbone. The model is trained for 20 epochs with a batch size of 4096. We randomly selected 20 million image-text pairs from Laion-2B as the pretraining dataset.
We first conduct ablation experiments to investigate the effects of sharing or separating different modules. As shown in Table 11, sharing both self-attention layers and FFN layers yields better results compared to not sharing. This suggests that utilizing a single Transformer can effectively align the semantic space of vision and language. Furthermore, it is more beneficial to separate the FFN layer instead of sharing it. This implies that separating the FFN layer enhances the model’s ability to extract modality-specific information, leading to more accurate representations. We also find that separating the self-attention layer and sharing the FFN layer yields the poorest results. We speculate that this is due to the self-attention layer playing a more significant role in aligning modalities compared to the FFN layer. Therefore, separating the self-attention layer lead to inferior performance. Figure 4 demonstrates the convergence performance of different architectures. Among all the architectures, the model with shared self-attention layers and separated FFNs exhibits the fastest convergence speed.
Effects of Intra-modal Denoising Contrastive Learning.
We examine the effects of intra-modal denoising contrastive learning (DCL). As shown in Table 12, applying DCL to language data (DCL-L) can enhance the model’s performance in text retrieval tasks. Furthermore, applying DCL to vision data (DCL-V) can improve the model’s cross-modal retrieval ability, as well as fine-tuning performance in image classification. By applying DCL to vision-language data (DCL-VL), ONE-PEACE achieves the best results in terms of all the evaluation metrics. These results demonstrate that intra-modal denoising contrastive learning can complement cross-modal contrastive learning. It not only enables ONE-PEACE to achieve excellent downstream fine-tuning performance but also enhances the model’s capability for zero-shot cross-modal retrieval.
Ablation on Different Denoising Losses.
We conduct a systematic comparison of different denoising losses, including the smooth L1 loss used in , the L2 loss used in , the cosine loss used in , and the denoising contrastive loss used in this paper. As shown in Table 13, different types of denoising loss can improve the performance of the model in both cross-modal retrieval and image classification tasks. Among all the denoising losses, the denoising contrastive loss has the greatest improvement in terms of all the metrics compared to other losses. For example, it increased by on COCO text retrieval R@1, increased by on COCO image retrieval R@1, and increased by on image classification. This indicates that denoising contrastive loss is more compatible with cross-modal contrastive loss than other denoising losses.
5 Emergent Zero-shot Retrieval
In our pretraining, we exclusively align other modalities with text which plays as an intermediary role. We assume that our model is able to align those modalities that are not paired in the training data. For example, ONE-PEACE should be able to align image and audio. Thus, we conduct experiments on the retrieval of those modalities to assess the emergent zero-shot capabilities .
To be more specific, we evaluate the audio-to-image, audio+image-to-image, and audio+text-to-image retrieval abilities and demonstrate case studies in Figure 5. The first two cases demonstrate the emergent capability of uni-modal retrieval, while the other cases show that of the retrieval of image based on multimodal inputs. Specifically, we find that ONE-PEACE is able to retrieve images that contain elements concerning inputs of different modalities, e.g., the model uses the text “snow” and the sound of bird chirping to retrieve the images of birds in the snow. These examples demonstrate that ONE-PEACE has strong potential in emergent zero-shot capabilities. This indicates that for a universal representation model, there is no need to learn all pairing relationships between modalities, but instead it is sufficient for modalities to be aligned to an intermediary one. We provide more quality examples in Appendix E.
Conclusion, Limitation and Future Work
In this work, we explore a scalable way for building a general representation model across different modalities. Based on the flexible architecture and modality-agnostic pretraining tasks, we release ONE-PEACE, a general representation model that can seamlessly align and integrate representations across vision, audio, and language modalities. We conduct a series of experiments across 3 modalities, 11 tasks, and 16 datasets. The experimental results demonstrate that ONE-PEACE achieves leading results in a wide range of tasks, including image classification, semantic segmentation, audio-text retrieval, audio classification, audio question answering, image-text retrieval, and visual grounding. Furthermore, we show that ONE-PEACE possesses a strong emergent zero-shot retrieval capability, enabling it to align modalities that are not paired in the training data.
Although ONE-PEACE achieves leading results in a wide range of tasks, it falls short of achieving state-of-the-art results in zero-shot image-text retrieval and vision-language understanding tasks. There are two possible reasons for this: 1). ONE-PEACE didn’t see enough image-text pairs during pretraining. We only trained on 6.4 billion image-text pairs, while previous works typically train on 12.8 billion image-text pairs or more. 2). ONE-PEACE didn’t use language pretrained models for initialization or introduce any pure text data. Both the vision and language modules of ONE-PEACE are completely randomly initialized, while previous works show that introducing pure text data or initialized with the language pretrained models can greatly enhance the model’s performance. In fact, as a highly extensible model, ONE-PEACE can combine with language pretrained models to achieve better results.
Future Work.
In the future, we will test ONE-PEACE on more downstream tasks, such as vision-audio-language tasks and extend to more modalities for pretraining like video, 3D point cloud, etc. Also, we pursue an active interaction with large language models (LLMs) to continue influencing broader areas. This includes:
With the help of LLMs, building a more powerful general representation model.
By combining LLMs, creating a more general multimodal language model.
Acknowledgments
We would like to thank Yang Zhang, Benjin Mei, Dongkun Li, Jin Wang, Wei Wang, and Yinghui Liu for their support to this project, and we would like to thank Yusong Wu for patiently answering our questions about the audio dataset. We would also thank the M6-Team for providing a supportive research environment.
References
Appendix A Pretraining Details
For image-text pairs, we use LAION-2B , a dataset obtained by web crawling that may contain some noisy pairs. To improve the data quality, we apply several pre-processing steps, including removing images with an aspect ratio greater than 3.5, removing images with the shortest side less than 128, and removing images with a CLIP score less than 0.3. We also remove texts containing non-English or emoji characters, as well as texts with lengths less than 3 or greater than 512. After these steps, we retain about 1.5 billion image-text pairs.
For audio-text pairs, we mainly use the environmental sound datasets processed by . Specifically, for some datasets that only contain tags, uses a pretrained language model T5 to rewrite these tags into captions. We also perform simple cleaning on the data, which involves removing samples with text lengths less than 3 or greater than 512, as well as texts containing non-English or emoji characters. Ultimately, we obtain about 2.4 million audio-text pairs, with a total duration of around 8,000 hours. Table 14 presents the environmental sound datasets utilized by ONE-PEACE.
A.2 Pretraining Settings
As mentioned in Sec 3.3, the pretraining of ONE-PEACE is divided into two stages: vision-language pretraining and audio-language pretraining.
For vision-language pretraining, we pretrain ONE-PEACE for 200K steps with a batch size of 32768. We use the AdamW optimizer with and . The peak learning rate is set to , with a linear warmup of steps and a cosine decay scheduler. The image resolution is set to . The maximum text sequence length is set to 70. For regulation, we use weight decay with and disable dropout. We employ drop path with a rate.
For audio-language pretraining, we keep the model parameters related to vision and language (e.g., self-attention layers) frozen and only update the parameters that pertain to audio, such as A-Adapter and A-FFN. In this stage, we pretrain ONE-PEACE for epochs with a batch size of . The peak learning rate is set to , with a linear warmup of epoch and cosine decay scheduler. The maximum audio duration is set to s. For audio with a duration of less than s, we first repeat the input and then truncate it to s. Other hyper-parameters remain the same as vision-language pretraining.
Appendix B Details of Downstream Tasks
Here we describe the implementation details of different vision tasks, including image classification , semantic segmentation , object detection , and video action recognition . All detailed hyperparameters are listed in Table 15.
We provide the fine-tuning results on ImageNet-1k . Following recent studies in self-supervised learning for computer vision, we use global pooling of all image tokens excluding the class token, and append a LayerNorm with a linear layer for classification. To further unleash the potential of ONE-PEACE, we perform intermediate fine-tuning on ImageNet-21k . We set the label smoothing as 0.3 and do not use random erasing, mixup, and cutmix data augmentations. For fine-tuning on ImageNet-1k, we use exponential moving average (EMA) for model parameters and set the EMA decay rate as 0.9998. For intermediate fine-tuning on ImageNet-21k, we do not use EMA. We also use Zero Redundancy Optimizer and set the stage as 1.
Semantic Segmentation
We provide the fine-tuning results on ADE20k . We use Mask2Former as the segmentation head. We first intermediate fine-tune segmentation head on coco-stuff dataset for 80k steps. The learning rate is set as 2e-5 and the rest hyperparameters are the same as ADE20K shown in Table 15. Then we fine-tune the model on ADE20K. Both experiments use the cosine learning rate decay scheduler.
Object Detection
We provide the fine-tuning results on COCO with ViTDet . We use large-scale jitter data augmentation and fine-tune for 50 epochs. We use the linear learning rate decay scheduler and decay the learning rate at 44 and 48 epochs respectively.
Video Action Recognition
To perform video action recognition, following AIM , we freeze the parameters of the pre-trained model and add spatial and temporal MLP adapters in each transformer layer. We conduct experiments on Kinetics 400 dataset. Due to the invalid video links, there are many different versions of the K400 dataset and we use the version released on AcademicTorrents. We use the cosine learning decay scheduler and set the backbone learning rate multiplier of 0.1.
B.2 Audio-(language) Tasks
We describe the implementation details of audio-text retrieval, audio classification, and audio question answering here. All detailed hyperparameters are listed in Table 16.
We evaluate ONE-PEACE on AudioCaps and Clotho datasets. To get better results, we merge the training set of AudioCaps , Clotho , and MACS as the fine-tuning dataset. Similar to image-text retrieval, we use A-Branch and L-Branch to extract the features of audio clips and texts respectively, and then calculate the cosine similarity between these features. The recall@k is employed as the evaluation metric.
Audio Classification
We conduct experiments on three datasets: ESC-50 , FSD50K , and VGGSound . ESC-50 is an environmental sound dataset that contains environmental audio recordings and labels. We directly use the pretrained ONE-PEACE model to perform zero-shot audio classification on ESC-50. Specifically, we use A-Branch to extract audio embeddings from the audio clips and use L-Branch to extract text embeddings from the label names. Then we determine the labels of the audio clips by calculating the similarity between the embeddings. For FSD50K and VGGSound, we input the original audio into the A-Branch and utilize multi-head attention pooling (MAP) to aggregate the features. FSD50K is a multi-label sound event dataset, for which we use BCELoss as the loss function and report the mean average precision on the test set. VGGSound is an audio-visual dataset, where each sample includes a video with audio. We extract the audio clips from the videos and excluded the visual information, using cross entropy as the loss function and reporting accuracy on the test set.
Audio Question Answering
We conduct experiments on the AVQA dataset . Each sample in this dataset consists of a video, a question, and four candidate answers. To perform the audio question answering task, we extract audio clips from the videos and excluded the visual information. During training, we concatenate each answer with the audio and question, and extracted the features through AL-Branch. We then minimize the pairwise hinge loss between the positive features and negative features.
B.3 Vision-language tasks
Here we describe the implementation details of different vision-language tasks, including image-text retrieval , visual grounding , visual question answering , and visual reasoning . All detailed hyperparameters are listed in Table 17.
We evaluate ONE-PEACE on MSCOCO and Flickr30K datasets, and report the results on the widely used Karpathy test split . We use V-Branch and L-Branch to extract the features of images and texts respectively, and then calculate the cosine similarity between these features. The recall@k is employed as the evaluation metric.
Visual Grounding
This task requires the model to locate an image region based on a text description. We conduct experiments on RefCOCO, RefCOCO+, and RefCOCOg datasets . The image and text are fed to the VL-Branch simultaneously, then we use multi-head attention pooling (MAP) to aggregate the features from all image patches. The pooled output is used to predict the continuous corner coordinates of the bounding box , where and denotes the normalized top left coordinates, and denotes the normalized bottom right coordinates. We report the standard metric Acc@0.5 on the validation and test sets.
Visual Question Answering
This task requires the model to answer the question based on an image. We perform experiments on the VQAv2 dataset . Following previous works , we use the training and validation set of VQAv2 for training, including additional question-answer pairs from Visual Genome . The image and question are fed to the VL-Branch simultaneously, then we use MAP to aggregate the features from all text tokens. The pooled output is fed into a classifier to predict the answer from the 3,129 most frequent answers. We report the final score on the test-dev and test-std sets.
Visual Reasoning
Given a text and a pair of images, this task requires the model to distinguish whether the text truly describes the images. We conduct experiments on the NLVR2 dataset . Following the common practice, We treat each sample as two image-text pairs, each containing a text and one image. Then we input these pairs into VL-branch respectively. The final pooled outputs are concatenated together and fed to a classifier to predict the label. We report accuracy on the dev and test-P sets.
Appendix C Effects of Pretrained Audio Feature Extractor
We conduct a systematic analysis of the impact of the pretrained audio feature extractor. We find that although the parameters of the feature extractor are only M, accounting for only about of the total parameters, it has a significant impact on the model performance. As shown in Table 18, the feature extractor with random initialization only achieves accuracy on the ESC-50 dataset, while using pretrained feature extractors results in better performance. Notably, using the WavLM feature extractor can lead to the largest improvement (+6.2). We attribute this to the fact that WavLM is trained on a more diverse audio dataset compared to Hubert and Wav2Vec 2.0, making its feature extractor more suitable for environmental sound tasks.
Appendix D Evaluate ONE-PEACE on One Piece
We further test the visual grounding ability of ONE-PEACE by using a more complex anime picture, One Piece. The model is fine-tuned on the RefCOCOg dataset. As shown in Figure 6, we ask ONE-PEACE to locate the characters based on their names. Although ONE-PEACE hasn’t seen any anime pictures in the RefCOCOg dataset, it still achieves a recognition accuracy of 56.6%.
Appendix E More Examples of Emergent Zero-shot Retrieval
In this section, we provide more examples to demonstrate the emergent zero-shot abilities of ONE-PEACE, including audio-to-image, audio+image-to-image, and audio+text-to-image retrieval. The audios are selected from ESC-50 , and the images are retrieve from ImageNet-1K and MSCOCO . By reading this section, we hope that readers can better perceive ONE-PEACE.