Meta-Transformer: A Unified Framework for Multimodal Learning
Yiyuan Zhang, Kaixiong Gong, Kaipeng Zhang, Hongsheng Li, Yu Qiao, Wanli Ouyang, Xiangyu Yue
Introduction
The human brain, which is considered as the inspiration for neural network models, processes information from various sensory inputs, e.g. visual, auditory, and tactile signals, simultaneously. Moreover, knowledge from one source can benefit the comprehension of another. However, in deep learning, designing a unified network capable of processing a wide range of data formats is a non-trivial task due to the significant modality gap .
Each data modality presents unique data patterns, which makes it difficult to adapt models trained on one modality to another. For instance, images exhibit a high degree of information redundancy due to densely packed pixels, which is not the case with natural language . Point clouds, on the other hand, have a sparse distribution in 3D space, making them more susceptible to noise and challenging to represent . Audio spectrograms are time-varying and non-stationary data patterns consisting of combinations of waves across frequency domains . Video data contains a sequence of image frames, which gives it the unique capability to capture both spatial information and temporal dynamics . Graph data represents entities as nodes and relationships as edges in a graph, modeling complex, many-to-many relationships between entities . Owing to the substantial differences inherent to various data modalities, it is common practice to utilize distinct network architectures to encode each modality separately. For instance, Point Transformer leverages vector-level position attention to extract structural information from 3D coordinates, but it cannot encode an image, a natural language paragraph, or an audio spectrogram slice. Therefore, designing a unified framework capable of utilizing a modality-shared parameter space to encode multiple data modalities remains a significant challenge. Recently, the development of unified frameworks such as VLMO , OFA , and BEiT-3 have improved the ability of the network for multimodal understanding, through large-scale multimodal pretraining on paired data , but they are more focused on vision and language, and unable to share the whole encoder across modalities
The transformer architecture and attention mechanism, proposed by Vaswani et al. in 2017 for natural language processing (NLP), have made a significant difference in deep learning . These advancements have been instrumental in enhancing perception across different modalities such as 2D vision (including ViT and Swin Transformer ), 3D vision (such as Point Transformer and Point-ViT ), and audio signal processing ( AST ), etc. These works have demonstrated the versatility of transformer-based architectures, inspiring researchers to explore whether it’s possible to develop foundation models capable of unifying multiple modalities, ultimately achieving human-level perception across all modalities.
In this paper, We explore the potential of transformer architecture to process 12 modalities including images, natural language, point cloud, audio spectrogram, video, infrared, hyperspectral, X-Ray, IMU, tabular, graph, and time-series data, as shown in Figure 1. We discuss the learning process with transformers for each modality and address the challenges associated with unifying them into a single framework. Consequently, we propose a novel unified framework named Meta-Transformer for multimodal learning. Meta-Transformer is the first framework to simultaneously encode data from a dozen of modalities using the same set of parameters, allowing a more cohesive approach to multimodal learning (as shown in Table 1). Meta-Transformer incorporates three simple and effective components: a modality-specialist (§ 3.2) for data-to-sequence tokenization, a modality-shared encoder (§ 3.3) for extracting representations across modalities, and task-specific heads for downstream tasks. Specifically, Meta-Transformer first transforms multimodal data into token sequences that share a common manifold space. Then, a modality-shared encoder with frozen parameters extracts representations, which are further adapted to individual tasks by updating the parameters of downstream task heads and lightweight tokenizers only. Finally, task-specific and modality-generic representations can be effectively learned by this simple framework.
We conduct extensive experiments on various benchmarks of 12 modalities. By utilizing images of LAION-2B dataset for pretraining exclusively, Meta-Transformer demonstrates remarkable performance in processing data from multiple modalities, achieving consistently superior outcomes over state-of-the-art methodologies in different multimodal learning tasks. More detailed experimental settings can be found in § D.
In conclusion, our contributions can be summarized as follows:
For multimodal research, we propose a novel framework, Meta-Transformer, which enables a unified encoder to simultaneously extract representations from multiple modalities with the same set of parameters.
For multimodal network design, we comprehensively examine the functions of transformer components such as embeddings, tokenization, and encoders in processing various modalities. Meta-Transformer provides valuable insights and sparks a promising new direction in developing a modality-agnostic framework capable of unifying all modalities.
Experimentally, Meta-Transformer achieves outstanding performance on various datasets regarding 12 modalities, which validates the further potential of Meta-Transformer for unified multimodal learning.
Related Work
The development of various neural networks facilitates the perception of machine intelligence .
Multi-Layer Perceptron for pattern recognition. At the beginning, support vector machine (SVM) and multi-layer perceptron (MLP) are applied to text , image , point cloud , and audio classification. These innovative works merit the feasibility of introducing AI to pattern recognition.
Recurrent & Convolutional Neural Network. Hopfield Network is the original form of recurrent networks, then LSTM and GRU further explore the advantages of RNNs in sequence modeling and application in NLP tasks , which is also widely applied in audio synthesis . Meanwhile, the success of CNNs including LeNet , AlexNet , VGG , GoogleNet and ResNet in image recognition greatly promote the application of CNNs in other fields such as text classification , point cloud understanding , and speech classification .
Transformer. Recently, transformer architecture has been adopted in various tasks such as text understanding and generation in NLP, classification , detection and segmentation in images, point cloud understanding , and audio recognition .
However, similar to applications of CNNs and RNNs, these networks are modified according to distinct properties of modalities. There is no common architecture for modality-agnostic learning. More importantly, information from different modalities can be complementary , it’s significant to design a framework that can encode data from different modalities and bridge these complicated representations via a shared parameter space.
2 Transformed-based Multimodal Perception
The advantages of transformers for perception are the global receptive field and similarity modeling, which prominently facilitate the development of multimodal perception. MCAN proposes the deep modular co-attention networks between vision and language, which performs the cross-modal alignment by concisely maximizing the cross-attention. Then it becomes a consensus to utilize a cross-attention mechanism to bridge different modalities. With the success of pretrain-finetune paradigm, more works are getting focused on how to effectively align representations extracted across modalities by pretraining. VL-BERT pioneers modality-aligned representations for generic vision-language understanding with the MLM paradigm. Then Oscar described the object semantics in both visual and textural contents. Frameworks such as Vinvl , Simvlm , VLMO , ALBEF , and Florence further explore the advantages of joint representations across vision-language modalities in terms of semantic consistency.
Multimodal models are also utilized for few-shot learning , sequence-to-sequence learning , contrastive learning . BEiT-v3 proposes to take images as a foreign language with a more fine-grained cross-modal mask-and-reconstruction process, sharing partial parameters. And MoMo further explores the training strategy and objective functions while using the same encoder for images and texts.
Despite these advances, there remain significant obstacles to designing unified multimodal networks due to differences between modalities. Additionally, most research in this area has focused on vision and language tasks, and may not directly contribute to challenges such as 3D point cloud understanding, audio recognition, or other modalities. The Flamingo model represents a powerful few-shot learner, but its transferability to point clouds is limited, and it remains a challenge to leverage prior knowledge from one modality to benefit the others. In other means, existing multimodal methods have limited extensibility on more modalities, although they have taken expensive training costs. Addressing these discrepancies is dependent on bridging different modalities using the same set of parameters, akin to how a bridge connects multiple river banks.
Meta-Transformer
In this section, we depict the proposed framework, Meta-Transformer, in detail. Meta-Transformer unifies the multiple pipelines of processing data from different modalities and fulfills encoding texts, images, point clouds, audio, and the other 8 modalities with a shared encoder. To achieve this, Meta-Transformer is composed of a data-to-sequence tokenizer to project data to a shared embedding space, a modality-agnostic encoder to encode the embedding of different modalities, and task-specific heads to perform downstream predictions, as shown in Fig. 2.
Formally, we denote the input space of modalities as , while are the corresponding label spaces. In addition, we assume there exists an effective parameter space for each modality, where any parameter can be utilized for processing data from that modality. We say that the essence of Meta-Transformer is to find a shared that satisfies:
The multimodal neural networks can be formulated as a unified mapping function , where is the input data coming from any modality and denotes the prediction of the network. Let’s denote as the ground truth labels, the multimodal pipeline can be formulated as:
2 Data-to-Sequence Tokenization
We propose a novel meta-tokenization scheme designed to transform data across various modalities into token embeddings, all within a shared manifold space. This approach is then applied to tokenization, taking into account the practical characteristics of modality, as illustrated in Figure 3. We take text, images, point clouds, and audio as examples. More details can be found in supplementary materials. In specific, we use , , , and to denote a data sample of text, image, point cloud, and audio spectrogram.
Natural Language. Following the common practice , we use WordPiece embeddings with a 30,000 token vocabulary. WordPiece segments original words into subwords. For example, the original sentence: “The supermarket is hosting a sale”, could be converted by WordPiece to: “_The _super market _is _host ing _a _sale”.
Note that we use the same operation for infrared images but the linear projection for hyperspectral images. In addition, we simply replace 2D convolution layers with 3D convolution for video recognition. More details can be found in B.1 and B.3.
Audio Spectrogram. Initially, we pre-process the audio waveform with the duration of seconds with log Mel filterbank . Then we employ the Hamming window with a stride of on the frequency of to split the original wave into intervals and further transform the original wave into -dimensional filterbank.
Subsequently, we split the spectrogram into patches from time and frequency dimensions with the same patch size of . Different from image patches, audio patches overlap on spectrograms. Following AST , we also choose to split whole spectrograms into patches by convolution, then we flatten patches into token sequences. Finally, we summarize the process:
where and denote time and frequency dimensions.
3 Unified Encoder
After transforming the raw inputs to token embedding space, we leverage a unified transformer encoder with frozen parameters to encode the sequences of token embeddings from different modalities.
Pretraining. We utilize ViT as the backbone network and pre-train it on the LAION-2B dataset with contrastive learning, which reinforces the ability for generic token encoding. After pretraining, we freeze the parameters of the backbone network. In addition, for text understanding, we utilize the pretrained text tokenizer of CLIP to segment sentences into subwords and transform subwords into word embeddings.
Modality-Agnostic Learning. Following common practice , we prepend a learnable token to the sequence of token embeddings, and the final hidden state of token () serves as the summary representation of the input sequence, which is usually utilized for performing recognition.
To reinforce positional information, we incorporate position embeddings into the token embeddings. Recall that we tokenize the input data to 1D embeddings, thus, we opt for standard learnable 1D position embeddings. In addition, we do not observe substantial performance improvements using more sophisticated 2D-aware position embeddings on image recognition. We simply fuse the position embeddings and the content embeddings with an element-wise addition operation, and the resulting embedding sequences are then fed into the encoder.
where denotes the token embeddings from proposed tokenizer and denotes the number of tokens. We augment patch embeddings and learnable embedding with position embeddings .
4 Task-Specific Heads
After obtaining learning representations, we feed representations to the task-specific heads , which consists mainly of MLPs and varies from modalities and tasks. The learning objective of Meta-Transformer can be summarized as:
where , , and denote the function of tokenizer, backbone, and heads, respectively.
Experiments
In this section, we perform experiments on each of the 12 modalities. We demonstrate the potential of Meta-Transformer for multimodal perception. A summary of our experimental design is shown in Table 2 and more experimental details can be found in § C.1.
Text understanding. For text understanding evaluation, we employ the General Language Understanding Evaluation (GLUE) benchmark which incorporates several different datasets, covering a wide range of natural language understanding tasks.
Image understanding. 1) Classification: we conduct experiments on ImageNet-1K which contains approximately 1.3 million images with 1000 categories. Following common practices , base-scale models are trained for 300 epochs, while large models are pre-trained on ImageNet-22K (14.2 million images) for 90 epochs and fine-tuned on ImageNet-1K for another 20 epochs. 2) Object Detection: we conduct experiments on the MS COCO dataset using Mask R-CNN as the detector and training each model for 12 epochs. 3) Semantic Segmentation: we train the segmentation head UperNet on ADE20K for 160k iterations, providing a fair comparison with previous CNN-based and transformer-based backbones.
Infrared, X-Ray, and Hyperspectral data understanding. We conduct experiments on infrared image, X-Ray scan, and hyperspectral data recognition with RegDB , Chest X-Ray , and Indian Pine https://github.com/danfenghong/IEEE_TGRS_SpectralFormer/blob/main/data/IndianPine.mat datasets, respectively.
Point cloud understanding. 1) Classification: to assess the performance of Meta-Transformer in 3D object classification, we use the ModelNet-40 benchmark, consisting of CAD models across 40 classes, with 9,843 training samples and 2,468 validation samples. 2) Semantic segmentation: to evaluate performance in 3D point cloud segmentation, we assess the model on both S3DIS and ShapeNetPart datasets. The S3DIS dataset encompasses 6 large indoor areas and 13 semantic classes, comprising 271 rooms. The ShapeNetPart dataset includes 16,880 object models across 16 shape categories.
Audio recognition. For audio recognition, we utilize the Speech Commands V2 dataset, which consists of 105,829 one-second recordings of 35 common speech commands.
Video recognition. For video understanding, we conduct experiments on the UCF101 dataset for action recognition, with more details presented in § B.1.
Time-series forecasting. For time-series forecasting, we conduct experiments on ETTh1 , Traffichttps://pems.dot.ca.gov/, Weatherhttps://www.bgc-jena.mpg.de/wetter/, and Exchange datasets. We use the tokenizer of Autoformer .
Graph understanding. We conduct experiments on the PCQM4M-LSC dataset , which is a large-scale dataset consisting of 4.4 million organic molecules with up to 23 heavy atoms with their corresponding quantum-mechanical properties. With the target of predicting molecular properties using machine learning, it has plenty of applications in drug discovery, and material science.
Tabular analysis. We conduct experiments on adult and bank marketing from UCI repository http://archive.ics.uci.edu/ml/. We use the tokenizer of TabTransformer to encode raw tabular data.
IMU recognition. To evaluate the ability of Meta-Transformer to understand the inertial motion systems, we conduct experiments of IMU sensor classification on the Ego4D dataset.
Settings of Networks: We follow the default settings of ViT . denotes Meta-Transformer with a base-scale encoder which contains 12 transformer blocks and 12 attention heads, and the image patch size is 16. For the base-scale encoder, the embedding dimension is 768 and the output dimension of MLP is 3,072. ‘F’ and ‘T’ denotes that parameters of the encoder are Frozen and further Tuned, respectively.
2 Results on Natural Language Understanding
Table 3 illustrates the experimental results on the GLUE benchmark for text understanding tasks, comparing various state-of-the-art methods such as BERT , RoBERTa , and ChatGPT. The comparison centers on paraphrasing, sentiment, duplication, inference, and answering tasks. When using frozen parameters pretrained on images, achieves scores of 54.6% in sentiment (SST-2), 81.1% in paraphrase (MRPC), 66.0% in duplication (QQP), 63.4% in inference (MNLI), and 56.3% in answering (QNLI) tasks. After finetuning, Meta-Transformer-B16T exhibits improved performance, with 81.3% in sentiment, 81.8% in paraphrase, 78.0% in duplication, 70.0% in inference, and 60.3% in answering tasks. Although the Meta-Transformer’s performance on the GLUE benchmark might not be as impressive as that of BERT, RoBERTa, or ChatGPT, it still demonstrates competitive performance, adaptability, and potential for understanding natural language.
3 Results on Image Understanding
As shown in Table 4, Meta-Transformer exhibits outstanding performance when compared with Swin Transformer series and InternImage on image understanding tasks. On image classification, with the help of CLIP text encoder, Meta-Transformer delivers great performances under zero-shot classification with the and , achieving 69.3% and 75.3%, respectively. At the same time, when the pretrained parameters are further tuned, Meta-Transformer can outperform existing advanced methods, with and achieving 85.4% and 88.1% accuracy, respectively. The latter outperforms both SwinV2-L/24‡ (87.6%) and InternImage-XL ‡ (88.0%) on ImageNet classification.
When it comes to object detection and semantic segmentation, Meta-Transformer also delivers excellent performances, which further proves its generic ability on image understanding. On object detection, and achieve APs of 31.7% and 43.5%, while and reach 46.4% and 56.3% AP, respectively. In semantic segmentation, the mIoUs for and are 33.4% and 41.2%, while and achieve 51.0% and 55.0%, respectively. In comparison, SwinV2-L/24‡ outperforms the Meta-Transformer in both object detection (58.8% AP) and semantic segmentation (55.9% mIoU). The model has a similar performance to InternImage-XL‡ in semantic segmentation (both achieving 55.0% mIoU), but outperforms it in object detection (56.3% AP compared to 55.3% AP). These results highlight that Meta-Transformer demonstrates a competitive performance in various image understanding tasks even compared to Swin Transformer and InternImage.
4 Results on Infrared, Hyperspectral, and X-Ray data
Table LABEL:tab:infrared presents the performance comparison of Meta-Transformer and other advanced methods on the RegDB dataset for infrared image recognition. demonstrates competitive results with a Rank-1 accuracy of 73.50% and an mAP of 65.19%. While it may not outperform the top-performing methods, Meta-Transformer proves to be a simple transferable approach for infrared image recognition tasks. These results indicate the potential of Meta-Transformer in handling the challenges associated with infrared images and contribute to advancements in this field.
In addition, Table LABEL:tab:hyper presents the performance of Meta-Transformer on the Indian Pine dataset for hyperspectral image recognition. SpectralFormer achieves impressive accuracy scores, with a patch-wise approach. Plain vision transformer also performs well in comparison when fully tuning all parameters. demonstrates competitive results on hyperspectral image recognition with lower overall accuracy. However, Meta-Transformer stands out for its significantly fewer trainable parameters (only 0.17M) compared to other methods. This reveals a promising development direction of applying the Meta-Transformer to remote sensing, environmental monitoring, and mineral exploration. For X-Ray images, similar to dealing with infrared images, we take the same image tokenizer as common visible images. From Table 4.4, we can observe that Meta-Transformer can achieve a competitive performance of 94.1% accuracy.
5 Results on 3D Point Cloud Understanding
Table 6 showcases the experimental results for point cloud understanding, comparing the performance of Meta-Transformer with other state-of-the-art methods on the ModelNet-40 , S3DIS , and ShapeNetPart datasets. The tasks include classification, semantic segmentation, and object part segmentation. When pretrained on 2D data, demonstrates competitive performance, achieving an overall accuracy (OA) of 93.6% on ModelNet-40 with only 0.6M trainable parameters, which is comparable to the best-performing models. On the S3DIS Area-5 dataset, Meta-Transformer outperforms other methods with a mean IoU (mIoU) of 72.3% and a mean accuracy (mAcc) of 83.5%, using 2.3M parameters. Moreover, Meta-Transformer excels in the ShapeNetPart dataset, achieving the highest scores on both instances mIoU () and category mIoU () with 87.0% and 85.2%, respectively, using 2.3M parameters. In summary, Meta-Transformer demonstrates remarkable advantages in point cloud understanding tasks, offering competitive performance with fewer trainable parameters compared to other state-of-the-art methods.
6 Results on Audio Recognition
In order to fairly compare Meta-Transformer with existing audio transformer series of similar scale, we conduct experiments on audio recognition using Meta-Transformer-B32.
Table 4.6 showcases the performance of Meta-Transformer in the audio domain. These models are compared to existing methods such as AST and SSAST in terms of accuracy, all parameters (A-Params), and trainable parameters (T-Params). With frozen parameters, Meta-Transformer-B32F achieves an accuracy of 78.3% while requiring only 1.1M parameters for tuning. On the other hand, the Meta-Transformer-B32T model exhibits a significantly higher accuracy of 97.0% when tuning the parameters, whereas the AST model only reaches an accuracy of 92.6%. When AST is pre-trained on ImageNet and supplemented with additional Knowledge Distillation (KD), it achieves an improved performance of 98.1%, but with a higher number of trainable parameters of 86.9M. SSAST models display accuracy scores ranging from 97.8% to 98.0% while requiring 89.3M parameters. These results highlight that the Meta-Transformer performs competitively in the audio domain, demonstrating its versatility and effectiveness across different fields.
7 Results on Video Recognition
Table 4.6 presents the performance comparison of the Meta-Transformer and existing advanced methods on the UCF101 dataset for video understanding. Several state-of-the-art video-tailored methods achieve accuracies of over 90%. Meta-Transformer only contains a negligible amount of trainable parameters of 1.1 million to obtain an accuracy of 46.6% while other methods have to train around 86.9 million parameters. Though Meta-Transformer is not able to beat other state-of-the-art video understanding models, Meta-Transformer stands out for its significantly reduced trainable parameter count, suggesting the potential benefit of unified multi-modal learning and less architectural complexity.
8 Results on Time-series Forecasting
To explore the ability of Meta-Transformer for time-series forecasting, we conduct experiments on several widely-adopted benchmarks for Long-term forecasting tasks including ETTh1 , Traffic, Weather, and Exchange , with results shown in Table 10.
From Table 10, we can have the following observations. 1) With most of the model parameters being fixed, Meta-Transformer can still outperform existing methods including Pyraformer , Informer , LogTrans , and Reformer on these datasets. 2) The number of trainable parameters of Meta-Transformer is very few. With only 19K trainable parameters, Meta-Transformer can still outperform Informer . When 2M parameters are trained, Meta-Transformer can directly outperform Pyraformer . Therefore, Meta-Transformers pretrained on perception tasks can also be applied to time-series forecasting tasks, which is inspiring for this area.
9 Results on Tabular Data Understanding
Table 4.9 provides the comparison results about the performances of different methods for tabular data understanding on Adult Census and Bank Marketing datasets.
achieves a slightly lower accuracy than other methods on Adult Census but performs better than all other methods on Bank Marketing dataset in terms of accuracy and F1 scores. It suggests that Meta-Transformer is also advantageous for tabular data understanding, especially on complex datasets such as Bank Marketing.
10 Results on Graph and IMU Data Understanding
We report the performance of utilizing Meta-Transformer for graph understanding in Table 12. We compare with various graph neural network models for graph data understanding on the PCQM4M-LSC dataset . Among all the methods, Graphormer shows the best performance with the lowest train and validation MAE scores of 0.0582 and 0.1234, respectively. In contrast, delivers the train and validation MAE scores of 0.8034 and 0.8863, which reveals the limited ability of current Meta-Transformer architecture for structural data learning. We will further improve this in the future. Besides, following ImageBind , we conduct classification on the Ego4D dataset , with input data, Meta-Transformer delivers an accuracy of 73.9%.
Limitation
From the perspectives of complexity, methodology, and further application, the limitations of the Meta-Transformer are summarized as follows:
Complexity: Meta-Transformer requires computation dealing with token embeddings . High memory cost and heavy computation burden make it difficult to scale up.
Methodology: Compared with Axial Attention mechanism in TimeSformer and Graphormer , Meta-Transformer lacks temporal and structural awareness. This limitation may affect the overall performance of Meta-Transformer in tasks where temporal and structural modeling plays a critical role, such as video understanding, visual tracking, or social network prediction.
Application: Meta-Transformer primarily delivers its advantages in multimodal perception. It’s still unknown about its ability for cross-modal generation. We will work on this in the future.
Conclusion
In the early stages of artificial intelligence development, pioneers introduced the Multi-Layer Perceptron (MLP) to address prediction tasks in machine learning. Later, recurrent and convolutional networks expanded AI capabilities in multimedia data processing, achieving significant success in extracting representations from texts, images, point clouds, and audio. MLPs have since been integrated into deep convolutional networks. In this paper, we explore the potential of plain transformers for unified multimodal learning, highlighting a promising trend toward developing unified multimodal intelligence with a transformer backbone. To some extent, this paper supports the dominant position of transformers in next-generation networks. Importantly, CNNs and MLPs are not left behind. They play essential roles in data tokenization and representation projection. This process exemplifies the law of succession in neural networks and the ongoing evolution of artificial intelligence.
References
Appendix A Summary
The appendix is organized as the following:
We first validate and discuss the potential of the Meta-Transformer on more modalities (video, infrared, X-Ray, and hyperspectral images) in addition to the modalities shown in the main paper, and we provide surprising experimental results on these modalities in § B.
Then we further demonstrate the performance and merits of Meta-Transformer in dealing with multi-modal tasks (involving inputs from more than one modality to perform predictions) in § C.
In addition, we introduce more details of experiments on text, image, point cloud, and audio in § D.
Last but not least, we discuss the impact of Meta-Transformer on the machine learning and computer vision community in § E.
Appendix B Extensibility on Single-Modality Perception
In the main body of this paper, we illustrate that Meta-Transformer can simultaneously uncover the underlying patterns of natural language, 2D images, 3D point clouds, and audio spectrograms with the same network architecture and network parameters. Furthermore, we explore its ability in perceiving other modalities, like video recognition, infrared, X-Ray, and hyperspectral image recognition. In specific, we conduct experiments on UCF101 (video), RegDB (infrared images), Chest X-Ray , and Indian Pine (hyperspectral images) datasets.
For video recognition, we follow VideoMAE to modify the tokenizer by replacing the 2D embedding layer with a 3D embedding layer to simultaneously encode the spatial-temporal information from input frames. After tokenization, by leveraging the modality-shared encoder and task-specific heads, Meta-Transformer is able to extract high-level semantic features from videos and achieve favorable performance in the action recognition task of the UCF101 dataset.
Dataset. The UCF101 dataset is a common-used benchmark dataset for action recognition tasks. It is an extended version of UCF50 and contains 13,320 video clips of 101 categories. These 101 categories can be divided into 5 groups: Body motion, Human-human interactions, Human-object interactions, Playing musical instruments and Sports. All the input frames are with a resolution of 320240 and a fixed frame rate of 25 FPS, collected from YouTube.
B.2 Infrared Image Recognition
Infrared and hyperspectral image recognition poses unique challenges due to their specific characteristics. For infrared images, the Meta-Transformer framework could be adapted to capture thermal information by encoding temperature values alongside visual features, where the tokenizer for infrared images is the same as common RGB images.
Dataset. The RegDB dataset focuses on evaluating the performance of infrared recognition algorithms in unconstrained and realistic scenarios. It includes variations in pose, expression, illumination, and occlusion. We conduct experiments on the RegDB dataset to evaluate the performance of Meta-Transformer on infrared recognition.
B.3 Hyperspectral Image Recognition
Similarly, for hyperspectral images, we expect that Meta-Transformer can also handle the high-dimensional spectral information by representing each spectral band in token embeddings. Compared with dealing with RGB images, the only modification is that we employ the new linear projection layer to replace the existing 2D convolution layer.
Dataset. The Indian Pine dataset is widely used in remote sensing and hyperspectral image analysis. It consists of pixels with 145 spectral bands, which are captured in Indiana.
B.4 X-Ray Image Recognition
In addition, we explore the potential of the Meta-Transformer in medical image analysis. We leverage the tokenizer for RGB images here to encode raw medical images. Specifically, we conduct experiments regarding X-ray image analysis on the Chest X-Ray dataset. It is a collection of medical images commonly used for the analysis and diagnosis of various thoracic conditions. It comprises 7,000 X-ray images of the chest. The dataset is annotated with labels indicating the presence or absence of abnormalities such as lung diseases, fractures, and heart conditions.
Appendix C Extensibility on Multi-Modality Perception
Since the modalities of text, image, point cloud, and audio are all involved in this paper, we did not conduct comprehensive multi-modal experiments as common practice such as Flamingo , OFA , or BEiT-3 . Instead, we conduct multi-modal experiments on a new and challenging task of Audio-Visual Segmentation , which is mainly focused on building an intelligent listener to align with fundamental visual tasks.
Audio-visual segmentation refers to the task of segmenting objects from different audio sources within a referring image. It aims to develop algorithms that analyze both audio and visual signals simultaneously to identify and delineate distinct sources or events. It finds applications in fields like video conferencing, surveillance, multimedia analysis, and augmented reality.
We conduct experiments on the AVSS dataset, which is recently released in the field of audio-visual research. It provides a comprehensive collection of audio and visual data captured in real-world scenarios. The dataset includes synchronized audio and visual recordings, featuring various events of human actions and natural sounds. In contrast to introducing multi-modal fusion modules as existing methods, Meta-Transformer directly concatenates visual and audio embeddings after Data-to-Sequence tokenization. After extracting representation, we employ a simple global average pooling layer to obtain the final representations of two modalities.
Table 13 illustrates the performance of Meta-Transformer and existing methods on the AVSS dataset for audio-visual segmentation. The evaluation metrics reported in this task are mIou and F-score. In comparison, Meta-Transformer outperforms all other methods with the highest mIou of 31.33% and the highest F-score of 0.387. It also stands out for its significantly lower parameter count, with only 86.5 million parameters compared to the approximate 80M to 180M parameters of other methods.
Meta-Transformer offers several advantages over other methods in the field.
Unified architecture. It relieves modality-specific encoders and reduces computation by leveraging a unified encode to process both audio and images, resulting in a more efficient and streamlined process.
Faster convergence. Thanks to the unified architecture for processing both audio and images, the encoder can deeply align the two modalities instead of only at the output end, which leads to faster convergence. Meta-Transformer only needs 4 training epochs to reach 31.33% of mIou.
Superior performance. Meta-Transformer achieves a significant improvement of compared to other methods of a similar parameter scale.
Efficiency. Despite its enhanced performance, Meta-Transformer achieves this with much fewer parameters, requiring only of the parameter amount, which makes forward and backward progress ease.
In summary, the benefits of employing the Meta-Transformer to deal with multi-modal tasks are appealing due to computational efficiency, rapid convergence, improved performance, and parameter efficiency. It reveals the significantly promising direction to apply Meta-Transformer to more multi-modal tasks.
Appendix D Experimental Details
Our code is built on open-source projects including MMClassificationhttps://github.com/open-mmlab/mmpretrain/tree/mmcls-1.x, MMDetectionhttps://github.com/open-mmlab/mmdetection, MMsegmentationhttps://github.com/open-mmlab/mmsegmentation, OpenPointshttps://github.com/guochengqian/openpoints, Time-Series-Libraryhttps://github.com/thuml/Time-Series-Library, Graphomer https://github.com/microsoft/Graphormer.
We sincerely thank their great contributions. More implementation details can be found in our source code.
Appendix E Further Impact Discussion
We hope that Meta-Transformer can introduce new insight into both multi-modal learning and multi-modal generation fields. Meta-Transformer enables the usage of a shared encoder to encode diverse modalities, e.g. natural language, 2D images, 3D point clouds, as well as audio spectrograms., and project them into a shared representation space. This naturally reduces the modality gap across modalities and mitigates the burden of cross-modal alignment. In addition, Meta-Transformer removes the need for paired training data (such as image-text pairs), thus endowing multi-modal learning with more training flexibility.
E.2 Application Prospects
We investigate the application of Meta-Transformer on a wide range of modalities including RGB images, text, point clouds, video understanding, remote sensing (hyper-spectral images), nighttime surveillance (infrared images), and medical analysis (X-Ray images).
In video understanding, Meta-Transformer reveals the potential of enhancing the analysis and interpretation of videos by integrating information from text, audio, and image with the shared encoder. This benefits tasks such as action recognition, event detection, and video summarization. Meta-Transformer’s capability to handle video-related modalities paves the way for improved video understanding applications in areas like video surveillance, video indexing, and content-based video retrieval.
In hyperspectral imaging for remote sensing, Meta-Transformer enables the analysis and understanding of hyperspectral data by extracting high-level semantic features. It enhances tasks such as classification, target detection, and land cover mapping, improving the accuracy and efficiency of remote sensing applications. The ability to process hyperspectral images using Meta-Transformer opens doors for advancements in environmental monitoring, agriculture, urban planning, and disaster management.
In medical applications, particularly X-ray image analysis, Meta-Transformer offers a promising approach to improving diagnostic accuracy and efficiency with multi-modal information. It can effectively capture and fuse information from X-ray images, clinical data, and other modalities to aid in disease detection, anomaly identification, and treatment planning by leveraging its unified learning framework. Meta-Transformer’s capability to handle multi-modal data enhances the potential for more accurate and comprehensive medical imaging analysis, leading to better patient care and outcomes.
For infrared images used in nighttime recognition and surveillance, Meta-Transformer’s ability to process infrared data helps extract crucial information for object detection, tracking, and recognition in low-light conditions, which opens an avenue for advancements in nighttime surveillance, security systems, and autonomous navigation in challenging environments with the cooperation between infrared cameras with RGB cameras.
E.3 Conclusion
In summary, we think that the ability of Meta-Transformer to unify multi-modal learning comes from that neural network architectures can learn modality-invariant patterns. The architecture of Meta-Transformer illustrates the advantages of length-variable token embeddings in multi-modal learning, which provides flexible but unified forms of multi-modal semantics. Then it’s time to think about designing algorithms to train networks that generalize on unseen modalities. Meanwhile, it’s also intriguing to design the architecture of a unified multi-modal decoder, which can decode representations into any form of a specific modality.
Although Meta-Transformer presents a surprising performance and shows a new promising direction in multi-modal perception, we are not sure whether the proposed architectures are also effective in generative tasks. And it remains mysterious how to develop modality-invariant generative models. We hope that this can inspire future research.