OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning
Siddharth Srivastava, Gaurav Sharma
Introduction
Extracting meaningful representations from data is a central task in machine learning. Majority of the approaches proposed are usually specialized for specific modalities and tasks. The development of methods capable of handling multiple modalities, in a holistic way, has been an active topic of research recently . Multi task learning has a large body of literature , but has been traditionally limited to tasks from single modality. Learning a unified network that trains shared parameters across diverse tasks in different modalities, like image, video, depth maps, audio, has been shown to be more robust and give better generalization and reduce overfitting to a single task or modality cf. unimodal networks. Such joint learning also enables more efficient use of available labeled data across various modalities, potentially reducing the need for extensive labeling in specific modalities for particular tasks.
In the present work, we extend such line of research and propose a multimodal multitask method which learns embeddings in a shared space across different modalities and then employs task specific sub-networks for solving specific tasks in specific modalities.
The method utilizes a common transformer based bottleneck block to map the input to embeddings in a shared space, thus incorporating knowledge from multiple tasks associated with different respective modalities. This structure leads to learning of very robust representations informed and regularized by all tasks and modalities together. The embeddings are then used by the task heads to make required predictions.
Previous research in generalized multimodal learning falls into three main categories. First, there are methods that process multiple heterogeneous modalities such as images, 3D, and audio, directly without using separate encoders for each modality, learning representations directly from these inputs . Second, some approaches use modality specific encoders and then learn generalized embeddings, for data from each modality, based on a unified objective in the latent space . Third, there are methods focused on knowledge sharing across different modalities, employing either a single common encoder or distinct encoders for each modality . Our work aligns more closely with the third type of approaches, while incorporating elements from the first. We employ modality specific tokenizers and encoders, and have a bottleneck shared transformer backbone. Tokenization is tailored to each modality, drawing inspiration from the Uni-Perceiver model but with key modifications detailed in Sec. 3. After tokenization, transformer based network is used to obtain initial representations for the modalities which are passed through fully connected layers and then fused together with cross attention module. The fused representation then passes through the transformer backbone. The features from the transformer are then individually fused with original modality features using cross attention and are in turn fed to the modality specific task head.
In summary, the contributions of the work are as follows. (i) We propose a multimodal multitask network based on transformer architectures with modality specific tokenizers, shared backbone, and task specific heads. (ii) We provide comprehensive empirical results on 25 benchmark datasets over 12 distinct modalities i.e. text, image, point cloud, audio and video along with applications to X-Ray, infrared, hyperspectral, IMU, graph, tabular, and time-series data. The method achieves better or close to state of the art performances on these datasets. (iii) We propose a novel multimodal pretraining approach that alternates between a pair of modalities to enable crossmodal knowledge sharing. (iv) We propose a multimodal and multitask supervised training approach to leverage knowledge sharing between modalities for robust learning, simplifying the complex processes proposed in previous works on modality integration, e.g. .
Related Works
In this section, we discuss similar works and various similar paradigms to our work.
Multi-modal methods. Contemporary multi-modal methods predominantly employ modality-specific feature encoders , focusing on fusion techniques within their architectural designs. These networks usually vary across modalities, necessitating architectural modifications for combined usage. They must address challenges related to feature fusion timing, fine-tuning, and pre-training etc. . Such complexities restrict the adaptability of universal frameworks like transformers for diverse domains, including point clouds, audio, and images.
Common network for multiple modalities. A growing body of research aims to learn from multiple modalities without modality-specific encoders . Notably, architectures like the perceiver employ cross-attention among latent queries to process multiple modalities together. The hierarchical perceiver expands on this by structuring the input while maintaining locality. Other approaches, such as data2vec , use modality-specific encoders. Omnivore , with a common encoder, is limited to visual modalities only. Contrarily, VATT employs a unified transformer backbone but processes each modality independently. These multi-modal methods have demonstrated enhanced robustness .
Multi-task learning. As explored in the preceding section, there has been a surge in methods that process multiple modalities. PerceiverIO extends the capabilities of Perceiver to facilitate learning multiple tasks with a singular network architecture. Although PerceiverIO is capable of multitasking, often separate networks are employed . Various techniques learn from raw representations of multiple modalities and are applicable to numerous tasks.
Multi-modal masked pretraining. Approaches such as implement masked pre-training. This technique has proven beneficial for improving the performance of deep networks across different modalities and tasks.
Comparison to similar works. We draw motivations from UniPerceiver , MetaFormer and OmniVec . Unlike UniPerceiver line of methods, we do not use a unified task head definition, while similar to it we use task specific task heads. This allows our method to learn more robust and leverage fine details from each task depending upon the complexity of the tasks, which is important as each modality has distinct definition of complexity. For ex., in vision task, classification is a relatively simpler task as compared to segmentation, as segmentation tasks enforces networks to learn pixel level attention and learning better neighbourhood relationships . Further, MetaFormer uses unified tokenizers, and instead, we utilize modality specific tokenizers. Our experiments indicate that modality specific tokenizers perform better than MetaFormer’s unified tokenizer when training on multiple modalities. Further, OmniVec uses separate encoders for each modaity, that makes the network heavy and computationally expensive. In contrast, we use modality specific tokenizers with a shared backbone. Additionally, unlike other works, we train on multiple modalities in a multi task manner, allowing the network to learn from multiple modalities with varying task complexities simultaneously.
Approach
where are the parameters of all the predictors, and is the training set provided for task of modality . This is the extension of multiple task learning to multiple modalities as well.
Along with the supervised multimodal joint training explained above, the learning also consists of two stages of unsupervised masked pretraining with the first stage being unimodal and the second stage being multimodal pretraining, to achieve knowledge sharing between tasks and modalities leading to regularized and robust predictors. We now present each of the components and the full training algorithm in detail.
1 Network components
We now go through the network components sequentially from input to output. Tokenizers. Each modality is tokenized using a modality specific tokenizer. The tokenizers are similar to those used in Uni-Perceiver , however, instead of attaching an embedding to the tokens, we provide transformer with one type of modality at a time. Further, Uni-Perceiver utilizes a combination of tokens from multiple modalities passed to a single transformer. This limits the Uni-perceiver to a limited set of modalities, i.e. text, image and videos. However, our method does not suffer from any such limitation. The details of specific tokenizers for the different modalities are provided in Supplementary. Feature transformation network. Once the features are tokenized, they are then passed through a transformer network. While the method can utilize any transformer backbone, in the current implementation we use a transformer based on BERT . Here, the multi head attention involves standard self-attention , and GeLU activation prior to the MLP layer. The output from the transformer network is passed to a fully connected neural network with three fully connected layers with ReLU activation. This transformer network along with the fully connected layers is denoted a in Fig. 1. The network could be used without the fully connected layers—we added the fully connected layers to reduce the dimensions of the features so that the computational complexity of the remaining part of the network could be reduced. Mixing features with cross attention. When training, we fuse the features from the two transformer streams, corresponding to two modalities, with cross attention module. The output fused features are then passed to another transformer network, denoted a in Fig. 1. The architecture of the transformer network is same as the transformers used in feature transformation network. Modality and task specific heads. The part of the network are the modality and task specific heads, denoted a in Fig. 1. These task heads take as input, features from respective modality streams fused with features from the above network, fused with cross attention module. The task heads consist of a vanilla ViT-Tiny networks .
2 Training
The training is done in three steps: (i) masked pretraining iterating over modalities but doing masked prediction with one modality at a time, (ii) multimodal masked pretraining where two modalities are simultaneously used to do masked prediction for each, and (iii) finally supervised task based training. Stage 1 masked pretraining. The first step in training is self supervised pretraining of the transformer in the feature transformation network. We follow earlier works and add a decoder for predicting masked tokens. Specifically, for an input modality with patches, we randomly mask patches, and feed non-masked patches and their positions to an encoder network attached in addition to the feature transformer. Further, we iterate between modalities while keeping the transformer network common, so that it learns to work with all modalities. Once this stage is complete we discard the decoder added, and keep only the encoder transformer. Stage 2 masked pretraining. We engage the full network, except the task specific prediction heads. We take two inputs from two different modalities and pass them through the network till just before the task prediction heads. Instead of task prediction heads we add decoders to predict the masked tokens for respective input modalities. This process involves decoding the modalities in parallel, utilizing the outputs from the cross-attention modules and the modality-specific feature vectors. This alternating approach is key to achieving effective multimodal masked pretraining. Here also, we randomly mask tokens for both the modalities. Task balancing is not employed in this pretraining stage. Such a multi task multi modality approach allows us to utilize unpaired data across modalities. As in stage 1 pretraining, once this stage of training is finished, we discard the decoders added and keep the trained network .
In the final part of the training, we train for two tasks at a time from two different modalities. This lets up stochastically minimize the loss function in Eq. 1, but minimizing sum of two losses at a time instead of minimizing the sum of all of them. When we use two modalities, we use the network as shown in Fig. 1 in a two stream configuration. With the two modality features being fused together in the middle, passed through a transformer and then fused back with themselves, before finally being input to the task prediction heads. Such fusion of the the features from two modalities leads to knowledge sharing between the tasks of different modalies and makes the learning robust and regularized.
Given the varying complexities of these task pairs, as underscored in previous research , we found it essential to balance the complexity of tasks in a multitask learning setting. Hence, the we train while employing standard task balancing techniques. We adjust the loss magnitude for each task based on its convergence rate. As our ablation studies will demonstrate, this approach allows for random pairing of modalities, in contrast to the need for selecting specific pairs as suggested in prior works . We give details of such task balancing in the Supplementary material.
2.2 Masked pretraining for different modalities
2.3 Inference
When doing prediction, the network is used as a single stream without the cross attention layers in Fig. 1. The input data is tokenized with the tokenizer for its modality, passed through the feature transformation network followed by the second transformer , and finally input to the task prediction head , i.e. the full forward pass is where is the output of the tokenizer.
Experimental results
Masked pretraining. We use AudioSet (audio) , Something-Something v2 (SSv2) (video) , English Wikipedia (text), ImageNet1K (image) , SUN RGB-D (depth maps) , ModelNet40 (3D point cloud) for pretraining the network. For Stage 1 of masked pretraining (Sec. 3.2), we use the samples from the training set of the respective datasets. For Stage 2 of masked pretraining, we randomly select two modalities, and sample data from them to pretrain the full network. Further, we randomly mask patches. For image, video and audio, we mask of the patches. For point cloud and text, we mask and of the patches respectively. We perform pretraining for epochs. We use fraction as . Downstream tasks. We train the model on downstream tasks and report results. The datasets used for single modality methods are iNaturalist-2018 (Image Recognition), Places-365 (Scene Recognition), Kinetics-400 (Video Action Recognition), Moments in Time (Video Action Recognition), ESC50 (Audio Event Classification), S3DIS (3D point cloud segmentation), DialogueSUM (Text summarization). Adaptation on unseen datasets. To assess our method’s adaptability to datasets not seen at training, we report comparisons with image classification on Oxford-IIIT Pets , action recognition in videos using UCF-101 and HMDB51 , 3D point cloud classification on ScanObjectNN , point cloud segmentation with NYU v2 seg , text summarization using the SamSum dataset . As the number of classes and labels differ in each dataset as compared to the datasets used during pretraining, we randomly sample data from each of the training set. Further, we extract the embeddings using the pretrained network, and train two fully connected layers with task specific loss functions. This allows us to demonstrate the ability of the proposed method to generate embeddings which can generalize across datasets. Cross domain generalization. We follow prior work and evaluate on video-text retrieval on two benchmark datasets i.e. YouCook2 , and MSR-VTT , for multiple modalities. Adaptation on unseen modalities. We also evaluate our method on unseen modalities. Specifically, we evaluate our method on the following (i) X-Ray scan, and hyperspectral data recognition, where we utilize the RegDB , Chest X-Ray , and Indian Pine datasetshttps://github.com/danfenghong/IEEE_TGRS_SpectralFormer/blob/main/data/IndianPine.mat. (ii) Time-series forecasting, where our experiments are based on the ETTh1 , Traffichttps://pems.dot.ca.gov/, Weatherhttps://www.bgc-jena.mpg.de/wetter/, and Exchange datasets . (iii) Graph understanding through the PCQM4M-LSC dataset , which comprises 4.4 million organic molecules with quantum-mechanical properties, focusing on predicting molecular properties with applications in drug discovery and material science. (iv)Tabular analysis, where we engage with the adult and bank marketing datasets from the UCI repositoryhttp://archive.ics.uci.edu/ml/, (v) IMU recognition, where we conduct experiments on IMU sensor classification using the Ego4D dataset , assessing the capability to understand inertial motion systems. We follow for the train test splits and evaluation metrics on these datasets. Further, we use modality specific tokenizers and follow similar network settings as for generalization on unseen datasets.
We provide more details on the tokenizers used for each modality, description of task heads, and formulations of loss functions in the supplementary material.
We performed masked pretraining followed by training on multiple modalities and task groups as described in Section 3 for comparing with existing methods. We discuss the comparison on each modality below. Image. Table 4 shows state of the art on iNaturalist 2018 and Places 365 datasets. On the iNaturalist 2018 dataset, our method achieves a top-1 accuracy of 94.6%, surpassing notable contenders such as OmniVec (93.8%), MetaFormer (87.5%), and MAE (86.8%). This superior accuracy demonstrates capability of the proposed method in accurately recognizing a diverse range of natural species. In the context of the Places 365 dataset, our method achieves an accuracy of 65.1%, notably outperforming OmniVec (63.5%), and significantly surpassing MetaFormer’s 60.7% and Omnivore’s 59.9%. The substantial margin of improvement, particularly in the challenging and variable environment of Places 365, underscores the robustness and adaptability of the proposed architecture. We also conduct experiments on ImageNet (classification), MSCOCO (object detection), and ADE-20K (semantic segmentation) datasets (detailed table is in supplementary). 89.3% (accuracy) on ImageNet, 60.1 (AP) on MSCOCO and an mIoU of 58.5 on ADE-20K. Video. Table 4 and Table 4 show comparison against state of the art methods on Kinetics-400 and Moments in Time datasets.We observe that we outperform all the competing methods on both the datasets achieving top-1 accuracy of and respectively. Audio. Table 4 shows our comparison with top-performing methods on the ESC50 dataset. We outperform competing methods, achieving an accuracy of 99.1%, significantly higher than the Audio Spectrogram Transformer (AST) at 85.7%, and OmniVec at 98.4%. Point Cloud. Table 7 and Table 7 compare against state of the art methods on ModelNet40-C and S3DIS datasets respectively. On ModelNet40-C, we evaluate a classification task, while on S3DIS we evaluate semantic segmentation. On both the datasets, we outperform the competing methods. On ModelNet-C, we achieve an error rate of 0.142, which is notably lower than the rates observed in other contemporary methods. This is particularly evident when compared against methods like OmniVec, which recorded an error rate of 0.156, and PCT + PCM-R, with an error rate of 0.163. On S3DIS, we achieve an mIoU of 77.1, which is the highest among all the methods evaluated c.f. 75.9 of OmniVec, and 74.5 of Swin3D. This demonstrates that the proposed method is able to obtain a robust performance with the shared backbone network across tasks. Text. Table 7 shows state of the art on DialogueSUM dataset for text summarization. Our method surpasses other methods in all the metrics. Despite utilizing significantly fewer datasets for text in comparison to visual tasks , our method demonstrates strong performance. This suggests proposed method’s capacity to bridge the modality gap across distinct domains in the latent space, even when the data distribution is skewed.
Table 9 illustrates the experimental results on the GLUE benchmark for text understanding tasks, comparing various state-of-the-art methods such as BERT , RoBERTa , and ChatGPT. The comparison centers on paraphrasing, sentiment, duplication, inference, and answering tasks. We achieve second best performance on three out of five tasks demonstrating its capability to perform reasoning and adaptability to natural language tasks. Comparison on pretraining datasets. We fine tune our pretrained network on the respective training sets with related task heads. We obtain an mAP of 55.8 and 56.4 on AudioSet(A) and AudioSet(A+V) respectively. Further, on SSv2, ImageNet-1K, SUN-RGBD, and ModelNet we achieve top-1 accuracies of 86.1%, 93.6%, 75.9% and 97.1% respectively. We outperform the competing state of the art methods on these datasets(detailed results are in supplementary).
2 Adaptation on unseen datasets
In Table 8 (rows 1-6), we observe that our method performs close to SoTA on all the datasets. Specifically, except on UCF-101, we outperform the SoTA (OmniVec) on all the datasets. We observe that on NYUv2, we obtain a performance improvement of , while on an average perform better by approx on other datasets. It must be noted that we freeze the base embeddings, and unlike other methods do not fine tune the full network, and use simpler task head for analysis on these datasets.
3 Cross domain generalization
Table 8 (rows 7,8) demonstrates results using our pretrained network on various tasks. On the YouCook2 dataset, our pretrained network surpasses the state of the art in zero-shot retrieval, achieving a Recall@10 of 69.9% compared to OmniVec’s 64.2% on pretrained network. Interestingly, we are very close to the full fine tuned OmniVec i.e. 70.8. This demonstrates that our method is able to leverage the cross domain information better potentially due to multi task pretraining while OmniVec sequentially trains on one modality at a time. On MSR-VTT, when compared with SM , our fine-tuned method has a Recall@10 of 89.4% cf. SM’s 80.0% (pretrained). It must be noted that SM uses internet-scale data while our method utilizes significantly less data.
4 Adaptation on Unseen Modalities
Infrared, Hyperspectral, and X-Ray data. Table LABEL:tab:infrared presents the performance comparison on the RegDB dataset for infrared image recognition. Our method achieves state of the art performance i.e. R@1 of 86.21 c.f. 83.86 of MSCLNet, and mAP of 84.24 c.f. 78.57 of SMCL. This demonstrates that our method can transfer knowledge across unseen modalities. Specifically, we significantly outperform Meta-Transformer, which pretrains on similar modalities as ours. This could be potentially due to separate tokenizers for each modality allowing better integration with the transformer encoder as compared to a common tokenizer in meta-transformer.
In addition, Table LABEL:tab:hyper presents the performance on the Indian Pine dataset for hyperspectral image recognition. We achieve an overall accuracy of 90.6%, which is better than the SpectralFormer (81.76%) and significantly better than Meta-Transformer(67.6%). For X-Ray images (table in supplementary), our method achieves an accuracy of 98.1%, significantly outperforming competing methods. Graph and IMU Data. We show results in Table 11. We achieve performance close to the state of the art methods i.e. validate MAE of 0.1397 c.f. 0.1234 of Graphormer. It is important to note that our method was not designed for graphical data, while competing methods are designed to exploit graphical data. Meta-Transformer, which is a unified learning mechanism like ours, significantly lies behind with 0.8863 MAE cf. 0.1397 of ours. Time series forecasting. We achieve an MSE of 0.399, 0.601, 0.210, 0.330 on ETTh1, Traffic, Weather and Exchange datasets respectively, outperforming all the competing methods such as Pyraformer , Informer , LogTrans , Meta-former and Reformer . The detailed results are in supplementary.
Tabular Data. We achieve an accuracy of 88.1 and 92.3 on Adult and Bank Marketing datasets respectively, outperforming the competing methods (details in supplementary). Our method has never seen tabular data or structured textual information demonstrating its generalization ability to adapt to unseen patterns within data while providing better performance than competing methods.
5 Ablations
We study the impact of various components of the network in Table 12 on image (iNaturalist), video (Kinetics-400) and audio (ESC50) modalities. Specifically, we study the impact of pretraining with a single modality only, using the full pretraining mechanism, and then fine tuning on the respective training set. We also study the impact of modality specific tokenizers compared to unified tokenizers of MetaFormer , and impact of utilizing multiple task heads as compared to unified task head design of UniPerceiver-v2 . For unimodal pretraining (Table 12-row 1), we train the network on a single modality following Step 1 of Masked pretraining (see Sec. 3.2). We use corresponding modality for each dataset i.e. for iNaturalist, we pretrain on ImageNet1K, for K400, we pretrain on SSv2 and for ESC50, we pretrain on AudioSet. For multimodal multitask pretraining (Table 12-row 2), we pretrain using the full pretraining discussed previously. For fine tuning, we utilize the respective train sets. Impact of unimodal vs. multimodal pretraining We can observe that multimodal multitask pretraining using our approach (row 5) provides a significant improvement in comparison to unimodal pretraining (row 1). Specifically, it outperforms unimodal pretraining by on iNaturalist and K400 datasets while is better by on ESC50. This demonstrates that the network is able to leverage the information from multiple modalities. Impact of modality specific tokenizer vs. unified tokenizer. We observe that the performance of unified tokenizer (row 3) lags behind that of a modality specific tokenizer (row 4) by an average of across all the tasks, while keeping unified heads. Similarly, while keeping task specific heads, and modality specific tokenizer (row 6) vs unified tokenizer (row 5), we observe an average performance gap of in favour of modality specific tokenizer. Multiple task heads vs unified task head. Comparing row 4 and row 6, we see that the the task specific heads contribute to an increase (average ) in performance while keeping a modality specific tokenizer.
Conclusion
We presented a novel multimodal multitask network and associated training algorithm. The proposed method utilizes modality specific tokenizers and then uses shared transformers based backbone feeding to task specific heads. The traning proceeds in three stages, (i) masked pretraining with one modality at a time, (ii) masked pretraining with pairs of modalities together, and (iii) supervised traning for tasks with pairs of modalities together. The pairwise pretraining and supervised training allows for knowledge sharing between tasks and modalities and leads to a robust and regularized network. We showed empirical results on 25 challenging benchmark datasets over 12 modalities obtaining better or close to existing state of the art results. The method can incorporate arbitrary number of modalities, with only the tokenizer and task heads being modality specific.