Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models
Jize Cao, Zhe Gan, Yu Cheng, Licheng Yu, Yen-Chun Chen, Jingjing Liu
Introduction
Recently, Transformer-based large-scale pre-trained models have prevailed in Vision-and-Language (V+L) research, an important area that sits at the nexus of computer vision and natural language processing (NLP). Inspired by BERT , a common practice for pre-training V+L models is to first encode image regions and sentence words into a common embedding space, then use multiple Transformer layers to learn image-text contextualized joint embeddings through well-designed pre-training tasks. There are two main schools of model design: () single-stream architecture, such as VLBERT and UNITER , where a single Transformer is applied to both image and text modalities; and () two-stream architecture, such as LXMERT and ViLBERT , in which two Transformers are applied to images and text independently, and a third Transformer is stacked on top for later fusion. When finetuned on downstream tasks, these pre-trained models have achieved new state of the art on image-text retrieval , visual question answering , referring expression comprehension , and visual reasoning . This suggests that substantial amount of visual and linguistic knowledge has been encoded in the pre-trained models.
There has been several studies that investigate latent knowledge encoded in pre-trained language models . However, analyzing multimodal pre-trained models is still an unexplored territory. It remains unclear how the inner mechanisms of cross-modal pre-trained models induce their empirical success on downstream tasks. Motivated by this, we present Value (Vision-And-Language Understanding Evaluation), a set of well-designed probing tasks that aims to reveal the secrets of these pre-trained V+L models. To investigate both single- and two-stream model architectures, we select one model from each category (LXMERT for two-stream and UNITER for single-stream, because of their superb performance across many V+L tasks). As illustrated in Figure 1, Value is designed to provide insights on: () Multimodal Fusion Degree; () Modality Importance; () Cross-modal Interaction (Image-to-text/Text-to-image); () Image-to-image Interaction; and () Text-to-text Interaction.
For () Multimodal Fusion Degree, clustering analysis between image and text representations shows that in single-stream models like UNITER, as the network layers go deeper, the fusion between two modalities becomes more intertwined. However, the opposite phenomenon is observed in two-stream models like LXMERT. For () Modality Importance, by analyzing the attention trace of the [CLS] token, which is commonly considered as containing the intended fused multimodal information and often used as the input signal for downstream tasks, we find that the final predictions tend to depend more on textual input rather than visual input.
To gain deeper insights into how pre-trained models drive success in downstream tasks, we look into three types of interactions between modalities. For () Cross-modal Interaction, we propose a Visual Coreference Resolution task to probe its encoded knowledge. For () Image-to-image Interaction, we conduct analysis via Visual Relation Detection between two image regions. For () Text-to-text Interaction, we evaluate the linguistic knowledge encoded in each layer of the tested model with SentEval tookit , and compare with the original BERT . Experiments show that both single- and two-stream models, especially the former, can well capture cross-modal alignment, visual relations, and linguistic knowledge.
To the best of our knowledge, this is the first known effort on thorough analysis of pre-trained V+L models, to gain insights from different perspectives about the latent knowledge encoded in self-attention weights, and to distill the secret ingredients that drive the empirical success of prevailing V+L models.
Related Work
For image-text representation learning, ViLBERT and LXMERT used two-stream architecture for pre-training, while B2T2 , VisualBERT , Unicoder-VL , VL-BERT and UNITER adopted single-stream architecture. VLP proposed a unified pre-trained model for both image captioning and VQA. Multi-task learning and adversarial training in VILLA have been studied to boost performance. On video+language side, VideoBERT applied BERT to learn joint embeddings of video frame tokens and linguistic tokens from video-text pairs. CBT introduced contrastive learning to handle real-valued video frame features, and HERO proposed hierarchical Transformer architectures to leverage both global and local temporal visual-textual alignments. However, except for some simple visualization of the learned attention maps , no existing work has systematically analyzed these pre-trained models.
There has been some recent studies on assessing the capability of BERT in capturing structural properties of language . Multi-head self-attention has been analyzed for machine translation in , which observed that only a small subset of heads is important, and the other heads can be pruned without affecting model performance. reported that BERT can rediscover the classical NLP pipeline, where basic syntactic information appears in lower layers, while high-level semantic information appears in higher layers. Analysis on BERT self-attention showed that BERT can learn syntactic relations, and a limited set of attention patterns are repeated across different heads. and demonstrated that BERT has surprisingly strong ability to recall factual relational knowledge and perform commonsense reasoning. A layer-wise analysis of Transformer representations in provided insights to the reasoning process of how BERT answers questions. All these studies have focused on the analysis of BERT, while investigating pre-trained V+L models is still an uncharted territory. Given their empirical success and unique multimodal nature, we believe it is instrumental to conduct an in-depth analysis to understand these models, to provide useful insights and guidelines for future studies. New probing tasks such as visual coreference resolution and visual relation detection are proposed for this purpose, which can lend insights to other evaluation tasks as well.
VALUE: Probing Pre-trained V+L Models
Key curiosities this study aims to unveil include:
() What is the correlation between multimodal fusion and the number of layers in pre-trained models? (Sec. 3.1)
() Which modality plays a more important role that drives the pre-trained model to make final predictions? (Sec. 3.2)
() What knowledge is encoded in pre-trained models that supports cross-modal interaction and alignment? (Sec. 3.3)
() What knowledge has been learned for image-to-image (intra-modal) interaction (i.e., visual relations)? (Sec. 3.4)
() Compared with BERT, do pre-trained V+L models effectively encode linguistic knowledge for text-to-text (intra-modal) interaction? (Sec. 3.5)
To answer these questions, we select one model from each archetypal model architecture for dissection: UNITER-base (12 layers, 768 hidden units per layer, 12 attention heads) for single-stream model, and LXMERT for two-stream model.Our probing analysis can be readily extended to other pre-trained models as well. Single-stream model (UNITER) shares the same structure as BERT , except that the input now becomes a mixed sequence of two modalities. Two-stream model (LXMERT) first performs self-attention through several layers on each modality independently, then fuses the inputs through a stack of cross-self-attention layers (first cross-attention, then self-attention). Therefore, the attention pattern in two-stream models is fixed, as one modality is only allowed to attend over either itself or the other modality at any time (there is no such constraint in single-stream models). A more detailed model description is provided in the Appendix.
Two datasets are selected for our probing experiments:
() Visual Genome (VG) : image-text dataset with annotated dense captions and scene graphs.
() Flickr30k Entities : image-text dataset with annotated visual co-reference links between image regions and noun phrases in the captions.
Each data sample consists of three components: () an input image; () a set of detected image regions;An image region is also called a visual token in this paper; these two terms will be used interchangeable throughout the paper. and () a caption. In Flickr30k Entities, the caption is relatively long, describing the whole image; while in VG, a few short captions are provided (called dense captions), each describing an image region.
The number of annotated image regions in an image is relatively small (5-6 for Flick30k Entities, and 2-4 for VG dense annotated graph); while in pre-trained V+L models, the number of image regions fed to the model is typically 36 . Therefore, we extract an additional set of image regions from a Faster R-CNN pre-trained on the VG dataset , and combine them with the original image regions provided in the dataset. The initial image representation is obtained from the Faster R-CNN as well. Finally, we feed both image regions and textual tokens into the pre-trained model for probing analysis. The following sub-sections describe analysis results and key observations from each proposed probing task.
By nature, visual features in images and linguistic clues in text have distinctive characteristics. However, it is unknown whether the semantic gap between their corresponding representations narrows (i.e., the contextualized representations of the two modalities become less differentiable) through the cross-modal fusion between intermediate layers.
To answer this question, we design a probing task to test the multimodal fusion degree of a model. First, we extract all the embedding features of both image regions and textual tokens from aforementioned two datasets (VG and Flickr30k Entities). For single-stream model (UNITER-base), we extract the output representation from each layer, as multimodal fusion is performed through all the layers. For two-stream model (LXMERT), we only consider the output embeddings from the last 5 layers (i.e., the layers constituting the cross-modality encoder), because the two modalities do not have any interaction prior to that. To gain quantitative measurement, for each data sample, we apply the -means algorithm (with = 2) on the representations from each layer to partition them into two clusters, and measure the difference between the formed clusters and ground-truth visual/textual clusters via Normalized Mutual Information (NMI), an unsupervised metric for evaluating differences between clusters. A larger NMI value implies that the distinction between two clusters is more significant, indicating a lower fusion degree . For example, when NMI is equal to 1.0, the two clusters represent the original visual tokens and textual tokens, respectively. We use the mean value of NMI for all the data samples to measure the level of multimodal fusion.
1.2 Results
Table 1 summarizes the probing results on multimodal fusion degree. For single-stream model (UNITER-base), the NMI scores gradually decrease, indicating that the representations from the two modalities fuse together deeper and deeper from lower to higher layers. This observation matches our intuition that the embedding of a modality from a higher layer better attends over the other modality than a lower layer. However, in two-stream model (LXMERT), as the layers go deeper (except for the last layer), the representations deviate from each other. One possible explanation is that single-stream model applies the same set of parameters to both image and text modalities, while two-stream model uses two separate sets of parameters (as part of its network design) to model the attention on the image and text modality independently. The latter makes it relatively easier to distinguish the two representations, leading to a higher NMI score. An t-SNE visualization is provided in the Appendix.
2 Who Pulls More Strings: Textual Modality Is More Dominant than Image
Following BERT , pre-trained V+L models add a special [CLS] token at the beginning of a sequence, and a special [SEP] token at the end. In practice, the [CLS] token representation from the last layer is used as the fused representation for both modalities in downstream tasks. Since the [CLS] token absorbs information from both modalities through self-attention, the degree of attention of the [CLS] token over each modality could be regarded as evidence to answer the following question: which modality is more dominant during inference? Note that in the two-stream model design, the [CLS] token is not allowed to attend to both image and textual modality simultaneously. Therefore, the probing experiments for modality importance is focused on single-stream model.
Formally, to quantitatively analyze the [CLS] attention trace, the Modality Importance (MI) of a head is defined as the sum of the attention values that the [CLS] token spent on the modality (visual or textual) at for the whole sequence ([CLS], , [SEP], ), where and are textual and visual tokens, respectively. That is,
Since the attention heads in the same layer perform attention on the same representation, it is natural to expand the above head-level analysis to a layer-level analysis, by considering the mean of MI scores of all the 12 heads as the MI measurement of that layer. In addition, we calculate an overall MI score via summing up the MI scores of all the 144 heads. The mean MI value of all the data samples is reported as the score for probing modality importance. Note that the MI score for visual/textual modality is calculated based on the attention weights on the visual/textual tokens; while the attention weights on the special [CLS] and [SEP] tokens are not considered. Therefore, the two mean MI scores from the two modalities do not sum to one.
2.2 Results
Experiments are conducted on the Flickr30k Entity dataset. Figure 2 provides the MI scores for all the 144 attention heads of UNITER-base. The heatmap on the textual MI is denser than that of the visual MI, showing that more attention heads are learning useful knowledge from the textual modality than the image modality. Figure 3(a) further shows the layer-level MI scores for each modality. The average MI score on the text modality is higher than that on the image modality, especially for intermediate layers, suggesting that the pre-trained model relies more on the textual modality for making decisions during inference time.
3 Winner Takes All: A Subset of Heads is Specialized for Cross-modal Interaction
The key difference between single-modal Transformer (such as BERT) and two-modal Transformer (such as UNITER and LXMERT) is that two-modal Transformer requires extra cross-modal interaction. To gain an in-depth understanding of cross-modal attention heads, which is instructive to prompt better model design and enhance model interpretability, we look into two special types of head: () image-to-text head; and () visual-coreference head.
We first analyze whether there exists any head specialized in learning cross-modal interaction. Formally, for a given image-text pair, visual and textual tokens are denoted as and , respectively. We define a head as image-to-text head if:
where denotes the attention weight from a visual token to a textual token. This defines whether a visual token pays more attention to the text modality. Specifically, if there exists one visual token that has higher attention weight on text than other tokens, we regard the corresponding head as performing cross-modal attention from image to text.
Based on the above definition, we count the number of occurrences of head as an image-to-text head for all data samples in the Flickr30k Entities dataset, and report the empirical probability of each head being an image-to-text head. Note that in the two-stream model, the image-to-text head is by design, therefore, we only conduct this analysis on single-stream model.
Figure 3(b) shows there is a specific set of heads in UNITER-base that perform cross-modal interaction. The maximum probability of a head performing image-to-text attention is 0.92, the minimum probability is 0, and only 15% heads have more than 0.5 probability to pay the majority attention weight on the image-to-text part. Interestingly, by training single-stream model, the attention heads are automatically learned to exhibit a “two-stream” pattern, where some heads control the message sharing from the visual modality to the textual modality.
3.2 Visual Coreference Resolution
One straightforward way to investigate the visual-linguistic knowledge encoded in the model is to evaluate whether the model is able to match an image region to its corresponding textual phrase in the sentence. Thus, we design a Visual Coreference Resolution task (similar to coreference resolution) to predict whether there is a link between an image region and a noun phrase in the sentence that describes the image. In addition, each coreference link in the dataset is annotated with a label (8 in total).
Through this task, we can find out whether the coreference knowledge can be captured by the attention trace. To achieve this goal, for each data sample in the Flickr30k Entity dataset, we extract the encoder’s attention weights for all the 144 heads. Note that noun phrases typically consist of two or more tokens in the sequence. Thus, we extract the maximum attention weight between the image region and each word of the noun phrase for each head. The maximum weight is then used to evaluate which head identifies visual coreference (i.e., performing multimodal alignment).
Results are summarized in Table 2. The columns labeled with “(Rand.)” are considered as ablation groups to identify whether the high attention weight of a certain coreference relationship is triggered by the relation between the image-text pair rather than the effect of one specific image/text token. For a link , we measure the maximum attention weight of the visual token to a random noun phrase, to obtain the results for (Rand.). Similarly, we use the maximum attention weight of a noun phrase to a random visual token to obtain the results for (Rand.).
Results show that the relation between a noun phase and its linked visual token is encoded in the attention pattern, especially for . The heads (9-3)Head (-) means the -th head at the -th layer., (9-12) and (3-1) have captured richer coreference knowledge than other heads, indicating that there exists a subset of heads in the pre-trained model that is specialized in coreference linking between the two modalities. Furthermore, for single-stream model, we observe that some heads encode the coreference information in both directions, serving as additional evidence that these heads perform cross-modal alignment specifically. On the other hand, the amount of learned coreference knowledge is limited in the text modality attention trace (), which provides indirect evidence that the text modality does not incorporate much visual information, even with the forced cross-attention design in the two-stream model.
3.3 Probing Combinations of Heads
Previous analysis mainly investigates whether cross-modal knowledge can be captured through individual attention head. It is also possible that such knowledge can be induced via the cooperation of multiple heads. To quantitatively analyze this, we further examine visual coreference through probing over combinations of heads.
1) Attention Prober To reveal the learned knowledge across different attention heads, we use a linear classifier based on the combination of attention weights, following . Specifically,
where given tokens and in the sequence, is the probability of the link label between these two tokens being . is the attention weight for token attending to token at head , and are two learnable scalars, and is the number of attention heads.
2) Layer-wise Embedding Prober Similarly, to further examine the knowledge encoded in the model, we can naturally extend the above attention-based prober into an embedding-based prober:
For visual coreference resolution, we probe the model on two sub-tasks: () Visual Coref Detection (VCD): determine whether a noun phrase is coreferenced to a specific visual token (binary classification); and () Visual Coref Classification (VCC): classify the label of the coreference relation between a noun phrase and a visual token (multi-class classification). With these tasks, we can examine whether certain type of knowledge is encoded, as well as the granularity of the knowledge encoded in the probed feature spaceSince noun phrase may contain several tokens, we use the maximum attention weight among the tokens in that phrase over an image region as the attention weight between the noun phase and the image region. The embedding of the noun phrase is the mean of all the representations of its textual tokens..
Results are shown in Table 3. Neither model performs well on the VCD task, suggesting that there is no attention pattern or embedding feature that can handle all the coreference relations. However, the results on VCC are encouraging. This aligns with our observation from Table 2 that some attention heads of the two models are significantly effective in certain coreference relations.Though both models’ embedding probers achieve higher than 94% accuracy on the VCC task, it is worth noting that text embedding input can potentially leak the link information. For instance, the phrase “A guard with a white hat” may already provide coreference information between person and the corresponding image region.
4 Secret Liaison Revealed: Cross-Modality Fusion Registers Visual Relations
To evaluate the encoded knowledge learned from the image modality, we adopt the visual relation detection task, which requires a model to identify and classify the relation between two image regions. This task can be viewed as examining whether the model captures visual relations between image regions.
The VG dataset is used for this task, which contains 1,531,448 first-order object-object relations. First-order object-object relation can be determined simply by the visual representations of two objects, independent to other objects in the image or text annotation. To reduce the imbalance in the number of relations per relation type, we randomly select at most 15,000 subject-object relation pairs per relation type. Furthermore, to de-duplicate cases where the same type of relation comes from the same text annotation, we select at most 5 same relation types from the same annotation. This probing task is performed on 32 most frequent relation pairs in the dataset.
1) Probing Individual Attention Head We apply the same analysis here similar to the visual coreference resolution task. The only difference is that we do not consider the directions of attention (i.e., , or ), as both directions correspond to the same visual modality. The average of maximum attention value in both and directions is reported for visual relation analysis.
Results are summarized in Table 4. We report the maximum attention weights of 7 of 30 relations. The columns with “(Rand.)” are ablation groups. As shown in the comparison, the learned attention heads encode rich knowledge about the relations between visual tokens with much higher attention weights. Moreover, similar to the observation in visual coreference resolution task, specific heads (10-1 head in the single-stream model, 7-4 head in the two-stream model) have captured richer visual relations than others.
2) Probing Combinations of Attention Heads The above analysis only reveals the behavior of individual attention heads. Similar to Sec. 3.3, we further examine visual relations by training a linear classifier on top of a combination of attention heads over two sub-tasks: () Relation Identification: determine whether two image regions have a relationship; and () Relation Classification: classify the relation label between two image regions.
Results on probing a combination of heads are summarized in Figure 4. Two baselines are considered in this task. () Original visual embeddings from Faster R-CNN . This setup evaluates how much correlation between two related visual representations has elevated or diluted by the V+L models. () Mismatched image-text representation. In this baseline, we construct a dataset where an image is associated with an unrelated dense annotation instead of a related one. This baseline evaluates how the correlation between the image regions changes based on the text modality. Note that for the two-stream model, there are 10 layers involved in visual relation reasoning. The first five layers perform self-attention across the visual modality only, and in the last five layers the visual representation interacts with the text modality.
As shown, both models perform much better than the baseline with related annotations (the dashed balck line in Figure 4), indicating there is a substantial amount of visual-relation knowledge encoded in these representations. On the other hand, the specific visual relation knowledge degrades a lot with mismatched image regions and dense annotations (the solid red and blue lines vs. the dashed lines in Figure 4). For single-stream model, the visual-relation knowledge decayed to that of the original visual embedding. This makes sense because no visual relation information can be obtained from the mismatched caption.
The visual relation knowledge of the two-stream model is greatly influenced by the unpaired caption, leading to a huge performance drop after Layer 5 in the VRC task. This result may contribute to different inductive bias between the two models. For two-stream model, since the visual modality has to attend to the text modality during cross attention, the visual representation will be greatly influenced by the text modality even when the two modalities are totally unrelated. On the other hand, because the visual representation in the single-stream model is capable of selectively choosing whether to attend over the text modality, the effect of unrelated caption on the visual representation is negligible.
5 No Lost in Translation: Pre-trained V+L Models Encode Rich Linguistic Knowledge
Besides looking into the knowledge learned from the visual modality, we are also interested in the encoded knowledge learned from the text modality. To achieve this goal, we probe the pre-trained models over nine tasks defined in the SentEval toolkit . Descriptions about the tasks are provided in the Appendix.
First, we extract contextualized word representations from the pre-trained models. For single-stream model, the input is the sequence of tokens with text only. For two-stream model, since the last 5 cross-attention layers require visual input, which is not covered by these linguistic benchmarks, we only evaluate the first 9 layers that are performing self-attention over pure text input.
We use the Google-pretrained BERT-base model as the baseline. For each task, we obtain the results from all the layers and report the best number. The results in Table 5 show that pre-trained V+L model generally performs worse than the original BERT-base model on these linguistic benchmarks, which is as expected. A full table with results from all the layers is provided in the Appendix. As shown in the table, the single-stream model performs better than the two-stream model across all the tasks. One possible reason is that LXMERT does not initialize the parameters of the language Transformer encoder from BERT-base, whereas UNITER does. Thus, using the parameters of BERT as initialization is potentially useful for the model to acquire rich linguistic knowledge, and helpful for tasks involving complex text-based reasoning.
To measure the gains due to learning, we have conducted all the above experiments (Secs. 3.1-3.5) on untrained baselines with random weights. These additional results are provided in the Appendix.
Conclusion and Key Takeaways
Intrigued by the five questions presented at the beginning of Section 3, we have provided a thorough analysis of UNITER-base and LXMERT models as a deep dive into Vision+Language pre-training. To summarize our key findings:
() In single-stream model, deeper layers lead to more intertwined multimodal fusion; while the opposite trend is observed in two-stream model.
() Textual modality plays a more important role than image in making final decisions, consistent across both single- and two-stream models.
() In single-stream model, a subset of heads organically evolves to pivot on cross-modal interaction and alignment, which on the other hand is enforced by model design in two-stream model.
() Visual relations are inherently registered in both single- and two-stream pre-trained models.
() Rich linguistic knowledge is naturally encoded, even though the models are specifically designed for multimodal pre-training.
We provide additional guidelines in the Appendix. For future work, we plan to perform model compression via pruning attention heads based on the analysis and observations in this work.
References
Appendix 0.A Details on Pre-trained V+L Models
A comparison between single-stream and two-stream V+L models is provided in Figure 5. We choose UNITER-basehttps://github.com/ChenRocks/UNITER as the representative model for single-stream, and LXMERThttps://github.com/airsplay/lxmert for two-stream. As shown in Figure 5(a), UNITER-base has the same model structure as the BERT-base model , which composes of 12 layers of self-attention Transformers. Each layer has 12 self-attention heads, and each hidden representation is a 768-dimensional vector. As shown in Figure 5(b), LXMERT is a two-stream model that performs intra-attention in the same modality first, then cross-attention. We denote , and as the Transformer modules that specifically model text-to-text, image-to-image and cross-modal interactions, respectively. In LXMERT, has 9 layers, has 5 layers, and has 5 layers. Each Transformer’s hidden representation is in dimension of 768. Note that each layer in contains one cross-attention layer between two modalities, followed by two self-attention layers for each modality.
Appendix 0.B Results on Untrained Baselines
To measure the gain from learning, we also conducted additional experiments on untrained single-stream (SS) and two-stream (TS) baselines with random weights.
Multimodal Fusion Probe The untrained SS model has NMI of 0.99 for all output layers, suggesting that the two modalities are completely separated. The untrained TS model has NMI of 0.56 for all output layers. This is because the cross-modality encoder layers force the two modalities to fuse, even in untrained setting.
Modality Importance Probe For the untrained model, the average attention of [CLS] token on the image/text modality is 0.66/0.28. Note that the number of tokens in a sentence is usually smaller than that of the visual tokens.
Visual Coreference and Relation Probe We provide additional untrained baselines for visual coreference and visual relation probes in Table 6. Compared to Table 3 and Figure 4(a), for VCD and VRI, untrained baselines for both SS and TS are equivalent to random guess. For VRC, both SS and TS models outperform the baseline by around 10%. For VCC, the SS model outperforms the baseline by 17%; while the TS model performs worse. This may be because after hard-designed multimodal fusion, the direct coreference relationship between a pair of image/text tokens is diluted after training.
Furthermore, we provide an additional evaluation on whether the head selected for a specific coreference relation of an image-text pair imposes higher attention scores for coreference relation than all other pairs. Results are summarized in Table 7, which suggests that these attention heads with maximum attention weight do pay more attention to the coreference image-text pair, compared to other unpaired ones.
Appendix 0.C Additional Guidelines for Future Model Design
In addition to the key takeaways in Sec. 4, we provide a set of guidelines for future model design based on our analysis and observations.
() Single-stream model is able to capture sufficient intra- and cross-modal knowledge, while the restricted attention structure in two-stream model does not bring additional benefit. For future work, we will further explore single-stream model design, which also exhibits better interpretability as observed.
() Initializing V+L model with BERT’s weights should be helpful, which can enhance V+L model’s capability in language understanding.
() It remains unclear how to measure a pre-trained model without evaluating on downstream tasks. Given that finetuning is time consuming, the probing tasks we propose can provide a convenient tool to quickly test intermediate model checkpoints during pre-training.
() Explicitly adding extra supervision to probing tasks during model training may lead to more interpretable and robust model.
Appendix 0.D Details on Linguistic Probe
We probe the pre-trained models over nine tasks defined in the SentEval toolkit , under three categories:
() Surface tasks: probe for the length of a sentence (SentLen);
() Syntactic tasks: predict the depth of a sentence’s syntax tree, consecutive token inversions (BShift), and the top constituents sequences (TopConst);
() Semantic tasks: test the tense (Tense), the number implied by the subject/object (SubjNum/ObjNum), the replacement of the noun/verb form (SOMO), and the inversion of coordinating conjunctions (CoordInv).
Appendix 0.E Additional Results
We provide additional results on multimodal fusion, visual coreference resolution, visual relation detection, and linguistic probing.
An t-SNE visualization of multimodal fusion degree of the first and last layer of UNITER (over one image-text pair) is provided in Figure 6. As the layer goes deeper, the two modalities become more intertwined.
E.2 Visual Coreference Resolution
Due to space limit, we only reported results using the embeddings from Layer 1, 5 and 12 in Table 3. A complete set of results is provided in Table 8. We observe that the attention probers work well for VCC, but not for VCD. Our assumption is that task granularity matters to the prober’s performance. Attention behavior varies a lot in different coreference relations, thus it performs well on VCC. The dataset for training VCC is built with positive examples from VCD only. Therefore, VCD’s settings naturally dilute the distinction between different coreference relations’ attentions, which makes it a more challenging task.
E.3 Visual Relation Detection
Results of the layer-wise embedding probers on the Visual Relation Classification and Identification (VRC and VRI) tasks are visualized in Figure 4(b) and (c), respectively. Detailed numbers corresponding to these two figures are provided in Table 9 and 10.
E.4 Linguistic Probing
For linguistic probing, we first obtain results from all the layers of a pre-trained model, then report the best number in Table 5. Detailed results for all the layers are provided in Table 11.
E.5 Visualization of Attention Maps
We show the learned attention maps of one specific relation in the probing tasks: Figure 7 and 8 for visual coreference resolution (Section 3.3.2 of the main paper), and Figure 9 and 10 for visual relation detection (Section 3.4 of the main paper).