Unifying Vision-and-Language Tasks via Text Generation
Jaemin Cho, Jie Lei, Hao Tan, Mohit Bansal
Introduction
Mirroring the success of the pretraining-finetuning paradigm with transformer language models (Devlin et al., 2019), recent vision-and-language transformers (Tan & Bansal (2019); Lu et al. (2019); Chen et al. (2020); Li et al. (2020b), inter alia) have also been adopted in a wide range of vision-and-language tasks. These models are firstly pretrained on large image-text corpus (e.g., COCO Caption (Chen et al., 2015)), then finetuned on downstream tasks (e.g., visual question answering (Goyal et al., 2019) and referring expression comprehension (Mao et al., 2016)), which outperformed many previous non-pretraining-finetuning methods.
For each pretraining or downstream task, existing vision-and-language transformers typically require designing task-specific, separately-parameterized architectures on top of the transformer encoder (e.g., multi-label sigmoid classifier for visual question answering, and softmax classifier for referring expression comprehension). However, the reasoning skills required by these tasks overlap significantly. Consider the example in Fig. 1. Both answering the question “What is the man jumping over?” and grounding an image region corresponding to the phrase “yellow fire hydrant” require recognizing the object “fire hydrant”. In addition, the labels for these tasks can be easily expressed in text. For instance, we can assign a region id (e.g., “
Hence, in order to alleviate these hassles of designing task-specific architectures, we propose a unified framework for vision-and-language learning via generating labels in text. Specifically, we extend off-the-shelf pretrained language models T5 (Raffel et al., 2020) and BART (Lewis et al., 2020) with visual understanding ability, named ‘VL-T5’ and ‘VL-BART’. In contrast to existing methods that train different architectures for each pretraining and downstream task, our models tackle all tasks with the same language modeling head. To learn a new task, we can simply rewrite its input and output in text, without the need of adding extra parameters or designing new architectures and objectives. In addition, we can leverage the text generation ability of pretrained language models when making predictions. This is especially helpful when we answer open-ended questions that require non-trivial answers, where discriminative methods can only answer from a predefined set of frequent candidates, while our models can generate open-ended natural language answers.
To evaluate the effectiveness of our generative modeling approach, we compare our models against recent vision-and-language transformers on a diverse set of 7 downstream benchmarks, including visual question answering on VQA (Goyal et al., 2019) and GQA (Hudson & Manning, 2019), referring expression comprehension on RefCOCOg (Mao et al., 2016), natural language visual reasoning on (Suhr et al., 2019), visual commonsense reasoning on VCR (Zellers et al., 2019), image captioning on COCO Caption (Chen et al., 2015), and multimodal machine translation on Multi30K (Elliott et al., 2016). Our unified generative method reaches comparable performance to recent state-of-the-art vision-and-language pretraining methods. This is especially interesting because we use the same unified language modeling architecture with the same maximum likelihood estimation (MLE) objective for all the tasks, while existing approaches use task-specific architectures and objective functions. In addition, we found that our generative models have better generalization ability compared to the discriminative versions in the rare-answer scenario on visual question answering, when ground truth answers for given questions are rarely seen during training. Finally, we also experiment with our unified framework under the multi-task learning setup on all 7 downstream tasks. With a single architecture and a single set of weights, our model achieves similar performance to separately optimized single-task models.
Related Works
Vision-and-Language pretraining: Large-scale language pretraining with transformers (Vaswani et al., 2017; Devlin et al., 2019; Liu et al., 2019; Lan et al., 2020; Clark et al., 2020; Yang et al., 2019; Raffel et al., 2020) have achieved remarkable success for many natural language understanding tasks (Rajpurkar et al., 2016; Zellers et al., 2018; Wang et al., 2018; Williams et al., 2017). Following this success, image+text pretraining models (Lu et al., 2019; Tan & Bansal, 2019; Chen et al., 2020; Huang et al., 2020; Li et al., 2020b; Cho et al., 2020; Radford et al., 2021; Zhang et al., 2021) and video+text pretraining models (Sun et al., 2019b, a; Li et al., 2020a; Zhu & Yang, 2020; Miech et al., 2020) have also shown to perform better than previous non-pretraining approaches (Yu et al., 2018a; Anderson et al., 2018; Kim et al., 2018; Yu et al., 2018b) in a wide range of discriminative (Goyal et al., 2019; Hudson & Manning, 2019; Lei et al., 2018; Mao et al., 2016; Xu et al., 2016; Zhou et al., 2018) and generative tasks (Chen et al., 2015; Xu et al., 2016; Zhou et al., 2018). In this work, we focus on image+text tasks. While existing image+text models mostly use task-specific architectures and objectives, we seek to design a unified framework across different tasks.
Unified frameworks: One line of work focus on solving natural language processing tasks in a unified format, such as question answering (Mccann et al., 2018), span prediction (Keskar et al., 2019), or text generation (Raffel et al., 2020; Brown et al., 2020; Khashabi et al., 2020). These unified frameworks provide efficient knowledge sharing among different tasks and make it easy to leverage pretrained language models. In relation to these works, we propose to unify previously separately modeled vision-and-language tasks in a single unified format, via text generation, conditioned on multimodal inputs from the image and the textual context.
Model
We propose a new framework that unifies vision-and-language problems as multimodal conditional text generation. We introduce VL-T5 and VL-BART based on two pretrained transformer language models: T5Base (Raffel et al., 2020) and BARTBase (Lewis et al., 2020). Specifically, we extend their text encoders to multimodal encoders by incorporating image region embeddings as additional input. The overall architecture of our framework is shown in Fig. 2. Since the architecture differences between VL-T5 and VL-BART are minor, we use VL-T5 as an example to illustrate our framework in details in the rest of this section.
We represent an input image with object regions from a Faster R-CNN (Ren et al., 2015) trained on Visual Genome (Krishna et al., 2016) for object and attribute classification (Anderson et al., 2018). As shown in Fig. 2 (b), each image region is encoded as a sum of four types of features: () RoI (Region of Interest) object features; () RoI bounding box coordinates; () image ids ; and () region ids . RoI features and bounding box coordinates are encoded with a linear layer, while image ids and region ids are encoded with learned embeddings (Devlin et al., 2019). Image ids are used to discriminate regions from different images, and is used when multiple images are given to the model (i.e., in (Suhr et al., 2019), models take two input images). The final visual embeddings are denoted as .
2 Text Embeddings
Instead of designing task-specific architectures, we add different prefixes to the original input text to adapt to different tasks, as shown in Table. 1Note that since we use simple prefixes (e.g., “vqa:” for VQA task), it is likely that engineering in text prompts (Gao et al., 2020) would improve the accuracy of our methods. As this is not the focus of this paper, we leave it as future works.. This augmented input text is then tokenized as and encoded as learned embedding . The embedding parameters are shared by the encoder, decoder, and language modeling head (Press & Wolf, 2017). Since the attention layers are permutation-invariant, BART learns positional embeddings (Vaswani et al., 2017; Devlin et al., 2019) for absolute token positions and adds them to the token embeddings. In contrast, T5 adds relative position bias to each self-attention layer (Shaw et al., 2018). Our models follow the positional embedding configurations of their text backbone models.
In addition to the original vocabulary of T5 and BART, we introduce visual sentinel tokens {
3 Encoder-Decoder Architecture
4 Task-Specific Methods vs. Our Unified Framework
We compare our unified framework with existing vision-and-language transformers on two popular tasks: visual question answering (Goyal et al., 2019) and referring expression comprehension (Mao et al., 2016).
Visual question answering requires a model to answer a question to a given context image. As shown in Fig.3 (a), existing methods (Tan & Bansal, 2019; Lu et al., 2019; Chen et al., 2020) typically introduce a multi-layer perceptron (MLP) multi-label classifier head on top of , which is trained together with the transformer backbone through a binary cross-entropy loss, and weighted with VQA score (Goyal et al., 2019): .
Referring expression comprehension requires models to localize a target region in an image that is described by a given referring expression. Previous methods tackle this task as multi-class (Chen et al., 2020) or binary (Lu et al., 2019) classification over image regions. For example, UNITER (Chen et al., 2020) introduces an MLP region scoring head on top of the output representations of regions, as shown in Fig. 3(b). This region scoring head is jointly trained with the encoder by minimizing negative log-likelihood of target region : .
In contrast to existing methods that develop task-specific architectures and objectives (e.g., equations above), our unified framework is free from extra model designs for new tasks. As shown in Fig. 3 (c,d) and Table 1, we formulate the task labels to corresponding text, and we learn these different tasks by predicting label text with the same language modeling objective (Eq. 1).
Pretraining
In this section, we describe how we pretrain our VL-T5 and VL-BART models (Sec. 3). We start with the details of the pretraining data and illustrate how we formulate diverse vision-and-language pretraining tasks as multimodal conditional text generation.
We aggregate pretraining data from MS COCO (Lin et al., 2014; Chen et al., 2015) and Visual Genome (VG; Krishna et al. (2016)) imagesExisting vision-and-language transformers are trained with different datasets and computational budgets, thus their results may not be directly comparable to each other. We show the number of their pretraining images in Table 2.. The captioning data from these two datasets are used in the multimodal language modeling task. The COCO captions are also used in the image-text matching task to learn cross-modal alignment. Besides the captions, we also use three visual question answering datasets (VQA v2.0 (Goyal et al., 2019), GQA balanced version (Hudson & Manning, 2019), and Visual7W (Zhu et al., 2016)) as in Tan & Bansal (2019), but only used them for the visual question answering task. Details of these pretraining tasks are in Sec. 4.2. Overall, our pretraining dataset contains 9.18M image-text pairs on 180K distinct images. We show more details of the pretraining data in appendix.
2 Pretraining Tasks
We pretrain our models under a multi-task setup with diverse pretraining tasks, including multimodal language modeling, visual question answering, image-text matching, visual grounding, and grounded captioning. Table 1 shows input and output examples of our pretraining tasks. The training data for each of these tasks are summarized in appendix. In the rest of this section, we explain these tasks in detail.
Multimodal language modeling: We follow Raffel et al. (2020) and Lewis et al. (2020) to construct the language modeling pretraining task. For VL-T5, we mask 15% of input text tokens and replace contiguous text span with sentinel tokens (e.g.,
Visual question answering: We include visual question answering in our pretraining tasks as in Tan & Bansal (2019). While previous methods (Tan & Bansal, 2019; Lu et al., 2019; Chen et al., 2020) tackle the task as classification over predefined answer candidates (illustrated in Fig. 3), we directly generate answers in their original text format.
Image-text matching: In this task, the model needs to verify whether a text corresponds to an image. We consider an image and its captionsWe only use captions from COCO for this task, since many short captions from VG and visual questions are nondistinctive descriptions of an image (e.g., ‘what is in the image?’). as positive pairs. With a probability of 50%, we randomly sample another training image’s caption to create a negative pair. The model then predicts the correspondence with “true” or “false” as shown in Table 1.
Visual grounding: We develop an object-text matching task to endow the model with grounding ability, which is required in several tasks (e.g., referring expression comprehension and VCR). We give the model a region description and let it predict the id of the related object region. With the help of the visual sentinel token (e.g.,
Grounded captioning: To teach the model with object-level information, we also use grounded captioning as an inverse task of visual grounding. As shown in Table 1, given a visual sentinel token (which indicates an image region) as text input, the model is asked to generate a corresponding textual description of the image region.
3 Pretraining Implementation Details
For both VL-T5 and VL-BART, it takes 4 days for 30-epoch pretraining with mixed precision training (Narang et al., 2018) on 4 RTX 2080 Ti GPUs. We use batch size 320 and 600 for VL-T5 and VL-BART, respectively. We use AdamW (Loshchilov & Hutter, 2019) with and learning rate 1e-4 with 5% linear warmup schedule. Our code is based on PyTorch (Paszke et al., 2017) and Huggingface Transformers (Wolf et al., 2019).
Downstream Tasks and Results
In this section, we compare our generative architectures VL-T5 and VL-BART on a diverse set of 7 downstream tasks (details in Appendix) with existing vision-and-language pretrained transformers (Tan & Bansal, 2019; Lu et al., 2019; Chen et al., 2020; Zhou et al., 2020; Li et al., 2020b; Xia et al., 2020). As summarized in Table 2, our unified generative approach (with the input-output format in Table 1) shows performance close to the task-specific models, most of which are discriminative. In the rest of this section, we provide detailed comparisons w.r.t. the baselines.
The visual question answering task requires models to answer a question to a given context image. Table 2 compares our models VL-T5 and VL-BART with existing methods on VQA (Goyal et al., 2019) and GQA (Hudson & Manning, 2019). For both tasks, our models achieve comparable performance to existing approaches.
Generative vs. Discriminative model: Modern approaches (Tan & Bansal, 2019; Lu et al., 2019; Chen et al., 2020; Zhou et al., 2020; Li et al., 2020b) are discriminative models, where they tackle visual question answering tasks as multi-label classification over a predefined set of answer candidates. This strategy achieves strong performance but not generalizes to real-world open-ended scenarios. To quantitatively compare the existing discriminative approaches and our generative approach, we break down VQA questions into in-domain and out-of-domain questions, in terms of whether the best answer for each question is included in the top-K () answer candidates . After this split, the in-domain subset contains 24,722 questions, and the out-of-domain subset contains 1,558 questions. Table 3 shows the performance. For discriminative baselines, we introduce a sigmoid MLP classifier on top of the decoder representation of start-of-sequence token , following LXMERT and UNITER. Comparing models with the same backbone, we notice the generative models improve upon the discriminative baselines across all the subsets. This improvement is more significant on the out-of-domain subset, where the generative VL-T5 and VL-BART achieve 6 and 6.2 points improvement over their discriminative counterparts, showing the effectiveness of using generative modeling. Compared to the strong discriminative baseline UNITERBase (pretrained with 4M extra images), our generative models still show comparable overall performance while significantly outperform it on the out-of-domain subset (about 3 points).
Dataset-specific prefixes: As shown in recent works (Gao et al., 2020; Shin et al., 2020; Li & Liang, 2021; Radford et al., 2021), different text prompts could result in different finetuning results. We thus experiment with a single prefix ‘vqa’ for both VQA and GQA in VL-T5 pretraining/finetuning. Interestingly, we found slight performance increases from the original dataset-specific prefix: VQA Karpathy-test (); GQA test-dev (). This shows that a single model can successfully handle multiple VQA tasks without dataset-specific prefixes (similar results were observed in text QA (Khashabi et al., 2020)).
The task of (Suhr et al., 2019) is to determine whether a natural language statement is true about two images. To apply our model to this task, we concatenate region features from the two images and use different image id embeddings to disambiguate the regions from the two images. Then our model learns to generate text labels “true” and “false”. This is similar to the Triplet setting described in UNITER (Chen et al., 2020). In Fig. 4, we illustrate three common encoding settings for .
Table 4 shows the model results on under different encoding settings: () Triplet: joint encoding of image pairs and text; () Pair: the concatenation of individual embedding of each image-text pair; () Pair-biattn: bidirectional attention added to Pair. UNITER shows that one can improve performance with a more complex encoding setting, i.e., Pair-biattn achieves better performance than Pair, which is again better than the simplest Triplet. Note that both the Pair and the Pair-biattn settings approximately double the computational cost compared to that of the Triplet setting. While there’s the gap between our models and baselines in Pair and Pair-biattn setting, VL-T5 shows comparable performance to UNITER in Triplet setting.
3 Referring Expression Comprehension: RefCOCOg
Referring expression comprehension requires a model to correctly localize an object described by a given phrase (e.g., ‘the car on the left’). In this work, we evaluate models on the RefCOCOg (Mao et al., 2016) dataset. Similar to the visual grounding pretraining task in Sec. 4, we give our model a referring phrase and candidate region features from the image, the model then generates the visual sentinel token (e.g.,
Table 2 compares our models with discriminative baselines. With pretraining, VL-T5 significantly outperforms the strong modular model MAttNet, and achieves a reasonable performance compared to the UNITER model that has been pretrained on a much larger corpus. While our method did not achieve state-of-the-art performance, these results suggest that referring expression comprehension can be effectively formulated as a text-generation task, rather than previously (Yu et al., 2018a; Chen et al., 2020) formulated classification task over a set of visual regions, allowing more flexible architecture design. We hope our work would inspire future works in this direction. We also observe that our experiments with VL-BART on RefCOCOg diverges. One reason might be the difference in positional encoding methods of T5 and BART. During training, BART adds learned absolute positional embedding to text token embedding, whereas T5 uses relative position biases in self-attention layers instead. We hypothesize that VL-BART found strong correspondence by memorizing the positions of each training object (we observe high training accuracy, but low validation accuracy).
4 Visual Commonsense Reasoning: VCR
Visual Commonsense Reasoning (VCR) (Zellers et al., 2019) is a multiple-choice question answering task that requires commonsense reasoning beyond object or action recognition. Each VCR question (Q) has 4 answers (A) and 4 rationales (R), and it can be decomposed into two multiple choice sub-tasks: question answering (QA), and answer justification (QAR). The overall task (QAR) requires a model to not only select the correct answer to the question, but also the correct rationale for choosing the answer. Similar to Nogueira et al. (2020) that leverages language model for document ranking, we concatenate context (image+question) with each candidate choice, and let our models generate “true” for the correct choice and generate “false” otherwise, as shown in Table 1, During inference, we use to rank the choices and select the one with the highest score.
UNITER (Chen et al., 2020) has shown that a second-stage in-domain pretraining (with the same pretraining objectives as generic-domain pretraining) on the VCR dataset would significant improve VCR task performance. This is likely due to the domain difference between VCR and the generic-domain pretraining corpus (e.g., COCO Captions), e.g., the input text (concatenation of multiple sentences: [Q]+[A]+[R]) in VCR is much longer than in generic-domain pretraining. In Table 6, we show the experiment results with second stage pretraining on VCR. On VCR val split, comparing to the base models that do not pretrain, we find that both Stage 1 generic-domain pretraining and Stage 2 in-domain pretraining help improve the VCR task performance, which is consistent with the findings in UNITER. On VCR test split, we notice that our best model VL-T5 achieves a comparable (slightly better) performance to UNITER, while significantly higher performance when compared to ViLBERT.
5 Image Captioning: COCO Caption
We evaluate automatic caption generation performance on MS COCO Caption dataset (Chen et al., 2015). We use Karparthy split (Karpathy & Fei-Fei, 2015), which re-splits train2014 and val2014 images (Lin et al., 2014) into 113,287 / 5000 / 5000 for train / validation / test. While some methods use reinforcement learning-based optimization on CIDEr, we only compare with methods using cross-entropy loss. Note that image captioning is the only task in our experiments where textual context is not meaningful, which results in a notable difference in pretraining and finetuning w.r.t. the input format. Inspired by Oscar (Li et al., 2020a), we also experiment with using object tags as additional text inputs during finetuning. We use BLEU (Papineni et al., 2002), CIDEr (Vedantam et al., 2015), METEOR (Banerjee & Lavie, 2005), SPICE (Anderson et al., 2016) as evaluation metrics using COCOEvalCaphttps://github.com/tylin/coco-caption.
In Table 7, we compare our models with baselines in different settings: use of vision-and-language pretraining and use of object tag as additional text inputs. With and without vision-and-language pretraining, our models show comparable performance to baselines. Since the use of object tags requires significant extra computation, we only use it for finetuning. Using tags gives a comparable or slightly improved performance for both models, and the improvement is significant (2.5) in CIDEr for VL-BART. We expect object tag augmentation during pretraining like Oscar would further boost the performance of our models.
6 Multimodal Machine Translation: Multi30K
We evaluate multimodal machine translation performance on Multi30K dataset (Elliott et al., 2016), where a model translates English text to German text given context images. We report BLEU score using SacreBLEU (Post, 2018)https://github.com/mjpost/sacrebleu. We compare our method with state-of-the-art transformer models: Multimodal self-attention (MSA) (Yao & Wan, 2020), MeMAD (Grönroos et al., 2018). Table 8 shows that our T5-based models outperform the baselines that use strong data augmentation (e.g., back-translation) on all three test splits. Our vision-and-language models improve the text-only backbones although we did not observe improvement with vision-and-language pretraining. This might be because the source text in Multi30K contains sufficient information for translation as discussed in Caglayan et al. (2019)
7 Multi-Task Finetuning
Single-task vs. Multi-task Finetuning: While our framework has unified the architecture for different downstream tasks, the parameters are separately optimized. To see whether we can go further, we finetune a single VL-T5 model for 20 epochs, where it tackles 7 different tasks with the same set of weights. At each finetuning step, we sample a mini-batch of examples from one of the 7 tasks in a round-robin fashion. For a fair comparison, we use single-task baselines without augmentations (e.g., no 2nd stage pretraining for VCR, no object tags for COCO Captioning). Table 9 shows that our multi-task model achieves comparable performance to the separately optimized single-task models on all 7 tasks with a single set of parameters.
Single shared head vs. Task-specific heads: We also experiment with the multi-task finetuning setup of ViLBERT-MT (Lu et al., 2020), where a task-specific head is fine-tuned for each of 7 downstream tasks while sharing backbone. The head parameters are initialized from the pretrained LM head and separately updated during finetuning. The 7 task-specific heads (7H) add parameters, which is 80% of original VL-T5’s 220M parameters (P), resulting around 400M parameters in total. Since the increased parameters make the training slow, we compare both models by 5th epoch checkpoints. Table 10 shows that VL-T5 with single shared head achieves almost equal performance with task-specific heads, while having much fewer total parameters.
Conclusion
In this work, we proposed VL-T5 and VL-BART which tackle vision-and-language tasks with a unified text generation objective. Experiments show VL-T5 and VL-BART can achieve comparable performance with state-of-the-art vision-and-language transformers on diverse vision-and-language tasks without hand-crafted architectures and objectives. Especially, we demonstrate our generative approach is better suited for open-ended visual question answering. In addition, we also showed it is possible to train seven different tasks simultaneously using a single architecture with single parameters without not losing much performance. It would be an interesting future work to further explore this direction by adding even more tasks.
Acknowledgments
We thank Hyounghun Kim, Zineng Tang, Swarnadeep Saha, Xiang Zhou, and anonymous reviewers for their comments and suggestions. This work was supported by NSF-CAREER Award 1846185, ARO-YIP Award W911NF-18-1-0336, DARPA MCS Grant N66001-19-2-4031, Google Focused Research Award, and Bloomberg Data Science Ph.D. Fellowship. The views, opinions, and/or findings contained in this article are those of the authors and not of the funding agency.
References
Appendix A Comparison with Baselines
In Table 11, we compare the baseline vision-and-language transformers with our VL-T5 and VL-BART in detail, including their pretraining datasets, architecture, etc.
Appendix B Implementation Details
In Table 12 and Table 13, we show the detailed statistics of our pretraining and downstream datasets and tasks. In Table 14, we show the hyperparameters that we used in our pretraining and downstream task experiments. We provide the links to download pretraining and downstream datasets.
Overall, our pretraining dataset contains 9.18M image-text pairs on 180K distinct images. We carefully split our pretraining data to avoid any intersection between our training data and the validation/test sets of the downstream tasks (e.g., COCO Captioning, RefCOCOg). In this process, around 10K images are excluded from the training sets of COCOhttps://cocodataset.org/#download and Visual Genomehttp://visualgenome.org/api/v0/api_home.html. We use COCO Karpathy val split (Karpathy & Fei-Fei, 2015) with 5,000 images as our validation set to monitor pretraining performance.
B.2 Downstream Tasks
For both VQA and COCO captioning tasks, we follow Karparthy split (Karpathy & Fei-Fei, 2015), which re-splits train2014 and val2014 COCO images (Lin et al., 2014) into 113,287 / 5,000 / 5,000 images for train / validation / test.
GQAhttps://cs.stanford.edu/people/dorarad/gqa/download.html
Following LXMERT (Tan & Bansal, 2019), we use GQA-balanced version. We use train and val splits for training and use test-dev split for validation. Train / val / test-dev splits consist of 943,000 / 132,062 / 12,578 questions, respectively.
Train / val / test-P splits consist of 86,373 / 6982 / 6967 sentences, respectively. We train our model on train split and use val split for validation.
VCRhttps://visualcommonsense.com/download/
Train / val / test splits consist of 212,923 / 26,534 / 25,263 questions, respectively. We train our model on train split and use val split for validation.
RefCOCOghttps://github.com/lichengunc/refer
We use umd split, which consists of train / val / test sets with 42,226 / 2,573 / 5,023 sentences, respectively. Following UNITER (Chen et al., 2020) and MAttNet (Yu et al., 2018a), we use ground truth COCO boxes for training, and use the detected boxes from an off-the-shelf Mask R-CNN https://github.com/lichengunc/MAttNet#pre-computed-detectionsmasks as candidates during inference.
Multi30K En-Dehttps://github.com/multi30k/dataset
The train / val / test2016 / test2017 / test2018 splits consist of 29,000 / 1,014 / 1,000 / 1,000 / 1,017 English-German sentence pairs, respectively.