VirTex: Learning Visual Representations from Textual Annotations
Karan Desai, Justin Johnson
Introduction
The prevailing paradigm for learning visual representations is first to pretrain a convolutional network to perform image classification on ImageNet , then transfer the learned features to downstream tasks . This approach has been wildly successful, and has led to significant advances on a wide variety of computer vision problems such as object detection , semantic and instance segmentation, image captioning , and visual question answering . Despite its practical success, this approach is expensive to scale since the pretraining step relies on images annotated by human workers.
For this reason, there has been increasing interest in unsupervised pretraining methods that use unlabeled images to learn visual representations which are then transferred to downstream tasks . Some recent approaches have begun to match or exceed supervised pretraining on ImageNet , and have been scaled to hundreds of millions or billions of images.
Continuing to scale unsupervised pretraining to ever-larger sets of unlabeled images is an important scientific goal. But we may also ask whether there are alternate ways of pretraining that learn high-quality visual representations with fewer images. To do so, we revisit supervised pretraining and seek an alternative to traditional classification pretraining that uses each image more efficiently.
In this paper we present an approach for learning Visual representations from Textual annotations (VirTex). Our approach is straightforward: first, we jointly train a ConvNet and Transformer from scratch to generate natural language captions for images. Then, we transfer the learned features to downstream visual recognition tasks (Figure 1).
We believe that using language supervision is appealing due to its semantic density. Figure 2 compares different pretraining tasks for learning visual representations. Captions provide a semantically denser learning signal than unsupervised contrastive methods and supervised classification. Hence, we expect that using textual features to learn visual features may require fewer images than other approaches.
Another benefit of textual annotations is simplified data collection. To collect classification labels, typically human experts first build an ontology of categories , then complex crowdsourcing pipelines are used to elicit labels from non-expert users . In contrast, natural language descriptions do not require an explicit ontology and can easily be written by non-expert workers, leading to a simplified data collection pipeline . Large quantities of weakly aligned images and text can also be obtained from internet images .
Our main contribution is to show that natural language can provide supervision for learning transferable visual representations with better data-efficiency than other approaches. We train models from scratch on the COCO Captions dataset , and evaluate the learned features on downstream tasks including image classification, object detection, instance segmentation, and low-shot recognition. On all tasks, VirTex matches or exceeds the performance of existing methods for supervised or unsupervised pretraining on ImageNet, despite using up to fewer images. Our code and pretrained models are available at https://github.com/kdexd/virtex
Related Work
Our work is related to recent efforts to move beyond supervised pretraining on ImageNet using alternate data sources or pretraining tasks.
Weakly Supervised Learning scales beyond supervised pretraining with a quantity over quality approach, and learns on large numbers of images with noisy labels from web services. Li et al. trains visual N-gram models on the YFCC-100M dataset , that provides 100M Flickr images with user-provided tags. Recent works also use JFT-300M dataset, curated by automatic labeling of images from web signals using Google’s internal tooling. Weakly-supervised learning has also been studied on up to 3.5B Instagram images, using hashtags as labels . These approaches learn visual representations with large quantities of images with low-quality labels; in contrast we focus on using fewer images with high-quality annotations.
Self-Supervised Learning focuses on learning visual representations by solving pretext tasks defined on unlabeled images. Early works on self-supervised learning proposed hand-crafted pretext tasks, such as context prediction , colorization , solving jigsaw puzzles , predicting rotation , inpainting , clustering , and generative modeling . Recent works are based on contrastive learning , encouraging similarity between image features under different random transformations on single input image . Other approaches use contrastive losses based on context prediction , mutual information maximization , predicting masked regions , and clustering .
These methods lack semantic understanding as they rely on low-level visual cues (color, texture), whereas we leverage textual annotations for semantic understanding. Unlike these methods, our approach can leverage additional metadata such as text, when scaled to internet images .
Vision-and-Language Pretraining attempts to learn joint representations of image-text paired data that can be transferred to multimodal downstream tasks such as visual question answering , visual reasoning , referring expressions , and language-based image retrieval . Inspired by the success of BERT in NLP, several recent methods use Transformers to learn transferable joint representations of images and text .
These methods employ complex pretraining pipelines: they typically (1) start from an ImageNet-pretrained CNN; (2) extract region features using an object detector fine-tuned on Visual Genome , following ; (3) optionally start from a pretrained language model, such as BERT ; (4) combine the models from (2) and (3), and train a multimodal transformer on Conceptual Captions ; (5) fine-tune the model from (4) on the downstream task. In this pipeline, all vision-and-language tasks are downstream from the initial visual representations learned on ImageNet. In contrast, we pretrain via image captioning, and put vision tasks downstream from vision-and-language pretraining.
Concurrent Work: Our work is closest to Sariyildiz et al. on learning visual representations from captions via image conditioned masked language modeling, with one major difference – we train our entire model from scratch, whereas they rely on pretrained BERT for textual features. Moreover, we evaluate on additional downstream tasks like object detection and instance segmentation. Our work is also closely related to Stroud et al. on learning video representations using paired textual metadata, however they solely operate and evaluate their method on video tasks.
Method
Given a dataset of image-caption pairs, our goal is to learn visual representations that can be transferred to downstream visual recognition tasks. As shown in Figure 2, captions carry rich semantic information about images, including the presence of objects (cat, plate, cake); attributes of objects (orange and white cat); spatial arrangement of objects (cat near a plate); and their actions (looking at apples). Learned visual representations that capture such rich semantics should be useful for many downstream vision tasks.
To this end, we train image captioning models to predict captions from images. As shown in Figure 3, our model has two components: a visual backbone and a textual head. The visual backbone extracts visual features from an input image . The textual head accepts these features and predicts a caption token by token, where and are fixed special tokens indicating the start and end of sentence. The textual head performs bidirectional captioning (bicaptioning): it comprises a forward model that predicts tokens left-to-right, and a backward model that predicts right-to-left. All model components are randomly initialized, and jointly trained to maximize the log-likelihood of the correct caption tokens
where , , and are the parameters of the visual backbone, forward, and backward models respectively. After training, we discard the textual head and transfer the visual backbone to downstream visual recognition tasks.
Language Modeling: Our choice of pretraining task is image captioning – a well-studied vision-and-language task, so far kept downstream from vision-based pretraining. We draw inspiration from recent work in NLP using language modeling as a pretraining task to learn transferable text representations. This involves training massive language models – either unidirectional or bidirectional , for predicting tokens one by one. However, following BERT , many large-scale models instead use masked language models (MLMs): some tokens are randomly masked and are predicted by the model.
We performed preliminary experiments with MLMs, but like we observed that MLMs converge more slowly than directional models. We note that MLMs have poor sample efficiency, as they only predict a subset of tokens for each caption, while directional models predict all tokens. Due to computational constraints, we focus on directional models and leave MLMs to future work.
Visual Backbone: The visual backbone is a convolutional network which computes visual features of images. It inputs raw image pixels, and outputs a spatial grid of image features. During pretraining, these features are used to predict captions. In downstream tasks, we either train linear models on features extracted from the visual backbone, or fine-tune the visual backbone end-to-end.
In principle we could use any convolutional network architecture for the visual backbone. In our experiments we use a standard ResNet-50 as the visual backbone to facilitate comparison with our baseline methods (Section 4). It accepts a image and produces a grid of -dimensional features after the final convolutional layer. During pretraining, we apply a linear projection layer to the visual features before passing them to the textual head to facilitate decoder attention over visual features. This projection layer is not used in downstream tasks.
Textual Head: The textual head receives features from the visual backbone and predicts captions for images. It provides a learning signal to the visual backbone during pretraining. Our overall goal is not to predict high-quality captions, but instead to learn transferable visual features.
The textual head comprises two identical language models which predict captions in forward and backward directions respectively. Following recent advances in language modeling, we use Transformers , which use multiheaded self-attention both to propagate information along the sequence of caption tokens, as well as to fuse visual and textual features. We closely follow the transformer decoder architecture from , but use GELU rather than ReLU, following . We briefly review the architecture here; refer to for a more complete description.
During training, the forward model receives two inputs: image features from the visual backbone, and a caption describing the image. Image features are a matrix of shape giving a -dimensional vector for each of the positions in the final layer of the visual backbone. As described earlier, the caption is a sequence of tokens, with and . It is trained to predict token-by-token, starting with . The prediction is causal – it only depends on past predictions and visual features. The backward model is similar; it operates right-to-left – trained to predict , given .
First, we convert the tokens of to vectors via learned token and positional embeddings, followed by elementwise sum, layer normalization and dropout . Next, we process these vectors through a sequence of Transformer layers. As shown in Figure 3, each layer performs masked multiheaded self-attention over token vectors, multiheaded attention between token vectors and image vectors, and applies a two-layer fully-connected network to each vector. These three operations are each followed by dropout, wrapped in a residual connection, and followed by layer normalization. Token vectors interact only through self-attention; the masking in this operation maintains causal structure of the final predictions. After the last Transformer layer, we apply a linear layer to each vector to predict unnormalized log-probabilities over the token vocabulary.
The forward and backward models consist of independent Transformer layers. However they share the same token embedding matrix (similar to ) which is also reused at the output layers of each model (similar to ).
Model Size: Several architectural hyperparameters control the size of our textual head. We can control the width of each Transformer layer by varying its hidden size , the number of attention heads used in multiheaded attention, and the feedforward size of the fully-connected network. We follow and always set and ; this allows us to control the width of our textual head by varying . We can also control the depth of our textual head by varying the number of transformer layers .
Tokenization: We tokenize captions with SentencePiece using the BPE algorithm . Prior to tokenization we lowercase and strip accents from captions. We build a vocabulary of 10K tokens, including boundary ([SOS], [EOS]) and out-of-vocab ([UNK]) tokens. Following we restrict subword merges between letters and punctuation to prevent redundant tokens such as dog? and dog!. Compared to basic tokenization schemes often used for image captioning that split on whitespace , BPE makes fewer linguistic assumptions, exploits subword information, and results in fewer out-of-vocab tokens.
Training Details: We train on the train2017 split of the COCO Captions dataset , which provides images with five captions each. During training we apply standard data augmentation: we randomly crop to 20-100% of the original image size, apply color jitter (brightness, contrast, saturation, hue), and normalize using the ImageNet mean color. We also apply random horizontal flips, also interchanging the words ‘left’ and ‘right’ in the caption.
We train using SGD with momentum 0.9 and weight decay wrapped in LookAhead with and 5 steps. Following , we do not apply weight decay to layer normalization and bias parameters in Transformers. We perform distributed training across 8 GPUs with batch normalization per GPU, following . We train with a batch size of 256 images (32 per GPU) for 500K iterations (1080 epochs). We use linear learning rate warmup for the first 10K iterations followed by cosine decay to zero. We found that the visual backbone required a higher LR than the textual head for faster convergence. The visual backbone uses a max LR of ; the textual head uses . We implement our models using PyTorch with native automatic mixed-precision .
We observe that performance on image captioning has a positive but imprecise correlation with performance on downstream visual recognition tasks (Refer Appendix A.4). We thus perform early stopping based on the performance of our visual backbone on downstream PASCAL VOC linear classification (see Section 4.1) since it is fast to evaluate and correlates well with our other downstream tasks.
Experiments
In our experiments, we aim to demonstrate the effectiveness of learning visual features via natural language supervision. As described in Section 3, we jointly train a VirTex model from scratch on the COCO Captions dataset. Here, we evaluate the features learned by visual backbone on six downstream vision tasks. We select these tasks based on two common mechanisms for transfer learning: where the visual backbone is either used as (a) frozen feature extractor, or (b) weight initialization for fine-tuning.
Our first set of evaluations involve training linear models on frozen visual backbones – we compare VirTex with various pretraining methods to test our two hypotheses:
Learning visual features via captions is cheaper than using other types of annotations, like labels and masks.
Using semantically dense captions helps with learning effective visual features using fewer training images.
We evaluate on two datasets: PASCAL VOC and ImageNet-1k . We choose these tasks based on their simplicity and evaluation speed. We briefly describe the setup here. Refer Appendix A.1 for more details.
PASCAL VOC: We follow same protocol as SwAV (highly similar to ); we train on VOC07 trainval split (9K images, 20 classes) and report mAP on test split. We train per-class SVMs on 2048-dimensional global average pooled features extracted from the last layer of the visual backbone. For each class, we train SVMs for cost values and select best by 3-fold cross-validation. Other SVM hyperparameters are same as .
ImageNet-1k: We follow similar protocol as MoCo and SwAV : we train on the ILSVRC 2012 train split and report top-1 accuracy on val split. We train a linear classifier (fully connected layer + softmax) on 2048-dimensional global average pooled features extracted from the last layer of the visual backbone. We train with batch size 256 distributed across 8 GPUs for 100 epochs. We use SGD with momentum and weight decay . We set the initial LR to and decay it to zero by cosine schedule.
Annotation Cost Efficiency: We believe that using captions is appealing due to a simple and cost-efficient collection pipeline. Here, we test our first hypothesis by comparing various pretraining methods on COCO, each drawing supervision from different annotation types (Figure 2):
MoCo-COCO (self-supervised): We train a MoCo-v1 model on COCO images with default hyperparameters.
Multi-label Classification (labels): We use COCO object detection annotations (80 classes), and train a ResNet-50 backbone to predict a -hot vector with values with a KL-divergence loss, similar to .
Instance Segmentation (masks): We use a pretrained Mask R-CNN from Detectron2 model zoo , and extract its ResNet-50 backbone for downstream tasks. This model is trained from scratch on COCO, following .
VirTex (captions): We train a VirTex model on COCO Captions, with ResNet-50 visual backbone and textual head. Note that COCO Captions provides five captions per image, which effectively increases image-caption pairs by five-fold. Hence for a fair comparison, we also train an additional VirTex model using only one randomly selected caption per image.
Results are shown in Table 1. We also compare annotation costs in terms of worker hours. For labels and masks, we use estimates reported by COCO . For captions, we estimate the cost based on nocaps We could not find estimates for COCO Captions in existing literature., that follows a similar data collection protocol as COCO. We observe that VirTex outperforms all methods, and has the best performance vs. cost tradeoff, indicating that learning visual features using captions is more cost-efficient than labels or masks.
Data Efficiency: We believe that the semantic density of captions should allow VirTex to learn effective visual features from fewer images than other methods. To test our hypothesis, we compare VirTex and ImageNet-supervised models (IN-sup) trained using varying amount of images from COCO Captions and ImageNet-1k respectively.
We train 4 VirTex models using of COCO Captions (118K images) and 7 ResNet-50 models using of ImageNet-1k (1.28M images). Similar to prior experiments, we also train 4 VirTex models using one randomly selected caption per image. All VirTex models use textual heads.
We show results in Figure 4. On VOC07, VirTex- outperforms IN-sup- (mAP 88.7 vs 87.6), despite using fewer images (118K vs. 1.28M). When using similar amount of images, VirTex consistently outperforms IN-sup (blue, orange vs green), indicating superior data efficiency of VirTex. We also observe that given the same number of captions for training, it is better to spread them over more images – VirTex- (1 caption) significantly outperforms VirTex- (5 captions) (mAP 79.4 vs 69.3).
Comparison with IN-sup on ImageNet-1k classification is unfair for VirTex, since IN-sup models are trained for the downstream task, using the downstream dataset. Even so, VirTex- outperforms IN-sup- (53.8 vs. 53.6, 118K vs. 128K images), and consistently outperforms it when both methods use fewer than 100K images.
Comparison with other methods: Here, we compare VirTex with recent pretraining methods that have demonstrated competitive performance on downstream tasks.
Self-supervised pretraining: We choose three recent methods based on their availability and compatibility with our evaluation setup – MoCo , PCL , and SwAV . We choose models trained with a similar compute budget as ours (8 GPUs, 200 ImageNet epochs).
ICMLM (Concurrent Work): We adapt numbers from Sariyildiz et al. ; evaluation may slightly differ. This model uses pretrained BERT for textual features.
Note on vision-and-language pretraining: Since we use captions, we also consider methods that learn multimodal representations for downstream vision-and-language tasks ). As described in Section 2, all these methods use an object detector trained on Visual Genome (with ImageNet-pretrained backbone) to extract visual features, made available by . These features are kept frozen, and do not learn from any textual supervision at all. Our comparison with ImageNet-supervised models subsumes this family of models.
Results are shown in Table 2. VirTex outperforms all methods on VOC07, despite being trained with much fewer images. On ImageNet-1k, comparison between self-supervised models and VirTex is unfair on both ends, as the former observes downstream images during pretraining, while the latter uses annotated images.
2 Ablations
The preceeding linear classification experiments demonstrate the effectiveness and data-efficiency of VirTex. In this section, we conduct ablation studies to isolate the effects of our pretraining setup and modeling decisions, and uncover performance trends to seed intuition for future work. We evaluate all ablations on PASCAL VOC and ImageNet-1k linear classification, as described in Section 4.1.
Pretraining Task Ablations: We choose bicaptioning task as it gives a dense supervisory signal per caption. To justify this choice, we form three pretraining tasks with sparser supervisory signal and compare them with bicaptioning:
Forward Captioning: We remove the backward transformer decoder and only perform left-to-right captioning.
Token Classification: We replace the textual head with a linear layer and perform multi-label classification (Table 1, row 2). We use the set of caption tokens as targets, completely ignoring the linguistic structure of captions.
Masked Language Modeling (MLM): We use a single bidirectional transformer in the textual head, and perform BERT-like masked language modeling. We randomly mask of input tokens, and train the model to predict ground-truth tokens of masked positions.
All textual heads with transformers have .
Results are shown in Figure 5(a). Bicaptioning outperforms forward captioning, indicating that denser supervisory signal from bidirectional modeling is beneficial. Bicaptioning and forward captioning both outperform token classification, demonstrating that learning to model the sequential structure of language improves visual features.
MLM performs quite worse than all three methods, possibly due to poor sample efficiency (discussed in Section 3) It may benefit from longer training schedules, however we leave this for future work due to computational constraints.
Visual Backbone Ablations: Bigger visual backbones tend to show improvements on many vision tasks . We investigate whether VirTex models with bigger visual backbones can improve downstream performance. We train three VirTex models with textual heads, and different visual backbones: (a) ResNet-50 (default), (b) ResNet-50 w2 ( channel width), and (c) ResNet-101 ( depth). We observe that bigger visual backbones better results on VOC07, however the trends are opposite on ImageNet (Figure 5(b)). We believe it to be an optimization issue. See Appendix A.2 for comparison on other tasks.
Transformer Size Ablations: Prior work in language modeling has shown that larger Transformers tend to learn better textual features . We investigate whether this holds for VirTex: do larger transformers in the textual head cause the visual backbone to learn better visual features? As discussed in Section 3, we may scale our textual head by increasing its width (hidden size ) or its depth (number of layers ). We investigate both, training VirTex models with:
Fixed , increasing .
Fixed , increasing .
Results are shown in Figure 5(c) – increasing transformer size, both width and depth, generally improves downstream performance. Performance degrades slightly with very deep transformers (), indicating overfitting. We hope that massive transformers with billions of parameters will help when scaling VirTex to large-scale, more noisy image-text paired datasets that are larger than COCO Captions.
3 Fine-tuning Tasks for Transfer
So far we have evaluated VirTex using features extracted from frozen visual backbones. Another common mechanisms for transfer learning is fine-tuning, where the entire visual backbone is updated for the downstream task.
We evaluate features learned using VirTex on four downstream tasks with fine-tuning: (a) Instance Segmentation on COCO ; (b) Instance Segmentation on LVIS ; and (c) Object Detection on PASCAL VOC ; (d) Fine-grained Classification on iNaturalist 2018 . In all these experiments, we use the VirTex model with ResNet-50 visual backbone and a textual head with .
Baselines: Our main baselines are ImageNet-supervised (IN-sup) and MoCo. We consider three variants of IN-sup pretrained with of ImageNet images (Figure 4). Similarly for MoCo, we consider both MoCo-IN (Table 2) and MoCo-COCO (Table 1). We also include Random Init baseline, trained from scratch on downstream task.
We follow the same evaluation protocol as MoCo for all four tasks. We use Detectron2 for tasks (a,b,c). Our IN-sup- results are slightly better than those reported in – we use pretrained ResNet-50 model from torchvision, whereas they used the MSRA ResNet-50 model from Detectron . We briefly describe implementation details that differ from default Detectron2 settings. Refer Appendix A.3 for full details.
COCO Instance Segmentation: We train Mask R-CNN models with ResNet-50-FPN backbones . We initialize backbone with pretrained weights, train on train2017 split, and evaluate on val2017 split. We fine-tune all layers end-to-end with BN layers synchronized across GPUs (SyncBN). We also use SyncBN in FPN layers. We train with batch size 16 distributed across 8 GPUs, following schedule (180K iterations with initial LR 0.02, multiplied by 0.1 at iterations 120K and 160K).
LVIS Instance Segmentation: The LVIS dataset provides instance segmentation labels for a long tail of 1203 entry-level object categories, and stresses the ability to recognize many object types from few training samples. We train Mask R-CNN models with ResNet-50-FPN backbones on train_v1.0 and evaluate on val_v1.0 split. Following MoCo settings, we keep BN parameters frozen for all IN-sup baselines. We train with schedule as COCO, use class resampling and test-time hyperparameters (0.0 score threshold and 300 detections per image) same as .
PASCAL VOC Detection: We train Faster R-CNN models with ResNet-50-C4 backbones on trainval07+12 split, and evaluate on test2007 split. Like COCO, we fine-tune all models with SyncBN. We train for 24K iterations, including linear LR warmup for first 100 iterations. We set the maximum LR as 0.02, that is divided by 10 at iterations 18K and 22K. We distribute training across 8 GPUs, with batch size 2 per GPU. We use gradient checkpointing to reduce the heavy memory footprint of these models and train them with desired batch size on our 12 GB GPUs.
iNaturalist 2018 Fine-grained Classification: The iNaturalist 2018 dataset provides labeled images for 8142 fine-grained categories, with a long-tailed distribution. We fine-tune the pretrained ResNet-50 with a linear layer end-to-end. We train on train2018 split and evaluate on val2018 split, following training setup from torchvision – we train for 100 epochs using SGD with momentum 0.9 and weight decay , and batch size 256 distributed across 8 GPUs. Fine-tuning uses LR 0.025 (and Random Init uses 0.1), which is multiplied by 0.1 at epochs 70 and 90.
Results: We show results in Table 3. VirTex matches or exceeds ImageNet-supervised pretraining and MoCo-IN on all tasks (row 2,5 vs. 7) despite using 10 fewer pretraining images. Moreover, VirTex significantly outperforms methods that use similar, or more pretraining images (row 3,4,6 vs. 7), indicating its superior data-efficiency. Among all tasks, VirTex shows significant improvements on LVIS, that shows the effectiveness of natural language annotations in capturing the long tail of visual concepts in the real world.
4 Image Captioning
Our goal is to learn transferable visual features via textual supervision. To do so, we use image captioning as a pretraining task. Although our goal is not to advance the state-of-the-art in image captioning, in Figure 6 we show quantitative and qualitative results of VirTex models trained from scratch on COCO. All models show modest performance, far from current state-of-the-art methods, that commonly involve some pretraining. However, captioning metrics are known to correlate weakly with human judgement – we surpass human performance on COCO.
We show some predicted captions by VirTex (R-50, ) model. We apply beam search on the forward transformer decoder (5 beams) to decode most likely captions. The decoder attention module in this transformer attends over a grid of image features through heads at each time-step for predicting a token. We average these attention weights over all the heads, and overlay them on input image (via bicubic upsampling).
In Figure 6, we show visualizations for some tokens. We observe that our model attends to relevant image regions for making predictions, indicating that VirTex learns meaningful visual features with good semantic understanding.
Conclusion
We have shown that learning visual representations using textual annotations can be competitive to methods based on supervised classification and self-supervised learning on ImageNet. We solely focus on downstream vision tasks – future works can explore other tasks that transfer both the visual backbone and the textual head. Finally, the usage of captions opens a clear pathway to scaling our approach to web-scale image-text pairs, that are orders of magnitude larger, albeit more noisy than COCO Captions.
Acknowledgments
We thank Harsh Agrawal, Mohamed El Banani, Richard Higgins, Nilesh Kulkarni and Chris Rockwell for helpful discussions and feedback on the paper. We thank Ishan Misra for discussions regarding PIRL/SwAV evaluation protocol; Saining Xie for discussions about replicating iNaturalist evaluation as MoCo; Ross Girshick and Yuxin Wu for help with Detectron2 model zoo; Georgia Gkioxari for suggesting the Instance Segmentation pretraining task ablation; and Stefan Lee for suggestions on figure aesthetics. We thank Jia Deng for access to extra GPUs during project development; and UMich ARC-TS team for support with GPU cluster management. Finally, we thank all the Starbucks outlets in Ann Arbor for many hours of free WiFi. This work was partially supported by the Toyota Research Institute (TRI). However, note that this article solely reflects the opinions and conclusions of its authors and not TRI or any other Toyota entity.
References
Appendix A Additional Experiments
In this section, we describe additional implementation details about our experiments in Section 4. Our evaluation protocol is consistent with prior works on pretraining visual representations – we report differences where applicable.
PASCAL VOC: We use standard data augmentation on images from both trainval and test split – we resize the shorter edge to 256 pixels, and take a center crop. We normalize images by ImageNet color (RGB mean = , std = ).
Prior works train per-class SVMs for (26 values), and choose best SVM based on 3-fold cross-validation. In our initial evaluations, we observed that the best performing SVMs are typically trained with cost values . Based on this observation, we only use these values for faster evaluation. For training SVMs, we use scikit-learn with LIBLINEAR backend, default parameters are: LinearSVC(penalty=‘l2’, dual=True, max_iter=2000, tol=1e-4, class_weight={1: 2, -1: 1}, loss=‘squared_hinge’).
ImageNet-1k: For data augmentation during training, we randomly crop 20–100 of the original image size, with a random aspect ratio in , resize to , apply random flip, and normalization by ImageNet color. During evaluation, we resize the shorter edge to 256 pixels and take a center crop. We initialize the weights of the linear layer as , and bias values as 0.
Note that we perform a small LR sweep separately for our VirTex model (ResNet-50 and ), and ImageNet-supervised models. For Figure 4, best LR values for VirTex models is 0.3 (as mentioned in Section 4.1, and ImageNet-supervised models is 0.1.
Annotation Cost Efficiency: Here, we provide details on our cost estimates for different methods in Table 1. For labels and masks, we use estimates reported by COCO , and for captions we use estimates reported by nocaps , collected in a similar fashion as COCO.
Labels: We consider total time of Category Labeling and Instance Spotting steps in (30K hours). This estimate corresponds to 328K images – we scale it for COCO Captions train2017 split (118K images).
Masks: As reported in , it takes 22 worker hours for collecting 1000 instance segmentation masks. We use this estimate to compute time for 860K masks in COCO train2017 split. The collection of masks is dependent on Category Labeling and Instance Spotting, we add the time for collecting labels in our total estimate.
Captions: We use the median time per caption (39.2 seconds) as reported in (151K captions) to estimate the cost of collecting (118K ) captions in COCO.
Data Efficiency: We train our ImageNet-supervised models on randomly sampled subsets of ImageNet (, , , , ). We sample training examples such that the class distribution remains close to ImageNet. For VirTex models, we randomly sample , , , and of COCO Captions – we do not use any class labels to enforce uniform class distribution. Note that this may put ImageNet-supervised models at an advantage.
We train our ImageNet-supervised models by following the exact setup used to train the publicly available ResNet-50 model in torchvision. We use SGD with momentum 0.9 and weight decay . We use a batch size of 256, and perform distributed training across 8 GPUs (batch size 32 per GPU). We train for 90 epochs, with an initial learning rate 0.1, that is divided by 10 at epochs 30 and 60. We keep the number of training epochs fixed for models trained on smaller subsets of ImageNet (else they tend to overfit). For VirTex models, we scale training iterations according to the size of the sampled training set.
Comparison: ImageNet vs. Cropped COCO. Note that the ImageNet images mostly contain a single object (commonly called iconic images). On the other hand, COCO dataset contains 2.9 object classes and 5.7 instances per image. It may seem that VirTex requires fewer images than ImageNet-supervised models as they contain multiple objects per image. Here, we make an additional comparison to control the varying image statistics between datasets.
Specifically, we crop objects from COCO images and create a dataset of 860K iconic images. We randomly expand bounding boxes on all edges by 0–30 pixels before cropping, to mimic ImageNet-like images. We train a ResNet-50 with same hyperparameters as ImageNet-supervised models, described above. It achieves 79.1 VOC07 mAP (vs. 88.7 VirTex). This shows that the data-efficiency of VirTex does not entirely stem from using scene images with multiple objects.
A.2 Ablations
Bicaptioning vs. Masked Language Modeling. In our pretraining task ablations (Section 4.2), we observed that Masked Language Modeling performs quite worse than all other pretraining tasks on downstream linear classification performance. This issue arises from the poor sample efficiency of Masked LM, discussed in Section 3.
For more evidence, we inspect VOC07 mAP of Masked LM, validated periodically during training. In Figure 7, we compare this with VOC07 mAP of Bicaptioning. Both models use textual heads. We find that Masked LM indeed converges slower than bicaptioning, as it receives weaker supervision per training caption – only corresponding to masked tokens. We believe that a longer training schedule may lead to MLM outperforming bicaptioning, based on its success in language pretraining .
Additional Evaluation: Backbone Ablations. In our backbone ablations (Figure 5), we observed that larger visual backbones improve VOC07 classification performance. However, the performance trend for ImageNet-1k linear classification is opposite. We think this is an optimization issue – the hyperparameters chosen for ResNet-50 may not be optimal for other backbones. To verify our claims, we evaluate these models on PASCAL VOC object detection.
In Table 4, we observe that the performance trends of PASCAL VOC object detection match with VOC07 classification. Hence, we conclude that using larger visual backbones can improve downstream performance.
A.3 Fine-tuning Tasks for Transfer
We described the main details for downstream fine-tuning tasks in Section 4.3. We provide config files in Detectron2 format to exactly replicate our downstream fine-tuning setup for COCO (Table 5), PASCAL VOC (Table 6), LVIS (Table 7). We apply modified hyperparameters on top of base config files available at: github.com/facebookresearch/detectron2 @ b267c6
iNaturalist 2018 Fine-grained Classification: We use data augmentation and weight initialization same as ImageNet-1k linear classification (Section A.1). Despite a long-tailed distribution like LVIS, we do not perform class balanced resampling, following the evaluation setup of MoCo .
LVIS v0.5 Instance Segmentation: In Section 4.3, we evaluated VirTex and baseline methods on LVIS Instance Segmentation task using LVIS v1.0 train and val splits. One of our baselines, MoCo, conducted this evaluation using LVIS v0.5 splits. For completeness, we report additional results on LVIS v0.5 split. The main changes in config (Table 7) following original LVIS v0.5 baselines are: NUM_CLASSES: 1230 and SCORE_THRESHOLD_TEST: 0.0
Results are shown in Table 8. We observe the VirTex significantly outperforms all baseline methods on LVIS v0.5 split, similar to evaluation on LVIS v1.0 split.
A.4 Selecting Best Checkpoint by VOC07 mAP
As described in Section 3, we observed that image captioning performance has an imprecise correlation with performance on downstream vision tasks. Hence, we select our best checkpoint based on VOC07 classification mAP.
In Figure 8, we compare validation metrics of our best VirTex model (ResNet-50, ). We observe the trends of VOC07 mAP and CIDEr score of the forward transformer decoder. We observe that an improvement in captioning performance indicates an improvement in downstream performance. However these are not strongly correlated – the best performing checkpoints according to these metrics occur at different iterations: 496K according to VOC07 mAP (88.7), and 492K according to CIDEr (105.8). Hence, we select the best checkpoint based on PASCAL VOC linear classification performance. We use this task as a representative downstream vision task for evaluation due to its speed and simplicity.
Appendix B Decoder Attention Visualizations for Caption Predictions
In Figure 9 and Figure 10, we show more qualitative examples showing decoder attention weights overlaid on input images, similar to Section 4.4. All captions are decoded from VirTex model using beam search. We normalize the attention masks to $$ to improve their contrast for better visibility.