A Corpus for Reasoning About Natural Language Grounded in Photographs
Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, Yoav Artzi
Introduction
Visual reasoning with natural language is a promising avenue to study compositional semantics by grounding words, phrases, and complete sentences to objects, their properties, and relations in images. This type of linguistic reasoning is critical for interactions grounded in visually complex environments, such as in robotic applications. However, commonly used resources for language and vision (Antol et al. 2015; Chen et al. 2016, e.g.,) focus mostly on identification of object properties and few spatial relations (Ferraro et al. 2015; Alikhani and Stone 2019, Section 4;). This relatively simple reasoning, together with biases in the data, removes much of the need to consider language compositionality Goyal et al. 2017. This motivated the design of datasets that require compositional In parts of this paper, we use the term compositional differently than it is commonly used in linguistics to refer to reasoning that requires composition. This type of reasoning often manifests itself in highly compositional language. visual reasoning, including NLVR Suhr et al. 2017 and CLEVR Johnson et al. 2017a; Johnson et al. 2017b. These datasets use synthetic images, synthetic language, or both. The result is a limited representation of linguistic challenges: synthetic languages are inherently of bounded expressivity, and synthetic visual input entails limited lexical and semantic diversity.
We address these limitations with Natural Language Visual Reasoning for Real (NLVR2), a new dataset for reasoning about natural language descriptions of photos. The task is to determine if a caption is true with regard to a pair of images. Figure 3 shows examples from NLVR2. We use images with rich visual content and a data collection process designed to emphasize semantic diversity, compositionality, and visual reasoning challenges. Our process reduces the chance of unintentional linguistic biases in the dataset, and therefore the ability of expressive models to take advantage of them to solve the task. Analysis of the data shows that the rich visual input supports diverse language, and that the task requires joint reasoning over the two inputs, including about sets, counts, comparisons, and spatial relations.
Scalable curation of semantically-diverse sentences that describe images requires addressing two key challenges. First, we must identify images that are visually diverse enough to support the type of language desired. For example, a photo of a single beetle with a uniform background (Table 2, bottom left) is likely to elicit only relatively simple sentences about the existence of the beetle and its properties. Second, we need a scalable process to collect a large set of captions that demonstrate diverse semantics and visual reasoning.
This paper includes four main contributions: (1) a procedure for collecting visually rich images paired with semantically-diverse language descriptions; (2) NLVR2, which contains examples of captions and image pairs, including unique sentences and images; (3) a qualitative linguistically-driven data analysis showing that our process achieves a broader representation of linguistic phenomena compared to other resources; and (4) an evaluation with several baselines and state-of-the-art visual reasoning methods on NLVR2. The relatively low performance we observe shows that NLVR2 presents a significant challenge, even for methods that perform well on existing visual reasoning tasks. NLVR2 is available at http://lil.nlp.cornell.edu/nlvr/.
Related Work and Datasets
Language understanding in the context of images has been studied within various tasks, including visual question answering (Zitnick and Parikh 2013; Antol et al. 2015, e.g.,), caption generation Chen et al. 2016, referring expression resolution (Mitchell et al. 2010; Kazemzadeh et al. 2014; Mao et al. 2016, e.g.,), visual entailment Xie et al. 2019, and binary image selection Hu et al. 2019. Recently, the relatively simple language and reasoning in existing resources motivated datasets that focus on compositional language, mostly using synthetic data for language and vision Andreas et al. 2016; Johnson et al. 2017a; Kuhnle and Copestake 2017; Kahou et al. 2018; Yang et al. 2018. A tabular summary of the comparison of NLVR2 to existing resources is available in Table 7, Appendix A. Three exceptions are CLEVR-Humans Johnson et al. 2017b, which includes human-written paraphrases of generated questions for synthetic images; NLVR Suhr et al. 2017, which uses human-written captions that compare and contrast sets of synthetic images; and GQA Hudson and Manning 2019, which uses synthetic language grounded in real-world photographs. In contrast, we focus on both human-written language and web photographs.
Several methods have been proposed for compositional visual reasoning, including modular neural networks (Andreas et al. 2016; Johnson et al. 2017b; Perez et al. 2018; Hu et al. 2017; Suarez et al. 2018; Hu et al. 2018; Yao et al. 2018; Yi et al. 2018, e.g.,) and attention- or memory-based methods (Santoro et al. 2017; Hudson and Manning 2018; Tan and Bansal 2018, e.g.,). We use FiLM Perez et al. 2018, N2NMN Hu et al. 2017, and MAC Hudson and Manning 2018 for our empirical analysis.
Data Collection
We require sets of images where the images in each set are detailed but similar enough such that comparison will require use of a diverse set of reasoning skills, more than just object or property identification. Because existing image resources, such as ImageNet Russakovsky et al. 2015 or COCO Lin et al. 2014, do not provide such grouping and mostly include relatively simple object-focused scenes, we collect a new set of images. We retrieve sets of images with similar content using search queries generated from synsets from the ILSVRC2014 ImageNet challenge Russakovsky et al. 2015. This correspondence to ImageNet synsets allows researchers to use pre-trained image featurization models, and focuses the challenges of the task not on object detection, but compositional reasoning challenges.
We identify a subset of the synsets in ILSVRC2014 that often appear in rich contexts. For example, an acorn often appears in images with other acorns, while a seawall almost always appears alone. For each synset, we issue five queries to the Google Images search engine https://images.google.com/ using query expansion heuristics. The heuristics are designed to retrieve images that support complex reasoning, including images with groups of entities, rich environments, or entities participating in activities. For example, the expansions for the synset acorn will include two acorns and acorn fruit. The heuristics are specified in Table 1. For each query, we use the Google similar images tool for each of the first five images to retrieve the seven non-duplicate most similar images. This results in five sets of eight similar images per query, At the time of publication, the similar images tool is available at the “View more” link in the list of related images after expanding the results for each image. Images are ranked by similarity, where more similar images appear higher. sets in total. If at least half of the images in a set were labeled as interesting according to the criteria in Table 2, the synset is awarded one point. We choose the synsets with the most points. We pick and remove one set due to high image pruning rate in later stages. The synsets are distributed evenly among animals and objects. This annotation was performed by the first two authors and student volunteers, is only used for identifying synsets, and is separate from the image search described below.
Image Search
We use the Google Images search engine to find sets of similar images (Figure 2a). We apply the query generation heuristics to the synsets. We use all synonyms in each synset Deng et al. 2014; Russakovsky et al. 2015. For example, for the synset timber wolf, we use the synonym set timber wolf, grey wolf, gray wolf, canis lupus . For each generated query, we download sets containing at most related images.
Image Pruning
We use two crowdsourcing tasks to (1) prune the sets of images, and (2) construct sets of eight images to use in the sentence-writing phase. In the first task, we remove low-quality images from each downloaded set of similar images (Figure 2b). We display the image set and the synset name, and ask a worker to remove any images that do not load correctly; images that contain inappropriate content, non-realistic artwork, or collages; or images that do not contain an instance of the corresponding synset. This results in sets of sixteen or fewer similar images. We discard all sets with fewer than eight images.
The second task further prunes these sets by removing duplicates and down-ranking non-interesting images (Figure 2c). The goal of this stage is to collect sets that contain enough interesting images. Workers are asked to remove duplicate images, and mark images that are not interesting. An image is interesting if it fits any of the criteria in Table 2. We ask workers not to mark an image if they consider it interesting for any other reason. We discard sets with fewer than three interesting images. We sort the images in descending order according to first interestingness, and second similarity, and keep the top eight.
2 Sentence Writing
In contrast to the collection process of NLVR, using real images does not allow for as much control over their content, in some cases permitting workers to write simple sentences. For example, a worker could write a sentence stating the existence of a single object if it was only present in both selected pairs, which is avoided in NLVR by controlling for the objects in the images. Instead, we define more specific guidelines for the workers for writing sentences, including asking to avoid subjective opinions, discussion of properties of photograph, mentions of text, and simple object identification. We include more details and examples of these guidelines in Appendix B.
3 Validation
4 Splitting the Dataset
We assign a random of the examples passing validation to development and testing, ensuring that examples from the same initial set of eight images do not appear across the split. For these examples, we collect four additional validation judgments to estimate agreement and human performance. We remove from this set examples where two or more of the extra judgments disagreed with the existing label (Section 3.3). Finally, we create equal-sized splits for a development set and two test sets, ensuring that original image sets do not appear in multiple splits of the data (Table 4).
5 Data Collection Management
We use a tiered system with bonuses to encourage workers to write linguistically diverse sentences. After every round of annotation, we sample examples for each worker and give bonuses to workers that follow our writing guidelines well. Once workers perform at a sufficient level, we allow them access to a larger pool of tasks. We also use qualification tasks to train workers. The mean cost per unique sentence in our dataset is 0.18. Appendix B provides additional details about our bonus system, qualification tasks, and costs.
6 Collection Statistics
We collect sets of related images and a total of images (Section 3.1). Pruning low-quality images leaves sets and images. Most images are removed for not containing an instance of the corresponding synset or for being non-realistic artwork or a collage of images. We construct sets of eight images each.
We crowdsource sentences (Section 3.2). We create two writing tasks for each set of eight images. Workers may flag sets of images if they should have been removed in earlier stages; for example, if they contain duplicate images. Sentence-writing tasks that remain without annotation after three days are removed.
During validation, sentences are reported as nonsensical. examples pass validation; i.e., the validation label matches the initial selection for the pair of images (Section 3.3). Removing low-agreement examples in the development and test sets yields a dataset of examples, unique images, and unique sentences. Each unique sentence is paired with an average of pairs of images. Table 3 shows examples of three unique sentences from NLVR2. Table 4 shows the sizes of the data splits, including train, development, a public test set (Test-P), and an unreleased test set (Test-U).
Data Analysis
We perform quantitative and qualitative analysis using the training and development sets.
Following validation, of the examples not reported during validation are removed due to disagreement between the validator’s label and the initial selection of the image pair (Section 3.3). The validator is the same worker as the sentence-writer for of examples. In these cases, the validator agrees with themselves of the time. For examples where the sentence-writer and validator were not the same person, they agree in of examples. We use the five validation labels we collect for the development and test sets to compute Krippendorff’s and Fleiss’ to measure agreement Cocos et al. 2015; Suhr et al. 2017. Before removing low-agreement examples (Section 3.4), and . After removal, and , indicating almost perfect agreement Landis and Koch 1977.
Synsets
Each synset is associated with examples. The five most common synsets are gorilla, bookcase, bookshop, pug, and water buffalo. The five least common synsets are orange, acorn, ox, dining table, and skunk. Synsets appear in equal proportions across the four splits.
Language
NLVR2’s vocabulary contains word types, significantly larger than NLVR, which has word types. Sentences in NLVR2 are on average tokens long, whereas NLVR has a mean sentence length of . Figure 3 shows the distribution of sentence lengths compared to related corpora. NLVR2 shows a similar distribution to NLVR, but with a longer tail. NLVR2 contains longer sentences than the questions of VQA Antol et al. 2015, GQA Hudson and Manning 2019, and CLEVR-Humans Johnson et al. 2017b. Its distribution is similar to MSCOCO Chen et al. 2015, which also contains captions, and CLEVR Johnson et al. 2017a, where the language is synthetically generated.
We analyze sentences from the development set for occurrences of semantic and syntactic phenomena (Table 5). We compare with the -example analysis of VQA and NLVR from Suhr et al. 2017, and examples from the balanced split of GQA. Generally, NLVR2 has similar linguistic diversity to NLVR, showing broader representation of linguistic phenomena than VQA and GQA. One noticeable difference from NLVR is less use of hard cardinality. This is possibly due to how NLVR is designed to use a very limited set of object attributes, which encourages writers to rely on accurate counting for discrimination more often. We include further analysis in Appendix C.
Estimating Human Performance
We use the additional labels of the development and test examples to estimate human performance. We group these labels according to workers. We do not consider cases where the worker labels a sentence written by themselves. For each worker, we measure their performance as the proportion of their judgements that matches the gold-standard label, which is the original validation label. We compute the average and standard deviation performance over workers with at least such additional validation judgments, a total of unique workers. Before pruning low-agreement examples (Section 3.4), the average performance over workers in the development and both test sets is . After pruning, it increases to . Table 6 shows human performance for each data split that has extra validations. Because this process does not include the full dataset for each worker, it is not fully comparable to our evaluation results. However, it provides an estimate by balancing between averaging over many workers and having enough samples for each worker.
Evaluation Systems
We evaluate several baselines and existing visual reasoning approaches using NLVR2. For all systems, we optimize for example-level accuracy. System and learning details are available in Appendix E.
We use two baselines that consider both language and vision inputs. The CNN+RNN baseline concatenates the encoding of the text and images, computed similar to the Text and Image baselines, and applies a multilayer perceptron to predict a truth value. The MaxEnt baseline computes features from the sentence and objects detected in the paired images. We detect the objects in the images using a Mask R-CNN model He et al. 2017; Girshick et al. 2018 pre-trained on the COCO detection task Lin et al. 2014. We use a detection threshold of . For each -gram with a numerical phrase in the caption and object class detected in the images, we compute features based on the number present in the -gram and the detected object count. We create features for each image and for both together, and use these features in a maximum entropy classifier.
Several recent approaches to visual reasoning make use of modular networks (Section 2). Broadly speaking, these approaches predict a neural network layout from the input sentence by using a set of modules. The network is used to reason about the image and text. The layout predictor may be trained: (a) using the formal programs used to generate synthetic sentences (e.g., in CLEVR), (b) using heuristically generated layouts from syntactic structures, or (c) jointly with the neural modules with latent layouts. Because sentences in NLVR2 are human-written, no supervised formal programs are available at training time. We use two methods that do not require such formal programs: end-to-end neural module networks (Hu et al. 2017, N2NMN;) and feature-wise linear modulation (Perez et al. 2018, FiLM;). For N2NMN, we evaluate three learning methods: (a) N2NMN-Cloning: using supervised learning with gold layouts; (b) N2NMN-Tune: using policy search after cloning; and (c) N2NMN-RL: using policy search from scratch. For N2NMN-Cloning, we construct layouts from constituency trees Cirik et al. 2018. Finally, we evaluate the Memory, Attention, and Composition approach (Hudson and Manning 2018, MAC;), which uses a sequence of attention-based steps. We modify N2NMN, FiLM, and MAC to process a pair of images by extracting image features from the concatenation of the pair.
Experiments and Results
We use two metrics: accuracy and consistency. Accuracy measures the per-example prediction accuracy. Consistency measures the proportion of unique sentences for which predictions are correct for all paired images Goldman et al. 2018. For training and development results, we report mean and standard deviation of accuracy and consistency over three trials as . The results on the test sets are generated by evaluating the model that achieved the highest accuracy on the development set. For the N2NMN methods, we report test results only for the best of the three variants on the development set. For reference, we also provide NLVR results in Table 11, Appendix D.
Table 6 shows results for NLVR2. Majority results demonstrate the data is fairly balanced. The results are slightly higher than perfect balance due to pruning (Sections 3.3 and 3.4). The Text and Image baselines perform similar to Majority, showing that both modalities are required to solve the task. Text shows identical performance to Majority because of how the data is balanced. The best performing system is the feature-based MaxEnt with the highest accuracy and consistency. FiLM performs best of the visual reasoning methods. Both FiLM and MAC show relatively high consistency. While almost all visual reasoning methods are able to fit the data, an indication of their high learning capacity, all generalize poorly. An exception is N2NMN-RL, which fails to fit the data, most likely due to the difficult task of policy learning from scratch. We also experimented with recent contextualized word embeddings to study the potential of stronger language models. We used a 12-layer uncased pre-trained BERT model Devlin et al. 2019 with FiLM. We observed BERT provides no benefit, and therefore use the default embedding method for each model.
Conclusion
We introduce the NLVR2 corpus for studying semantically-rich joint reasoning about photographs and natural language captions. Our focus on visually complex, natural photographs and human-written captions aims to reflect the challenges of compositional visual reasoning better than existing corpora. Our analysis shows that the language contains a wide range of linguistic phenomena including numerical expressions, quantifiers, coreference, and negation. This demonstrates how our focus on complex visual stimuli and data collection procedure result in compositional and diverse language. We experiment with baseline approaches and several methods for visual reasoning, which result in relatively low performance on NLVR2. These results and our analysis exemplify the challenge that NLVR2 introduces to methods for visual reasoning. We release training, development, and public test sets, and provide scripts to break down performance on the examples we manually analyzed (Section 4) according to the analysis categories. Procedures for evaluating on the unreleased test set and a leaderboard are available at http://lic.nlp.cornell.edu/nlvr/.
Acknowledgments
This research was supported by the NSF (CRII-1656998), a Google Faculty Award, a Facebook ParlAI Research Award, an AI2 Key Scientific Challenge Award, Amazon Cloud Credits Grant, and support from Women in Technology New York. This material is based on work supported by the National Science Foundation Graduate Research Fellowship under Grant No. DGE-1650441. We thank Mark Yatskar, Noah Snavely, and Valts Blukis for their comments and suggestions, the workers who participated in our data collection for their contributions, and the anonymous reviewers for their feedback.
References
Appendix A Frequently Asked Questions
Composition of reasoning skills including counting, comparing, and reasoning about sets is critical for robotic agents following natural language instructions. Consider a robot on a factory floor or in a cluttered workshop following the instruction get the two largest hammers from the toolbox at the end of the shelf. Correctly following this instruction requires reasoning compositionally about object properties, comparisons between these properties, counts of objects, and spatial relations between observed objects. The language in NLVR2 reflects this type of linguistic reasoning. While the task we define does not use this kind of application directly, our data enables studying models that can understand this type of language.
How can I use NLVR2 to build an end application?
The task and data are not intended to directly develop an end application. Our focus is on developing a task that drives research in vision and language understanding towards handling diverse set of reasoning skills. It is critical to keep in mind that this dataset was not analyzed for social biases. Researchers who wish to apply this work to an end product should take great care in considering what biases may exist.
Doesn’t using a binary prediction task limit the ability to gain insight into model performance?
Because our dataset contains both positive and negative image pairs for each sentence, we can measure consistency Goldman et al. 2018, which requires a model to predict each label correctly for each use of the sentence. This metric requires generalization across at most four image pair contexts.
Why collect a new set of images rather than use existing ones like COCO Lin et al. 2014?
Our goal was to achieve similar semantic diversity to NLVR, but using real images. Like NLVR, we use a sentence-writing task where sets of similar images are compared and contrasted. However, unlike NLVR, we do not have control over the image content, so cannot guarantee image sets where the content is similar enough (e.g., where the only difference is the direction in which the same animal is facing) such that the written sentence does not describe trivial image differences (e.g., the types of objects present). In addition to image similarity within sets, we also prioritize image interestingness, for example images with many instances of an object. Existing corpora, including like COCO and ImageNet Russakovsky et al. 2015, were not constructed to prioritize interestingness as we define it, and are not comprised of sets of eight very similar images as required for our task.
We select a set of ImageNet synsets which often appear in visually rich images.
We generate search queries which result in visually rich images, e.g., containing multiple instances of a synset.
We use a similar images tool to acquire sets of images with similar image content, for example containing the same objects in different relative orientations.
We prune images which do not contain an example of the synset it was derived from.
We apply a re-ranking and pruning procedure that prioritizes visually rich and interesting images, and prunes set which do not have enough interesting images.
These steps result in a total of sets of eight similar, visually rich images.
Why use pairs of images instead of single images?
We use pairs of images to elicit descriptions that reason over the pair of images in addition to the content within each image. This setup supports, for example, comparing the two images, requiring that a condition holds in both images or in one but not the other, and performing set reasoning about the objects present in each image. This is analogous to the three-box setup in NLVR.
Why allow workers to select the pairs themselves during sentence writing?
We found that for some image pair selections, it was too difficult for workers to write a sentence which distinguishes the pairs. Allowing the workers to choose the pairs avoids this feasiblity issue.
Why get multiple validations for development and test splits?
This ensures the test splits are of the highest quality and have minimal noise, as required for reliable measure of task performance. The additional annotatiosn also allow us to measure agreement and estimate human performance.
How does the NLVR2 data compares to the NLVR data?
NLVR and NLVR2 share the task of determining whether a sentence is true in a given visual context. In NLVR, the visual input is synthetic and includes a handful of shapes and properties. In NLVR2, each visual context is a pair of real photographs obtained from the web. Grounding sentences in image pairs rather than single images is related to NLVR’s use of three boxes per image.
How does the NLVR2 data collection process compare to NLVR?
We adapt the NLVR sentence-writing and validation tasks. However, rather than using four related synthetic images for writing, we use four pairs of real images. The pairing of images encourages set comparison. This was accomplished in NLVR through careful control of the generated image content, something that is not possible with real images. The NLVR image generation process is also controlled for the type of differences possible between images and the visual complexity, by ensuring the objects present in the selected and unselected images were the same. This guarantees that the only differences are in the object configurations and distribution among the three boxes in each image. Neither form of control is possible with real images. Instead, we rewrite the guidelines and develop a process to educate workers to follow them. In our process, we use the similar images tool to identify images that require linguistically-rich descriptions to distinguish. While using the similar images tool does not guarantee that the objects in the selected images are also present in the unselected images, our process successfully avoids this issue; in practice, only around of examples take advantage of this by mentioning objects only present in the selected images.
Can you summarize the key linguistic differences between NLVR2 and NLVR?
NLVR contains significantly Using a test with . more examples of hard cardinality, existential quantifiers, spatial relations, and prepositional attachment ambiguity. NLVR2 contains significantly11 more examples of soft cardinality, universal quantifiers, coordination, coreference, and comparatives. NLVR2’s descriptions are longer on average than NLVR ( vs. tokens), and the vocabulary is much larger ( vs. word types). This demonstrates both the lexical diversity and challenges of understanding a wide range of image content in NLVR2 that are not present in NLVR. However, NLVR allows studying compositionality in isolation from lexical diversity, an intended feature of the dataset’s design. NLVR has also been used as a semantic parsing task, where images are represented as structured representations Goldman et al. 2018, a use case that is not possible with NLVR2. NLVR remains a challenging dataset for visual reasoning; recent approaches have shown moderate improvements over the initial baseline performance, yet remain far from human accuracy, which we compute in Table 11.
How does NLVR2 compare to existing visual reasoning datasets?
Table 7 compares NLVR2 with several existing, related corpora. In the last several years there has been an increase in the number of datasets released for vision and language research. One trend includes building datasets for compositional visual reasoning (SHAPES, CLEVR, CLEVR-Humans, ShapeWorld, NLVR, FigureQA, COG, and GQA), all of which use synthetic data either for at least one of the inputs. While NLVR2 requires related visual reasoning skills, it uses both real natural language and real visual inputs.
How does NLVR2 compare to recent attempts to avoid biases in vision and language datasets?
Recently, several approaches were proposed to identify unintended biases present in vision-and-language tasks, such as the ability to answer a question without using the paired image Zhang et al. 2016; Goyal et al. 2017; Li et al. 2017; Agrawal et al. 2017; Agrawal et al. 2018. The data collection process of NLVR2 is designed to automatically pair each sentence with both labels in different visual contexts. This makes NLVR2 robust to implicit linguistic biases. This is illustrated by our initial experiments with BERT, which have been shown to be extremely effective at capturing language patterns for various tasks Devlin et al. 2019. With our balanced data, using BERT does not help identifying and using language biases.
Are the differences in the linguistic analysis between the datasets significant?
We measure significance using a test with . Our qualitative linguistic analysis shows several differences from VQA Antol et al. 2015 and GQA Hudson and Manning 2019. NLVR2 contains significantly more examples of hard cardinality, soft cardinality, existential quantifiers, universal quantifiers, coordination, coreference, spatial relations, comparatives, negation, and preposition attachment ambiguity than both GQA and VQA. However, VQA and GQA both contain significantly more examples of presupposition than NLVR2.
Given your linguistic analysis, how does GQA compare to VQA?
We found that the distribution of phenomena in VQA and GQA are roughly similar, with notable differences being significantly11 more examples of hard cardinality and coreference in VQA, and significantly11 more examples of universal quantifiers, coordination, and coordination and subordinating conjunction attachment ambiguity in GQA.
Appendix B Data Collection Details
We consider the images of each search query in the order of the search results. For each result associated with a set of similar images, we save the URL of the result image and the URLs of the fifteen most similar images, giving us a set of sixteen images. We skip and ignore URLs from a hand-crafted list of stock photo domains; images from these domains include large, distracting watermarks. We stop after observing result images, saving sets of image URLs, or observing five consecutive results that do not have similar images. For collective nouns and the numerical phrase two
After downloading a set of URLs of related images (Section 3.1), we automatically prune the images. We remove any broken URLs or any URLS that appeared in other previously-downloaded sets from the same search query. We remove downloaded images smaller than pixels. We apply basic duplicate removal by removing any images which are exact duplicates of a previously-downloaded image in the set. This automatic pruning may result in image sets consisting of fewer than images. We discard any sets after this stage with fewer than images.
Sentence Writing
Data Collection Management
We use two qualification tasks. For the set construction and sentence writing tasks, we qualify workers by first showing six tutorial questions about the guidelines and task. We then ask them to validate guidelines for nineteen sentences across two sets of four pre-selected image pairs, and to complete a single sentence-writing task for pre-selected image pairs. We validate the written sentence by hand. We qualify workers for validation with eight pre-selected validation tasks.
We use a bonus system to encourage workers to write linguistically diverse sentences. We conduct sentence writing in rounds. After each round, we sample twenty sentences for each worker from that round. If at least of these sentences follow the guidelines, they receive a bonus for each sentence written during the last round. If between and follow our guidelines, they receive a slightly lower bonus. This encourages workers to follow the guidelines more closely. In addition, each worker initially only has access to a limited pool of sentence-writing tasks. Once they successfully complete an evaluation round where at least of their sentences followed the guidelines, they get access to the entire pool of tasks.
Table 9 shows the costs and number of workers per task. The final cost per unique sentence in our dataset is 0.18.
Appendix C Additional Data Analysis
Figure 4 shows the counts of examples per synset in the training and development sets.
Image Pair Reasoning
We use a -sentence subset of the sentences analyzed in Table 5 to analyze what types of reasoning are required over the two images (Table 10). We observe that sentences commonly use the pair structure used to display the images: of sentences require that a property to hold in both images, simply require that a property holds in at least one image, and of sentences require a property to be true in the left or right images specifically. The pair is also used for comparison, with of sentences requiring comparing properties of the two images. Finally, of sentences simply state a property that must be true across the image pair, e.g., One sliding door is closed.
Appendix D Results on NLVR
Table 11 shows previously published results using raw images in NLVR from Suhr et al. 2017 and more recent approaches. Not all previously evaluated methods report consistency. We also report results for visual reasoning systems originally developed for CLEVR. We compute human performance for each split of the data using the procedure described in Section 5; a threshold of covers of annotators. NMN Andreas et al. 2016, N2NMN, and FiLM achieve the best results for methods that were not developed using NLVR. However, both perform worse than CNN-BiATT Tan and Bansal 2018 and CMM Yao et al. 2018, which were developed originally using NLVR. Consistency for CNN-BiATT was taken from the NLVR leaderboard.
Appendix E Implementation Details
The caption’s representation is computed using an RNN encoder. We use -dimensional GloVe vectors trained on Common Crawl as word embeddings Pennington et al. 2014. We encode the caption using a single-layer long short-term memory (Hochreiter and Schmidhuber 1997, LSTM,) RNN of size . The hidden states of the caption are averaged and processed with the MLP described above to predict the truth value.
Image
The image pair’s representation is computed by extracting features from a pre-trained model. We resize and pad each image with whitespace to a size of pixels, which is the size of the image displayed to the workers during sentence-writing. Each padded image is resized to and passed through a ResNet-152 pre-trained model He et al. 2016. The features from the final layer before classification are extracted for each image and concatenated. This representation is processed with the MLP described above to predict a truth value.
E.2 Image and Text Baselines
The caption and image pair are encoded as described in Appendix E.1, then concatenated and passed through the MLP described above to predict a truth value.
MaxEnt
We use -grams where . We train a maximum entropy classifier with Megam. https://www.umiacs.umd.edu/~hal/megam
E.3 Module Networks
We use the publicly available implementation. https://github.com/ronghanghu/n2nmn The model parameters used for NLVR2 are the same as those used for the original experiments on VQA. We use GloVe vectors of size to embed words Pennington et al. 2014. The model parameters used for NLVR are the same as those used for the original N2NMN experiments on CLEVR. This includes learning word embeddings from scratch and embedding images using the pool5 layer of VGG-16 trained on ImageNet Simonyan and Zisserman 2014; Hu et al. 2017. The two paired images are resized and padded with white space to size , then concatenated horizontally and resized to a single image of pixels. The resulting image is embedded using the res5c layer of ResNet-152 trained on ImageNet He et al. 2016; Hu et al. 2017.
FiLM
We use the publicly available implementation. https://github.com/ethanjperez/film For NLVR2, we first resize and pad both images with whitespace to images of size . The two images are concatenated horizontally and resized to a single image of pixels. This image is passed through a ResNet-101 pretrained model and the features from the conv4 layer are extracted He et al. 2016; Perez et al. 2018. For NLVR, we resize images to and use the raw pixels directly. The parameters of the models are the same as described in Perez et al. Perez et al. 2018’s experiments on featurized images, except for the following: RNN hidden size of , classifier projection dimension of size , final MLP hidden size of , and feature maps. Using the original parameters did not result in significant differences in accuracy, while updates using our parameters were computed faster and the computation graph used less memory.
E.4 MAC
We use the implementation provided online. https://github.com/stanfordnlp/mac-network For experiments on NLVR2, we adapt the image processing procedure. Both images are resized and padded with white space to images of size , then concatenated horizontally and resized to pixels. We use the same image featurization approach used in Hudson and Manning 2018. For experiments on NLVR, we use the NLVR configuration provided in the repository.
E.5 Training
For the Text, Image, and CNN+RNN methods on NLVR2, we perform updates using Adam Kingma and Ba 2014 with a global learning rate of . The weights and biases are initialized by sampling uniformly from . All fully-connected and output layers use a learned bias term. For MAC, we use the same training setup as described in Hudson and Manning 2018, stopping early based on performance over the development set. For all other experiments, we use early stopping with patience, where patience is initially set to a constant and multiplied at each epoch the validation accuracy improves over a global maximum. We use of the training data as a validation set, which is not used to update model parameters. We choose a validation set such that unique sentences do not appear in both the validation and training sets. For FiLM and N2NMN, we set the initial patience to . For Text, Image and CNN+RNN baselines, initial patience was set to . For MaxEnt, we use at most epochs.
Appendix F Additional Examples
Table 12 includes additional examples sampled from the training and development sets of NLVR2, as well as license information for each image. All images in this paper were sampled from websites known for hosting non-copyrighted images, for example Wikimedia.
Appendix G Lisence Information
Tables 13, 14, 15, and 15 detail license and attribution information for the images included in the main paper.