Reasoning About Pragmatics with Neural Listeners and Speakers

Jacob Andreas, Dan Klein

Introduction

We present a model for describing scenes and objects by reasoning about context and listener behavior. By incorporating standard neural modules for image retrieval and language modeling into a probabilistic framework for pragmatics, our model generates rich, contextually appropriate descriptions of structured world representations.

This paper focuses on a reference game RG played between a listener LL and a speaker SS.

Figure 1 shows an example drawn from a standard captioning dataset [Zitnick et al., 2014].

In order for the players to win, SS’s description dd must be pragmatic: it must be informative, fluent, concise, and must ultimately encode an understanding of LL’s behavior. In Figure 1, for example, the owl is wearing a hat and the owl is sitting in the tree are both accurate descriptions of the target image, but only the second allows a human listener to succeed with high probability. RG is the focus of many papers in the computational pragmatics literature: it provides a concrete generation task while eliciting a broad range of pragmatic behaviors, including conversational implicature [Benotti and Traum, 2009] and context dependence [Smith et al., 2013]. Existing computational models of pragmatics can be divided into two broad lines of work, which we term the direct and derived approaches.

Direct models (see Section 2 for examples) are based on a representation of SS. They learn pragmatic behavior by example. Beginning with datasets annotated for the specific task they are trying to solve (e.g. examples of humans playing RG), direct models use feature-based architectures to predict appropriate behavior without a listener representation. While quite general in principle, such models require training data annotated specifically with pragmatics in mind; such data is scarce in practice.

Derived models, by contrast, are based on a representation of LL. They first instantiate a base listener L0\textrm{L}0 (intended to simulate a naïve, non-pragmatic listener). They then form a reasoning speaker S1\textrm{S}1, which chooses a description that causes L0\textrm{L}0 to behave correctly. Existing derived models couple hand-written grammars and hand-engineered listener models with sophisticated inference procedures. They exhibit complex behavior, but are restricted to small domains where grammar engineering is practical.

The approach we present in this paper aims to capture the best aspects of both lines of work. Like direct approaches, we use machine learning to acquire a complete grounded generation model from data, without domain knowledge in the form of a hand-written grammar or hand-engineered listener model. But like derived approaches, we use this learning to construct a base model, and embed it within a higher-order model that reasons about listener responses. As will be seen, this reasoning step allows the model to make use of weaker supervision than previous data-driven approaches, while exhibiting robust behavior in a variety of contexts.

Our goal is to build a derived model that scales to real-world datasets without domain engineering. Independent of the application to RG, our model also belongs to the family of neural image captioning models that have been a popular subject of recent study [Xu et al., 2015]. Nevertheless, our approach appears to be:

the first such captioning model to reason explicitly about listeners

the first learned approach to pragmatics that requires only non-pragmatic training data

Following previous work, we evaluate our model on RG, though the general architecture could be applied to other tasks where pragmatics plays a core role. Using a large dataset of abstract scenes like the one shown in Figure 1, we run a series of games with humans in the role of LL and our system in the role of SS. We find that the descriptions generated by our model result in correct interpretation 17% more often than a recent learned baseline system. We use these experiments to explore various other aspects of computational pragmatics, including tradeoffs between adequacy and fluency, and between computational efficiency and expressive power.Models, human annotations, and code to generate all tables and figures in this paper can be found at http://github.com/jacobandreas/pragma.

Related Work

As an example of the direct approach mentioned in the introduction, ?) collect a set of human-generated referring expressions about abstract representations of sets of colored blocks. Given a set of blocks to describe, their model directly learns a maximum-entropy distribution over the set of logical expressions whose denotation is the target set. Other research, focused on referring expression generation from a computer vision perspective, includes that of ?) and ?).

Derived pragmatics

Derived approaches, sometimes referred to as “rational speech acts” models, include those of ?), ?), ?), and ?). These couple template-driven language generation with probabilistic or game-theoretic reasoning frameworks to produce contextually appropriate language: intelligent listeners reason about the behavior of reflexive speakers, and even higher-order speakers reason about these listeners. Experiments [Frank et al., 2009] show that derived approaches explain human behavior well, but both computational and representational issues restrict their application to simple reference games. They require domain-specific engineering, controlled world representations, and pragmatically annotated training data.

An extensive literature on computational pragmatics considers its application to tasks other than RG, including instruction following [Anderson et al., 1991] and discourse analysis [Jurafsky et al., 1997].

Representing language and the world

In addition to the pragmatics literature, the approach proposed in this paper relies extensively on recently developed tools for multimodal processing of language and unstructured representations like images. These includes both image retrieval models, which select an image from a collection given a textual description [Socher et al., 2014], and neural conditional language models, which take a content representation and emit a string [Donahue et al., 2015].

Approach

Our goal is to produce a model that can play the role of the speaker SS in RG. Specifically, given a target referent (e.g. scene or object) rr and a distractor r′r^{\prime}, the model must produce a description dd that uniquely identifies rr. For training, we have access to a set of non-contrastively captioned referents {(ri,di)}\{(r_{i},d_{i})\}: each training description did_{i} is generated for its associated referent rir_{i} in isolation. There is no guarantee that did_{i} would actually serve as a good referring expression for rir_{i} in any particular context. We must thus use the training data to ground language in referent representations, but rely on reasoning to produce pragmatics.

Our model architecture is compositional and hierarchical. We begin in Section 3.2 by describing a collection of “modules”: basic computational primitives for mapping between referents, descriptions, and reference judgments, here implemented as linear operators or small neural networks. While these modules appear as substructures in neural architectures for a variety of tasks, we put them to novel use in constructing a reasoning pragmatic speaker.

Section 3.3 describes how to assemble two base models: a literal speaker, which maps from referents to strings, and a literal listener, which maps from strings to reference judgments. Section 3.4 describes how these base models are used to implement a top-level reasoning speaker: a learned, probabilistic, derived model of pragmatics.

Formally, we take a description dd to consist of a sequence of words d1,d2,…,dnd_{1},d_{2},\ldots,d_{n}, drawn from a vocabulary of known size. For encoding, we also assume access to a feature representation f(d)f(d) of the sentence (for purposes of this paper, a vector of indicator features on nn-grams). These two views—as a sequence of words did_{i} and a feature vector f(d)f(d)—form the basis of module interactions with language.

Referent representations are similarly simple. Because the model never generates referents—only conditions on them and scores them—a vector-valued feature representation of referents suffices. Our approach is completely indifferent to the nature of this representation. While the experiments in this paper use a vector of indicator features on objects and actions present in abstract scenes (Figure 1), it would be easy to instead use pre-trained convolutional representations for referring to natural images. As with descriptions, we denote this feature representation f(r)f(r) for referents.

2 Modules

All listener and speaker models are built from a kit of simple building blocks for working with multimodal representations of images and text:

These are depicted in Figure 2, and specified more formally below. All modules are parameterized by weight matrices, written with capital letters W1W_{1}, W2W_{2}, etc.; we refer to the collection of weights for all modules together as WW.

The referent and description encoders produce a linear embedding of referents and descriptions in a common vector space.

Choice ranker

The choice ranker takes a string encoding and a collection of referent encodings, assigns a score to each (string, referent) pair, and then transforms these scores into a distribution over referents. We write R(ei∣e−i,ed)R(e_{i}|e_{-i},e_{d}) for the probability of choosing ii in contrast to the alternative; for example, R(e2∣e1,ed)R(e_{2}|e_{1},e_{d}) is the probability of answering “2” when presented with encodings e1e_{1} and e2e_{2}.

(Here ρ\rho is a rectified linear activation function.)

Referent describer

The referent describer takes an image encoding and outputs a description using a (feedforward) conditional neural language model. We express this model as a distribution p(dn+1∣dn,d<n,er)p(d_{n+1}|d_{n},d_{<n},e_{r}), where dnd_{n} is an indicator feature on the last description word generated, d<nd_{<n} is a vector of indicator features on all other words previously generated, and ere_{r} is a referent embedding. This is a “2-plus-skip-gram” model, with local positional history features, global position-independent history features, and features on the referent being described. To implement this probability distribution, we first use a multilayer perceptron to compute a vector of scores ss (one sis_{i} for each vocabulary item): s=W6ρ(W7[dn,d<n,ei])s=W_{6}\rho(W_{7}[d_{n},d_{<n},e_{i}]). We then normalize these to obtain probabilities: pi=esi/∑jesjp_{i}=e^{s_{i}}/\sum_{j}e^{s_{j}}. Finally, p(dn+1∣dn,d<n,er)=pdn+1p(d_{n+1}|d_{n},d_{<n},e_{r})=p_{d_{n+1}}.

3 Base models

From these building blocks, we construct a pair of base models. The first of these is a literal listener L0\textrm{L}0, which takes a description and a set of referents, and chooses the referent most likely to be described. This serves the same purpose as the base listener in the general derived approach described in the introduction. We additionally construct a literal speaker S0\textrm{S}0, which takes a referent in isolation and outputs a description. The literal speaker is used for efficient inference over the space of possible descriptions, as described in Section 3.4. L0\textrm{L}0 is, in essence, a retrieval model, and S0\textrm{S}0 is neural captioning model.

Both of the base models are probabilistic: L0\textrm{L}0 produces a distribution over referent choices, and S0\textrm{S}0 produces a distribution over strings. They are depicted with shaded backgrounds in Figure 3.

Given a description dd and a pair of candidate referents r1r_{1} and r2r_{2}, the literal listener embeds both referents and passes them to the ranking module, producing a distribution over choices ii.

That is, pL0(1∣d,r1,r2)=R(e1∣e2,ed)p_{\textrm{L}0}(1|d,r_{1},r_{2})=R(e_{1}|e_{2},e_{d}) and vice-versa. This model is trained contrastively, by solving the following optimization problem:

Here r′r^{\prime} is a random distractor chosen uniformly from the training set. For each training example (ri,di)(r_{i},d_{i}), this objective attempts to maximize the probability that the model chooses rir_{i} as the referent of did_{i} over a random distractor.

This contrastive objective ensures that our approach is applicable even when there is not a naturally-occurring source of target–distractor pairs, as previous work [Golland et al., 2010, Monroe and Potts, 2015] has required. Note that this can also be viewed as a version of the loss described by ?), where it approximates a likelihood objective that encourages L0\textrm{L}0 to prefer rir_{i} to every other possible referent simultaneously.

Literal speaker

As in the figure, the literal speaker is obtained by composing a referent encoder with a describer, as follows:

As with the listener, the literal speaker should be understood as producing a distribution over strings. It is trained by maximizing the conditional likelihood of captions in the training data:

These base models are intended to be the minimal learned equivalents of the hand-engineered speakers and hand-written grammars employed in previous derived approaches [Golland et al., 2010]. The neural encoding/decoding framework implemented by the modules in the previous subsection provides a simple way to map from referents to descriptions and descriptions to judgments without worrying too much about the details of syntax or semantics. Past work amply demonstrates that neural conditional language models are powerful enough to generate fluent and accurate (though not necessarily pragmatic) descriptions of images or structured representations [Donahue et al., 2015].

4 Reasoning model

As described in the introduction, the general derived approach to pragmatics constructs a base listener and then selects a description that makes it behave correctly. Since the assumption that listeners will behave deterministically is often a poor one, it is common for such derived approaches to implement probabilistic base listeners, and maximize the probability of correct behavior.

The neural literal listener L0\textrm{L}0 described in the preceding section is such a probabilistic listener. Given a target ii and a pair of candidate referents r1r_{1} and r2r_{2}, it is natural to specify the behavior of a reasoning speaker as simply:

At a first glance, the only thing necessary to implement this model is the representation of the literal listener itself. When the set of possible utterances comes from a fixed vocabulary [Vogel et al., 2013] or a grammar small enough to exhaustively enumerate [Smith et al., 2013] the operation max⁡d\max_{d} in Equation 7 is practical.

For our purposes, however, we would like the model to be capable of producing arbitrary utterances. Because the score pL0p_{\textrm{L}0} is produced by a discriminative listener model, and does not factor along the words of the description, there is no dynamic program that enables efficient inference over the space of all strings.

We instead use a sampling-based optimization procedure. The key ingredient here is a good proposal distribution from which to sample sentences likely to be assigned high weight by the model listener. For this we turn to the literal speaker S0\textrm{S}0 described in the previous section. Recall that this speaker produces a distribution over plausible descriptions of isolated images, while ignoring pragmatic context. We can use it as a source of candidate descriptions, to be reweighted according to the expected behavior of L0\textrm{L}0. The full specification of a sampling neural reasoning speaker is as follows:

Draw samples d1,…dn∼pS0(⋅∣ri)d_{1},\dots d_{n}\sim p_{\textrm{S}0}(\cdot|r_{i}).

Score samples: pk=pL0(i∣dk,r1,r2)p_{k}=p_{\textrm{L}0}(i|d_{k},r_{1},r_{2}).

While primarily to enable efficient inference, we can also use the literal speaker to serve a different purpose: “regularizing” model behavior towards choices that are adequate and fluent, rather than exploiting strange model behavior. Past work has restricted the set of utterances in a way that guarantees fluency. But with an imperfect learned listener model, and a procedure that optimizes this listener’s judgments directly, the speaker model might accidentally discover the kinds of pathological optima that neural classification models are known to exhibit [Goodfellow et al., 2014]—in this case, sentences that cause exactly the right response from L0\textrm{L}0, but no longer bear any resemblance to human language use. To correct this, we allow the model to consider two questions: as before, “how likely is it that a listener would interpret this sentence correctly?”, but additionally “how likely is it that a speaker would produce it?”

Formally, we introduce a parameter λ\lambda that trades off between L0\textrm{L}0 and S0\textrm{S}0, and take the reasoning model score in step 2 above to be:

This can be viewed as a weighted joint probability that a sentence is both uttered by the literal speaker and correctly interpreted by the literal listener, or alternatively in terms of Grice’s conversational maxims [Grice, 1970]: L0\textrm{L}0 encodes the maxims of quality and relation, ensuring that the description contains enough information for LL to make the right choice, while S0\textrm{S}0 encodes the maxim of manner, ensuring that the description conforms with patterns of human language use. Responsibility for the maxim of quantity is shared: L0\textrm{L}0 ensures that the model doesn’t say too little, and S0\textrm{S}0 ensures that the model doesn’t say too much.

Evaluation

We evaluate our model on the reference game RG described in the introduction. In particular, we construct instances of RG using the Abstract Scenes Dataset introduced by ?). Example scenes are shown in Figure 1 and Figure 4. The dataset contains pictures constructed by humans and described in natural language. Scene representations are available both as rendered images and as feature representations containing the identity and location of each object; as noted in Section 3.1, we use this feature set to produce our referent representation f(r)f(r). This dataset was previously used for a variety of language and vision tasks (e.g. ?), ?)). It consists of 10,020 scenes, each annotated with up to 6 captions.

The abstract scenes dataset provides a more challenging version of RG than anything we are aware of in the existing computational pragmatics literature, which has largely used the tuna corpus of isolated object descriptions [Gatt et al., 2007] or small synthetic datasets [Smith et al., 2013]. By contrast, the abstract scenes data was generated by humans looking at complex images with numerous objects, and features grammatical errors, misspellings, and a vocabulary an order of magnitude larger than tuna. Unlike previous work, we have no prespecified in-domain grammar, and no direct supervision of the relationship between scene features and lexemes.

We perform a human evaluation using Amazon Mechanical Turk. We begin by holding out a development set and a test set; each held-out set contains 1000 scenes and their accompanying descriptions. For each held-out set, we construct two sets of 200 paired (target, distractor) scenes: All, with up to four differences between paired scenes, and Hard, with exactly one difference between paired scenes. (We take the number of differences between scenes to be the number of objects that appear in one scene but not the other.)

We report two evaluation metrics. Fluency is determined by showing human raters isolated sentences, and asking them to rate linguistic quality on a scale from 1–5. Accuracy is success rate at RG: as in Figure 1, humans are shown two images and a model-generated description, and asked to select the image matching the description.

In the remainder of this section, we measure the tradeoff between fluency and accuracy that results from different mixtures of the base models (Section 4.1), measure the number of samples needed to obtain good performance from the reasoning listener (Section 4.2), and attempt to approximate the reasoning listener with a monolithic “compiled” listener (Section 4.3). In Section 4.4 we report final accuracies for our approach and baselines.

To measure the performance of the base models, we draw 10 samples djkd_{jk} for a subset of 100 pairs (r1,j,r2,j)(r_{1,j},r_{2,j}) in the Dev-All set. We collect human fluency and accuracy judgments for each of the 1000 total samples. This allows us to conduct a post-hoc search over values of λ\lambda: for a range of λ\lambda, we compute the average accuracy and fluency of the highest scoring sample. By varying λ\lambda, we can view the tradeoff between accuracy and fluency that results from interpolating between the listener and speaker model—setting λ=0\lambda=0 gives samples from pL0p_{\textrm{L}0}, and λ=1\lambda=1 gives samples from pS0p_{\textrm{S}0}.

Figure 5 shows the resulting accuracy and fluency for various values of λ\lambda. It can be seen that relying entirely on the listener gives the highest accuracy but degraded fluency. However, by adding only a very small weight to the speaker model, it is possible to achieve near-perfect fluency without a substantial decrease in accuracy. Example sentences for an individual reference game are shown in Figure 5; increasing λ\lambda causes captions to become more generic. For the remaining experiments in this paper, we take λ=0.02\lambda=0.02, finding that this gives excellent performance on both metrics.

On the development set, λ=0.02\lambda=0.02 results in an average fluency of 4.8 (compared to 4.8 for the literal speaker λ=1\lambda=1). This high fluency can be confirmed by inspection of model samples (Figure 4). We thus focus on accuracy or the remainder of the evaluation.

2 How many samples are needed?

Next we turn to the computational efficiency of the reasoning model. As in all sampling-based inference, the number of samples that must be drawn from the proposal is of critical interest—if too many samples are needed, the model will be too slow to use in practice. Having fixed λ=0.02\lambda=0.02 in the preceding section, we measure accuracy for versions of the reasoning model that draw 1, 10, 100, and 1000 samples. Results are shown in Table 1. We find that gains continue up to 100 samples.

3 Is reasoning necessary?

Because they do not require complicated inference procedures, direct approaches to pragmatics typically enjoy better computational efficiency than derived ones. Having built an accurate derived speaker, can we bootstrap a more efficient direct speaker?

To explore this, we constructed a “compiled” speaker model as follows: Given reference candidates r1r_{1} and r2r_{2} and target tt, this model produces embeddings e1e_{1} and e2e_{2}, concatenates them together into a “contrast embedding” [et,e−t][e_{t},e_{-t}], and then feeds this whole embedding into a string decoder module. Like S0\textrm{S}0, this model generates captions without the need for discriminative rescoring; unlike S0\textrm{S}0, the contrast embedding means this model can in principle learn to produce pragmatic captions, if given access to pragmatic training data. Since no such training data exists, we train the compiled model on captions sampled from the reasoning speaker itself.

This model is evaluated in Table 3. While the distribution of scores is quite different from that of the base model (it improves noticeably over S0\textrm{S}0 on scenes with 2–3 differences), the overall gain is negligible (the difference in mean scores is not significant). The compiled model significantly underperforms the reasoning model. These results suggest either that the reasoning procedure is not easily approximated by a shallow neural network, or that example descriptions of randomly-sampled training pairs (which are usually easy to discriminate) do not provide a strong enough signal for a reflex learner to recover pragmatic behavior.

4 Final evaluation

Based on the following sections, we keep λ=0.02\lambda=0.02 and use 100 samples to generate predictions. We evaluate on the test set, comparing this Reasoning model S1\textrm{S}1 to two baselines: Literal, an image captioning model trained normally on the abstract scene captions (corresponding to our L0\textrm{L}0), and Contrastive, a model trained with a soft contrastive objective, and previously used for visual referring expression generation [Mao et al., 2015].

Results are shown in Table 2. Our reasoning model outperforms both the literal baseline and previous work by a substantial margin, achieving an improvement of 17% on all pairs set and 15% on hard pairs. For comparison, a model with hand-engineered pragmatic behavior—trained using a feature representation with indicators on only those objects that appear in the target image but not the distractor—produces an accuracy of 78% and 69% on all and hard development pairs respectively. In addition to performing slightly worse than our reasoning model, this alternative approach relies on the structure of scene representations and cannot be applied to more general pragmatics tasks. Figures 4 and 6 show various representative descriptions from the model.

Conclusion

We have presented an approach for learning to generate pragmatic descriptions about general referents, even without training data collected in a pragmatic context. Our approach is built from a pair of simple neural base models, a listener and a speaker, and a high-level model that reasons about their outputs in order to produce pragmatic descriptions. In an evaluation on a standard referring expression game, our model’s descriptions produced correct behavior in human listeners significantly more often than existing baselines.

It is generally true of existing derived approaches to pragmatics that much of the system’s behavior requires hand-engineering, and generally true of direct approaches (and neural networks in particular) that training is only possible when supervision is available for the precise target task. By synthesizing these two approaches, we address both problems, obtaining pragmatic behavior without domain knowledge and without targeted training data. We believe that this general strategy of using reasoning to obtain novel contextual behavior from neural decoding models might be more broadly applied.

References