Learning in the Rational Speech Acts Model

Will Monroe, Christopher Potts

1 Pragmatic language use

In the Gricean view of language use , people are rational agents who are able to communicate efficiently and effectively by reasoning in terms of shared communicative goals, the costs of production, prior expectations, and others’ belief states. The Rational Speech Acts (RSA) model is a recent Bayesian reconstruction of these core Gricean ideas. RSA and its extensions have been shown to capture many kinds of conversational implicature and to closely model psycholinguistic data from children and adults .

Both Grice’s theories and RSA have been criticized for predicting that people are more rational than they actually are. These criticisms have been especially forceful in the context of language production. It seems that speakers often fall short: their utterances are longer than they need to be, underinformative, unintentionally ambiguous, obscure, and so forth . RSA can incorporate notions of bounded rationality , but it still sharply contrasts with views in the tradition of , in which speaker agents rely on heuristics and shortcuts to try to accurately describe the world while managing the cognitive demands of language production.

In this paper, we offer a substantially different perspective on RSA by showing how to define it as a trained statistical classifier, which we call learned RSA. At the heart of learned RSA is the back-and-forth reasoning between speakers and listeners that characterizes RSA. However, whereas standard RSA requires a hand-built lexicon, learned RSA infers a lexicon from data. And whereas standard RSA makes predictions according to a fixed calculation, learned RSA seeks to optimize the likelihood of whatever examples it is trained on. Agents trained in this way exhibit the pragmatic behavior characteristic of RSA, but their behavior is governed by their training data and hence is only as rational as that experience supports. To the extent that the speakers who produced the data are pragmatic, learned RSA discovers that; to the extent that their behavior is governed by other factors, learned RSA picks up on that too. We validate the model on the task of attribute selection for referring expression generation with a widely-used corpus of referential descriptions (the TUNA corpus; ), showing that it improves on heuristic-driven models and pure RSA by synthesizing the best aspects of both.

2 RSA as a speaker model

RSA is a descendent of the signaling systems of and draws on ideas from iterated best response (IBR) models , iterated cautious response (ICR) models , and cognitive hierarchies (see also ). RSA models language use as a recursive process in which speakers and listeners reason about each other to enrich the literal semantics of their language. This increases the efficiency and reliability of their communication compared to what more purely literal agents can achieve.

For instance, suppose a speaker and listener are playing a reference game in the context of the images in Figure 1(a). The speaker SS has been privately assigned referent r1r_{1} and must send a message that conveys this to the listener. A literal speaker would make a random choice between beard and glasses. However, if SS places itself in the role of a listener LL receiving these messages, then SS will see that glasses creates uncertainty about the referent whereas beard does not, and so SS will favor beard. In short, the pragmatic speaker chooses beard because it’s unambiguous for the listener.

RSA formalizes this reasoning in probabilistic Bayesian terms. It assumes a set of messages MM, a set of states TT, a prior probability distribution PP over states TT, and a cost function CC mapping messages to real numbers. The semantics of messages is defined by a lexicon L\mathcal{L}, where L(m,t)=1\mathcal{L}(m,t)=1 if mm is true of tt and otherwise. The agents are then defined as follows:

The model that is the starting point for our contribution in this paper is the pragmatic speaker s1s_{1}. It reasons not about the semantics directly but rather about a pragmatic listener l1l_{1} reasoning about a literal speaker s0s_{0}. The strength of this pragmatic reasoning is partly governed by the temperature parameter λ\lambda, with higher values leading to more aggressive pragmatic reasoning.

Figure 1 tracks the RSA computations for the reference game in Figure 1(a). Here, the message costs CC are all , the prior over referents is flat, and λ=1\lambda=1. The chances of success for the literal speaker s0s_{0} are low, since it chooses true messages at random. In contrast, the chances of success for s1s_{1} are high, since it derives the unambiguous system highlighted in gray.

The task we seek to model is a language generation task, so we present RSA from a speaker-centric perspective. It has been explored more fully from a listener perspective. In that formulation, the model begins with a literal listener reasoning only in terms of the lexicon L\mathcal{L} and state priors. Models of this general form have been shown to capture a wide range of pragmatic behaviors and to increase success in task-oriented dialogues .

RSA has been criticized on the grounds that it predicts unrealistic speaker behavior . For instance, in Figure 1, we confined our agents to a simple message space. If permitted to use natural language, they will often produce utterances expressing predicates that are redundant from an RSA perspective—for example, by describing r1r_{1} as the man with the long beard and sweater, even though man has no power to discriminate, and beard and sweater each uniquely identify the intended referent. This tendency has several explanations, including a preference for including certain kinds of descriptors, a desire to hedge against the possibility that the listener is not pragmatic, and cognitive pressures that make optimal descriptions impossible. One of our central objectives is to allow these factors to guide the core RSA calculation.

3 The TUNA corpus

In Section 0.6, we evaluate RSA and learned RSA in the TUNA corpus , a widely used resource for developing and testing models of natural language generation. We introduce the corpus now because doing so helps clarify the learning task faced by our model, which we define in the next section.

In the TUNA corpus, participants were assigned a target referent or referents in the context of seven other distractors and asked to describe the target(s). Trials were performed in two domains, furniture and people, each with a singular condition (describe a single entity) and a plural condition (describe two). Figure 2 provides a (slightly simplified) example from the singular furniture section, with the target item identified by shading. In this case, the participant wrote the message “blue fan small”. All entities and messages are annotated with their semantic attributes, as given in simplified form here. (Participants saw just the images; we include the attributes in Figure 2 for reference.)

The task we address is attribute selection: reproducing the multiset of attributes in the message produced in each context. Thus, for Figure 2, we would aim to produce {\{[[size:small]], [[colour:blue]], [[type:fan]]}\}. This is less demanding than full natural language generation, since it factors out all morphosyntactic phenomena. Section 0.6 provides additional details on the nature of this evaluation.

4 Learned RSA

We now formulate RSA as a machine learning model that can incorporate the quirks and limitations that characterize natural descriptions while still presenting a unified model of pragmatic reasoning. This approach builds on the two-layer speaker-centric classifier of , but differs from theirs in that we directly optimize the performance of the pragmatic speaker in training, whereas apply a recursive reasoning model on top of a pre-trained classifier. Like RSA, the model can be generalized to allow for additional intermediate agents, and it can easily be reformulated to begin with a literal listener.

To build an agent that learns effectively from data, we must represent the items in our dataset in a way that accurately captures their important distinguishing properties and permits robust generalization to new items . We define our feature representation function ϕ\phi very generally as a map from state–utterance–context triples <t,m,c>\left<t,m,c\right> to vectors of real numbers. This gives us the freedom to design the feature function to encode as much relevant information as necessary.

As noted above, in learned RSA, we do not presuppose a semantic lexicon, but rather induce one from the data as part of learning. The feature representation function determines a large, messy hypothesis space of potential lexica that is refined during optimization. For instance, as a starting point, we might define the feature space in terms of the cross-product of all possible entity attributes and all possible utterance meaning attributes. For mm entity attributes and nn utterance attributes, this defines each ϕ(t,m,c)\phi(t,m,c) as an mnmn-dimensional vector. Each dimension of this vector records the number of times that its corresponding pair of attributes co-occurs in tt and mm. Thus, the representation of the target entity in Figure 2 would include a 11 in the dimension for clearly good pairs like colour:blue ∧\wedge [[colour:blue]] as well as for intuitively incorrect pairs like size:small ∧\wedge [[colour:blue]].

Because ϕ\phi is defined very generally, we can also include information that is not clearly lexical. For instance, in our experiments, we add dimensions that count the color attributes in the utterance in various ways, ignoring the specific color values. We can also define features that intuitively involve negation, for instance, those that capture entity attributes that go unmentioned. This freedom is crucial to bringing generation-specific insights into the RSA reasoning.

Literal speaker.

Learned RSA is built on top of a log-linear model, standard in the machine learning literature and widely applied to classification tasks .

This model serves as our literal speaker, analogous to s0s_{0} in (1). The lexicon of this model is embedded in the parameters (or weights) θ\theta. Intuitively, θ\theta is the direction in feature representation space that the literal speaker believes is most positively correlated with the probability that the message will be correct. We train the model by searching for a θ\theta to maximize the conditional likelihood the model assigns to the messages in the training examples. Assuming the training is effective, this increases the weight for correct pairings between utterance attributes and entity attributes and decreases the weight for incorrect pairings.

To find the optimal θ\theta, we seek to maximize the conditional likelihood of the training examples using first-order optimization methods (described in more detail in Learning, below). This requires the gradient of the likelihood with respect to θ\theta. To simplify the gradient derivation and improve numerical stability, we maximize the log of the conditional likelihood:

where the first two equations can be derived by expanding the proportionality constant in the definition of S0S_{0}.

Pragmatic speaker.

We now define a pragmatic listener L1L_{1} and a pragmatic speaker S1S_{1}. We will show experimentally (Section 0.6) that the learned pragmatic speaker S1S_{1} agrees better with human speakers on a referential expression generation task than either the literal speaker S0S_{0} or the pure RSA speaker s1s_{1}.

The parameters for L1L_{1} and S1S_{1} are still the parameters of the literal speaker S0S_{0}; we wish to update them to maximize the performance of S1S_{1}, the agent that acts according to S1(m∣t,c;θ)S_{1}(m\mid t,c;\theta), where

This corresponds to the simplest case of RSA in which λ=1\lambda=1 and message costs and state priors are uniform: s1(m∣t,L)∝l1(t∣m,L)∝s0(m∣t,L)s_{1}(m\mid t,\mathcal{L})\propto l_{1}(t\mid m,\mathcal{L})\propto s_{0}(m\mid t,\mathcal{L}).

In optimizing the performance of the pragmatic speaker S1S_{1} by adjusting the parameters to the simpler classifier S0S_{0}, the RSA back-and-forth reasoning can be thought of as a non-linear function through which errors are propagated in training, similar to the activation functions in neural network models . However, unlike neural network activation functions, the RSA reasoning applies a different non-linear transformation depending on the pragmatic context (sets of available referents and utterances).

For convenience, we define symbols for the log-likelihood of each of these probability distributions:

The log-likelihood of each agent has the same form as the log-likelihood of the literal speaker, but with the value of the distribution from the lower-level agent substituted for the score θTϕ\theta^{T}\phi. By a derivation similar to the one in (6) above, the gradient of these log-likelihoods can thus be shown to have the same form as the gradient of the literal speaker, but with the gradient of the next lower agent substituted for the feature values:

The value JS0J_{S_{0}} in (12) is as defined in (5).

Training.

The stochastic gradient descent (SGD) family of first-order optimization techniques can be used to approximately maximize J(θ)J(\theta) by obtaining noisy estimates of its gradient and “hill-climbing” in the direction of the estimates. (Strictly speaking, we are employing stochastic gradient ascent to maximize the objective rather than minimize it; however, SGD is the much more commonly seen term for the technique.)

The exact gradient of this objective function is

The learning rate α\alpha determines how “aggressively” the parameters are adjusted in the direction of the gradient. Small values of α\alpha lead to slower learning, but a value of α\alpha that is too large can result in the parameters overshooting the optimal value and diverging. To find a good learning rate, we use AdaGrad , which sets the learning rate adaptively for each example based on an initial step size η\eta and gradient history. The effect of AdaGrad is to reduce the learning rate over time such that the parameters can settle down to a local optimum despite the noisy gradient estimates, while continuing to allow high-magnitude updates along certain dimensions if those dimensions have exhibited less noisy behavior in previous updates.

5 Example

In Figure 3, we illustrate crucial aspects of how our model is optimized, fleshing out the concepts from the previous section. The example also shows the ability of the trained S1S_{1} model to make a specificity implicature without having observed one in its data, while preserving the ability to produce uninformative attributes if encouraged to do so by experience.

As in our main experiments, we frame the learning task in terms of attribute selection with TUNA-like data. In this toy experiment, the agent is trained on two example contexts, consisting of a target referent, a distractor referent, and a human-produced utterance. It is evaluated on a third test example. This small dataset is given in the top two rows of Figure 3. The utterance on the test example is shown for comparison; it is not provided to the agent.

Our feature representations of the data are in the third row. Attributes of the referents are in small caps; semantic attributes of the utterances are in [[square brackets]]. These representations employ the cross-product features described in Section 0.4; in TUNA data, properties that the target entities do not possess (e.g., ¬\negglasses) are also included among their “attributes.”

The RSA reasoning yields gradients that express both lexical and contextual knowledge. From the first training example, the model learns the lexical information that [[person]] and [[glasses]] should be used to describe the target. However, this knowledge receives higher weight in the association with glasses, because that attribute is disambiguating in this context. As one would hope, the overall result is that intuitively good pairings generally have higher weights, though the training set is too small to fully distinguish good features from bad ones. For example, after seeing both training examples and failing to observe both a beard and glasses on the same individual, the model incorrectly infers that [[beard]] can be used to indicate a lack of glasses and vice versa. Additional training examples could easily correct this.

The distributions in Figure 3(b) show that the linear classifier correctly learns that human-produced utterances in the training data tend to mention the attribute [[person]] even though it is uninformative. However, for the referent that was not seen in the training data, the model cannot decide among mentioning [[beard]], [[glasses]], both, or neither, even though the messages that don’t mention [[glasses]] are ambiguous in context. The pure RSA model, meanwhile, chooses messages that are unambiguous, but because it has no mechanism for learning from the examples, it does not prefer to produce [[person]] without a manually-specified prior.

Our pragmatic speaker S1S_{1} gives us the best of both models: the parameters θ\theta in learned RSA show the tendency exhibited in the training data to produce [[person]] in all cases, while the RSA recursive reasoning mechanism guides the model to produce unambiguous messages by including the attribute [[glasses]].

6 Experiments

We report experiments on the TUNA corpus (Section 0.3 above). We focus on the singular portion of the corpus, which was used in the 2008 and 2009 Referring Expression Generation Challenges. We do not have access to the train/dev/test splits from those challenges, so we report five-fold cross-validation numbers. The singular portion consists of 420 furniture trials involving 176 distinct referents and 360 people trials involving 228 distinct referents.

Evaluation metrics.

The primary evaluation metric used in the attribute selection task with TUNA data is multiset Dice calculated on the attributes of the generated messages:

Experimental set-up.

We evaluate all our agents in the same pragmatic contexts: for each trial in the singular corpus, we define the messages MM to be the powerset of the attributes used in the referential description and the states TT to be the set of entities in the trial, including the target. The message predicted by a speaker agent is the one with the highest probability given the target entity; if more than one message has the highest probability, we allow the agent to choose randomly from the highest probability ones.

Features.

We use indicator features as our feature representation; that is, the dimensions of the feature representation take the values 0 and 1, with 1 representing the truth of some predicate P(t,m,c)P(t,m,c) and 0 representing its negation. Thus, each vector of real numbers that is the value of ϕ(t,m,c)\phi(t,m,c) can be represented compactly as a set of predicates.

The baseline feature set consists of indicator features over all conjunctions of an attribute of the referent and an attribute in the candidate message (e.g., P(t,m,c)=\textscred(t)∧P(t,m,c)=\textsc{red}(t)\wedge{}[blue]∈m{}\in m). We compare this to a version of the model with additional generation features that seek to capture the preferences identified in prior work on generation. These consist of indicators over the following features of the message:

attribute type (e.g., P(t,m,c)=P(t,m,c)= “mm contains a color”);

pair-wise attribute type co-occurrences, where one can be negated (e.g., “mm contains a color and a size”, “mm contains an object type but not a color”); and

message size in number of attributes (e.g., “mm consists of 3 attributes”).

For comparison, we also separately train literal speakers S0S_{0} as in (4) (the log-linear model) with each of these feature sets using the same optimization procedure.

Results.

The results (Table 1) show that training a speaker agent with learned RSA generally improves generation over the ordinary classifier and RSA models. On the more complex people dataset, the pragmatic S1S_{1} model significantly outperforms all other models. The value of the model’s flexibility in allowing a variety of feature designs can be seen in the comparison of the different feature sets: we observe consistent gains from adding generation features to the basic cross-product feature set. Moreover, the two types of features complement each other: neither the cross-product features nor the generation features in isolation achieve the same performance as the combination of the two.

Of the models in Table 1, all but the last exhibit systematic errors. Pure RSA performs poorly for reasons predicted by —for example, it under-produces color terms and head nouns like desk, chair, and person. This problem is also observed in the trained S1S_{1} model, but is corrected by the generation features. On the people dataset, the S0S_{0} models under-produce beard and hair, which are highly informative in certain contexts. This type of communicative failure is eliminated in the S1S_{1} speakers.

The performance of the learned RSA model on the people trials also compares favorably to the best dev set performance numbers from the 2008 Challenge , namely, .762 multiset Dice, although this comparison must be informal since the test sets are different. (In particular, the Accuracy values given in are unfortunately not comparable with the values we present, as they reflect “perfect match with at least one of the two reference outputs” [emphasis in original].) Together, these results show the value of being able to train a single model that synthesizes RSA with prior work on generation.

7 Conclusion

Our initial experiments demonstrate the utility of RSA as a trained classifier in generating referential expressions. The primary advantages of this version of RSA stem from the flexible ways in which it can learn from available data. This not only removes the need to specify a complex semantic lexicon by hand, but it also provides the analytic freedom to create models that are sensitive to factors guiding natural language production that are not naturally expressed in standard RSA.

This basic presentation suggests a range of potential next steps. For instance, it would be natural to apply the model to pragmatic interpretation (the listener’s perspective); this requires no substantive formal changes to the model as defined in Section 0.4, and it opens up new avenues in terms of evaluating pragmatic models in standard classification tasks like sentiment analysis, topic prediction, and natural language reasoning. In addition, for all versions of the model, one could consider including additional hidden speaker and listener layers, incorporating message costs and priors into learning, to capture a wider range of pragmatic phenomena.

References