CHAI: A CHatbot AI for Task-Oriented Dialogue with Offline Reinforcement Learning

Siddharth Verma, Justin Fu, Mengjiao Yang, Sergey Levine

Introduction

Constructing fluent and intelligent dialogue agents could pave the way for intuitive interfaces and automation of human-interactive tasks. However, this requires dialogue agents that both generate fluent, natural responses and effectively pursue the goals of the given dialogue task. A predominant approach to training dialogue agents is through supervised learning, where an agent is tasked with imitating language provided by humans. While this can provide for fluent responses, it becomes difficult to ensure that such agents systematically pursue the goals of the dialogue interaction. If we instead view dialogue as a control problem, frameworks such as reinforcement learning (RL) could allow agents to automatically optimize dialogue with respect to a task goal through a trial-and-error process and improve over human behavior.

However, implementing an RL system in practice, where an agent learns online from interacting with real humans, can be prohibitively expensive and time-consuming. This is in stark contrast to supervised learning approaches, where we can cheaply construct datasets for training imitation agents by simply logging conversations. Therefore, existing RL approaches for dialogue often rely on interacting with a learned model of a human (Li et al., 2016b; He et al., 2018), from which experience can be generated inexpensively. However, naïve training in this manner can result in the dialogue agent exploiting the model, which can degenerate into non-intelligible language. To mitigate this, algorithms must typically enforce strong priors to keep generated language similar to those seen in the dataset (Li et al., 2016b; Jaques et al., 2019), or adopt dialogue management and template-based approaches which directly re-use language seen in the dataset (He et al., 2018).

Issues such as model exploitation and distribution shift when training on static datasets are a primary concern of offline RL (Levine et al., 2020), and can provide a formalized approach to tackling these problems. While offline RL is motivated by scaling RL to large datasets, annotated datasets for dialogue are still small compared to the large amount of raw text datasets available today. Therefore, we propose an offline, model-free approach to dialogue generation that leverages language models. Because the size of unlabeled language datasets dwarfs that of curated datasets for dialogue, using a pre-trained language model as a central component of our method allows it to learn aspects of language fluency from unlabeled datasets, while learning higher-level strategies for goal-directed dialogue via RL on a smaller annotated datasets. This combined approach enables us to utilize the large amounts of existing language data that standard RL methods cannot.

The main contribution of this work is CHAI (CHatbot AI), an algorithm for learning task-oriented dialogue that utilizes a language model in conjunction with offline RL. We show that this leads the policy to generate goal-oriented dialogue that is both realistic and functional, and does not require training against a simulated model of human language. We evaluate our method on a negotiation task, which requires the model to both reason about strategic aspects of conversation along with generating fluent language. We show that CHAI consistently bargains for better prices and with higher rates of successful negotiation than prior RL approaches to goal-oriented dialogue.

Related Work

Recent developments in deep learning have led to end-to-end approaches to dialogue using supervised learning, such as sequence-to-sequence models (Dušek and Jurcicek, 2016; Eric and Manning, 2017), hierarchical models (Serban et al., 2017), attention (Mei et al., 2017; Chen et al., 2019), and Transformer-based models (Wu et al., 2021; Hosseini-Asl et al., 2020; Peng et al., 2020; Adiwardana et al., 2020). However, supervised learning only allows an agent to imitate behaviors, requires optimal data, and does not allow agents to exceed human performance. Supervised learning for dialogue generation also has well-known issues such as outputting commonplace responses (e.g., I do not know) regardless of the inputs Li et al. (2016a). Therefore, additional training of the dialogue agent is required for performing goal-oriented tasks.

Task-oriented dialogue has been formulated as a sequential decision making problem in a Markov Decision Process (MDP) since the 1990s (Smith and Hipp, 1994; Singh et al., 1999; Williams and Young, 2007; Young et al., 2013; Paek, 2006; Henderson et al., 2008; Gao et al., 2018; Pieraccini et al., 2009; Young et al., 2013; Su et al., 2015; Chen et al., 2020). Dialogue is converted into abstract states and actions from which an agent is trained using reinforcement learning (RL) (Eckert et al., 1997; Levin et al., 2000; Chung, 2004; Georgila et al., 2006; Schatzmann et al., 2007; Heeman, 2009; Georgila and Traum, 2011; Su et al., 2016; Fatemi et al., 2016; Asri et al., 2016; Zhao et al., 2019; Zhang et al., 2020; Wang et al., 2020). These methods differ in how the abstract states/actions are designed and whether the simulated environment for training the policy is hand created, learned as a fixed model, or is an agent itself. For instance, Eckert et al. (1997); Levin et al. (2000) learn a fixed transition model from human conversations and Georgila and Traum (2011) learn negotiation agents where each agent is the user simulator for the other agent. These methods also differ in how the decision making policy is trained, e.g., online (Gašić et al., 2011) or off-policy/offline (Yu et al., 2016; Pietquin et al., 2011) using actor-critic (Su et al., 2017), policy gradient (He et al., 2018), or fitted Q-iteration (Pietquin et al., 2011). Regardless of the RL method used, since policies are trained on abstract states and actions, these methods lack the ability to generate natural language (i.e., response is created via templates depending on an abstract action).

To overcome these limitations, recent work has trained policies directly on text, using a recurrent neural network to output language tokens, and using self-play for policy training while interacting with another learned agent (Li et al., 2016b; Lewis et al., 2017; Liu et al., 2018). To further improve the generated language quality, hierarchical methods decouple the strategic high-level dialogue decisions from generation (Yarats and Lewis, 2018; He et al., 2018; Saleh et al., 2020). These model-based approaches require accurate estimation of the environment/human (e.g., the trained self-play agent needs to mimic complex human behavior), which is beyond current capability of model-based reinforcement learning algorithms. Similar to our proposal, Jaques et al. (2019) use offline RL based on KL-control for text generation in open-domain dialogue. Our work differs in that our model is able to utilize large amounts of unsupervised data through the use of pre-trained language models, and that our work focuses on task-oriented (as opposed to open-domain) dialogue tasks. Goal-oriented tasks have clearly defined objectives that can be quantified, allowing us to provide an objective comparison between our method and prior approaches.

Preliminaries

In this section, we describe our evaluation task and cover the necessary background and notation.

We evaluate our approach on the CraigslistBargain task (He et al., 2018). CraigslistBargain consists of 6682 advertisements scraped from Craigslist, along with dialogues for each advertisement collected via Amazon Mechanical Turk where two users play the role of buyer/seller. An example advertisement from this dataset is shown in Fig. 1, along with a sample conversation between a human and CHAI.

During each round of interaction, the buyer and seller can execute four possible response types. A message allows one player to send an utterance to the other. An offer allows one player to propose a price at which to conduct the transaction. Once an offer is made, the other player can either accept or reject the offer, which ends the episode. A reward is then computed based on the transaction price. Our bot receives a reward equal to the normalized price the item is sold for at the end of an episode (normalized by the list price), scaled by a constant factor of 10. Additionally, we penalize the bot by a constant of -20 for episodes resulting in a reject to incentivize the agent to make deals.

We selected this task because it provides a clear objective, allowing us to illustrate our approach with quantifiable metrics. Of course, practical applications of CHAI to goal-directed dialogue may tackle other problems, including non-adversarial problms such as helping a user to answer a question or fulfill a request. However, our choice of tasks was constrained by the limited availability of public datasets for dialogue tasks that are goal-directed and have objective task goals.

2 Reinforcement Learning Setup

We formulate the task-oriented dialogue problem as an RL problem, where the agent serves the role of seller in the CraigslistBargain problem, and the environment serves the role as the buyer. We consider a Markov decision process defined as a tuple (S,A,T,r,γ)(\mathcal{S},\mathcal{A},\mathcal{T},r,\gamma). The state and action spaces, S\mathcal{S} and A\mathcal{A}, consist of three main components: the action type type (one of the four described in Sec. 3.1), an utterance uu (used only for message responses), and a normalized price price, expressed as a fraction of the list price. Additionally, the context component scontexts_{\texttt{context}} contains the advertisement listing description. The price component is used in two ways. First, because we do not wish to represent prices in the dialogue as discrete tokens, we instead replace prices in the utterance with a placeholder token understood to be substituted with the price component. Second, the price component is used in offer response type to communicate the desired transaction price. In all other cases, it is ignored. An example of the state space is shown in Fig. 2.

We write individual states as s={su,stype,sprice,scontext}s=\{s_{u},s_{\texttt{type}},s_{\texttt{price}},s_{\texttt{context}}\} and actions as a={au,atype,aprice}a=\{a_{u},a_{\texttt{type}},a_{\texttt{price}}\}, where su,aus_{u},a_{u} denote the utterances, or sequences of words generated by the environment and agent. {stype,atype}\{s_{\texttt{type}},a_{\texttt{type}}\} denote the action types, {sprice,aprice}\{s_{\texttt{price}},a_{\texttt{price}}\} denote the prices, and scontexts_{\texttt{context}} denotes the listing context. The transition distribution T(s′∣s,a)\mathcal{T}(s^{\prime}|s,a) governs the distribution over responses generated by the buyer agent (environment), and the reward R\mathcal{R} defines the task objective. The goal of RL is to find a policy π(a∣s)\pi(a|s) that maximizes the expected returns:

Offline Reinforcement Learning with Language Models

Offline RL potentially allows reinforcement learning methods to leverage large datasets for policy learning. However, it still requires datasets to be annotated with rewards, and for conversations to come from the task at hand. Because of this, annotated dialogue datasets, such as the CraigslistBargain dataset presented in Section 3.1, are many orders of magnitude smaller than unlabeled datasets collected for unsupervised and language modeling tasks. In order to utilize these large unlabeled datasets, we propose an algorithm that combines offline RL with fine-tuned language models.

Our approach begins with training a language model, such as GPT-2 (Radford et al., 2018), and fine-tuning it on our task-specific dialogue corpus (Sec. 3.1). We use LM(u∣s1:t){\textrm{LM}}(u|s_{1:t}) to denote a distribution over utterances uu produced by the language model given the dialogue history, denoted s1:ts_{1:t}. We then train a critic or Q-function as described in Sec. 4.1, which is responsible for scoring good and bad responses and is used to select responses from a pool of candidates generated from the language model. Our approach can be viewed as using a Q-function to steer a language model (which has no concept of a task) towards producing language that accomplishes some task-specific goal.

In this section we describe how to train a Q-function that can score candidate responses based on their potential to maximize returns. We implement and evaluate three different training procedures, each utilizing different offline RL methods. In the overall Q-learning framework, we sample a batch of transitions (consisting of states, actions, rewards, and successor states) from our dataset and perform updates based on minimizing a modified Bellman loss:

where target value Qtarget(s,a)Q_{\textrm{target}}(s,a) is typically computed via the Bellman operator defined as

However, using this update directly can lead to problems if we only have access to offline datasets. A widely studied issue in offline RL is the challenge of handling out-of-distribution actions: when the maximization over the action in the target value is not constrained in any way, it is easy to obtain actions for which the Q-value predictions are erroneously high (Levine et al., 2020). In dialogue, this issue is greatly exacerbated, since the Q-function is only trained on responses in the dataset, and therefore is unlikely to make accurate predictions for arbitrary strings. The following modifications to Eqn.1 and Eqn.2 address this issue.

Proposal sampling (CHAI-prop) In the proposal sampling approach, the target value Qtarget(s,a)Q_{\textrm{target}}(s,a) is computed via a modified Bellman operator that utilizes a proposal distribution based on the language model, μ(at∣s1:t)\mu(a_{t}|s_{1:t}), to generate NN response proposals, and then uses the target Q-function, Qˉ\bar{Q}, to score those responses and selects the highest one:

This sampling scheme serves a dual purpose: it both constrains the responses to be naturalistic, and it also prevents out-of-distribution inputs to the Q-function in the target value calculation. This approach resembles a number of prior offline RL methods that also employ proposal distributions (Kalashnikov et al., 2018; Kumar et al., 2019; Fujimoto et al., 2019; Wu et al., 2019). Similarly to several prior works, we use samples from a proposal distribution for the target value, without an explicit actor (Kalashnikov et al., 2018; Ghasemipour et al., 2020). Unlike these approaches, our method leverages a pretrained and finetuned language model LM, which additionally makes use of extensive unsupervised prior datasets during the pretraining stage and enables our method to handle the complex and combinatorial action space of dialogue generation. In addition, following prior work, we use a separate target network Qˉ\bar{Q} whose weights are updated to track those of Q using a soft update rule as done in prior methods (Lillicrap et al., 2016; Haarnoja et al., 2018).

The proposal distribution μ(at∣s1:t)\mu(a_{t}|s_{1:t}) represents a distribution over actions a={au,atype,aprice}a=\{a_{u},a_{\texttt{type}},a_{\texttt{price}}\}. We use the language model in order to sample utterances au∼ LM(⋅∣s1:t)a_{u}\sim~{}{\textrm{LM}}(\cdot|s_{1:t}). To make training more computationally efficient, we pre-generate a batch of 5 utterances per transition in the dataset using the language model, and resample these during training as an approximation of directly sampling from the language model. For the prices apricea_{\texttt{price}}, we uniformly sample a value between 70%70\% to 100%100\% of the previously offered price, which roughly matches the distribution of the seller’s offers in the dataset. Finally, we infer the message type based on the utterance sampled using a simple heuristic, as the CraigslistBargain task requires us to specify a type for each response. During language model fine-tuning, we replaced each offer, accept, or reject action with the utterances “offer”, “accept”, and “reject”, respectively. We then check if the language model generated any of these tokens and return the corresponding action type, and label the action as a message otherwise. This simplifies our method and allows us to use the language model to generate the action types as well as the utterances.

Conservative Q-learning (CHAI-CQL) Conservative Q-learning (CQL) (Kumar et al., 2020) proposes a complimentary approach to reducing the harmful effect of out-of-distribution Q-values by explicitly penalizing the Q-value of actions not seen in the dataset. We adapt CQL as an additional regularizer on the Q-value in addition to the proposal sampling scheme. Specifically, we use the CQL(H)\textrm{CQL}(\mathcal{H}) variant, which add an additional regularizer fCQLf^{CQL} to the Q-learning objective:

Our adaptation of CQL differs from Kumar et al. (2020) in that they propose an actor-critic method which trains an explicit actor. Rather, the remainder of our method is identical to the proposal sampling variant, specifically in regards to the computation of QtargetQ_{\textrm{target}} and language model sampling.

Behavior-regularized Q-learning (CHAI-BRAC) Behavior-regularized actor-critic (BRAC) (Wu et al., 2019) proposes an alternative method for regularizing the Q-function such that out-of-distribution Q-values are penalized. Adapting this method to the setting of dialogue with language models, we use this approach to regularize the price proposal mechanism. Rather than uniformly sampling prices as described for proposal sampling, we train an additional price proposal network πϕ(aprice∣s1:t)\pi_{\phi}(a_{\texttt{price}}|s_{1:t}) that outputs a Gaussian distribution over prices given the conversation states. Using the notation a′∼πϕ,μa^{\prime}\sim\pi_{\phi},\mu to denote sampling prices from the proposal network, utterances from the language model, and action types uniformly, the proposal network is trained according to the objective

The prior proposal network, πB\pi_{B}, is estimated as a univariate conditional Gaussian of the current offer given the previous offer, where the mean and standard deviations are linear functions of the previous offer. The target value is then computed as:

This adaption of the behavior regularized objective controls out-of-distribution queries on the Q-value by regularizing the price towards those seen within the dataset. This prevents the target Q-value from being queried in low-data regimes which can cause inaccuracies during training.

2 Dialogue Generation

Once the Q-function has been trained, dialogue generation from our model is a three-phase process. The first phase is sampling: given a response from the buyer, we query the language model to sample 5 candidate utterances, and sample 5 candidate prices. The next step is scoring: we then take the cross-product of these sets, and score each potential action with the Q-function. Finally, in the selection phase, several methods are considered in order to select an action. A straightforward method is to return the action that had the highest Q-value. However, we found that this approach resulted in behavior with low diversity. Instead, we opted to follow the approach in soft Q-learning (Haarnoja et al., 2018) and sample actions from a softmax distribution over the Q-values, p(a∣s)∝exp⁡{Q(s,a)}p(a|s)\propto\exp\{Q(s,a)\}, which increases diversity in the responses as sub-optimal actions are occasionally sampled. This decoding process is depicted in Fig. 3.

3 Architecture Details

For our language model, we use an off-the-shelf implementation of GPT2-medium (Radford et al., 2018). This language model is finetuned on a transcript of each scenario in the dataset containing the context (title, description) and spoken dialogue. The prices in the dialogue are masked out with a special price token allowing us use GPT to generate templates which we can substitute prices into. The input to the language model is a concatenation of the scenario context and the dialogue history.

The Q-function is parameterized as a feedforward network that maps states and actions into a single scalar representing the Q-value. To process the utterances into the state and action, we separately compute state and action embeddings by taking the average of the masked GPT2 attention embeddings of the entire dialogue history up to the current utterance. These embeddings are then concatenated with the prices (represented as a fraction of the list price) and message types (represented as a one-hot vector) to produce a single vector that is given to the critic as input. The critic is parametrized using a 2-layer feedforward network with hidden sizes of 256 and ReLU nonlinearities. Additional details about our model architectures, language model sampling method, and how inputs to the Q-function are structured, can be found in Appendix A.1.

Experiments

Our experimental evaluation aims to compare our proposed goal-directed offline RL dialogue method to both prior dialogue management approaches and language modeling baselines. We conduct two studies: an objective evaluation against other dialogue agents to measure each method’s performance in negotiation, and a subjective human study to measure the overall end-to-end performance of the system in a similar manner to prior work (He et al., 2018; Jaques et al., 2019). Qualitative results showing actual dialogue generated from our method can be found in Appendix A.4. We consider 4 baseline approaches. The first is the current state-of-the-art approach for the CraigslistBargain task proposed by He et al. (2018) (referred to as the retrieval-based baseline). This is a hierarchical approach to dialogue generation that separately handles language generation and dialogue management. This method parses utterances into coarse “dialogue acts,” which represent high-level categorizations of the utterance such as greetings, offers, or counter-offers. An RL agent is then trained against a learned model of the environment to select “dialogue acts”, and a retrieval-based generator is then used to convert “dialogue acts” back into text. In contrast, our method directly generates text, and does not require any manually designed categorizations of natural dialogue into dialogue acts. Since our method utilizes a modern language model, we also include a pure language modeling baseline, which consists of the same GPT-2 language model Radford et al. (2018) finetuned on the CraigslistBargain dataset using the same method as done in CHAI. This baseline allows us to determine whether any improvement from our method is due to the language model, or to the use of offline RL. Finally, we also evaluate the end-to-end approaches described by Lewis et al. (2017), which include a dialogue agent trained via supervised learning, and an RL agent optimized for the task objective.

To evaluate the effectiveness of our offline RL goal-directed dialogue system, we first conducted a systematic study including the 3 variations of CHAI outlined in Sec. 4: the proposal sampling method (CHAI-prop), CQL method (CHAI-CQL), and behavior regularized method (CHAI-BRAC). In order to ensure that the results are not overfit to a single strategy, we run each method against a suite of 5 evaluation buyer agents, based on the retrieval agents presented by He et al. (2018). We choose these agents because they have been evaluated by humans as being human-like and have the strongest performance on the CraigslistBargain benchmark task. Specifically, we use the rule-based and RL agents proposed by He et al. (2018) (trained using “utility”, “fairness”, and conversation “length” as rewards). To introduce additional variety in negotiation styles, we additionally modify the rule-based agent to offer 25% of the difference between offers rather than splitting the difference between offers, which we refer to as the “Stingy” rule-based agent. We record the percentage of negotiations that result in an accept and the average normalized revenue generated per negotiation, which totals the average sale price (rejections have zero revenue) normalized by the listing price of the advertisement. Our results are presented in Table 1.

Overall, we find that among the variations of CHAI, the conservative Q-learning variant performs the best by a small margin, but results are very comparable between all 3 variations. This suggests that the particular choice of offline regularizer is far less important than the CHAI framework of utilizing a pre-trained language model with Q-function scoring. On average, CHAI-CQL performs significantly higher on acceptance rate and similarly on revenue to the next best agent, the retrieval agent (He et al., 2018) using conversation length as reward. Computing statistical significance between these two methods, we find that p<1.96∗10−9p<1.96*10^{-9} using a chi-squared test for acceptance rate, indicating that the difference in acceptance rates is statistically significant. We find that p<0.946p<0.946 using a t-test for revenue, indicating that the difference in revenue is not significant. We also note that the performance of CHAI has significantly less variation across evaluations against different buyer agents than the retrieval-based agents. For example, the retrieval agent with utility reward scores near-zero on 3 evaluations but scores near-perfectly on the other two. This suggests that the CHAI framework produces dialogue agents that are more consistent and less susceptible to exploitation. We also note that implementing the retrieval method (He et al., 2018) requires hand-designing high-level dialogue actions, topic categories, and rules for parsing or labeling these components. These designs are specifically tailored to the CraigslistBargain task, whereas such hand-engineering for CHAI does not exist outside of the interface requirements to the task itself. Thus, CHAI has significantly weaker assumptions, generates dialogue end-to-end via RL, and yet is able to narrowly outperform prior methods. Among prior methods with similar assumptions to CHAI (the language modeling baseline and (Lewis et al., 2017)), CHAI outperforms by a wide margin on both acceptance rates and revenue.

We ran an additional ablation study on the choice of reward in Appendix A.3. We find that this has a significant effect on performance, and we based our reward design on balancing between maximizing acceptance rate (through a rejection penalty) and revenue (through the utility function).

2 Human User Study

To evaluate the effectiveness and naturalness of our offline RL goal-directed dialogue system, we conducted a user study with 16 individuals, who were each asked to carry out 2 negotiations with each of three agents. The users were then asked to rate the conversation on fluency, coherency, on-topicness, and human-likeness on a 5-point Likert scale. Fluency specifically refers to the frequency of grammatical and word-choice errors. Coherency measures whether the agent’s responses are coherent. On-topic measures how well the agent was aligned with performing the task at hand. Finally, human-likeness measures how similar the agent’s responses were to a human. Because of the cost of human evaluations, we were limited in our ability to evaluate as many baselines. Therefore, we chose methods that were the most directly comparable - the simplest variation of CHAI (CHAI-prop) optimized for utility, evaluated against the utility-optimized agent from He et al. (2018), and a language model baseline that shares the same finetuning procedure as CHAI. Additional details of the user study, including the questions posed to the users and statistical significance tests, are included in Appendix A.3.1.

Results are shown in Table 2. We ran a one-way repeated measures ANOVA test, and found that the type of agent used leads to statistically significant rating differences for all metrics (with at least p<0.01p<0.01). CHAI outperforms both baselines on almost all metrics, except for fluency, where both CHAI and the language model perform similarly. The fact that fluency is similar between the two models makes sense, since both methods use a GPT-2 model to generate utterances. However, the langauge modeling baseline lacks an understanding of the task goal, and therefore makes unreasonable offers or responds in illogical ways (see Appendix A.4 for examples). It therefore scores lower on other metrics as compared to CHAI. This result suggests that the ability of language models to execute goal-directed dialogues is limited by a lack of awareness of the task objectives, and that offline RL potentially addresses this issue, producing dialogue that is perceived as more coherent, task-oriented, and human-like.

In Fig. 4 we show a comparative example between CHAI, the retrieval-based agent (He et al., 2018), and the language modeling baseline on the same scenario with human responses. CHAI and the language modeling baseline tend to produce more specific responses to the prompt due to the use of a language model, rather than generating text via the usage of templates. For example, when presented with a questioned about utilities and viewing time, both methods are able to answer the question, whereas the retrieval-based agent gives a non-sequitur answer. However, the language modeling baseline struggles in understanding prices, and offers \2200whenthebuyerrequestedwhen the buyer requested\30003000. CHAI is able to demonstrate understanding both language and the flow of the negotiation by offering reasonable counter-offers to the user, such as responding to a low offer with “I can’t go that low” and offering a higher counter-offer that the Buyer accepts. Additional examples from our human evaluation can be found in Appendix A.4.

Discussion and Future Work

We presented a system for goal-directed dialogue based on combining offline RL with finetuned language models. CHAI learns with RL, but does so from offline datasets of human dialogue. The language model allows CHAI to benefit from large-scale unsupervised pre-training, and the offline RL component enables CHAI to select responses that are more likely to lead to a successful task outcome. Quantitatively, CHAI achieves higher acceptance rate at higher revenue than prior dialogue management systems designed for this task.

Goal-oriented dialogue agents have many potentially useful applications such as building personal assistants, improving accessibility to technology for the disabled or the elderly, and simply saving time by automating menial tasks. Of course, as with any natural language generation technology, this kind of method can be used both beneficially and maliciously, for example by users who aim to create intentionally deceptive and realistic agents.

While CHAI provides a proof-of-concept that offline RL can successfully learn complex human-interactive tasks such as dialogue, it also has limitations. The goal of RL is to maximize reward, which can lead to unintended responses – for example, without additional objective terms, there is no reason for CHAI to be truthful. Similar issues affect language models more broadly, though we anticipate that it would be easier to address such issues in RL by employing better reward design. Although reward design can itself be a difficult problem, it does provide a more direct lever for influencing the agent’s behavior than what is available in standard language models, which must be directed either through the choice of training data or other indirect mechanisms. In CHAI, we only investigated a single task due to architectural constraints, and the exact same architecture presented in this paper (e.g. a price prediction head) may not be directly transferable to other domains. However, using a value-based selection mechanism is more generally applicable to any goal-oriented task.

An exciting direction for future work is to extend offline RL to address a wider range of human-interactive tasks, particularly tasks with longer-range dependencies and delayed rewards, where complex task goals can lead to the emergence of dialogue that is ultimately more useful to human users.

Acknowledgements

We thank Natasha Jacques, Daniel Fried, Dilek Hakkani-tur, Yang Liu, Alexandros Papangelis, and Wei Wei for insightful discussions. We would also like to thank all of the anonymous participants in the user study. This research was supported by the Office of Naval Research.

References

Appendix A Appendix

For our language model, we use an off-the-shelf implementation of GPT2-medium (Radford et al., 2018). This language model is finetuned on a transcript of each scenario in the dataset containing the context (title, description) and spoken dialogue. The prices in the dialogue are masked out with a special price token allowing us use GPT to generate templates which we can substitute prices into. To produce language samples, we concatenate the context of the scenario along with the dialogue history and feed it to the language model. The language model then generates the next utterance. This process is repeated in order to get multiple samples.

The Q-function is a feedforward network that maps the states and actions into a single scalar representing the Q-value. To process the utterances into the state and action, we separately compute state and action embeddings by taking the average of the masked GPT2 attention embeddings of the entire dialogue history up to the current utterance. State and action embeddings are then concatenated with the prices (represented as a percentage of the list price) and message types (represented as a one-hot vector) to produce a single vector that is given to the critic as input. The critic is parametrized as a 2-layer feedforward network with hidden sizes of 256 and ReLU nonlinearities.

The equations below describe the inputs of the Q-function. “STRCAT” is a custom string formatting function that concatenates the context and utterances in the dialog and prefixes each utterance with the string “Buyer:” or “Seller:”. “EMBED” calculates the masked attention embeddings from GPT2. dialog represents the dialog history up to the current state, [su,1,au,1,…,au,t−1,su,t][s_{u,1},a_{u,1},\dots,a_{u,t-1},s_{u,t}], and candidate denotes a candidate utterance au,ta_{u,t} generated by the language model that is being considered for scoring.

A.2 Experiment Details

Hyperparameter Selection. For our Q-learning algorithm, we used default hyperparameters from an SAC implementation and did not vary the parameters. We used a critic learning rate of 3∗10−43*10^{-4}, and a soft target update rate of 0.050.05. For the language model architecture, we finetuned 2 GPT models architectures (GPT2-small, and GPT2-medium). In order to select which model to use, we manually rated samples from the language generated in their quality across both model types and checkpoints, and selected the best performing model. We selected the GPT2-medium architecture at training epoch 2000.

Compute Resources. We finetuned our language models on TPUs (TPU v3-8, 16GB memory per core with 8 cores) within a GCP instance. We trained our Q-learning models on an internal compute cluster using an Nvidia 1080 GPU (12 GB memory).

A.3 Reward Ablation Study

Because we are limited in the number of evaluations possible in a human study, we use a simulated evaluation against another chatbot to run an ablation study measuring the effect of using different reward functions. Specifically, we instantiate a rule-based dialogue manager proposed by (He et al., 2018) as the “buyer” and negotiate with it on randomly sampled scenarios from the dataset. This is done for both our method and the baselines, and the results are tabulated in Table 3.

We evaluated 5 variations of CHAI. CHAI(final) is the method used in our paper, which uses 2 components to the reward: positive reward for the price the item was sold at, and a penalty for the episode ending in a rejection. The specific reward used was 10 * the price sold (normalized by the list price) if the offer was accepted, or a penalty of -20 if the offer was rejected. CHAI(penalty) uses the same reward, except with an increased rejection penalty. CHAI(accept) is given a positive reward of +20 for episodes ending in an accept and negative reward of -20 for episodes ending in a rejection, without regard to the price. CHAI(utility) is rewarded solely for the price an item is sold at at 10 * the price sold. Finally, CHAI(fair) is rewarded for negotiating to a midpoint price between the buyer and seller’s target prices.

We report the acceptance rate, average revenue, and average offers made and offers accepted. We see that CHAI(accept), CHAI(fair) and CHAI(penalty) achieve higher acceptance rates, but offer lower prices on average. In contrast, CHAI(utility) and CHAI(final) offer higher prices with lower acceptance rates, with the pure utility optimizing agent CHAI(utility) offering the highest prices with the lowest acceptance rates. Thus, we can see how changing the reward function can significantly affect the behavior of the resulting agent.

Setup. We conducted our user study through a web-interface, where an advertisement from the test set of CraigslistBargain is displayed to the user. This is shown in Fig. 5. Users were instructed to “type any message to speak with the bot and negotiate”, and a chatbot agent replied to each message. After the user is satisfied with the negotiation, they are asked to indicate whether they want to accept or reject the current offer on the item.

After interacting with the chatbot agent, users were given a survey and asked to rate the bot on a 5-point Likert scale: strongly agree (5), agree (4), neutral (3), disagree (2), strongly disagree (1). The ratings were:

The bot was fluent (did not make grammatical or word choice errors).

The flow of the conversation was coherent.

The bot demonstrated human-like behavior.

These questions correspond to the fluency, coherency, on-topicness, and human-likeness scores reported in our paper, respectively.

Manipulated factors. We evaluate 3 chatbot agents. The first is CHAI. The second is a retrieval-based baseline (RLutility(act)\textrm{RL}_{\textrm{utility}}(\textrm{act}) as proposed in (He et al., 2018)). The this is a language modeling baseline based on finetuning GPT2-medium on the CraigslistBargain dataset.

Dependent measures. We measure the performance of each chatbot agent according to 4 subjective metric (fluency, coherency, on-topicness, and human-likeness), which correspond to the survey questions given to participants described above. Each metric is rated on a 5-point Likert scale. We also measure the price that was agreed upon, and whether the negotiation resulted in an acceptance or a rejection.

Risks The risks presented to participants were minimal, and participants were never placed in the way of physical harm. There was a small probability the chatbot could generate offensive or otherwise inconsiderate langauge, as the language generation for CHAI and the language model baseline were unconstrained. However, we did not observe this behavior prior to the study, and did not observe this behavior during the study. We minimized the risk of confidentiality breaches by anonymizing all data stored as user IDs – participant names were not stored on our servers.

Subject allocation. We recruited 10 male and 6 female participants, with an average age of 24. Participation was voluntary, and subjects were not compensated for their participation. Participants were asked for consent before participating, and consent for including examples from their interactions for this paper. All examples contained in this paper are anonymized and contain no personally identifiable information.

Prior to participation in the study, each user was provided with instructions for how to use the interface, and a practice evaluation against the retrieval agent to familiarize them with the user interface and rating system. Each user interacted with all 3 dialogue agents (CHAI, retrieval, and language model) twice, with the order of interaction randomized per user. The advertisement displayed to the user is also uniformly sampled from the test set for each trial, independently from the agent being used.

Analysis. We ran a one-way repeated measures ANOVA test for each metric reported (fluency, coherency, on-topic, human-like) to examine the effect of chatbot agent on the metric.

Results showed that the type of agent used lead to statistically significant differences in the ratings. We found that:

For coherency, f(2,30)=16.9518,p<0.0001f(2,30)=16.9518,p<0.0001

For on-topicness, f(2,30)=10.1840,p<0.0004f(2,30)=10.1840,p<0.0004

For human-likeness, f(2,30)=20.1592,p<0.0001f(2,30)=20.1592,p<0.0001

A.4 Additional Qualitative Results

In this section, we include additional examples collected from human dialog for CHAI-prop, the retrieval method of (He et al., 2018), and the pure language modeling baseline.

A.4.2 Retrieval

A.4.3 Language Model