Are Neural Open-Domain Dialog Systems Robust to Speech Recognition Errors in the Dialog History? An Empirical Study
Karthik Gopalakrishnan, Behnam Hedayatnia, Longshaokan Wang, Yang Liu, Dilek Hakkani-Tur
Introduction
Neural modeling approaches are prominent in research on both task-oriented and open-domain dialog. Traditional sequence-to-sequence models have been used for encoding the dialog history and predicting domains, intents, slot types, spans and more generally decoding full-fledged system responses. In recent years, large pre-trained Transformer-based models for natural language understanding (NLU) and natural language generation (NLG) have become ubiquitous , leading to tremendous advances by fine-tuning towards these dialog tasks .
In task-oriented speech-based dialog systems, the effect of ASR hypotheses has been widely studied and techniques have been devised to minimize the resulting downstream NLU errors . More recently, end-to-end spoken language understanding approaches have been attempted to sidestep this problem . On the other hand, research in open-domain dialog is increasingly focusing on large, monolithic end-to-end neural models like Google’s Meena that are built using written data and evaluated on written interactions. Several written textual datasets have been created recently to tackle various problems in open-domain dialog, including persona-grounding, knowledge-grounding and reasoning, and state-of-the-art chatbots have been built using them. However, it is not clear whether these written text-based open-domain chatbots would seamlessly integrate with ASR models to serve the speech modality, which is popular due to the ubiquity of voice assistants like Alexa, Google Assistant and Siri.
Collecting large-scale written text-based dialog datasets is cheaper and more practical than collecting audio-based dialog datasets in many ways. But speech-robustness should be a factor of consideration when designing any (task-oriented or open-domain) dialog system intended to be deployed to the speech modality, even in the absence of audio-based training data. To bring attention to this important aspect in the open-domain dialog community, we empirically study the effects of various types of synthetic and actual ASR hypotheses in the dialog history on TransferTransfo (TF2) , a state-of-the-art neural open-domain dialog system based on the Generative Pre-trained Transformer (GPT) from the NeurIPS ConvAI2 Conversational Intelligence Challenge . We build off the Topical-Chat dataset and perform two augmentations in our study: one creating simulated ASR hypotheses for the entire dataset, and another creating actual ASR hypotheses with a smaller audio-based analogue of the Topical-Chat test sets.
We observe that TF2 trained on written textual data is very sensitive to synthetic and actual ASR hypotheses introduced to the dialog history during inference time, with the sensitivity being particularly prominent for the task of response selection. As a baseline mitigation strategy, we introduce synthetic ASR hypotheses to the dialog history during training and observe marginal improvements, demonstrating the need for further research into techniques to make end-to-end open-domain chatbots fully speech-robust. Figure 1 shows a sample snippet with responses from TF2 models trained on written text and synthetic ASR hypotheses when fed speech-distorted dialog history.
A close work to ours in spirit is , which shows that Transformer-based generative dialog models are insensitive to unrealistic perturbations like token-shuffling of the dialog history. Our work is more focused on evaluating the effects of introducing realistic perturbations to the dialog history in the form of synthetic and actual ASR hypotheses. Our augmentation of Topical-Chat, dubbed the Topical-Chat ASR dataset, is open-sourced https://github.com/alexa/Topical-Chat/tree/master/TopicalChatASR/ to enable open-domain dialog researchers to perform speech-robustness evaluation and fuel research into novel techniques to make monolithic neural open-domain dialog models more speech-robust.
Preliminaries
Let denote a dialog history containing a sequence of turns. Let denote a flattened sequence of all tokens in . Let be a truncate parameter for a flattened dialog history , which retains at most tokens from the end in . When applied to , we denote it as . Finally, we denote a dialog example pair by , where is a candidate response that is either the ground-truth response at turn or a distractor response.
2 Data
For our experiments, we use the dialogs from Topical-Chat . This is one of the largest and most diverse knowledge-grounded text-based open-domain dialog datasets publicly available today. Each dialog in Topical-Chat contains 20+ turns alternating between two Turkers and 19 tokens per turn on average. Topical-Chat has two types of test sets: test freq and test rare.
We performed two augmentations of Topical-Chat. First, a simulated augmentation wherein simulated errors are introduced at a corpus-level target word error rate (WER). For each data split (train/valid/test), we simulated ASR hypotheses for each turn in a dialog with four WER settings (0.1, 0.15, 0.2 and 0.3), each with a single seed for train and five random seeds for valid and test. We used the ASR error simulator method based on -gram confusion matrix and trained the simulator on transcribed ASR output from an internal user study.
Second, an actual augmentation wherein two new test sets were created: test freq audio and test rare audio, which are smaller speech-based analogues of the original test sets. 40 dialogs corresponding to corpus-level distinct entity triplets were picked from each of the original test sets, and English-speaking human subjects of various ethnicities were asked to verbally read the dialogs with their own audio setup and record their audio, resulting in phonetically rich test sets. Two automated transcription systems (A and B) by Amazon were independently used to transcribe the collected audio, and each dialog transcription was aligned with the text of the original dialog based on edit distance followed by manual re-alignment to obtain the turn-level transcriptions. The Word Error Rate (WER) for the dialogs in test freq audio and test rare audio were in the range 0.1-0.39 by system A and in the range 0.08-0.35 by system B. For all experiments, we use the transcripts by system A.
Since our focus is on studying the effect of synthetic and actual ASR hypotheses for user utterances, we need to label the partners in a dialog example as user and system. For a dialog history where turn is the target turn, we treat the Turker corresponding to turn as the system and the Turker corresponding to turn as the user. This leads to system/user label assignments for all turns in . Figure 2 shows this process.
3 Models
TransferTransfo (TF2) is a state-of-the-art neural open-domain dialog system by Hugging Face based on the Generative Pre-trained Transformer (GPT) model that won 1st place in automated evaluation and 2nd place in human evaluation at the NeurIPS ConvAI2 Conversational Intelligence Challenge .
In this system, GPT is fine-tuned on a dialog dataset in a multi-task learning fashion with the language modeling and next utterance classification tasks. Starter code for TransferTransfo along with the pre-trained model and BPE-based vocabulary is provided on GitHub by Hugging Face.
Language modeling: Trained with pairs, this task is for learning to generate a response given a dialog history.
Next utterance classification: This task is for learning to select the ground-truth response from a set of candidate responses given a dialog history and is trained with pairs, . The set of candidate responses contains distractors alongside the ground-truth. In our experiments, we used and randomly sampled turns from other dialogs in the same corpus (train/valid/test) as the one containing to serve as distractors.
To remove confounding factors that may affect our study of the effect of synthetic and actual ASR hypotheses in the dialog history , we avoid conditioning TF2 on knowledge from the reading sets provided to Turkers in Topical-Chat.
4 Training/Inference
Like , we use special tokens to identify speakers with their associated segments and initialize them with random embeddings to be learned during fine-tuning. We fine-tuned for a fixed number of 3 epochs with equal weight of to the losses for both tasks. We wanted to fine-tune with a large dialog history and set . We used a train batch size of 2, performed gradient accumulation for 8 steps and gradient clipping with a max norm of , used the Adam optimizer and linearly decayed the learning rate from 6.25e-5 to 0 during the course of training. For simplicity, we set during inference time, i.e., we evaluate TF2’s ability to generate and select given the last user turn. We used top-, top- nucleus sampling with temperature for decoding, where , and . We also set a maximum decode length of 40 tokens.
Experimental Setup
Unlike written text (which we denote by GOLD), raw ASR hypotheses typically do not contain punctuation and casing. Since TF2 performs lowercasing during tokenization, we exclude casing from consideration, and evaluate our models on the following three variations of the dialog history by using synthetic and actual ASR hypotheses:
NO-PUNC: No punctuation in user turns in .
WER-SIM: Simulated errors are introduced to user turns in at a corpus-level target WER.
REAL: Actual ASR errors are introduced to user turns in . This case is applicable for evaluation on test freq audio and test rare audio only.
We train these TF2 models as described in Section 2.4:
TF2-GOLD: Trained on written text dialog examples.
TF2-NO-PUNC: Trained on a NO-PUNC version of written text dialog examples.
TF2-WER-SIM: For each WER, TF2 is trained on a WER-SIM version of written text dialog examples.
For the language modeling task, we use the standard automated metrics of perplexity (PPL) and unigram F1 between the generated and ground-truth response. For the next utterance classification task, we use recall with and , which measures the accuracy of selecting the ground-truth response from a set of candidate responses.
How Robust is TF2-GOLD?
To answer this question, we evaluate TF2-GOLD in various settings and empirically compare them against the GOLD setting.
Table 1 shows the results when evaluating TF2-GOLD in the NO-PUNC and GOLD settings on test freq and test rare. In order to better understand the effect of sentence segmentation, we report results separately for three versions of each test set:
single: a subset of dialog examples where contains just a single sentence
multi: a subset of dialog examples where contains more than one sentence
all: single multi, i.e., all dialog examples
We observe a clear degradation in all metrics in the NO-PUNC setting relative to the GOLD setting, showing that TF2-GOLD relies on punctuation in the dialog history during both generation and selection. The degradation in PPL and when evaluating with the multi version of the test sets is greater than when evaluating with the single version, showing that the presence of sentence segmentation in dialog history comprised of multi-sentence utterances has a large impact on performance.
2 WER-SIM
Table 2 shows the results when evaluating TF2-GOLD in the WER-SIM and GOLD settings on all dialog examples of test freq and test rare. For each WER, we compute metrics using ASR hypotheses corresponding to all five seeds separately and report mean and standard deviation. We observe a significant degradation in all metrics in the WER-SIM setting relative to the GOLD setting, showing that TF2-GOLD is very sensitive to simulated ASR hypotheses in the dialog history. We also observe a significant degradation in metrics as the WER increases. The magnitude of degradation in shows that response selection relies heavily on the written form.
3 REAL
Table 3 shows the results when evaluating TF2-GOLD in the REAL and GOLD settings on all dialog examples of test freq audio and test rare audio. We observe a significant degradation in all metrics in the REAL setting relative to the GOLD setting, demonstrating that TF2-GOLD is very sensitive to actual ASR hypotheses in the dialog history.
Does Synthetic Training Help?
We now study the efficacy of synthetic training in making TF2 robust to both synthetic and actual ASR hypotheses.
We evaluate TF2-NO-PUNC in the NO-PUNC setting. Results for the NO-PUNC setting are in Table 4 and can be interpreted jointly with Table 1. We observe that TF2-NO-PUNC is more effective than TF2-GOLD at handling the NO-PUNC setting during inference time. There is a larger improvement in PPL and when evaluating with the multi version of the test sets than the single version, showing that NO-PUNC training is especially useful when the dialog history is comprised of multi-sentence utterances.
2 WER-SIM
We evaluate TF2-WER-SIM in the WER-SIM setting. Results for the WER-SIM setting are in Table 5 and can be interpreted jointly with Table 2. We observe that TF2-WER-SIM is more effective than TF2-GOLD at handling the WER-SIM setting during inference time.
3 REAL: Automated Evaluation
We evaluate TF2-NO-PUNC and TF2-WER-SIM in the REAL setting. Results are in Table 6 and can be interpreted jointly with Table 3. We observe that TF2-NO-PUNC and TF2-WER-SIM are generally more effective than TF2-GOLD at handling the REAL setting during inference time.
Our evaluation shows that synthetic training provides reasonable improvements when the nature of errors matches during training and inference and very marginal improvements otherwise. An analysis of our simulated and actual ASR augmented data showed there are more insertion and deletion errors in the simulated data, and only about 10% of substitution errors in actual appear in simulated ASR augmented data. But regardless of the existence of training-inference match/mismatch, the perfomance gap with TF2-GOLD in the GOLD setting (particularly prominent in ) demonstrates that there is still a need for creative error-specific and/or error-agnostic techniques to make monolithic neural open-domain dialog models demonstrably robust to non-trivial target errors.
4 REAL: Human Evaluation
We also performed a human evaluation of the language modeling task from the two tasks in TF2, specifically generating responses from TF2-GOLD, TF2-NO-PUNC and TF2-WER-SIM in the REAL setting on the Topical-Chat audio test sets. 100 unique snippets were prepared from each test set, where each snippet contained a written-text dialog history and responses from all four models. But, the responses were obtained by feeding the models the error-distorted version of the dialog history. For each test set, a pair of annotators were asked to annotate snippets 1-60 and 40-100 respectively, thus providing 20 overlapping annotations to compute inter-annotator agreement. The annotators were asked to annotate on a 5-point nominal scale (1: Strongly Disagree, 2: Disagree, 3: Neither Agree Nor Disagree, 4: Agree, 5: Strongly Agree) whether each response was appropriate (AP) for the provided written-text dialog history.
The goal of this human evaluation was to see if there’s any perceptible difference in responses generated from TF2-GOLD and the top 3 synthetic models with the least PPL in the REAL setting from Section 5.3 when fed error-distorted dialog history.
We bucketed ratings 1-2 and 4-5 when computing Fleiss’ kappa annotator agreement and got scores of 0.54 and 0.38 on the two test sets. We observe in Table 7 that TF2-NO-PUNC is marginally better than TF2-GOLD for both test sets and TF2-WER-SIM is marginally better than TF2-GOLD on one test set.
Conclusion
We empirically studied the effects of synthetic and actual ASR hypotheses in the dialog history on TF2, a large state-of-the-art text-based neural open-domain dialog system from the NeurIPS ConvAI2 challenge. We observed that TF2 trained on written data is very sensitive to such hypotheses introduced to the dialog history during inference time, demonstrating that text-based neural open-domain chatbots may not be very effective at serving the speech modality as-is. We observed that training TF2 with synthetic ASR hypotheses makes it more robust to both synthetic and actual ASR hypotheses during inference time, with considerable room for improvement left for future work. Our augmentation of Topical-Chat, dubbed the Topical-Chat ASR dataset, is open-sourced and we hope our work sparks discussion and further research into modality-centric and modality-agnostic open-domain dialog systems.