Aligning Language Models to Explicitly Handle Ambiguity
Hyuhng Joon Kim, Youna Kim, Cheonbok Park, Junyeob Kim, Choonghyun Park, Kang Min Yoo, Sang-goo Lee, Taeuk Kim
Introduction
Large Language Models (LLMs) (Ouyang et al., 2022; Team et al., 2023; Achiam et al., 2023) have demonstrated remarkable capabilities in text generation, proving particularly effective for question-answering (QA) tasks (Zhang et al., 2023; Etezadi and Shamsfard, 2023). QA systems in the wild are frequently confronted with unexpected inputs from users, such as unanswerable (Kim et al., 2023b; Yin et al., 2023) or ambiguous questions (Cole et al., 2023; Lee et al., 2023; Kim et al., 2023a). To build a reliable user-friendly model, it is essential for the model to robustly handle such inputs. In this work, we seek to extend the scope of research to effectively handle invalid inputs. Specifically, we focus on managing "ambiguity" (Gleason, 1963; Mackay and Bever, 1967), which poses a significant challenge in Natural Language Processing (NLP) (Jurafsky, 1996).
Ambiguity refers to cases where an expression conveys multiple denotations (Wasow et al., 2005). Users may pose queries with clear intentions that, possibly due to insufficient domain knowledge, result in ambiguous requests. If the model arbitrarily responds to such ambiguity, there is a risk of misinterpreting the user’s original intent, potentially harming the model’s reliability. This is especially pronounced in sensitive domains such as legal (Schane, 2002; Choi, 2024) or medical (Stevenson and Guo, 2010; Gyori et al., 2022) domains, where misinterpretations can lead to serious drawbacks. Despite the importance, approaches to robustly manage ambiguity are still significantly unexplored. In this paper, we endeavor to utilize the model’s intrinsic knowledge to align the model in a manner that effectively handles ambiguity.
Properly processing ambiguous inputs is challenging primarily due to the following two hurdles. Firstly, models are not directly trained to explicitly express ambiguity. Even if a model perceives ambiguity, it is challenging to verify the recognition without explicit feedback. The second challenge is that the degree of ambiguity for the query can vary depending on the intrinsic knowledge of the model. Consider the scenario depicted in Figure 1. The initial query is ambiguous as the phrase "national championship" poses various denotations, such as "national tennis championship" or "national golf championship". If a model possesses comprehensive knowledge across the possible denotations, it is plausible for the model to recognize the ambiguity (left). However, if the model’s knowledge is limited to "national tennis championship", it would perceive the query as unambiguous (right). Therefore, it is essential to verify whether the input is deemed ambiguous from the model’s point of view.
To overcome these issues, this paper proposes a method to align models to explicitly handle ambiguous queries. Specifically, we design a proxy task that guides the model to self-disambiguate a given query by utilizing its intrinsic knowledge. Then, we quantify the information gain from the disambiguation as an implicit measure of the extent to which the models perceive their inputs as ambiguous. This measure serves as a cue for selecting samples deemed ambiguous from the model’s perspective, which are then utilized for alignment. Experimental results from several QA datasets demonstrate that the alignment process enables the model to properly clarify ambiguous inputs while maintaining its inherent capabilities. The findings underscore the value of assessing the perceived ambiguity, rather than relying solely on the ground-truth ambiguity. Furthermore, to provide a comprehensive framework for assessing ambiguity, we construct a new dataset dubbed AmbigTriviaQA. The dataset facilitates a more extensive evaluation of models’ robustness in addressing ambiguity, thus contributing to the further expansion of related research.
Related Work
An expression is defined as ambiguous if it has two or more distinct denotations (Wasow et al., 2005). Ambiguity challenges NLP applications by obscuring the intended meaning of expressions, leading to difficulties in accurately performing specific tasks. Efforts addressing this issue span across various domains, including machine translation (Pilault et al., 2023), coreference resolution (Poesio and Artstein, 2005; Yuan et al., 2023), and natural language inference (Liu et al., 2023).
The challenge intensifies in the scope of QA as ambiguous questions may yield various answers, potentially not aligning with the user’s initial intent. Min et al. (2020) introduce the AmbigQA dataset to tackle ambiguity in open-domain QA and Stelmakh et al. (2022) expands it to long-form generation. Furthermore, Cole et al. (2023) discovered that quantifying sampling repetition presents a reliable uncertainty measure for ambiguity, while Kim et al. (2023a) generates tree-of-clarification (ToC) that refines ambiguity within the inputs. As we share the goal of handling ambiguity, we adopt a novel approach of directly aligning the model to address ambiguity.
Alignment of LLMs
LLMs are fundamentally trained through causal language modeling, a process essential for understanding and generating text of high fluency and consistency. To better harness these models, approaches have been developed to align them with human preferences (Leike et al., 2018; Ji et al., 2023b). This has taken various forms, notably through Reinforcement Learning from Human Feedback (RLHF) (Ouyang et al., 2022; Bai et al., 2022a; Chakraborty et al., 2024), as well as Supervised Fine-tuning (SFT) (Dong et al., 2023; Yang et al., 2023; Zhou et al., 2024).
Previous works focused on preferences such as helpfulness (Ding et al., 2023; Köpf et al., 2023; Xu et al., 2024) and safety (Bai et al., 2022b; Ji et al., 2023a; Liu et al., 2024b). Recent studies have concentrated on the factuality (Yang et al., 2023; Tian et al., 2024), avoiding hallucinations. Building on this foundation, our research extends the scope and aims to align models to effectively understand and handle ambiguities, which is a relatively unexplored area within the field of model alignment. This stands in contrast to previous methods which typically bypass the nuanced interpretation of ambiguous contexts inherent in language.
Data Quality Control for Alignment
Data-centric AI (Chu et al., 2016; Majeed and Hwang, 2023; Kumar et al., 2024) emphasizes the importance of data quality in training models. In the context of the instruction following, LIMA (Zhou et al., 2024) demonstrates that models can be effectively aligned even with 1,000 human-curated samples. Similarly, AlpaGasus (Chen et al., 2024) shows effective alignment can be achieved by utilizing only a small subset of the Alpaca dataset (Taori et al., 2023) selected by ChatGPT. Various approaches for data selection have been explored, including those based on pre-defined quality factors such as length and complexity (Liu et al., 2024a), and utilizing gradient similarity from validation sets as a selection criterion (Xia et al., 2024). In this paper, we distinctly define data quality to effectively align models to handle ambiguity. To do so, we make use of the model’s perceived ambiguity as an implicit cue for data quality.
Methodology
The objective of our approach is to align models to explicitly handle potentially ambiguous inputs leveraging intrinsic model knowledge. To this end, we propose a four-stage alignment pipeline, depicted in Figure 2. In this section, we first formulate the problem and describe each stage in detail.
The goal of a QA task is to generate a factually correct answer , given an unambiguous input , a pre-defined inference template , and a language model . As we expand our input scope to ambiguous queries , the model is expected to generate a clarification request We have considered various approaches to handle ambiguity but were concluded to be impractical. Arbitrarily offering one of the valid answers may fail to reflect the user’s intent, and presenting all possible answers is often impractical due to the potentially vast number of valid answers. for to resolve the ambiguity, where the user is best positioned to clarify their intent.
1 Explicit Prediction Stage
This initial stage involves assessing whether the model can appropriately handle each sample and identifying samples that the model currently fails to explicitly manage. By comparing the model’s prediction with the ground-truth label, samples are categorized based on their response accuracy. We collect correct samples, which the model can properly handle, as and incorrect samples are classified as .
2 Implicit Ambiguity Detection Stage
The objective of this stage is to identify samples that the model perceives as ambiguous from . Given that it is challenging for the model to explicitly express ambiguity, we construct a proxy task to estimate the ambiguity from the model’s point of view.
The proxy task is designed to self-disambiguate and implicitly measure the perceived ambiguity. Specifically, the model is first prompted to generate a disambiguation for the input . In this process, the model leverages its intrinsic knowledge related to and generates further details. If lacks specifications and the model possesses related knowledge necessary to compensate, then would yield a higher certainty (lower entropy) for the model. On the other hand, if requires no specification or the model lacks the necessary knowledge, would exhibit a similar level of uncertainty to . To quantify the uncertainty associated with and , we employ the model’s average entropy (Malinin and Gales, 2021; Abdar et al., 2021). Formally, the entropy of an output distribution is defined as follows:
where is the probability of the token of a sentence from the full vocabulary set . The average entropy for can be defined as:
where is composed of -tokens. We quantify the changes in input uncertainty by the difference in average entropy, which we define as information gain (Infogain). The Infogain from the disambiguation can be defined as the following:
If the disambiguation results in a meaningful specification by utilizing intrinsic knowledge, a substantial Infogain would be measured, suggesting that the model considers as ambiguous. Conversely, a negligible Infogain indicates that the model does not perceive as ambiguous. Samples with Infogain greater than the threshold are classified as ambiguous, denoted as .
3 Data Construction Stage
In this stage, we construct datasets for the alignment process. This involves labeling samples identified as ambiguous and constructing an ambiguous dataset . serves as the ground-truth label for ambiguous samples, which are randomly selected from pre-defined clarification requests stipulated in Appendix D. To prevent the potential loss of the model’s existing knowledge, we also incorporate for training. We balance the number of samples from both datasets so that . The final training dataset is thus established as .
4 Supervised Fine-tuning (SFT) Stage
Utilizing the dataset , the model is trained to generate ground-truth label for input , employing the identical inference template . The model with parameter is trained as follows:
Experimental Setting
The capability of the model to perform within the trained domain is pivotal. However, its ability to generalize to out-of-distribution (OOD) is essential for real-world applicability, as queries that deviate from the training data are frequently confronted in the wild. To this end, we employ one training dataset and three OOD test sets to evaluate in diverse domains. All the datasets include both ambiguous and unambiguous queries.
Introduced by Min et al. (2020), AmbigQA is a derivative of the Natural Questions dataset (Kwiatkowski et al., 2019), designed to verify data points deemed ambiguous. The dataset covers diverse sources of ambiguity such as event and entity references. We set AmbigQA as the in-domain dataset and utilize it for training.
SituatedQA
SituatedQA (Zhang and Choi, 2021) specifically focuses on temporal and geographic ambiguity from the input query. As the cause of ambiguity and its construction process are distinct, we assess performance on the temporal split and the geographic split separately, denoted as Temp and Geo, respectively.
AmbigTriviaQA
Since there are limited datasets for evaluating ambiguity in open-domain QA, we construct a new dataset, namely AmbigTriviaQA. By taking questions from the widely-used TriviaQA dataset (Joshi et al., 2017), we prompt gpt-3.5-turbo to ambiguate the initial query and verify the results. More details on dataset construction are described in Appendix B.
2 Baselines
To assess the effectiveness of our approach, we establish two sets of baselines: inference-only methods and trained methods. Further implementation details are described in Appendix A.
Inference-only methods address ambiguity by directly prompting the model. We employ naïve prompting (Naïve) as a fundamental baseline, applying a simple QA prompt. Furthermore, we explore ambiguity-aware prompting (Ambiguity-aware), which additionally provides instructions on handling ambiguity. We also examine Sample Repetition (Cole et al., 2023) by measuring the consistency of the sampled generations. Finally, we evaluate Self-Ask (Amayuelas et al., 2023), where the model initially generates an answer and subsequently determines the ambiguity based on the generation.
Trained Methods
Given the lack of directly comparable prior work, we compare a fine-tuned baseline wherein the model is trained with the in-domain training set. Full-set applies the full in-domain training dataset and Subset is trained on a randomly selected subset of size , which is equivalent in size to our approach. Additionally, we compare Honesty-tuned (Yang et al., 2023), which takes a similar approach utilizing an implicit measure named "expected accuracy" to estimate the model’s factuality. The expected accuracy is measured as the average accuracy of sampled prediction for a single input query. We have revised the method so that incorrect samples based on "expected accuracy" are selected from the ground-truth ambiguous samples. Although Honesty-tuned shares similarities with our approach in the way of estimating the model’s knowledge with an implicit measure, a notable distinction lies in our main focus on handling ambiguity beyond factuality.
3 Evaluation Metrics
As we expand the input scope to possibly ambiguous questions, the model should be capable of handling both unambiguous and ambiguous queries simultaneously. Therefore, we employ two widely used evaluation metrics to assess performance on both types of inputs. A successful alignment should preserve the model’s capability to handle unambiguous inputs while successfully managing ambiguous queries. All evaluations are conducted by comparing the greedy generation to the ground truth.
While expanding the task scope, it remains crucial for the model to preserve the ability to handle unambiguous inputs. Thus, our analysis persists in exclusively evaluating the model’s accuracy in processing unambiguous queries. We measure the quality of the generation by employing RougeLhttps://huggingface.co/spaces/evaluate-metric/rouge (Lin and Och, 2004) with all the possible answers, where the prediction is regarded as correct if the score is above 0.3.
Ambiguity Detection F1-score (Ambig. F1)
The model should be capable of detecting ambiguity and generating clarification requests for ambiguous inputs. However, especially for trained methods, models may exhibit biased predictions toward clarification requests. Taking these aspects into account, we evaluate the model’s ambiguity detection capability with F1-score, which captures both the precision and recall of prediction, offering a balanced view of the model’s ambiguity detection performance. Further details on the detection process are described in Appendix C.
4 Implementation Details
For our experiments, we utilize Llama2 7B 13B (Touvron et al., 2023) and Mistral 7B (Jiang et al., 2023). We utilized QLoRA (Dettmers et al., 2023) to facilitate efficient training. Implementation details are stipulated in Appendix D.
Experimental Results
The main results of our experiments are presented in Table 1. Inference-only methods exhibit a pronounced deficiency in handling ambiguous queries. Specifically, Naïve establish poor performance in responding to ambiguous queries, resulting in a notably low F1-score. Ambiguity-aware demonstrates a strong bias towards clarification requests, as it achieves a relatively high F1-score at the expense of accuracy. Similarly, Sample Repetition exhibits a substantial trade-off between F1-score and accuracy. Self-Ask displays a subpar F1-score, indicating that it is challenging to resolve ambiguity by explicitly "self-asking" the model.
Trained methods exhibit enhanced performance overall compared to inference-only approaches. Honesty-tuned struggles to handle ambiguity, as it also demonstrates biased detection performance. This is likely because the implicit measure from Honesty-tuned can be influenced by various factors but not specifically ambiguity. The results underscore the necessity of distinct methods for perceiving ambiguity. Compared to Honesty-tuned, Subset exhibits relatively balanced performance across both metrics. Full-set demonstrates the most superior performance among the baselines, particularly in the in-domain setting, as it has access to the ground-truth ambiguity.
Our approach yields comparable results in in-domain and demonstrates superior OOD performances. Despite employing identical inference templates as Naïve, our method demonstrates equal or improved unambiguous accuracy. This indicates the effectiveness of our alignment in managing ambiguity while preserving the inherent capabilities of the model. We can also observe an improvement in unambiguous accuracy, particularly in SituatedQA splits. It is especially surprising given that our method was trained on , which the model is already capable of handling. Compared to Full-set, we note a slight decline in the in-domain performance, an expected result given that Full-set is optimized with ground-truth ambiguity of the in-domain data. However, our method outperforms Full-set across OOD datasets in F1-score up to 17 points. This discrepancy underscores the effectiveness of utilizing perceived ambiguity for alignment, facilitating superior generalization and robustness. The efficacy of leveraging only the data perceived ambiguous (about 32% in the Llama2 family and 13% in Mistral) emphasizes the importance of data quality over quantity (Zhou et al., 2024; Chen et al., 2024).
Ablation Study
For a deeper analysis of the influence of Infogain for data selection within our pipeline, we conduct an ablation study by varying the criteria for selecting ambiguous data. While maintaining the same for unambiguous samples, we alter the selection of samples labeled as ambiguous. We compare the following data selection strategies:
Random Selection (Random) We randomly select ground-truth ambiguous samples, without any consideration of Infogain.
Implicit Measure-based Selection (Implicit) We select top- samples with the largest Infogain among those that are ground-truth ambiguous. It differs from our approach as our method utilizes samples perceived as ambiguous, allowing the potential inclusion of unambiguous samples.
Table 2 is the ablation results on Llama2 7B. For AmbigQA, baselines leveraging ground-truth ambiguity slightly outperform our method, which is a similar tendency from Section 5 where Full-set exhibit better in-domain performance. However, across OOD datasets, our approach demonstrates significantly superior performance. Specifically, Random demonstrates a notable drop in F1-score by up to 10 points, illustrating the limitations of simply utilizing the ground-truth ambiguity, which might not align with the model’s perceived ambiguity. Furthermore, Implicit surpasses Random by up to 5 points in F1-score, validating the effectiveness of Infogain as a cue for data selection. Finally, our approach outperforms Implicit across the majority of metrics, even with unambiguous samples selected for training, again highlighting the effectiveness of Infogain as the data selection measure.
2 Analysis on Sample-level Prediction Change
Our method is designed to align the model to generate clarification requests for ambiguous queries. However, the process may lead to a potential trade-off, where the model erroneously generates clarification requests for unambiguous inputs that were previously well-handled. To assess this balance, we introduce three metrics: Valid Alignment Rate (VAR) measures the proportion of ambiguous samples incorrectly handled before alignment that are correctly addressed post-alignment and Misaligned Clarification Rate (MCR) measures the rate of correct unambiguous samples before training that erroneously generates clarification requests after alignment. A high VAR is desirable, whereas a low MCR is preferred simultaneously. Inspired by Yang et al. (2023), we additionally define Overall Alignment Performance (OAP) that measures the balance between VAR and MCR.
Table 3 compares the results on Llama2 7B. Honesty-tuned exhibits high VAR but poor MCR, implying a tendency to misinterpret known knowledge as ambiguous. This aligns with the previous results where Honesty-tuned displays biased generation towards clarification requests. On the other hand, Full-set and Subset demonstrate a good MCR and a relatively low VAR. Our method performs superior OAP, successfully addressing ambiguities (high VAR) while preserving existing capabilities (low MCR).
Self-disambiguation Case Study
Table 4 demonstrates examples of initial query and its disambiguation generated by Llama2 7B. The first example is when is inherently ambiguous, yet the model perceives it as unambiguous. Specifically, the model generates hallucination ("in the 1960s") where the song "don’t mess around with jim" was originally released in 1972. This non-factual generation would not provide any information gain to the model, classifying as ambiguous. In such a case, should be considered "unknown" with no related knowledge within the model. The second and third examples are correctly classified, as the model properly applies its intrinsic knowledge to perceive ambiguity. Regardless of the quantity of additional context generated, the model is capable of verifying its ambiguity. The last example is a misclassification as ambiguous. Despite disambiguation provides factually correct information ("1932 novel" and "by Aldous Huxley") for "brave new world", we speculate that the misclassification may arise from the existence of various media, such as movies and songs, sharing the title "brave new world", leading to an erroneous integration of knowledge.
Conclusion
In this paper, we present a novel alignment pipeline designed to enhance the ability of LLMs to address ambiguities within queries, leveraging the model’s intrinsic knowledge. Our method employs an implicit measure, dubbed Infogain, to quantify ambiguity as perceived by the model. Through alignment based on the measure, the model learns to explicitly handle ambiguous as well as unambiguous queries. Experimental results demonstrate the effectiveness of our alignment, particularly pronounced in out-of-distribution scenarios. Results indicate the importance of alignment based on the model’s perceived ambiguity. Future work may explore the extension of this methodology to broader domains and more complex types of ambiguities, further solidifying the role of LLMs in managing the inherent uncertainty present in NLP tasks.
Limitations
For the experiments, we explore the most widely used models for evaluation, specifically Llama2 and Mistral. Despite this, a more comprehensive evaluation encompassing a broader consideration of LLMs could have enriched our findings, providing insights across different architectures and capabilities. Larger models in scale could demonstrate different tendencies and should be explored for future work. Furthermore, our work mainly focuses on supervised fine-tuning (SFT) as the alignment method. However, alternative methods, such as Reinforcement Learning from Human Preference (RLHF) (Ouyang et al., 2022) or Direct Preference Optimization (DPO) (Rafailov et al., 2023) could offer distinct advantages towards our objective. Finally, the experiments are mainly focused on short-form QA tasks. The research scope could be expanded to long-form generation tasks such as detailed reasoning.
References
Appendix A Baseline Details
In this section, we describe detailed implementation of the baselines.
We make a direct inference using template from Table 5. We evaluate the greedy generation result with temperature 0.
Ambiguity-aware
We utilize the template from Table 6, where we explicitly describe how to handle ambiguity. Identically, we use the greedy generations for evaluation.
Sample Repetition
Template from Table 5 is used to generate a single greedy generation and 10 sampled generations with sampling temperature 1.0. We measure the rate of sampled generations that match the greedy generation, which is reported to be best-calibrated (Cole et al., 2023). Samples with the measure below a specific threshold is considered ambiguous. We empirically select a threshold that demonstrates the best accuracy and F1-score with the least trade-off.
Self-Ask
We initially prompt the model with the template from Table 5 and generate a greedy generation. Then, the initial query and the generated answer is utilized with the template from Table 7 and prompt the model to verify the ambiguity of the query. We modified the prompt from Amayuelas et al. (2023) so that the model can specifically focus on ambiguity. The ambiguity detection is determined based on the model’s final verification.
Honesty-tuned
The approach involves measuring the "expected accuracy" of sampled generations. The expected accuracy is measured by generating 10 samples with temperature 1.0 and measuring the average accuracy among the generations. Ground-truth ambiguous samples with expected accuracy below the specific threshold, in this case 0.1 from Yang et al. (2023), are labeled ambiguous. The samples classified as ambiguous is re-labeled with . The model is trained to generate the ground-truth given the input query and the inference template from Table 5.
Full-set
The full training set is utilized for training. The ground-truth ambiguous samples are labeled with . The model is trained to generate given input question and inference template from Table 5.
Subset
We randomly select samples from the training data, with the same number () of ambiguous and unambiguous samples. Subset is trained in a same way as Full-set.
Appendix B AmbigTriviaQA Construction Details
AmbigTriviaQA is constructed by ambiguating the widely-used TriviaQA dataset (Joshi et al., 2017). We first prompt gpt-3.5-turbo to ambiguate the original question with the template from Table 8. To further validate the generation and control the quality of the dataset, we prompt gpt-3.5-turbo again for a secondary verification. We utilize the template in Table 9 and collect samples verified as ambiguous. This process yielded a total of 4,374 question pairs to examine the model’s capability to interpret and generate responses to intentionally ambiguous queries. Examples from AmbigTriviaQA are demonstrated in Table 10.
Appendix C Ambiguity Detection Details
For ambiguous questions, we expect the model to generate clarification requests. Since there are various ways to express clarification requests, we use the following list of phrases to detect the requests. The presence of pre-defined ambiguity-related phrases in the model’s output is treated as a successful detection. The pre-defined phrases are the follows: [ambiguous, ambig, unclear, not clear, not sure, confused, confusing, vague, uncertain, doubtful, doubt, questionable, clarify, not clear]
Appendix D Implementations Details
For explicit prediction (Stage 1), we utilize the same inference template as Naïve (Table 5) and the disambiguation is generated with the template from Table 11. We use the greedy generation for the disambiguated output. The threshold is set to 0.1 for filtering ambiguous inputs. For balancing training set size, if , we randomly select samples from , where and . If , we select samples from with the largest Infogain. Finally, for , we randomly select from the following pre-defined phrases : [The questions is ambiguous. Please clarify your question. Your question is ambiguous. Can you clarify your question? Your question is not clear. Can you clarify your question please?]
D.2 Training Details
For training, we applied AdamW optimizer (Loshchilov and Hutter, 2019) with a batch size of 32. We selected the model with the best performance from learning rates {1e-3, 5e-4, 1e-4} and training epochs {1, 2, 3}. All the experiments were implemented with Pytorch (Paszke et al., 2019) and Huggingface Transformers library (Wolf et al., 2020). For efficient training, we applied QLoRA from Huggingface PEFT library (Mangrulkar et al., 2022) with r=4 and alpha=16. The training takes about half an hour on a single Tesla V100 GPU.