Successive Prompting for Decomposing Complex Questions
Dheeru Dua, Shivanshu Gupta, Sameer Singh, Matt Gardner
Introduction
Compositional reading comprehension datasets like HotpotQA Yang et al. 2018 and DROP Dua et al. 2019 have inspired a range of model architectures that learn to answer complex questions with weak supervision from the final answer. One recent direction is to leverage large language models (LMs) to solve compositional tasks with very few examples by generating latent reasoning steps before answering the question Wei et al. 2022; Nye et al. 2021; Karpas et al. 2022. Given a complex question, this approach first finds nearest-neighbor training examples from a dataset of (question, reasoning, answer) triples and then concatenates them to create an input for the LM. A large LM is then prompted with this input to generate the intermediate reasoning steps needed, while answering the complex question in a single pass.
While promising, this approach discards many of the benefits of prior approaches to this task Khot et al. 2021; Karpas et al. 2022 by coupling the supervision for question decomposition to the supervision for performing the intermediate steps. Moreover, its non-modular nature does not allow using alternate symbolic reasoning engines in cases where they perform better than LMs. Additionally, the model gets exposed to only a single set of in-context examples, selected based on their proximity to the complex question, which may not contain optimal supervision for the intermediate steps that need to be taken.
We propose “Successive Prompting”, where we iteratively decompose the complex question into the next simple question to answer, answer it, and then repeat until the complex question is answered (Figure 1). Each of these steps is performed with separate a query to the LM. Since the decomposition and answering steps are performed separately, we can decouple the supervision of each step, providing two primary benefits. First, when performing in-context learning, we get multiple opportunities to select different in-context examples, which can be tailored to the particular decomposition or answering step being performed, instead of selecting a single set of examples based only on the complex question. Second, when fine-tuning (with or without in-context examples Chen et al. 2022), we can provide training examples for each step independently, so the model only has to learn to perform one step at a time.
This decoupling additionally allows us to judiciously inject synthetic data into the learning process, e.g., to help the model answer a particular kind of simple question that it could not previously answer, or a new reasoning composition it did not know how to decompose. Because the steps are separate, we can isolate model failures and develop synthetic approaches to fill in the gaps. It also allows us to replace the LM with other, purpose-built components to perform symbolic reasoning when appropriate Khot et al. 2021; Segal et al. 2020; Jin et al. 2021.
We demonstrate the utility of successive prompting using a few-shot variant of the DROP dataset Dua et al. 2019, selecting 300 examples for training (either fine-tuning or in-context example selection). These 300 examples are manually annotated with simple QA pairs as decompositions. We find that performance of all models is quite low in this few-shot setting, so we develop a synthetic data generator that produces complex questions with their decompositions from semi-structured Wikipedia tables Yoran et al. 2021. This synthetic data provides not just complex question supervision, but also supervision for the intermediate steps. We augment this data with the 300 (complex) training examples and their decompositions from DROP. In this few-shot setting, our best performing successive prompting model shows a 5% improvement in F1 when compared to state-of-the-art model on DROP. The code and data are available at https://github.com/dDua/succesive_prompting
Decomposing Complex Questions
The goal of compositional question answering is to answer a complex question in the context of a passage (together denoted as ) by reasoning through latent sequential decisions to reach the final answer, . Many models have been proposed to accomplish this with varying amounts of supervision and interpretability. In prompting methods like Chain-of-Thought (Wei et al. 2022, CoT,) the latent steps are supervised, interpretable sentences; in other models these latent steps might be a program Gupta et al. 2020; Chen et al. 2020 or even just the (unsupervised) hidden states in the model Segal et al. 2020; Andor et al. 2019
There are three kinds of model outputs in this general form: intermediate questions , intermediate answers , and the final answer . We refer to the first kind of output as question decomposition (QD) and the second kind as question answering (QA). We treat final answer prediction as a special case of question decomposition, where the model decides that no more decomposition is necessary and outputs a final answer, so we iteratively alternate between question decomposition and question answering until the model terminates.
2 Training paradigm
During in-context learning, a small number of training examples are provided directly in the prompt that is given to a large LM, before the test input. These examples are selected from an index based on their similarity with the test input. For successive prompting, we create two indices: , for looking-up relevant demonstrations for QD, and , for looking-up relevant demonstrations for QA. The index contains partially decomposed chains at each step , demonstrating the next question to be produced for every complex question in the training data. The index contains all the simple QA pairs in the training data from all the complex questions.
In the QD stage, the index is queried with the complex test question, and current step number, , to select demonstrations regarding how to generate the next question for the held-out example. In the QA stage, the index is queried with the simple question generated during QD to select relevant simple QA pairs. Figure 2 shows a demonstration of how in-context learning is executed step-by-step in each stage until QD outputs the special phrase “There are no more questions left to ask”, along with a final answer.
Successive prompting allows the QA stage to access simple questions derived from complex questions that would not have been retrieved by Chain-of-Thought prompting because on the surface they are not similar to the held-out complex question, even though they share similar sub-questions.
Model Fine-tuning
For model fine-tuning, we use T5 Raffel et al. 2020 based sequence-to-sequence models. Such models are typically trained with control codes in a multi-task setting Ma et al. 2021; Rajagopal et al. 2022 to switch between QD and QA tasks with shared model parameters. We adapt and extend the control codes introduced by text modular networks (Khot et al. 2021, TMNs,) for training with our synthetic data. TMNs are limited in terms of the operations they can handle as they do not go beyond first order reasoning. We use synthetically generated data, which allows us to deal with higher-order reasoning questions in DROP. Because we are fine-tuning the model, we can use special tokens to denote question decomposition and other separators, instead of the natural language prompts shown in Figure 2, though the content is the same. The specific tokens used for each step are listed in Appendix A.
Specialized Modules
Successive prompting also allows us to use specialized sub-modules for solving different QA tasks because we no longer perform QD and QA in an end-to-end manner. Solving arithmetic operations like counting, difference, sorting, etc., can be challenging for language models. As a result, we follow Khot et al. 2021 and construct a simple mathematical sub-module for QA which parses the generated simple question for symbolic operation type and its arguments and then executes them in a deterministic way. If the generated simple question cannot be parsed as a mathematical operation, we apply the language model to solve it.
Synthetic Dataset
Any method that prompts LMs to produce intermediate reasoning steps to answer complex questions needs some amount of supervision for those reasoning steps. This kind of annotation can be expensive to collect and often requires expert knowledge. Prior work has typically relied on a small handful of manually-written example decompositions. We find that such small collections lead to very poor performance on a dataset as varied as DROP, even for large models.
To mitigate these data issues, we propose a way to synthetically generate complex questions and their decompositions using semi-structured data which is easy to parse. We show that we can bootstrap model learning with this out-of-domain, synthetically generated data so it can adapt better when fine-tuned with limited in-domain supervision.
Inspired by Yoran et al. 2021, we use semi-structured data from tables in English Wikipedia which are available in plenty. We employ curated templates to convert the rows in the tables into paragraphs. We use single column headers to create first order simple questions and a combination of columns for higher order complex questions. We synthesize data for 10 simple operations: COUNT, TOP(k), BOTTOM(k), FILTER, SUM, COMPARISON, DIFFERENCE, NEGATION, GATHER, and INTERSECTION.
We generate higher order combinations of first-order operations, wherever possible. Figure 3 shows examples of higher order combinations of the atomic operation COUNT with a few other simple operations using Table 1 as context. The complete list of all decompositions is provided in Appendix A. Depending on the model, we use either symbolic or natural language version of the arithmetic operations. If we are using an LM to perform arithmetic operations, we output natural language; if we are using a separate symbolic reasoning engine, we output symbolic operations. We generate approximately 141K total complex questions which result in 525K examples for QD and 257K examples for QA. See Appendix A for more dataset statistics.
Experiments and Results
The DROP dataset contains a variety of reasoning compositions which are not uniformly distributed. In order to get a fair representation of DROP examples, we first embed the examples using a sentence embedding method trained on the QQP dataset Reimers and Gurevych 2019. We then use cosine similarity to get the top-50 nearest neighbor questions for each training example. The connection graph between each training question to its neighbors is then used to obtain 300 questions that cover the majority of the training data, via the vertex cover algorithm. We manually annotate these 300 examples with decomposed QA pairs in the same format as our synthetic data (Figure 3). For synthetic examples, since we know the reasoning types, we uniformly sample example demonstration from each reasoning type.
We use faiss https://ai.facebook.com/tools/faiss/ index with the QQP-based sentence embedding Reimers and Gurevych 2019 for indexing all the questions. We use GPT-J (6B) https://github.com/EleutherAI which is the largest freely available model we could use with prompts containing 6 in-context examples.
Results
In Table 2, we compare performance of language models without any prompting (Standard), with chain-of-thought prompting (CoT) and successive prompting. We observe that successive prompting performs better than CoT by 3.5% when only synthetic data is available, and 4.3% better with synthetic data and 300 annotations from DROP. The best successive prompting version on the dev set (Synthetic+DROP) has a test set performance of 30.6% F1. We also perform an ablation where the symbolic calculator is replaced by language model and observe that the performance drops by 1.5% F1. This further shows that modular approach is better over a single model that tries to solve all the tasks.
2 Model Fine-tuning
We employ a shared question decomposition (QD) and answering model (QA) based on T5-large version of UnifiedQA Khashabi et al. 2020, trained in a multi-task manner. We use the format described in Appendix A for prompting UnifiedQA. For symbolic questions, we use a simple calculator that parses the operator and arguments in the generated question and executes the discrete operator on the detected arguments.
To deter the model from learning incorrect steps, we use contrastive estimation Smith and Eisner 2005. In particular, we first train the model for two epochs with cross-entropy loss while generating the output sequence (simple question or answer). Then we continue training by adding an auxiliary loss term which increases the likelihood of the intermediate sub-question that would produce a correct sub-answer at the cost of one that does not Dua et al. 2021. We sample up to 3 negative samples at each step. We use HuggingFace transformers https://github.com/huggingface/transformers to train our models, with a learning rate of 5e-5 and maximum input length of 768.
Due to variance in the types of context tables present in Wikipedia, the synthetic dataset distribution is not uniform across different reasoning types. To have a balanced representation of questions pertaining to different reasoning types, we employ dynamic sampling Gottumukkala et al. 2020, where at the beginning of each epoch we select 80,000 instances from across all reasoning types in proportion to the drop in their current performance with respect to previous epoch on held-out synthetic data. For the first epoch we sample in proportion to original the size of each reasoning type. During inference, we use beam search with size 5 to generate decompositions, switching between QD and QA stages until QD reaches end of decomposition (“EOQ”) or maximum number of steps which we set as 10.
Baseline models
We compare against a number of different baselines, both symbolic and non-symbolic. As non-symbolic baselines, we use UnifiedQA Khashabi et al. 2020, which is pre-trained on a number of existing question answering datasets, and PReasM Yoran et al. 2021, which is pre-trained on synthetically generated compositional QA pairs. We also include a baseline with symbolic components, TASE Segal et al. 2020. This model (and others like it Jin et al. 2021; Andor et al. 2019) are capable of performing a combination of continuous and discrete operations, which is essential for DROP. TASE does not require expressing decomposition in a specific grammar and can work with natural language. We chose this model as it is close to state of the art on the full DROP dataset and has publicly available code.
Results
In Table 3, we use the DROP dev set to compare the performance of different symbolic and non-symbolic models in three settings: (1) using no training data from DROP (0-shot), (2) using only question-answer supervision from the 300 DROP examples, and (3) using both question-answer supervision and the decompositions for the 300 DROP examples. In each of these settings, we can train the model with or without the synthetic data that we generated.
We observe that our out-of-domain synthetic data universally improves model performance, and the improvement is most pronounced in TASE, nearing a 20% absolute improvement. Without synthetic data, PReasM is the best performing baseline, but TASE overtakes PReasM when synthetic data is available. Additionally, and unsurprisingly, increasing the amount of supervision from 0-shot to complex QA pairs to decompositions universally improves model performance.
Finally, our method, which is a fine-tuned successive prompting model combined with a symbolic reasoning engine, achieves the best performance, giving an improvement of 5.4 F1 over the state-of-the-art model with similar supervision, i.e. TASE+Synthetic w/ decomp. We follow the standard practice of using test set for only our final best performing model (SP w/ decomp). We observe that our best model with a test set performance of 50.2 F1 is better than the state-of-the-art model with similar supervision (45.1 F1) by 5.1% F1.
Overall, methods that learn to decompose complex questions into simple QA pairs adapt well to complex questions in new domain with little (SP w/ decomp: 51.3 F1) to no in-domain supervision for decomposition (SP 0-shot: 49.8). If we have limited complex QA supervision (without any decompositions), un-interpretable symbolic models result in the best performance (TASE + Synthetic w/o decomp: 44.1). This is because of two reasons. First, such models can capture domain specific answer priors which may result it decent held-out performance Dua et al. 2020; Agrawal et al. 2018. Second, depending on the context, sometimes it may not be straight-forward to decompose the complex questions into QA pairs.
3 In-context vs Fine-Tuning
To understand the gap in performance between successive prompting with in-context learning and fine-tuning, we perform ablations across in-context and fine-tuned version of QD and QA modules. We observe that in-context learning is unable to do well on answering simple questions that result in a list of answers—which is especially important for DROP as symbolic aggregations are generally applied on a list of answers. On using a fine-tuned QA model we see an improvement of 10% in F1 with an in-context QD model. Moreover, since the final answer performance is dependent on how well the QA model performs, using a better QD model (fine-tuned) does not help the overall performance much unless the QA model can handle the decompositions produced by the QD model.
4 Qualitative Examples
To evaluate the correctness of decomposed QA pairs, we manually analyze a subset of predictions on the dev set with in-context (DROP-only) learning and model fine tuning (few shot). We do this by randomly sampling 50 correct predictions to determine how often the incorrect decompositions result in correct answer. We observe that QD stage has an accuracy of 88% for in-context and 96% for fine-tuned model. The incorrect decompositions are mainly because the decomposed question is identical to the original question. For instance, "Who made the longest field goal?" can sometimes be answered correctly without decomposing the question if the passage contains a single field goal mention.
We also sample 50 incorrect predictions to ascertain the reason for incorrect predictions in both in-context and fine-tune setup. We observe that the final predictions are incorrect due to three main categories of errors: incorrect QA model prediction, incorrect next question prediction (QD) and out-of-scope reasoning type. The QA model outputs incorrect answers to simple question 40% and 22% of the times for in-context and fine-tuned respectively. The second class of errors, due to incorrect decomposition, occur 30% of the times for both in-context and fine-tuned. The final class of errors, due to compositional questions that are not covered by synthetically generated annotations, occur 28% (in-context) and 46% (fine-tune) of the times.
In Figure 4, we show a few examples of correct and incorrect predictions and point out the strengths and weaknesses of successive prompting. The main strength of successive prompting is that, by breaking down the question, we are able to get improved supervision for QA. As a result, it is able to correctly identify the goals kicked in the first half while answering the question “How many field goals did both teams kick in the first half?", unlike CoT that returns goals for the entire game.
One of the limitations of in-context learning, when compared with fine-tuning (irrespective of the type of prompting), is that examples are chosen based on the question alone, overlooking the context. For instance, DROP has questions like “How many people were not Germans, in terms of percentage?” where we first need to answer “How many people were Germans, in terms of percentage?" and then perform a negation operation (i.e, subtract from 100). The word “not" influences the example lookup to choose decomposition that involves a negation even when the question being answered requires a different operation.
A limitation of successive prompting is that it is sometimes challenging to decompose a question, especially when it involves implicit reasoning from the passage. For instance, for “Which port did the Korean immigrants leave first Chemulpo or Veracruz?”, it is difficult to explicitly define a comparison style decomposition from the sentence, “After which they took a train to Veracruz”.
Related Work
Prompting was introduced as a way to test the reasoning capabilities of large language models Brown et al. 2020. In follow-up works Schick 2022; Chowdhery et al. 2022; Marasović et al. 2021 prompting techniques have been used as a mechanism to supervise the model decision with few demonstrations as a conditioning context to guide its predictions on an unseen example. Works like Chain-of-Thought reasoning Wei et al. 2022; Zelikman et al. 2022 especially focus on compositional questions where they provide a chain of reasoning as demonstrations. In concurrent work, Least-to-Most prompting Zhou et al. 2022 takes a similar view as ours to break down the problem into sub-problems. However, in Successive Prompting the question decomposition and answering stages are interleaved, unlike Least-to-Most where the problem is first reduced into sub-problem and then executed in a sequence. In our method, the next question prediction has access to previously answered sub-questions, which is useful in questions that need long chain referencing. Other contemporaneous works Press et al. 2022; Khot et al. 2022 use very large language models (more than twice the size we used) and show better few-shot generalization. Works like Perez et al. 2021 have shown the importance of having the right in-context examples for downstream performance leading to works that learn to retrieve relevant in-context examples Rubin et al. 2021.
Non-symbolic methods
Most non-symbolic methods are sequence-to-sequence models trained on a large amount of question answering data Khashabi et al. 2020; Yoran et al. 2021.
Symbolic methods
Neural module networks like approaches parse complex questions into a pre-specified grammar and learn neural components to handle symbolic mathematical operations Gupta et al. 2020; Chen et al. 2020; Nye et al. 2021 which are recursively executed. State-of-the-art models on DROP, however, use a combination of BERT-based contextual models along with a calculator that performs discrete operations Andor et al. 2019; Segal et al. 2020; Hu et al. 2019. Works like Text Modular networks Khot et al. 2021 and MRKL Karpas et al. 2022 are closest to our work. However, they are limited in the terms of types of simple questions they can answer (single-span only) and the complexity of reasoning they can do (single-order only). TMNs, additionally, use a classifier that scores the generated chains module and filters out incorrect question decompositions, while we use contrastive estimation to learn a better question decomposer and as a result do not need a chain scorer.
Conclusion
We present a way to successively decompose complex questions into simple QA pairs, which allows for modular QD and QA systems that can be trained and queried independently. When performing in-context learning, we showed that successive prompting yields an improvement of 4.6 F1 over chain-of-thought prompting. When replacing just the in-context QA module with a fine-tuned one, which is adept at handling list type questions, we further improve the overall performance by 9.5 F1. We believe that modular systems that decompose and delegate tasks to the most appropriate model, whether that is a large LM or a tailored component, are more effective at solving complex tasks than trying to have a large LM solve the entire task on its own. Successive prompting shows one way this decomposition and delegation can be done.
Acknowledgements
We would like to thank Anthony Chen, Catarina Belem and the anonymous reviewers for the discussions and feedback. This material is based upon work sponsored in part by the DARPA MCS program under Contract No. N660011924033 with the United States Office Of Naval Research, in part by funding by AI2 and NSF IIS-1817183. We would also like to thank Hasso Plattner Institute(HPI) for supporting the first author through UCI-HPI fellowship. The views in this work are of authors and not the sponsors.
Limitations
We propose a way to decompose complex questions into interpretable simple QA pairs as latent steps that get successively asked and answered by large pretrained models. The notion of performing complex tasks by iteratively finding and then filling information needs is very general, but we have only shown the applicability of one specific version of this idea in one specific setting. There are many potential challenges in applying successive prompting more broadly. The biggest is that it requires at least some decomposition data, which may be hard or even impossible to obtain. Some complex questions are not easily decomposed, and some domains can be very challenging to write synthetic data generators for. We were able to generate synthetic data that covered most of the reasoning types in DROP, but other kinds of complex questions would not be covered by our generator (e.g., questions that require commonsense or causal reasoning).
There is also significant difficulty in choosing a level of granularity for decomposition. If a large pretrained model can directly answer a question as complex as “What was Barth’s second field goal?”, we should let the model answer the question instead of trying to decompose it further. The right granularity for the decomposition thus depends on the capabilities of the underlying model, and those capabilities are rapidly changing as newer and larger pretrained models are released. There is the possibility that newer model iterations will not need any decomposition to answer the complex questions covered by our synthetic data generator, making that generator obsolete. However, it seems unlikely that pretrained models will be able to handle all complex scenarios in the near future, so the ideas of successive prompting and generating synthetic data to bridge reasoning gaps should still be applicable even when our particular application of them becomes obsolete.
This method also increases the computational requirements for answering complex questions, as instead of making one query to a large LM, successive prompting makes many queries to answer a single question.
Ethics Statement
This work focuses on improving complex question answering with limited data. It uses existing training data and conventional methods of testing model performance. This work does not deal with any social impacts or biases in natural language processing systems.
References
Appendix A Appendix
To generate a simple question given complex question and previously generated latent steps, we append the input with control code “QS:”
{} QI: A: QI: A: QS:
To answer the simple question we prompt the model again with only the simple question this time and control code “A:” to generate answer,
We alternate between the two stages till we reach the end of decomposition marker “EOQ"
{} QI: A: QI: A: QS: EOQ