Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning

Yanda Chen, Chandan Singh, Xiaodong Liu, Simiao Zuo, Bin Yu, He He, Jianfeng Gao

Introduction

Pre-trained large language models (LLMs) have shown impressive proficiency in a range of complex natural language processing tasks, significantly advancing the field and opening new frontiers for applications Brown et al. (2020); Touvron et al. (2023); OpenAI (2023). However, the opaqueness of these models’ decision making process has hindered their use in high-stakes applications such as healthcare, and raised issues related to regulatory pressure, safety, and alignment Goodman and Flaxman (2016); Amodei et al. (2016); Gabriel (2020). Moreover, this lack of interpretability has heavily limited the use of LLMs in fields such as social science and data analysis Ziems et al. (2023), where trustworthy interpretation (rather than model deployment) is itself the end goal.

One growing avenue into interpretability is natural-language explanations produced by LLMs. These explanations are potentially very powerful, helping users predict model behavior Johnson-Laird (1980); Bansal et al. (2019), which is useful to calibrate a model’s capacity and limitations, e.g. limiting its demographic bias Vig et al. (2020). However, these natural-language explanations are limited by the inherent inconsistency of LLMs. For example, one recent work finds that modern LLMs often generate inconsistent explanations and answers on very related questions Chen et al. (2023b). In fact, LLMs often even struggle to consistently answer rephrasings of the same question Sclar et al. (2023); Zhang et al. (2023). It is unclear if popular methods for adapting LLMs, e.g. supervised finetuning or reinforcement learning from human feedback, are able to solve this issue.

We address this issue by introducing explanation-consistency finetuning (EC-finetuning). EC-finetuning finetunes an LLM on synthetic data that is precisely constructed to contain consistent explanations. We start with a question-explanation pair (e.g., “Can sparrows fly?”, “all birds can fly”), generate a set of related questions (e.g., “Can penguins fly?”), and then answer the related questions to be consistent with the initial explanation (e.g., “all birds can fly so penguins can fly”). We generate the synthetic data by prompting LLMs, which can be the same as or different from the explanation LLM.

We apply EC-finetuning to question-answering datasets and find that it improves the consistency of natural-language explanations of LLaMA2-13B by 10.0% relative on four finetuning datasets, and also generalizes to seven out-of-distribution datasets unseen during finetuning (+4.5% relative). This suggests that EC-finetuning may be generally useful for helping users build mental models of an LLM from its explanations (see Fig. 1).

Related work

A great deal of recent work has studied improving LLM controllability, e.g. through supervised finetuning Liu et al. (2019), reinforcement learning from human feedback Ouyang et al. (2022) or learning from human explanations Stiennon et al. (2020). Two related works study the consistency in the generations made by an LLM, either between the generation and validation of LLMs Li et al. (2023) or LLM predictions on implications of an original question Akyürek et al. (2024). In contrast to EC-finetuning, these works do not focus directly on improving an LLM’s explanation capabilities.

Many works extend and analyze explanations given by chain-of-thought prompting Wei et al. (2022), e.g. by evaluating counterfactuals introduced into the chain of thought Gat et al. (2023), testing their robustness to mistakes introduced into the reasoning chain Lanham et al. (2023), or using contrastive chain-of-thought to induce reliance on the reasoning chain Chia et al. (2023). These methods do not alter the underlying LLM, and thus can be used in conjunction with EC-finetuning.

Evaluating natural-language explanations

We summarize three existing orthogonal metrics for explanations: consistency, plausibility, and faithfulness. Consistency, which we focus on in this work, measures if the model generates consistent explanations on similar examples Hase and Bansal (2020); Chen et al. (2023b). Plausibility evaluates humans’ preference of an explanation based on its factual correctness and logical coherence Herman (2017); Lage et al. (2019); Jacovi and Goldberg (2020). It is different from faithfulness, which measures whether an explanation is consistent with the model’s internal decision process Harrington et al. (1985); Jacovi and Goldberg (2020).

Method: EC-finetuning

EC-finetuning is an intuitive method that augments data in a manner that enhances explanation consistency (Fig. 2). Specifically, it prompts LLMs to augment data in two steps. In the first step, a question-explanation pair is given to an LLM (e.g., “Can sparrows fly?”, “all birds can fly”), with the task of generating follow-up questions related to the explanation of the initial question (e.g., “Can penguins fly?”). This is achieved by explicitly prompting the LLM to generate questions that are answerable given the initial explanation.

In the second step, another LLM generates answers and explanations for the follow-up questions. To ensure these answers and explanations are consistent with the explanation in the initial question, the initial question-explanation pair is presented in the prompt, alongside explicit instructions to keep the new explanation consistent with the initial (e.g., “all birds can fly so penguins can fly”.) Finally, these augmented questions, along with their explanations and answers are used for finetuning an LLM to generate more consistent explanations.

We use different LLMs for the two data augmentation steps (here, GPT-4 OpenAI (2023) for the first step and Claude-2https://www.anthropic.com/index/claude-2 for the second step) to avoid issues with LLMs that favor their own outputs Zheng et al. (2023). Precise prompts are given in Sec. A.1. Note that the two data augmentation LLMs can be, but are not required to be, identical to the explanation LLM.

Measuring consistency

Evaluating the consistency of model explanations is challenging. Here, we follow the metric proposed by Chen et al. 2023b, which measures explanation consistencyWhat we call “consistency”, Chen et al. call “counterfactual simulatability precision”. as the fraction of answers on follow-up questions that match a human user’s expectation based on an explanation (similar to Fig. 1); the metric ranges from 0 to 1, with 1 being a perfect score. Additionally, following Chen et al. 2023b we use LLMs as a user simulator, as it was found to reliably emulate a user’s predictions on follow-up questions.

We evaluate consistency on two types of follow-up questions: related questions and rephrased questions. For related questions, similar to how we generate EC data, we again use GPT-4 to generate. For rephrased questions, we prompt GPT-4 to generate exact paraphrases of the original questions.

To ensure this metric from Chen et al. 2023b is reliable, we conduct two sanity checks. First, we check if the metric is stable with respect to how the metric is computed. We run various perturbations (e.g. varying how the followup questions are generated or the explanation format) and find the metric to be quite stable (see Table A3). Second, we study the correlation between explanation consistency and explanation length to see if the metric can be easily hacked by generating shorter/longer explanations. We do not observe a correlation on any of the 7 unseen datasets (see Table A1).

Results

We perform EC-finetuning on the LLaMA-2 13-billion parameter model Touvron et al. (2023). For finetuning, we use 4 datasets: StrategyQA Geva et al. (2021), MedMCQA Pal et al. (2022), and two versions of MedQA Zhang et al. (2018): MedQA-Sim contains related questions on diagnosis and treatment (similar to the original questions), whereas MedQA-Diff contains related questions on medical facts derived from the original questions.

We additionally evaluate consistency on 7 datasets that are not used for finetuning: BoolQ Clark et al. (2019), Natural Questions (NQ) Kwiatkowski et al. (2019), MS-Marco Nguyen et al. (2016), OBQA Mihaylov et al. (2018), MMLU-Medical Hendrycks et al. (2020), PubMedQA Jin et al. (2019) and ARC-Easy Clark et al. (2018). For a cleaner evaluation, these 7 datasets are all converted to have a shared yes-no answer format. We show each dataset’s domain and the skills it tests in Table 1. The testing datasets introduce a distribution shift as they cover new domains (science) and new skills (commonsense reasoning and quantitative reasoning) not seen during finetuning. Table A2 shows the size of each dataset.

2 Main result: EC-finetuning improves explanation consistency

Table 2 shows the main results for EC-finetuning. EC-finetuning can effectively improve consistency, yielding an average relative improvement of 10.0% for tasks seen during finetuning and 4.5% for unseen tasks. An improvement is seen for every dataset studied here and for both types of followup questions. The largest gain in consistency after EC-finetuning is for MedQA-Diff; this suggests that EC-finetuning can also improve the LLM’s explanation consistency on related questions that are more different from the original questions. These consistency improvements also come with modest accuracy improvements (5.2% relative for finetuning tasks and 4.3% relative for unseen tasks). There is no significant correlation between improvement in consistency and the improvement in accuracy (Pearson correlation coefficient ρ=0.001\rho=0.001). This suggests that the consistency improvement derived from EC-finetuning differs from the improvement attained by standard supervised finetuning.

We explore a simplified setting, where EC-finetuning is run using only the LLaMA-2 13-billion parameter, both for synthetic data generation and explanation finetuning. This setting tests whether EC-finetuning can be used with smaller LLMs and whether those LLMs can improve their own explanation consistency. We find that when running EC-finetuning on StrategyQA, EC-finetuning yields a 4.4% relative improvement but decreases accuracy by 5.4%. This suggests that EC-finetuning may succeed in improving explanation consistency in today’s relatively small models, but can incur some tradeoffs as a result, i.e. decreasing accuracy.

3 Analysis

Table 3 shows examples of explanations before/after EC-finetuning. The consistency of the explanation in both examples increases after EC-finetuning, but in different ways. In the first example, EC-finetuning encourages the model to generate more precise explanations that are not overgeneralized/vague. On the other hand, in the second example, EC-finetuning does not change the explanation the model generates for the initial question, but instead changes the model’s predictions on related questions to be more consistent with the explanation on the initial question.

Inconsistent explanations suggest incorrect predictions.

Do LLMs generate more consistent explanations on correct predictions? We study the correlation between explanation consistency and prediction accuracy across different examples of the same dataset. We find that the baseline model shows a positive correlation of 0.099 (Pearson), and this correlation increases to 0.185 after EC-finetuning (dataset-level breakdown in Table 4). This indicates that inconsistent explanations suggest wrong predictions, and we may calibrate LM’s predictions based on the consistency of its explanations Chen et al. (2023a).

EC-finetuning improves consistency more on correct predictions.

We compare the consistency improvement from EC-finetuning on correct versus incorrect predictions. EC-finetuning improves explanation consistency on correct predictions by 5.7% relative but only 1.2% relative on incorrect predictions (see full breakdown in Table 5). This is possibly because examples where the model makes mistakes are likely challenging ones beyond the model’s capability, so the model can produce inconsistent hallucinations.

Conclusion

Forming reliable mental models of LLMs will be critical to understanding and deploying them as their capabilities continue to grow. EC-finetuning is a step in this direction, helping to align models to not only be more accurate, but also to provide more self-consistent explanations. A big part of future work will be i) scaling up EC-finetuning to larger models, ii) exploring EC-finetuning on more complicated tasks (e.g., multimodal understanding), iii) studying how LLMs represent consistency in their parameters and if EC-finetuning improves this representation, and iv) proposing other methods to improve explanation consistency.

Acknowledgements

This research is supported in part by Open Philanthropy. We thank Kathleen McKeown, Zhou Yu, Sarah Wiegreffe, Ruiqi Zhong and Qing Lyu for valuable discussions and feedback on this paper.

References

Appendix A Appendix

We prompt LLMs to generate the EC training data in two steps 1) generate follow-up questions related to an initial explanation, 2) answer the related questions conditioned on the initial explanation. We include a few demonstration examples for both steps (in-context learning), which are sampled from each dataset. We include below the prompts we use for the StrategyQA dataset, and include the demonstration examples for other datasets at https://github.com/yandachen/explanation-consistency-finetuning.

Human: In the questions below, you will be asked to read a starter question and its answer. After that you will be asked to write a follow-up question that can be answered based on the starter QA, and write your answer to the follow-up question based on the starter QA. Your follow-up question should be self-contained even without the starter question.

Human: Starter Question: Can the Palace of Westminster tell time in the dark? Answer to the Starter Question: The Palace of Westminster is a building. Buildings cannot tell time.

Assistant: here is my response. Follow-up Question: Can the White House tell time? Answer to the Follow-up Question: We know from the starter QA that buildings cannot tell time. Since the White House is also a building, the White House cannot tell time either. So the answer is no.

Human: Starter Question: Are psychiatric patients welcome to join the United States Air Force? Answer to the Starter Question: The United States Air Force is a branch of the United States military. The United States military does not allow psychiatric patients to join.

Assistant: here is my response. Follow-up Question: Are psychiatric patients welcome to join the United States Army? Answer to the Follow-up Question: We know from the starter QA that the United States military does not allow psychiatric patients to join. Because United States Army is a branch of the United States military, psychiatric patients are not welcome to join the United States Army either. So the answer is no.

Human: Starter Question: Does The Hague border multiple bodies of water? Answer to the Starter Question: The Hague is a city in the Netherlands. The Netherlands borders the North Sea.

Assistant: here is my response. Follow-up Question: Is Amsterdam near the North Sea? Answer to the Follow-up Question: We know from the starter QA that the Netherlands borders the North Sea. Because Amsterdam is a city in the Netherlands, Amsterdam is also near the North Sea. So the answer is yes.

Human: Starter Question: Can an Asian black bear use chopsticks? Answer to the Starter Question: Asian black bears are omnivores. Omnivores can use chopsticks.

Assistant: here is my response. Follow-up Question: Can pigs use chopsticks? Answer to the Follow-up Question: We know from the starter QA that omnivores can use chopsticks. Because pigs are omnivores, pigs can use chopsticks as well. So the answer is yes.

Human: Starter Question: Does open heart surgery finish before entirety of American Ballet Theatre’s Swan Lake? Answer to the Starter Question: The American Ballet theatre’s Swan Lake has a run time of 45 minutes. The National Heart, Lung, and Blood Institute states that a coronary artery bypass takes 3 to 6 hours.

Assistant: here is my response. Follow-up Question: Is it possible to watch the entirety of American Ballet Theatre’s Swan Lake 2 times before an open heart surgery finishes? Answer to the Follow-up Question: We know from the starter QA that the run time of Swan Lake is 45 minutes and an open heart surgery takes 3 to 6 hours. Therefore it is possible to watch Swan Lake twice before the surgery finishes (90 minutes). So the answer is yes.

Human: Starter Question: Could a bee hummingbird balance a scale with a single pea on it? Answer to the Starter Question: The bee hummingbird is the smallest bird in the world. The bee hummingbird weighs about 2 grams. A pea weighs about 1 gram.

Assistant: here is my response. Follow-up Question: Could a pea balance a scale with a dollar bill on it? Answer to the Follow-up Question: We know from the starter QA that a pea weighs about 1 gram. Because a dollar bill weighs about 1 gram as well, a pea can balance a scale with a dollar bill on it. So the answer is yes.

Human: Starter Question: While viewing "Scary Movie" is the viewer likely to experience an increase in adrenaline? Answer to the Starter Question: The movie "Scary Movie" is a comedy. Comedies are not scary.

Assistant: here is my response. Follow-up Question: Is "Superbad" scary? Answer to the Follow-up Question: We know from the starter QA that comedies are not scary. "Superbad" is also a comedy, so "Superbad" is not scary either. So the answer is no.

Human: Starter Question: Is Mixed martial arts totally original from Roman Colosseum games? Answer to the Starter Question: Mixed Martial arts in the UFC takes place in an enclosed structure called The Octagon. The Roman Colosseum games were fought in enclosed arenas where combatants would fight until the last man was standing. Mixed martial arts contests are stopped when one of the combatants is incapacitated. The Roman Colosseum was performed in front of crowds that numbered in the tens of thousands. Over 56,000 people attended UFC 193.

Write 10 follow-up yes or no questions. Start each question by "1. ", "2. ", etc. Skip writing answers to your follow-up questions. Remember, your task is to write follow-up questions that can be answered based on the starter QA. Your follow-up questions should be self-contained even without the starter question.

Follow-up Questions Answering

Human: In the questions below, you will be asked to read a starter question and its answer. After that you will be asked to read a follow-up question and judge whether the starter QA directly helps choosing a single answer for the follow-up question. If not, end your answer with "So the answer is unknown.". If yes, use the starter QA to answer the follow-up question, explain your reasoning as clearly and as detailed as possible using all relevant information in the starter QA, end your answer with "So the answer is yes/no.", and do NOT explicitly mention "the starter QA" or "According to the starter QA" in your answer. Stick to the starter QA when you answer the follow-up question, even if the reasoning or claims in the starter QA are wrong.

Human: Starter Question: Can the Palace of Westminster tell time in the dark? Answer to the Starter Question: The Palace of Westminster is a building. Buildings cannot tell time. Follow-up Question: Can the White House tell time? Judge whether the starter QA directly helps choosing a single answer for the follow-up question. If not, end your answer with "So the answer is unknown.". If yes, use the starter QA to answer the follow-up question, explain your reasoning as clearly and as detailed as possible using all relevant information in the starter QA, end your answer with "So the answer is yes/no.", and do NOT explicitly mention "the starter QA" or "According to the starter QA" in your answer. Stick to the starter QA when you answer the follow-up question, even if the reasoning or claims in the starter QA are wrong. Answer to the Follow-up Question:

Assistant: here is my response. The White House is a building. Buildings cannot tell time. So the answer is no.

Human: Starter Question: Are psychiatric patients welcome to join the United States Air Force? Answer to the Starter Question: The United States Air Force is a branch of the United States military. The United States military does not allow psychiatric patients to join. Follow-up Question: Are psychiatric patients welcome to join the United States Army? Judge whether the starter QA directly helps choosing a single answer for the follow-up question. If not, end your answer with "So the answer is unknown.". If yes, use the starter QA to answer the follow-up question, explain your reasoning as clearly and as detailed as possible using all relevant information in the starter QA, end your answer with "So the answer is yes/no.", and do NOT explicitly mention "the starter QA" or "According to the starter QA" in your answer. Stick to the starter QA when you answer the follow-up question, even if the reasoning or claims in the starter QA are wrong. Answer to the Follow-up Question:

Assistant: here is my response. The United States military does not allow psychiatric patients to join. The United States Army is a branch of the United States military. So the answer is no.

Human: Starter Question: Does The Hague border multiple bodies of water? Answer to the Starter Question: The Hague is a city in the Netherlands. The Netherlands borders the North Sea. Follow-up Question: Is Amsterdam near the North Sea? Judge whether the starter QA directly helps choosing a single answer for the follow-up question. If not, end your answer with "So the answer is unknown.". If yes, use the starter QA to answer the follow-up question, explain your reasoning as clearly and as detailed as possible using all relevant information in the starter QA, end your answer with "So the answer is yes/no.", and do NOT explicitly mention "the starter QA" or "According to the starter QA" in your answer. Stick to the starter QA when you answer the follow-up question, even if the reasoning or claims in the starter QA are wrong. Answer to the Follow-up Question:

Assistant: here is my response. Amsterdam is a city in the Netherlands. Netherlands borders the North Sea. So the answer is yes.

Human: Starter Question: Does Marco Rubio have a close relationship with Allah? Answer to the Starter Question: Marco Rubio is a Christian. Christians do not have a close relationship with Allah. Follow-up Question: Is Barack Obama a Christian? Judge whether the starter QA directly helps choosing a single answer for the follow-up question. If not, end your answer with "So the answer is unknown.". If yes, use the starter QA to answer the follow-up question, explain your reasoning as clearly and as detailed as possible using all relevant information in the starter QA, end your answer with "So the answer is yes/no.", and do NOT explicitly mention "the starter QA" or "According to the starter QA" in your answer. Stick to the starter QA when you answer the follow-up question, even if the reasoning or claims in the starter QA are wrong. Answer to the Follow-up Question:

Assistant: here is my response. The starter QA does not include information that directly answer this follow-up question. So the answer is unknown.

Human: Starter Question: Can an Asian black bear use chopsticks? Answer to the Starter Question: Asian black bears are omnivores. Omnivores can use chopsticks. Follow-up Question: Can pigs use chopsticks? Judge whether the starter QA directly helps choosing a single answer for the follow-up question. If not, end your answer with "So the answer is unknown.". If yes, use the starter QA to answer the follow-up question, explain your reasoning as clearly and as detailed as possible using all relevant information in the starter QA, end your answer with "So the answer is yes/no.", and do NOT explicitly mention "the starter QA" or "According to the starter QA" in your answer. Stick to the starter QA when you answer the follow-up question, even if the reasoning or claims in the starter QA are wrong. Answer to the Follow-up Question:

Assistant: here is my response. Pigs are omnivores. Omnivores can use chopsticks. So the answer is yes.

Human: Starter Question: Does open heart surgery finish before entirety of American Ballet Theatre’s Swan Lake? Answer to the Starter Question: The American Ballet theatre’s Swan Lake has a run time of 45 minutes. The National Heart, Lung, and Blood Institute states that a coronary artery bypass takes 3 to 6 hours. Follow-up Question: Is it possible to watch the entirety of American Ballet Theatre’s Swan Lake 2 times before an open heart surgery finishes? Judge whether the starter QA directly helps choosing a single answer for the follow-up question. If not, end your answer with "So the answer is unknown.". If yes, use the starter QA to answer the follow-up question, explain your reasoning as clearly and as detailed as possible using all relevant information in the starter QA, end your answer with "So the answer is yes/no.", and do NOT explicitly mention "the starter QA" or "According to the starter QA" in your answer. Stick to the starter QA when you answer the follow-up question, even if the reasoning or claims in the starter QA are wrong. Answer to the Follow-up Question:

Assistant: here is my response. The run time of Swan Lake is 45 minutes. An open heart surgery takes 3 to 6 hours. 45 minutes * 2 = 90 minutes < 3 hours. So the answer is yes.

Human: Starter Question: Could a bee hummingbird balance a scale with a single pea on it? Answer to the Starter Question: The bee hummingbird is the smallest bird in the world. The bee hummingbird weighs about 2 grams. A pea weighs about 1 gram. Follow-up Question: Could a pea balance a scale with a dollar bill on it? Judge whether the starter QA directly helps choosing a single answer for the follow-up question. If not, end your answer with "So the answer is unknown.". If yes, use the starter QA to answer the follow-up question, explain your reasoning as clearly and as detailed as possible using all relevant information in the starter QA, end your answer with "So the answer is yes/no.", and do NOT explicitly mention "the starter QA" or "According to the starter QA" in your answer. Stick to the starter QA when you answer the follow-up question, even if the reasoning or claims in the starter QA are wrong. Answer to the Follow-up Question:

Assistant: here is my response. A pea weighs about 1 gram. A dollar bill weighs about 1 gram. So the answer is yes.

Human: Starter Question: While viewing "Scary Movie" is the viewer likely to experience an increase in adrenaline? Answer to the Starter Question: The movie "Scary Movie" is a comedy. Comedies are not scary. Follow-up Question: Is "Superbad" scary? Judge whether the starter QA directly helps choosing a single answer for the follow-up question. If not, end your answer with "So the answer is unknown.". If yes, use the starter QA to answer the follow-up question, explain your reasoning as clearly and as detailed as possible using all relevant information in the starter QA, end your answer with "So the answer is yes/no.", and do NOT explicitly mention "the starter QA" or "According to the starter QA" in your answer. Stick to the starter QA when you answer the follow-up question, even if the reasoning or claims in the starter QA are wrong. Answer to the Follow-up Question:

Assistant: here is my response. The movie "Superbad" is a comedy. Comedies are not scary. So the answer is no.

Human: Starter Question: Do gorillas fight with panda bears? Answer to the Starter Question: Gorillas are herbivores. Pandas are herbivores. Herbivores do not fight with other herbivores. Follow-up Question: Do herbivores fight with carnivores? Judge whether the starter QA directly helps choosing a single answer for the follow-up question. If not, end your answer with "So the answer is unknown.". If yes, use the starter QA to answer the follow-up question, explain your reasoning as clearly and as detailed as possible using all relevant information in the starter QA, end your answer with "So the answer is yes/no.", and do NOT explicitly mention "the starter QA" or "According to the starter QA" in your answer. Stick to the starter QA when you answer the follow-up question, even if the reasoning or claims in the starter QA are wrong. Answer to the Follow-up Question:

Assistant: here is my response. The starter QA does not include information that directly answer this follow-up question. So the answer is unknown.

Human: Starter Question: Can citrus grow in Ulaanbaatar? Answer to the Starter Question: Citrus trees can grow in Ulaanbaatar. Thus, citrus can grow in Ulaanbaatar. Follow-up Question: Can palm trees grow in Ulaanbaatar? Judge whether the starter QA directly helps choosing a single answer for the follow-up question. If not, end your answer with "So the answer is unknown.". If yes, use the starter QA to answer the follow-up question, explain your reasoning as clearly and as detailed as possible using all relevant information in the starter QA, end your answer with "So the answer is yes/no.", and do NOT explicitly mention "the starter QA" or "According to the starter QA" in your answer. Stick to the starter QA when you answer the follow-up question, even if the reasoning or claims in the starter QA are wrong. Answer to the Follow-up Question:

Assistant: here is my response. The starter QA does not include information that directly answer this follow-up question. So the answer is unknown.

Human: Starter Question: Is Mixed martial arts totally original from Roman Colosseum games? Answer to the Starter Question: Mixed Martial arts in the UFC takes place in an enclosed structure called The Octagon. The Roman Colosseum games were fought in enclosed arenas where combatants would fight until the last man was standing. Mixed martial arts contests are stopped when one of the combatants is incapacitated. The Roman Colosseum was performed in front of crowds that numbered in the tens of thousands. Over 56,000 people attended UFC 193. Follow-up Question: Is the UFC Octagon considerably smaller than the Roman Colosseum? Judge whether the starter QA directly helps choosing a single answer for the follow-up question. If not, end your answer with "So the answer is unknown.". If yes, use the starter QA to answer the follow-up question, explain your reasoning as clearly and as detailed as possible using all relevant information in the starter QA, end your answer with "So the answer is yes/no.", and do NOT explicitly mention "the starter QA" or "According to the starter QA" in your answer. Stick to the starter QA when you answer the follow-up question, even if the reasoning or claims in the starter QA are wrong. Answer to the Follow-up Question: