Head-to-Tail: How Knowledgeable are Large Language Models (LLMs)? A.K.A. Will LLMs Replace Knowledge Graphs?
Kai Sun, Yifan Ethan Xu, Hanwen Zha, Yue Liu, Xin Luna Dong
Introduction
Pre-trained large language models (LLMs), such as ChatGPThttps://openai.com/blog/chatgpt and LLaMA Touvron et al. (2023a), have demonstrated impressive capabilities in internalizing knowledge and responding to common inquiries Ouyang et al. (2022); OpenAI (2023). Nevertheless, these models often lack knowledge of nuanced, domain-specific details and are susceptible to hallucinations Bang et al. (2023), underscoring the significant challenges of increasing the factuality of LLMs and minimizing hallucinations from LLM responses. Conversely, the rise of LLMs has sparked debates on whether Knowledge Graphs (KGs), which store real-world factual knowledge in triplet form (subject, predicate, object), will be replaced with LLMs. This paper tries to answer these questions from a new angle: How knowledgeable are LLMs?
Entities. An important contribution of the Head-to-Tail benchmark is the bucketing of head, torso, and tail entities, decided by the popularity of the entities (we will also discuss how the popularity of the predicates affect results in Section 3.3). We use two ways to approximate popularity: traffic and density. When there is traffic information, such as views and votes, we conveniently use traffic to measure the popularity; otherwise, we use density as a proxy, such as the number of facts or works about the entity. We often observe a correlation between density and traffic (e.g., the more popular a person is, the more we know about her), but as we will see soon from the benchmark statistics, they can still lead to slightly different distributions of head, torso, and tail. We now give details on how we decide the popularity of different types of entities from each data source.
IMDb (traffic): The number of votes (i.e., numVotes) the title (e.g., movie, short, TV series, etc.) has received; we do NOT consider whether the vote is high or low in the counting. For person entities, we use the total number of votes received by the titles the person is known for.
Goodreads (traffic): The count of ratings (i.e., ratings_count) the book has received; similarly, we do NOT take into consideration whether the rating is high or low.
MAG (traffic): The number of citations (i.e., CitationCount) the entity (i.e., scholarly article, conference, or journal) has received.
DBLP (density): The number of works the scholar has authored.
DBpedia (density): The number of relational triples in DBPedia that contain the entity.
We bucketed head, torso, and tail entities in three steps. First, we sorted the entities by their popularity, measured as above. Second, for each entity, we computed the cumulative popularity score up to the top-1 entity in the sorted list. Third, we bucketed the entities such that head entities comprise entities whose cumulative popularity score is up to 1/3 of that of all entities, torso entities comprise entities with cumulative scores ranging from 1/3 to 2/3, and tail entities from 2/3 to 1. We determined the partitioning separately for different entity types for each domain.
To make the popularity score fair, we filtered out entities that are likely too new to have sufficient statistical data for popularity measurement. For IMDb, MAG, DBLP, and Goodreads, we kept only entities by the years 2020, 2020, 2020, and 2015, respectively. The cut-off years are all before the cut-off time of the LLM training data, so the benchmark avoids questions that require recent knowledge. We did not perform similar filtering for DBpedia because the year attribute is unavailable or non-applicable for most entities, and our pilot study shows that very few (if at all) of the questions generated from DBpedia require knowledge after 2020.
Table 1 summarizes the distribution of head, torso, and tail entities. The distribution follows the power law, where very small percentages of entities fall in the head and torso buckets, and the majority of entities fall in the tail bucket; for example, over 99.9% of movies fall in the tail bucket, according to IMDb vote counts. We also observe that this phenomenon is more pronounced when we measure by traffic than by density; for the latter, the torso buckets are often larger (15% of entities), and the tails are slightly smaller (82%).
Questions. We generated questions using a template-based approach, where each generated question asks for an attribute of an entity. We filtered out the following types of attributes: (i) unspecific (e.g., seeAlso in DBpedia), (ii) dynamic (e.g., lastLaunchRocket in DBpedia), (iii) data source specific (e.g., averageRating in IMDb), and (iv) non-textual (e.g., picture in DBpedia). For each specific domain (Movie, Book, Academics), we manually designed the question template for each attribute. DBpedia contains a large set of attributes, so we first employed ChatGPT to draft the templates (using Prompt 1 in Appendix A.1), then proofread them manually and made necessary edits.
The answer for each question is the object of the relevant triple; when there are multiple answers (e.g., a book may have multiple authors), we included all in the answer. When necessary, we included extra information for an entity to avoid potential ambiguities (e.g., we included the publication year for a book to distinguish books of highly similar names).
We generated an equal number of questions for randomly sampled head, torso, and tail entities using each template. For each specific domain, we generated 1K questions for each of the head, torso, and tail buckets. As DBPedia contains more domains and relationship types, we generated 3K questions for each bucket. Table 2.1 summarizes the overall statistics of Head-to-Tail in the number of questions and templates.
We evaluated representative state-of-the-art LLMs of various sizes and architectures, including ChatGPT, LLaMA (7B, 13B, 33B, 65B) Touvron et al. (2023a), Vicuna (7B, 13B) Chiang et al. (2023), Flan-T5 (3B, 11B) Chung et al. (2022), RWKV (7B) Peng et al. (2023b), Falcon (7B, 40B), and Falcon-Instruct (7B, 40B) Almazrouei et al. (2023). We employed the most deterministic settings (i.e., temperature=0 or top_k=1) for all models.
We interacted with ChatGPT through OpenAI APIhttps://platform.openai.com/docs/api-reference. The employed ChatGPT version is gpt-3.5-turbo-0301. We used Transformers Wolf et al. (2020) to interact with the other LLMs on A100 (80GB) GPUs, and we used 16-bit floating point formats (i.e., float16 for Flan-T5 and RWKV, bfloat16 for LLaMA, Vicuna, Falcon, and Falcon-Instruct). We employed the original LLaMA, Flan-T5, Falcon, and Falcon-Instruct versions. The employed version of RWKV and Vicuna is v4 Raven and v1.1, respectively.
Table 9 in Appendix A.3 gives detailed results of all LLMs. We note that our goal is NOT to compare different LLM models; rather, by examining the metrics by different LLMs, we make sure to report the common patterns among the representative LLMs.
2 RQ1: How reliable are LLMs in answering factual questions?
We present in Table 3 the overall performance of ChatGPT and LLaMA-33B, which perform the best in most metrics on Head-to-Tail among all LLMs introduced in Section 3.1. Both models correctly answer only 20% of questions (measured by ALM).
Interestingly, for questions that are not answered correctly, ChatGPT and LLaMA-33B show opposite patterns: ChatGPT gives unsure or empty answers for the majority of them, and the hallucination rate is only 15% (still non-negligible), while LLaMA-33B mostly provides hallucinated answers, resulting with high hallucination rate (80%). We suspect fine-tuning and reinforcement learning of these models may explain the different patterns when the model is unsure of the answers. Figure 1 shows examples of counterfactual answers given by ChatGPT.
Finally, for both models, the overall performance is mostly similar to the performance on the open domain (recall that the open domain accounts for only half of the question-answer pairs), but the performance varies substantially across different specific domains. Both models perform the best in the Movie domain and worst in the Academics domain, likely due to the relatively low popularity of the Academics domain, as we will discuss soon.
Head-to-Tail predicates. We investigated whether the performance still correlates with the head-to-tail order regarding the popularity of predicates instead of entities. We sorted the predicates from DBpedia by popularity (measured by the number of relational triples with the predicate) and partitioned the sorted predicates into head, torso, and tail in a similar fashion. We then re-partitioned the open-domain questions into head, torso, and tail predicate buckets, each containing , , and questions, respectively. Since the number of questions in the head bucket is low, we merged the head and torso buckets.
Table 4 compares the performance on head & torso vs. on tail. We observe no consistent correlation among different LLMs between the performance and the head-to-tail predicate ordering, and the differences in accuracy are not very high. This is not too surprising for two reasons. First, the semantics of each predicate is mostly consistent with the semantics of the predicate names, which can be well understood by LLMs. Second, when facts are present for tail predicates, they are often about the head entities, and factual information for head entities is likely to be more abundant in the training data.
Table 5 compares LLMs in different sizes and with or without instruction tuning. First, we observe that an increased model size does not automatically translate to a better grasp of factual knowledge. For example, LLaMA-33B modestly outperforms LLaMA-65B across the head, torso, and tail subsets ( in ALM and in HLM on average) while they share the same training dataset and hyperparameters. This provides additional evidence for our hypothesis that once the model is sufficiently large, the abundance of training data plays a more critical role in the factuality of the LLMs.
Second, compared with LLaMA and Falcon, the instruction-tuned counterparts (i.e., Vicuna and Falcon-Instruct) have lower accuracy, as they learned to be more conservative in providing factual answers and thus generate “unsure” more often (e.g., Vicuna-13B is higher in M than LLaMA-13B). Despite so, they still have high hallucination rate.
Finally, we evaluate the robustness of our evaluation methodology.
Correlations between rule- and LLM-based metrics. For each combination of popularity (head, torso, tail) and domain (movie, book, academics, open), we calculate Spearman’s rank and Pearson correlation coefficients between rule- and LLM-based metrics over all LLMs. We report the aggregated results (minimum, mean) in Table 6. The correlation scores suggest that ALM (resp. HLM) strongly correlates with AEM, AF1, and ARL (resp. HEM, HF1, and HRL), indicating that rule-based metrics are good alternatives for lower-cost or faster evaluation.
Effect of brief and “unsure”. We randomly sampled K questions and tested the stability of answers if we call ChatGPT to regenerate answers. When not requiring brief or “unsure” answers, for of questions, ChatGPT regenerated different answers. Adding the requirement for brief answers (Prompt 6 in Appendix A.1) reduced the percentage to , and further asking “unsure” answers with few-shot examples (Prompt 3) reduced the percentage to . In addition, according to manual evaluation on randomly sampled questions, removing “unsure” as an option increases ChatGPT’s hallucination rate by percentage points.
Robustness of prompts. We explore two other prompts. Compared with the original prompt that conducts few-shot learning (Section 3.1), denoted as Few-shot, the Zero-shot prompt does not provide examples and thus is zero-shot learning (Prompt 4 in Appendix A.1), and the In-domain prompt has the answerable example swapped out for an in-domain example generated by the same question template as the target question (Prompt 5 in Appendix A.1).
As shown in Table 7, Few-shot and Zero-shot show very similar results, but performance differences are noticeable between Few-shot and In-domain. In particular, in-domain examples help get more correct answers (, , in ALM for head, torso, tail) but at the cost of more hallucinations (, , in HLM for head, torso, tail). We suspect that the in-domain examples boost the confidence of ChatGPT in answering a question, so it answers questions even when the real confidence is not that high, causing both higher accuracy and higher hallucination rate.
Despite the fluctuation, our original prompt template (Few-shot) appears to be better at approximating the (confident) factuality of LLMs with the QA accuracy, and the relative performance among the head, torso, and tail remains stable over different prompts.
The experimental analysis indicates that although LLMs have incorporated factual knowledge within their parameters, the amount of this encoded knowledge remains limited. Knowledge of long-tail entities is already sparse in KGs and is even more deficient in LLMs.
Nevertheless, LLMs have been revolutionizing the way people seek information and calling for reconsideration of the best representation of factual knowledge. We term the forthcoming generation of KGs as Dual Neural KGs: knowledge can reside explicitly as triples (similar to KGs) and implicitly as embeddings (like in LLMs); the symbolic form caters to human understanding and explainability, while the neural form benefits machine comprehension and seamless conversations. A piece of knowledge can exist in both formats or in the one that is more appropriate. The harmonious blend of the two forms, capitalizing on the latest LLM innovations, is an exciting research area as we elaborate next.
Head knowledge. This involves popular entities where training data are ample. Ideally, LLMs could be taught such knowledge for efficient retrieval, meaning head knowledge shall exist in both forms. Currently, LLMs still have a low QA accuracy for popular entities (see Table 3.2), so a critical research area is to infuse head knowledge into LLMs through model training or fine-tuning. Early work in this line includes knowledge infusion Liu et al. (2021); Wang et al. (2021); Zhen et al. (2022).
Torso-to-tail and recent knowledge. This involves non-popular entities and emerging knowledge, where training data are typically sparse or absent. This type of knowledge might be best represented as triples. Serving such knowledge requires effectively deciding when external knowledge is essential, efficiently retrieving the relevant knowledge, and seamlessly integrating it into the answers. Early attempts in this direction involve knowledge-augmented LLMs Asai et al. (2023); Nakano et al. (2022); Shi et al. (2023); Borgeaud et al. (2022).
2 Limitations and extensions
Taxonomy. Our work does not discuss the effectiveness of LLMs in capturing taxonomy or type hierarchies, which could be an extension of this study. Specifically, we hypothesize that LLMs can effectively incorporate type relationships (e.g., hypernyms and synonyms), even for the fine-granularity sub-types. Hence, it may no longer be worth manually constructing a very deep and complex hierarchy in the future.
More recent LLMs. Although we performed the study using various representative models, we cannot exhaustively benchmark every recent model in this fast-moving field. Notably, we did not employ GPT-4 OpenAI (2023) and Llama 2 Touvron et al. (2023b), which became publicly accessible after we finished the first version of this paper.
Benchmarks. Most works studied the factuality of LLMs using existing QA benchmarks such as WebQuestions Berant et al. (2013), TriviaQA Joshi et al. (2017), LC-QuAD Trivedi et al. (2017); Dubey et al. (2019), QALD-9 Usbeck et al. (2018), Natural Questions Kwiatkowski et al. (2019), and EntityQuestions Sciavolino et al. (2021). A recent line of work has been constructing new QA benchmarks to assess LLMs’ factuality, especially for long-tail knowledge Mallen et al. (2023); Kim et al. (2023). Compared with these benchmarks, Head-to-Tail is the first to specifically assess how well LLMs incorporate head, torso, and tail factual information.
LLM Evaluation. Recent years have seen a proliferation of research on assessing the factuality of LLMs Roberts et al. (2020); Petroni et al. (2021); Shuster et al. (2021); Mielke et al. (2022); Tan et al. (2023); Hu et al. (2023); Peng et al. (2023a); Omar et al. (2023); Kandpal et al. (2023); Mallen et al. (2023). Most of these works focus on a single knowledge source, such as Freebase or Wikipedia, and they have yet to systematically perform the evaluation explicitly regarding head/torso/tail entities or attributes. One work close to ours is Omar et al. (2023), which evaluated ChatGPT using facts collected from diverse knowledge sources; however, their evaluation was carried out manually on only QA instances.
There are three works that also showed the correlation between the QA accuracy of language models and fact popularity Mallen et al. (2023); Kandpal et al. (2023); Kim et al. (2023). Our work, conducted in parallel, focuses on a different angle—how knowledgeable are LLMs? For this purpose, we systematically designed experimental methodology, including the definition of head, torso, and tail entities, the design of metrics, and the evaluation method. Our benchmark is comprehensive in containing different knowledge sources, different domains, and rich relations. Compared with these three works, we gave more quantified answers for research questions RQ1–RQ3.
We introduce Head-to-Tail, the first benchmark designed to assess the ability of LLMs to internalize head, torso, and tail facts. Alongside the dataset, we present a new evaluation methodology with appropriate metrics for automatically evaluating LLMs’ factuality. Our evaluation shows that even the most advanced LLMs have notable limitations in representing factual knowledge, particularly for the torso and tail entities. Accordingly, we suggest new research areas to seamlessly blend knowledge in the symbolic form and neural form.
Appendix A Appendix
A.2 Impact of less naturally occurring questions
When constructing Head-to-Tail, we include all predicates that allow reasonable factual questions. Table 8, instead, shows metrics on predicates that users are more likely to ask about. In general we observed higher performance on the Movie and Book domains, but the accuracy is still fairly low and we observe similar patterns regarding head, torso, and tail entities.