Do Language Models Know When They're Hallucinating References?

Ayush Agrawal, Mirac Suzgun, Lester Mackey, Adam Tauman Kalai

Introduction

Language models (LMs) famously hallucinateThough it is an anthropomorphism, we use the term hallucinate due to its widespread adoption, following the use-theory of meaning . We use the terms hallucinate and fabricate interchangeably., meaning that they fabricate strings of plausible but unfounded text. As LMs become more accurate, their fabrications become more believable and therefore more problematic. A primary example is “hallucinated references” to non-existent articles with titles readily fabricated by the LM. For instance, a real New York Times article entitled “When A.I. Chatbots Hallucinate” leads with a ChatGPThttps://openai.com/blog/chatgpt-fabricated New York Times article titled “Machines Will Be Capable of Learning, Solving Problems, Scientists Predict” . In this work, we study the problem of hallucinated references.

The hallucinated reference problem is worth study for multiple reasons. First, as we discuss, hallucinated references can easily be evaluated and debunked. Second, hallucinated references impact applications, as LMs help generate literature reviews for the exploration and citation of related work and may assist in writing of paper reviews . Third, due to the deployment of these models, the problem of hallucinated references has pushed beyond an academic curiosity to attract the attention of masses [e.g., 5, 25, 15, 23, 19] and has been highlighted as a problem in the medical domain where hallucinations could be extremely harmful. Finally, the insights gained from studying hallucinated references may apply to hallucination in domains beyond references.

A motivating question for this work is, why do LMs hallucinate, and what can be done about it? Is it a problem of LM representation, a problem of training (maximizing next-word likelihood), or a problem due to the way they are used for generation? Specifically, we investigate whether an LM itself can be used to detect whether or not an output it has produced is a hallucination, without any external resources. While this does not provide a complete answer to the questions of why and what to do, it does inform the discussion. In particular, to the extent that LMs can be used to detect their own hallucinations, this suggests that the hallucination problem is not inherently one of training or representation but is rather one of generation because the models contain enough information to at least reduce the hallucination rate.

In this work, by hallucinations we are referring to fabricated text that has little or no grounding in the training data. Note that this has been referred to as open-domain hallucination to distinguish it from closed-domain hallucination [see, e.g., 10], which is often studied in summarization and machine translation, where the fabrications are defined relative to a specific source document to be summarized or translated (as opposed to the training data). The two types of hallucinations are different in terms of what is often considered a hallucination: background information based on the training corpus is often defined to be a hallucination in the study of closed-domain hallucinations (assuming it is not in the source document, e.g., the text to be translated). However, open-domain hallucination has attracted significant recent attention within scientific communities and journalism. In this work, when we refer to hallucinations we are referring to absolute (i.e., open-domain) hallucinations.

Groundedness versus correctness The opposite of fabrication is groundedness in the sense of being based on the training corpus rather than accuracy in the sense of being a true fact (a genuine publication, in the case of references). We define hallucination to be fabricated text, meaning text that is not grounded in this training set. In contrast, correctness is evaluated with respect to ground-truth answers. This distinction is called honesty versus truthfulness by Evans et al. . For example, the common misconception that “people use 10% of their brains” is grounded because it is almost surely mentioned in the training data, either exactly or in various paraphrased versions. However, it is not scientifically correct. Much work on hallucination conflates groundedness and accuracy, often equating hallucination with fallacy and evaluating hallucinations using accuracy on fact-based assessments, without regard to the training data . We adopt the groundedness definition of hallucination even though it may often be less clear-cut and more difficult to evaluate than factuality.

Evaluating groundedness Perfectly evaluating hallucinations would require access to the LM’s training data. An advantage of the hallucinated reference problem is ease of (approximate) evaluation in that exact-match Web search is a reasonable heuristic for groundedness. This is because the vast majority of article titles present in the training data are included in Web search results—articles are meant to be published and shared, and publishers aim to make their work discoverable by search. Furthermore, references generally have titles that are specific enough not to spuriously occur on the Web. Regarding other types of hallucinations, besides article names, which cannot be as easily evaluated, we still hope that our methodology and findings would apply, even if evaluating those types of hallucinations would require access to the training data.

Direct queries Our work builds upon and is inspired by two recent works that show how to use black-box generative LMs to assess confidence in generations, without consulting external references or inspecting weights. In particular, Kadavath et al. introduce multiple direct black-box strategies for using an LM to extract confidence estimates by querying the language models on question-answer problems. Manakul et al. apply a similar direct self-consistency check called SelfCheckGPT to identify relative hallucinations in a summarization context. These queries are direct true/false correctness queries. We test similar approaches in the context of hallucinated references. Black-box generative approaches stand in contrast to the work that either introspects the weights on LMs or that consults existing databases .

Indirect queries In addition, we suggest a new approach using what we call indirect queries. A direct query may ask, Is the following paper real? while an indirect query may ask, Who are the authors of this paper?, as illustrated in Fig. 1. Answers are then generated to the indirect query in i>1i>1 independent sessions, and tested for consistency. The motivation for indirect queries comes from investigative interviews, where detectives are advised to interview individuals separately and ask open-ended questions. For instance, consistency may be better evaluated by asking multiple witnesses to “Describe in detail what the suspect was holding” rather than asking, “Was the suspect holding a gun in their right hand?” . In the context of reference hallucination, our hypothesis is that the likelihood of multiple generations agreeing on the same authors for a hallucinated reference would be smaller than the likelihood of multiple responses to a direct query indicating that the reference exists.

Contributions There are several contributions of this work. First, we perform a systematic LM study of hallucinated references, enabling us to compare hallucination rates across LMs. Second, we introduce indirect queries for evaluating hallucinations. Third, we compare these to direct queries, inspired by studies in LM question-answering and summarization-based hallucinations . A conclusion of our work for reducing hallucination is the recognition that changing the generation pipeline can certainly help, while it is less clear if training or representation changes are necessary.

Related Work

Open-domain hallucination were discussed in the context of GPT-4 , due to their prevalence and potential danger, Bubeck et al. [4, page 82] write:

Open domain hallucinations pose more difficult challenges, per requiring more extensive research, including searches and information gathering outside of the session.

We show that open domain hallucinations can in fact be addressed, at least in part, without consulting external resources.

As mentioned, there are multiple definitions of hallucination. In this work, we use the term hallucinations to mean fabricated text that is not grounded in the training data. Factually incorrect generations can be decomposed into two types of errors: grounded errors which may be due to fallacies in the training data (e.g., that people use only 10% of their brains) and ungrounded errors. These two types of errors may need different techniques for remedy. The grounded errors may be reduced by curating a training set with fewer errors or other techniques such as RLHF . However, the ungrounded errors which we studyOne can also imagine ungrounded correct generations, such as a generated paper title that exists but is not in the training data, but we find these to be quite rare. are a fascinating curiosity which still challenge the AI community and one which is not clearly addressable by improving training data. The distinction is further elucidated by Evans et al. .

There is comparatively little prior work studying open-domain groundedness like ours. Some work [e.g., 9] in attribution aims to understand which training examples are most influential in a given output. In recent independent work in the health space, Athaluri et al. did an empirical evaluation of hallucinated references within the medical domain. Similar to our approach, they used a Google search for exact string match as a heuristic for evaluating hallucinations. Our study of hallucinated references enables us to estimate the hallucination rates of different models, and, as discussed in prior work, the hallucination problem interestingly becomes more pressing as models become more accurate because users trust them more .

Related recent works include black-box techniques for measuring confidence in LM generations. Although these works are targeted at factual confidence, the approaches are highly related to our work. While Kadavath et al. use probability estimates drawn from LMs, it is straightforward to extend their procedures to generation-only LMs like ChatGPT using sampling. Lin et al. show that LMs can be used to articulate estimates by generating numbers or words as we do. Finally, Manakul et al. perform self-checks in the context of summarizing a document. All of these works use direct queries which influenced the design of our direct queries.

Due to space limitations, we do not discuss the work studying closed-domain hallucination (e.g., in translation or summarization) but instead refer the reader to recent survey of Ji et al. .

Methodology

We now give an overview of our methodology followed by further details on our direct and indirect queries. Note that this full pipeline is run separately for each of our LMs, so there is no mixing across LMs. We first describe how we generate lists of candidate reference titles.

Generating references The input to our evaluation is a set of topics from which we generate kk references each using the LM by prompting it with temperature 11 as illustrated in Fig. 2. The procedure is re-run if the LM fails to generate a list of kk candidate titles. We then run our classification procedures, described below, on each of the candidate titles.

Hallucination estimation procedures Each of our procedures takes three inputs:

A candidate reference title. Given that there is generally less ambiguity in the title of a reference than in the spelling or abbreviation of its authors names, for each reference we chose to use only its title as input.

A black-box LM capable of completing a prompt. This is the most general model which includes dialogue-based models, such as ChatGPT that offer an API without probabilities.

A number of queries. This parameter is slightly different for direct and indirect queries.

Direct queries: parameter j≥1j\geq 1 which determines how many judgments to make.

Indirect queries: parameter i≥1i\geq 1 determining how many indirect responses to request.

In our experiments, the candidate title will have been generated using the LM, though this is not a requirement. The procedure detects (possibly) hallucinated references by querying the LM to check the existence of the reference. It does so by making black-box completion queries to the same LM. Finally, the procedure outputs a real-valued prediction in $oftheprobabilitythetitleisgrounded(G)orahallucination(H).Weconsiderbothperformingasinglejudgmentof the probability the title is grounded (G) or a hallucination (H). We consider both performing a single judgmentj=1perpapertitleandper paper title andj>1$ to implement a version of the procedure that outputs probabilities rather than just G/H judgments. Since we do not have access to the probability distribution of the LM completions for all models, the above procedure effectively simulates probabilities using sampling at temperature 1. (Note that each query is run independently “from scratch” in a new prompt; one would expect an artificially high degree of consistency if one were to ask the same query repeatedly within a single dialogue.)

Labeling For labeling, we use exact match in a search engine as a heuristic for labeling G/H. These labels are treated as ground truth (though like all labels they have some error and ambiguities). Final receiver operating characteristic (ROC) curves and false discovery rates (FDR) are determined by comparing the ground truth labels to the classifications. Note that we also experimented with academic reference APIs such as Semantic Scholar. While these gave thorough details about each paper in its index, many grounded references (even for real papers) did not appear in their indexes, and we found search engine results to be significantly more complete.

The direct query (DQ) procedures simply query whether or not the given title exists following the format shown in Fig. 3. We created three query templates (DQ1, DQ2, and DQ3) based on the multiple direct query approaches advocated by Kadavath et al. , Manakul et al. . The first query asks whether the reference exists directly. However, as discussed in prior work, some LMs can be strongly biased in answering the question when phrased this way, e.g., it may be presumed real without any context about where the reference came from. DQ2 and DQ3 establish the context indicating that the reference was generated by an assistant or LM. DQ3 goes further by giving additional comparisons, as advocated for in prior work. For DQ3, all kk queries from our generation step (using the same LM) are shown.

For each query, we generate j≥1j\geq 1 completions to approximate the probability distribution of the model. These strings are converted to binary judgements as follows: We calculate how many completions contained the word yes and divide it by the total number of completions to get the estimates of groundedness. This means that empty or otherwise invalid answers were assigned no. We do not assume that this score is calibrated as our analysis considers arbitrary probability thresholds.

We sampleFor models that support probability computations, they could be used directly for greater accuracy and efficiency. However, for uniformity, since models such as ChatGPT that we employ does not offer probabilities, we employ sampling. jj completions for each direct prompt. Temperature 11 is used when j>1j>1 and temperature is used when j=1j=1 to approximate the most likely LM completion.

2 Indirect queries details

The indirect queries proceed in two steps.

Step 1: Interrogation Separately for each reference, an indirect query is made of the LM i>1i>1 times at temperature 1, as shown in Fig. 4 (top). Responses were truncated to 100 characters.

Step 2: Overlap estimation. The LM is used to evaluate overlap between the ii responses. For each pair of answers, an estimate is computed by calling the overlap query, as shown in Fig. 4 (bottom). The leading number is extracted, or, if no number is given, then a 0 is used. (We divide by 100 and clip the answer to the interval $toconvertthepercentagestofractions.)ItisworthnotingthatLMsmayreturnananswerthatdoesnotconsistofalistofauthors,suchasalongresponsebeginningwith“Icouldnotfindaspecificreferencetitledto convert the percentages to fractions.) It is worth noting that LMs may return an answer that does not consist of a list of authors, such as a long response beginning with “I could not find a specific reference titled\ldots$”. Thus the overlap estimation prompt clarifies that an answer of 0 should be given if either response is not a list.

The rationale for this approach is that we expect consistent responses to indirect questions to indicate the existence of a grounded reference title, while inconsistent responses may be taken as an warning sign for hallucination. Our method does not rely on external resources and uses the same language model for hallucination detection end-to-end. Of course, parsing and string-matching could be used in place of a LM for the overlap step, though this would require name matching which is known to be a thorny problem and one which is well suited for pretrained LMs.

3 Ground Truth Labelling

A Web search the reference title surrounded by quotes (e.g., “Language models are few-shot learners”) using web search. If no results are retrieved, we label the reference title as hallucinated and vice versa. We perform a manual inspection of results to determine the efficacy of this proxy for groundedness of reference titles.

Results and Discussion

In this section, we describe our experiment details, discuss the performance of the indirect and direct methods using quantitative metrics, and present interesting qualitative findings. The code and data generated in our experiments will be made available upon publication.

Models We use the Azure OpenAI APIhttps://azure.microsoft.com/en-us/products/cognitive-services/openai-service for our LMs. We use the three most powerful models, GPT-3 (text-davinci-003), ChatGPT (gpt-35-turbo), and GPT-4 (gpt-4), for evaluation and generating the datasets. We also experimented with smaller models, but the accuracy with these models was extremely poor, as in the work of Kadavath et al. . As can be seen in our results, even the performance of the GPT-3 model was of limited accuracy.

Topics We use the ACM Computing Classification Systemhttps://dl.acm.org/ccs (CCS) for topics. CCS contains 12 high level categories, 84 second level concepts, and 543 subconcepts at the third level of granularity. For generating the dataset, we sample 200 of the 543 subconcepts uniformly at random, describing each by a topic string of the form concept: subconcept (e.g., Information retrieval: Retrieval models and ranking). For each topic, we generate k=5k=5 references. In this manner, we generate 200×5=1000200\times 5=1000 candidate paper titles using each LM.

Parameters We selected i=3i=3 indirect query results and took the average of the overlapping evaluations to compute the final score for each indirect query experiment. For direct query experiments, we sampled j=10j=10 judgments at temperature 1.0 and reported the fraction of yes responses as a final score.

Search engine labels The Bing search engine APIhttps://www.microsoft.com/en-us/bing/apis/bing-web-search-api is used for searching for the candidate title string on the Web. Note that even with exact string match, some flexibility beyond capitalization and punctuation is allowed. A manual inspection of 120 random examples, given in Appendix C, finds the use of the search engine to be a reliable method for detecting hallucinations.

2 Quantitative metrics

First, Table 1 shows the rates of hallucination for the three models studied. As expected, references produced by the newer models (which achieve higher scores on other benchmarks ) also exhibit a higher grounding rate or, equivalently, a lower hallucination rate. While this is expected, it is a positive indicator of the validity of our approach of using search engine results to measure hallucination. (This is discussed further in Appendix C.)

Since each of our querying strategies outputs a real-valued score, one can trade-off accuracy on G (i.e., how often truly grounded references are labeled G) and H (how often truly hallucinated references are labeled H) by thresholding the score to form a G or H classification. The standard receiver operating characteristic (ROC) curves based on these thresholded scores are shown for each approach and model in Figs. 5(b), LABEL:, 5(c), LABEL:, and 5(a). These figures enable one to explore different points on this trade off for each classifier. For the GPT-3 and ChatGPT models, the IQ procedure performs best as quantified via the area under the ROC curve (AUC). For GPT-4 (Fig. 5(c)), both the IQ and DQ approaches work well for classifying hallucination and groundedness with the IQ (AUC: 0.8780.878) and DQ1 (AUC: 0.8870.887) performing the best. The performance of each procedure generally improves as the model size increases. For smaller models, where the procedures perform worst, others have found that users are less likely to believe the generated text due to its inaccuracy . We additionally display 95%95\% confidence bands for each ROC curve using 100100 bootstrap replicates and a 95%95\% confidence interval for the AUC using the DeLong et al. estimate of AUC standard error.

Each groundedness classifier can also be used as a filter to generate a list of likely grounded references for a literature review based on the raw generations of an LM. Aside from relevance, which we do not study in this work, two primary quantities of interest to a user of this filter would be the fraction of references preserved (more references provide a more comprehensive review) and the fraction of preserved references which are actually hallucinations. Fig. 7 shows how these two quantities can be traded off. As one varies the threshold of G/H classification and returns only those references classified as grounded, the false discovery rate (FDR) captures the fraction of references produced which are hallucinations. Users may have a certain rate of tolerance for hallucinations, and one would like to maximize the number of generated references subject to that constraint. For GPT-3 and ChatGPT, the IQ method achieves significantly lower FDR and a provides a substantially better FDR-preservation rate trade-off than the other approaches. For GPT-4, both IQ and DQ methods offer low FDR with comparable trade-offs. Fig. 7 also displays 95% FDR prediction intervals (lighter bands) computed from the quantiles of 100100 bootstrap replicates and 95% confidence intervals (darker bands) for expected FDR computed from the bootstrap mean ±1.96\pm 1.96 times the bootstrap standard error.

Overall, our hypothesis that indirect queries would be more reliable than direct queries appears to hold for ChatGPT and GPT-3; for GPT-4 the direct queries were similarly effective. Finally, we now observe that one can improve accuracy for all models using a combination of direct and indirect queries.

We find that classification performance increases when we take ensemble of different approaches, as illustrated by ROC curves in Fig. 6. For creating the ensemble of the approaches, we simply compute the mean of the scores and use them as thresholds. The ensemble of three direct query approaches (computed using the equally-weighted mean of the DQ1, DQ2, and DQ3 scores), which we refer to as simply DQ, performs slightly better than the best performing direct query approach. The ensemble of IQ and DQ (computed using the 50-50 mean of IQ and the DQ mean), referred to as IQ+DQ performs the best for every model.

The compute costs, which involve ≈\approx6.6 million tokens and $412, are discussed in Appendix B.

3 Qualitative findings

A qualitative examination of the titles generated by the LMs and their classifications according to the Bing search API revealed several interesting observations:

Many hallucinated titles were combinations of multiple existing titles.

The Bing quoted search heuristic is more lenient than exact match, ignoring more than just capitalization and punctuation. However, presumably since Bing quoted search is designed to facilitate title searches, it works well.

Some hallucinations were “plausible sounding” such as A survey on X for topic X, even when such a survey did not exist.

Direct methods may fail to identify hallucinations on “plausible sounding” titles such as surveys or book chapters. The indirect method also sometimes failed to identify a hallucination because the LM would consistently produce a “likely author” based on the title, for a given non-existent paper. For example, GPT-4 hallucinated the title Introduction to Operations Research and Decision Making, but there is a real book called Introduction to Operations Research. In all three indirect queries, it hallucinated the authors of the existing book, Hillier Frederick S., Lieberman Gerald J.. Similarly, for the hallucinated title Exploratory Data Analysis and the Role of Visualization, 2 of 3 indirect queries produced John W. Tukey, the author of the classic, Exploratory Data Analysis.

The indirect method may sometimes fail to identify a grounded paper title which it can recognize/generate, as it may simply not be able to generate authors not encoded in its weights.

Since, in many applications, identifying potential hallucinations is more important than recognizing all grounded citations, errors due to falsely marking an H as a G are arguably more problematic than classifying a G as an H. A manual examination of 120 examples is given in Appendix C.

Conclusions, Limitations, and Future Work

This work investigates the hallucinated reference problem in LMs and provides a methodology by which LMs can be used for self-detection of hallucinations. Both direct and indirect queries were found to be effective for language models, and combining multiple methods led to further improvements in accuracy.

There are several limitations of this work. First, as discussed earlier, because we used LMs with inaccessible training data, we cannot conclude what is truly grounded versus hallucination. Second, while we consider a binary notion of hallucination in this work, as is done in much prior work, the notion of hallucination is not entirely black and white. Third, LMs are notoriously sensitive to prompt wording, and some of our findings comparing direct and indirect queries may be sensitive to the specific wording in the prompt. Since we use ACM Computing Classification System for our topics, the results are heavily biased towards computer science references, though it would be straightforward to re-run the procedure on any given list of topics. Also note that LMs have been shown to exhibit gender and racial biases which may be reflected in our procedure–in particular: our procedure may not recognize certain names as likely authors, or it may perform worse at matching names of people in certain racial groups where there is less variability in names. Since our work compares LMs and hallucination estimation procedures, the risk is lower compared to a system that might be deployed using our procedures to reduce hallucination. Before deploying any such system, one should perform a more thorough examination of potential biases against sensitive groups and accuracy across different research areas.

There are several directions for future work. Of course, an important consequence of our work is the recognition that reducing hallucination may be a problem at generation time. Thus, inventing improved (non-black-box) generation procedures is thus a crucial direction for future work.

There are also several more immediate ways in which our work may be extended. First, one may improve accuracy by adding more indirect questions such as year or venue. These pose additional challenges as a paper with the same title and authors may often appear in multiple venues (e.g., arXiv, a workshop, a conference, and a journal) in different years. Second, it would be very interesting to see if the methods we employ could be used to identify other types of open-domain hallucinations beyond references. Even though hallucinated references are often given as a blatant example of hallucination, perhaps due to the ease with which they can be debunked, these other types of hallucination are also important. Following the investigative interviewing analogy, one way to aim to discover general hallucinations would be to query the LM for “notable, distinguishing details” about the item in question. One could then use an LM to estimate the consistency between multiple answers. However, as mentioned for other domains besides references, it may be impossible to determine whether or not a generation is a hallucination without access to the training set (and unclear even with such access).

In summary, open-domain hallucination is an important but slippery concept that is difficult to measure. By studying it in the context of references using search engine results, we can quantitatively compare hallucinations across LMs and we can also quantitatively compare different black-box detection methods. Of course, for the sole purpose of detection, one could achieve higher accuracy by directly consulting curated publication indexes. However, we hope that our study of black-box self-detection of hallucinated references sheds light on the nature of open-domain hallucination more broadly, where detecting hallucinations is more challenging. It suggests that hallucination is not entirely a problem of training but rather one that can be addressed using only the same internal model representation with different generation procedures. While our direct and indirect query methods are only partially reliable and impractically expensive, we hope they may pave the way towards more efficient methods that generate text with fewer hallucinations and thereby reduce potential harms of language models.

References

Appendix A Licenses and Terms of Use

According to the OpenAI terms of use Sharing and Publication policy,https://openai.com/policies/sharing-publication-policy they “welcome research publications related to the OpenAI API.” Following the Bing Search API Legal Informationhttps://www.microsoft.com/en-us/bing/apis/legal, we do not store the results of the search queries but rather only whether or not there were any results. According to the ACM,https://www.acm.org/publications/class-2012 “The full CCS classification tree is freely available for educational and research purposes.” (This section will be included with any published version of our paper.)

Appendix B Computation and cost

We use OpenAI API for running the experiments on GPT-4, ChatGPT and GPT-3. We show the average tokens consumed for prompt and completion for each of the approaches and data generation per candidate query in Tables 2, 3, and 4. We estimate the cost based on the pricing details available as of May 2023.https://openai.com/pricing For GPT-4, around 2.2M tokens were used amounting to roughly 74toevaluateallapproaches.ForChatGPT,around2.3Mtokenswereusedamountingtoroughly74 to evaluate all approaches. For ChatGPT, around 2.3M tokens were used amounting to roughly5. For GPT-3, around 2.1M tokens were used amounting to roughly 258.ForBingSearch,weuseanS1instanceoftheBingSearchAPIhttps://www.microsoft.com/en−us/bing/apis/pricing.Wemade3,000queriesinalltothisendpointamountingto258. For Bing Search, we use an S1 instance of the Bing Search API https://www.microsoft.com/en-us/bing/apis/pricing. We made 3,000 queries in all to this endpoint amounting to75. Summing these costs gives a total of $412. The compute requirements of combining these results were negligible. While the exact model sizes and floating point operations are not publicly available for these models, the total cost gives a rough idea on the order of magnitude of computation required in comparison to the hourly cost of, say, a GPU on the Azure platform.

Appendix C Examples of hallucinations and references

Tables 5, 6, 7, and 8 each display a careful inspection of 30 random candidate paper titles classified as H and G as determined by whether the Bing Search API returned any results. A manual search for each suggested title indicated that the vast majority of Hs are in fact hallucinations and the vast majority of Gs are in fact real references. We show the titles classified as H by Bing search along with closest manually discovered match for ChatGPT (Table 5) and GPT-4 (Table 7). We show the titles classified as G by Bing search along with the web links to the matched titles for ChatGPT (Table 6) and GPT-4 (Table 8). We also list the score assigned by the IQ method for all the sampled candidate titles. Interestingly, for both models there was a case in which the IQ method assigned the score of 1 to an H title. These H titles were Design and Implementation of Digital Libraries: Technological Challenges and Solutions for ChatGPT (Table 5) and Enterprise Modeling: Tackling Business Challenges with the 4EM Approach for GPT-4 (Table 7). In both of these cases, the titles were very similar to the closest manually discovered matched titles - Design and Implementation of Digital Libraries and Enterprise Modeling with 4EM: Perspectives and Method, respectively.