Origin Tracing and Detecting of LLMs
Linyang Li, Pengyu Wang, Ke Ren, Tianxiang Sun, Xipeng Qiu
Introduction
Using LLMs such as ChatGPT and GPT4 for various daily routines and copilot in work is becoming a new trend that draws worldwide attention not only in the machine learning community. Starting from GPT and BERT , pre-trained models have developed for several years, and trustworthy and security concerns have been constantly discussed . The performances of LLMs are sensational, therefore, the usage of LLMs should be strictly supervised by users as well as service providers. One trend in ensuring the safety of LLMs is to build detection tools that can discriminate whether an AI system generates a certain text . AI-generated context detection is useful in releasing texts that require strict censoring or originality such as official documents, consultation, and student submissions to avoid abuse of AI systems. Further, a more critical and applicable field is to trace the origin of LLMs. While more and more companies and institutions are releasing their original LLMs, it is of great importance to trace whether an LLM is trained from a previous model, or is copied or distilled from another LLM. Since LLMs can produce massive generated data, future LLMs might be trained from these generated data, the human-written texts might be contaminated with different LLMs. Therefore, tracing the origin of the text is a major challenge in future LLM industries.
In this work, we first introduce the concept of origin tracing. Then we provide an effective tool named Sniffer and its evaluation benchmark to study the origin tracing problem.
Origin tracing is to further categorize the origin of a given text. Specifically, we categorize the origins by the LLM abilities and by LLM service providers. We first trace the origin of GPT2 level LLMs such as GPT2 from OpenAI, GPT-Neo/J from EleutherAI, we can also trace the GPT3 level LLMs such as text-davinci-003, ChatGPT (GPT3.5 turbo) , LLaMA . With origin tracing, we can avoid AI abuse or potential model theft since the texts can be traced to a specific service provider. While more and more companies and institutions are releasing their LLMs, origin tracing is the keystone of anthropology of LLMs. Then we introduce Sniffer, the first origin tracing tool. In Sniffer, we use contrastive features across open-source LLMs such as GPT2, GPT-Neo/J, and LLaMA. Specifically, we design heuristic features that capture the model-wise discrepancies which can help trace the origin of given texts. Then we utilize a simple linear classifier to project the extracted features to specific origins including known origins that have open-source models and unknown origins that are black boxes to users. The core motivation of the origin tracing tool Sniffer utilizes the discrepancies between LLMs as features to help trace the origins. Compared with previous methods, Sniffer utilizes model-wise features, which is different from supervised learning methods such as fine-tuning a RoBERTa model; further, Sniffer is able to trace text origins and generalize to unknown origins, while previous methods that use model-wise features cannot.
With Sniffer, we further introduce a test dataset that contains collected texts from different origins as a benchmark to study the origin tracing problem. We provide plenty of experiments and through the experimental results, we have several non-trivial observations that help future LLM studies. In general, we find that: (1) we are able to trace the origins of generated texts when we can possess the models; (2) it grows harder to detect and trace origins when the LLMs are stronger; (3) LLM providers need to be more cautious as the origin of generated texts might be known only by the providers. (4) We can trace the origins of distilled LLMs such such as Alpaca and Dolly.
To summarize, in this paper, we: (1) raise the concern of origin tracing of AI-generated contexts; (2) build a feature-wise tool to trace the origins of various open-source models by releasing a diversified benchmark for AI-generated contexts detection and origin tracing; (3) conduct various experiments to analyze the ability of LLMs when they are being traced and hope that future works can pay more attentions to the origin tracing of LLMs.
Related Work
Pre-trained models are supposed to be harmless, harmful, and honest to users , however, there are various aspects that challenge LLM securities such as social bias, stereotypes, privacy leak or adversarial examples .
Detection of AI-generated contexts is also a rapidly growing field that requires attention since the abuse of AI might be a major challenge in LLM applications . Still, current methods only consider detection as a binary task, that is, whether a given text is generated by an AI without considering tracing the origins of texts. Detection methods can be categorized into two lines:
The most straightforward method to detect AI-generated contexts is to construct a text classification task . Therefore, DetectGPT introduces a strong RoBERTa-trained baseline that uses the dataset released by OpenAI https://github.com/openai/gpt-2-output-dataset to train a classifier.
Model-wise Detection Unlike semantic-wise detection which discriminates the semantic difference between human-written and AI-generated texts, a more direct way is to explore the model-wise features. That is, AI-generated texts show discrepancies compared with humans when fed into AI models. These discrepancies are not easily noticed by humans as they might be a subtle difference that is possessed by a certain origin, therefore, many works focus on utilizing model-wise features such as log-likelyhood of model outputs, neural features, or bag-of-word features to detect AI-generated contexts.
In the era of LLMs, there are similar works that work on using model-wise features of LLMs such as watermarking , and backdoor plantings . In the computer vision field, model-wise features are also used in fake image detection but are less related to LLM origin tracing due to the continuous nature of images.
Methods
We aim to detect whether a context is generated by an LLM system and trace the origin of the texts. Therefore, we design a simple method Sniffer This name is inspired by the fact the sniffer dogs are able to trace the scents that cannot be easily noticed by humans. that is applicable in both white-box and black-box settings and only requires limited supervised data.
The core idea is to utilize the contrastive features between different accessible language models such as GPT-2, GPT-Neo, GPT-J, and LLaMA. We first obtain the perplexity of a target sample based on different models denoted as , then we craft several heuristic features and construct a simple linear classifier to classify the origin of the given sample. Through such a feature engineering process, we can trace the origin of the target sample down to a known model , we can also generalize the contrastive features to show differences between unknown source models or human-written texts.
Compared with previous detection methods, as seen in Table 1, our proposed method is the first to allow origin tracing. Model-wise detection methods such as DetectGPT are designed to detect whether a text is generated by a certain model, which can not be well generalized to origin tracing. Supervised learning methods such as RoBERTa finetuning, on the other hand, require a large amount of training data and are fixed to pre-defined labels, which is also limited when LLMs are developing drastically.
The process of Sniffer includes: (1) Obtain and align token-level perplexity between different models; (2) Extract contrastive features; (3) Train features for origin tracing.
Given target text and a known model , we obtain the encoded tokens , the perplexity of token in text given model is the log-likelyhood . Given a list of known models , we obtain a list of perplexities of the same text . Since the tokenization process of each LLM might be different, we use a general word-level tokenization of : and align calculated perplexities of tokens in to the general words. If the word is aligned to multiple tokens in , we use the averaged perplexity; if a token in is aligned to multiple words, we assign these words with the same value. The details of the aligning process can be seen in the Appendix.. Further, the aligned perplexity is conditioned on the specific model-training process, therefore, the perplexity should be normalized for comparison between models. We apply normalization strategies including dataset-wise-normalization and L1-normalization for these aligned features. Dataset-wise normalization is to normalize the perplexity with an averaged perplexity on all data. Therefore, as seen in Figure 2, we obtain lists of word-level perplexities that are aligned across models.
After obtaining aligned perplexities , which is a list of token-wise perplexities, we aim to search features within the list of token-wise perplexities. Instead of introducing assumptions of the difference between human-written texts and machine-generated texts , the core idea of Sniffer is to find different features between models. Therefore, we pose a simple hypothesis:
Hypothesis 1 A human written text tends to have a similar perplexity list across models, and a generated text tends to show discrepancy across models.
To further generalize the extracted features for tracing unknown models, we pose another hypothesis:
Hypothesis 2 The discrepancy of perplexity curves across models can reveal features of both known and unknown models.
With these hypotheses, we can utilize the contrastive features between known models to trace the origin of various models denoted as that includes known and unknown models.
Given known models, we have pairs and we can obtain scores. We collect all these scores as heuristic features that can be used to trace the text origins. Plus, we also include the sentence-level perplexities and the Pearson and Spearman correlation coefficient between and to construct a feature vector as the final representation for origin tracing. For instance, given 4 known LLMs (), we have 4 sentence-level perplexity features, pct-scores, and correlation coefficient score, therefore the final representation vector is a 22-dim vector.
After extracting the contrastive features from different known models, we can train a simple linear classifier to project the extracted features to different model origins. To train the classifier , we collect a small amount of human-written texts that cover various aspects and use different known models to generate texts as similar or parallel data compared to human-written texts to train the linear classifier.
The trained classifier has several unique features:
(1) Low-resource required: Unlike semantic-wise classification tasks that require a large amount of data and a strong natural language encoder, after feature extraction, the final representation is a low-dimension vector, which requires only a small amount of data since the feature contains abundant model-wise heuristic knowledge for origin tracing;
(2) Generalization ability to unknown models and stronger LLMs: In the classifier studying process, we can collect unknown-model-generated texts to obtain their features based on known models and these features can reveal different traces compared with known models and human-written texts. Different from supervised-learning methods that rely on semantic-level features to detect text origins, model-wise features are NOT influenced as the quality of generated texts improves. That is, in the supervised learning methods, stronger LLMs are harder to detect since the generated texts are human-like, but model-wise features cannot be easily optimized.
(3) Extend Ability: In our proposed method, the linear classifier can be easily modified to trace various model origins. Given a new open-source model, we can easily use the feature extraction strategy above to re-train the classification model; given a new unknown model, we can easily collect a few generated texts and extract the features based on known models, and train a new classifier with the unknown model origin features to trace the new unknown origin. With such extended ability, our proposed method can be used in both black-box and white-box settings.
Experiments
In the origin tracing of LLMs, one major challenge is that LLMs are almost omniscient to world knowledge since they are trained with various and huge amounts of data. We collect a wide range of texts from different origins for the proposed origin tracing tool Sniffer to serve as a general detector.
We collect texts from domains including News articles, social media posts, web texts, scientific articles or academic papers, and technical documentation. We use public datasets including XSum dataset that contains news articles; IMDB dataset that contains social media reviews of movies; web texts that contains texts from common-crawled online pages; PubMed and Arxiv dataset that contains academic topics; Wikipedia corpus used in SQuAD dataset that contains general world knowledge. For each dataset, we randomly collect 1,000 documents and shuffle them into a human-written text dataset containing 6k documents. Then we generate AI-generated contexts from the texts we collected as parallel data to study origin tracing.
When generating texts from language models such as GPT-2, we use the first 10 words as the prompt to generate a document. When generating texts from instruction-tuned models such as GPT-3.5(text-davinci-003) and ChatGPT (turbo), we give several instructions including re-write instruction and story generation instruction to obtain a similar AI-generated text. (We show instruction details in the Appendix.) With these simple instructions, we collect AI-generated texts from instruction-tuned LLMs.
For known origin models, we collect AI-generated texts from GPT2 (powered by OpenAI https://openai.com/), GPT-J and GPT-Neo (powered by EleutherAI https://www.eleuther.ai/), LLaMA (powered by Meta AI https://ai.facebook.com/blog/large-language-model-llama-meta-ai/). For unknown origin models, we collect AI-generated texts from GPT3.5-text-davinci-003, which is an instruction-tuned model with 175B parameters. Therefore, we collect 6k human written texts, and every 6k texts from GPT2, GPT-Neo, GPT-J, and 12k texts from GPT3 models, which is 36k texts in total. We divide the 36k texts into a train/test split with a 90%/10% partition.
We name the collected dataset SnifferBench, which can be further used in origin tracing and AI-generated contexts detection tasks. Further, the dataset collection process can be extended to different LLMs in the future with more different scenarios and prompts/instructions which further challenges the origin tracing ability.
2 Baseline Methods
As we compare Sniffer with previous detection methods in the form of the methods, we construct experiments to explore how previous methods fail to run origin tracing.
We first implement a method which is also used in the original version of GPTZero https://gptzero.me/; then we implement the DetectGPT method which introduces a discrepancy score to discriminate AI-generated texts by adding multiple perturbations (we try 40 perturbations for each sample). We collect the sentence-level perplexity and the DetectGPT discrepancy score of the collected dataset based on a certain known model and draw a histogram showing the score distributions of different origin texts. Then we select a threshold manually as the discrimination boundary in and DetectGPT. In Figure 3, we plot the histogram of the discrepancy between different texts’ origins in method and DetectGPT method. In , we plot the perplexity score of different text origins using a specific model such as GPT-2 or GPT-Neo, and in the DetectGPT method, we use their proposed z-score. As seen in Figure 3, though there are multiple peaks showing that there are differences between different text origins, the overlap is too large to successfully separate different text origins. we can conclude that it is extremely difficult to discriminate text origins from features or z-score features used in DetectGPT. Therefore, it is important to introduce strong features to trace the texts’ origins.
3 Implementations of Sniffer
In Sniffer, we select several open-source (L)LMs as known models: we use GPT2-xl(1.5B), GPT-Neo(2.7B), GPT-J(6B) and LLaMA(7B) as known models. In the SnifferBench, we collect texts from origins including GPT2(OpenAI), GPT-Neo and GPT-J(EleutherAI), LLaMA (MetaAI), ChatGPT(GPT3.5-turbo from OpenAI), and human-written texts, therefore, the unknown origin is the ChatGPT(GPT3.5-turbo) model since they are not open-source models. The goal of origin tracing in Sniffer is to trace both known and unknown origins. We construct an inference server for each known model based on NVIDIA4090 GPUs and set the max sequence length given the maximum GPU allowance. We align all texts with a white-space tokenizer to obtain uniform tokenizations.
Besides Sniffer which uses 4 known models and utilizes a linear classifier for origin tracing, we introduce a new variant of Sniffer: Sniffer(+GPT3) As we are able to obtain logits of the generated texts in the GPT-3 API provided by OpenAI, we are able to treat GPT-3 model as a white-box model. Therefore, we collect a subset in the SnifferBench to test the origin tracing with GPT-3 models as white boxes. Therefore, in our implementation of Sniffer, we have 5 models in total which result in a 35-dimension vector in obtaining the sniffer feature. Here, we align the tokens with the tokenizer provided by OpenAI.
4 Metrics of Origin Tracing
In the SnifferBench, we collect texts from different origins, therefore, we calculate the precision and recall of each text origin tracing prediction. We can only calculate the known model and the human-written texts’ precision and recall in baseline methods since these methods cannot be used in the origin tracing task.
5 Results of Sniffer Origin Tracing
In Table 2, we can observe that model-wise detection methods such as perplexity-based method and perturbation-based method DetectGPT cannot be used in origin tracing task. When predicting whether the texts are generated by the specific model GPT-2, and DetectGPT can obtain reasonable results but fail to discriminate the mixture of texts generated by EleugherAI-released models including GPT-J and GPTNeo, not to mention that these methods cannot discriminate other models.
We can observe that Sniffer can obtain impressive results in tracing the text origins in the known models including GPT-2, GPT-J/Neo; further, when Sniffer does not use the GPT-3 models as white boxes, Sniffer can generalize to trace GPT-3 model origins and can properly discriminate them from human-written lines. Such an ability cannot be easily obtained as the GPT-3 models are stronger and similar to humans, meanwhile remaining unknown to the public.
Plus, we can observe that recently released LLaMA models are more difficult to trace, though it is a white box in Sniffer, indicating that stronger LLMs are harder to detect. When LLMs are being studied by a wide range of researchers, it is also important to be alert that stronger models may be harder to detect, and harder to control which can cause potential harm to society.
6 Different Sniffers
In Table 3, we list the results of using Sniffer(+GPT3) to trace the origins of texts generated by GPT3 models while we treat GPT3 models as white boxes. We use the OpenAI service that returns token logits therefore we use the text-davinci-003 model and test on a dataset compared with the full dataset. Here, we replace texts generated from ChatGPT with davinci-003 outputs therefore the testset is different from the one used in tracing ChatGPT texts.
As seen, when the Sniffer feature can use GPT-3 logits, the results grow significantly higher compared with generalizing features from GPT2/J/Neo and LLaMA models to trace GPT-3. Therefore, we can conclude that although Sniffer is able to generalize its features to trace unknown origins, it is better to possess the LLMs, indicating that the actual LLM providers must be more cautious when releasing LLMs since they are more capable to avoid abuse or malicious usage of LLMs.
As we illustrated, in methods using model-wise features, we can use limited data to construct a powerful detector. Therefore, we use different numbers of training data to train Sniffer. As seen, when we use limited data to train the Sniffer model, the performances are not significantly harmed, indicating that the model-wise features are high-quality features that reveal obvious traces of texts for the origin tracing classification.
In the extracted features, we observe that using L1-norm can help obtain a higher performance in tracing human texts, but is rather weak in tracing different model origins, especially when tracing LLaMA. Therefore, we use datase-wise norm for the rest experiments.
To further analyze how the extracted features help trace the text origins, we run a simple ablation test that uses different features proposed in Sniffer. As seen in Table 2, the perplexity score is one key metric but can be significantly improved by the percent-of-perplexity score (pct-score), showing that we can trace text origins by analyzing the discrepancies between models when testing same texts as Sniffer did.
As discussed, the model-wise features trace the text origins by analyzing the model discrepancies, and the semantic-wise features are used as a fine-tuning task. We compare Sniffer with a fine-tuned RoBERTa model and measure the f1-score. Further, we combine Sniffer features and RoBERTa [CLS] feature to train a linear classifier as Sniffer+RoBERTa to test the performances of combinations of model-wise and semantic-wise features. As seen in Figure 4, the supervised learning method fails to solve the problem when the data is limited. Still, when the data is abundant, the performances are stronger than model-wise features, which is also discussed in Mitchell et al. , Souradip et al. .
We need to notice that the supervised learning method fails to detect GPT-2/GPT-J/Neo generated texts, which is different from Sniffer results as GPT-2 texts are known to be less fluent compared with stronger models such as GPT-3 models. Therefore, we can assume that the RoBERTa models may capture some specific patterns since the GPT-3 generated texts only use several static instructions and the semantic features are easier to detect. The failure of GPT-2 detection indicates that semantic features are limited in origin tracing, calling for better algorithms to utilize model-wise features. In Sniffer+RoBERTa, we obtain promising results in tracing all origins, indicating that a proper combination of two types of features can help better trace text origins.
7 Difficulty of Tracing Different Types of Generated Texts
As illustrated, we use several different instructions to instruct GPT3.5 (turbo) models and GPT3.5-text-davinci-003 models, we further discuss the tracing difficulty when the texts are generated by different instructions.
As seen in Figure 5(a), we find that when the texts are generated by rephrasing instructions, the texts are rather easy to trace while texts generated from a summary are harder to trace. Such an observation indicates that the difficulty of tracing texts generated by strong LLMs is also different when the instructions are different. The rephrased texts are more similar to human-written texts, therefore, are more difficult to detect.
In our origin tracing experiment setup, we divide GPT-J/Neo generated texts into the same origin since these models are provided by the same facility. That is, LLM origins can have different levels of classifications. As LLMs can be trained with distill texts from strong LLMs such as ChatGPT, it is harder to trace the origin if the model is trained based on a base model such as GPT-J or LLaMA but is instructed by ChatGPT-generated outputs. It is hard to tell whether the base model or the instruct model has a larger impact on the mixed-origin model. Therefore, we construct experiments to first separate GPT-J/Neo origins using the proposed datasets. Further, we test mixed-origin models such as Alpaca (ChatGPT instructions tuned based on LLaMA) and Dolly https://huggingface.co/databricks/dolly-v2-12b (ChatGPT instructions tuned based on GPT-J) by instructing these models to generate 400 samples and testing them with Sniffer.
In Figure 5(b), we show the f1-score of origin tracing that separates GPT-J and GPT-Neo models using the full data of SnifferBench. As seen, Sniffer is able to successfully divide GPT-J and GPT-Neo, indicating that the text origin divide can be of different levels. We can trace the text origins of a specific model or a certain party.
In Figure 5(c), we list the averaged probability of the Sniffer inference results of supervise fine-tuned model Alpaca and Dolly. As seen, texts from both Alpaca and Dolly tend to be categorized as ChatGPT-generated texts, indicating that the align process has a more significant impact on the generated texts compared with the base model. Therefore, Sniffer can be used as a detector for testing whether a model is trained from ChatGPT-generated instructions, helping protect the originality of LLMs.
Conclusion and Future Work
In this paper, we first introduce the concept of origin tracing, an important direction in the era of LLMs. Then we discuss two lines of AI-generated context detection and origin tracing methods and point out the necessity of studying model-wise features for origin tracing. We further design a simple Sniffer method as well as a benchmark to test the origin tracing challenge. Through extensive experiments, we find that the current origin tracing field is full of challenges including tracing texts from models with mixed origins; combining semantic features and model-wise features; tracing texts from LLMs that are given various instructions. Therefore, we can hope for a continuous line of works that study the origin tracing of LLMs and hope to improve the trustworthy and safe usage of LLMs.
References
Appendix A Appendix
As mentioned, we align tokens with different tokenization methods with a uniform tokenization strategy to compare the perplexity at the same level. As shown in Figure 6, when the texts are tokenized by different tokenizers, we project different tokens to the byte level, and from the byte-level projections, we align different tokenized texts to a list of tokens with the same tokenizations.
We also conduct a Chinese version of Sniffer as well as a collected testset for studying different types of LLMs.
We sample 1k samples from each open-source dataset of various domains including a long-document corpus of stories ; various web texts from Iflytek classification dataset ; Chinese academical documents from CSL dataset ; Chinese Wikipedia corpus from CMRC dataset ; review corpus from Xiecheng APP https://huggingface.co/datasets/seamew/ChnSentiCorp. We collect these datasets that cover various domains with relatively long documents (more than 200 Chinese characters per document). The dataset construction setup is similar to English dataset setups. We also use ChatGPT (GPT turbo) model to generate Chinese texts with the Chinese instructions as ChatGPT (GPT turbo) is a strong multi-lingual model.
As for open-source model selections, we adopt several open-source Chinese LLMs including Wenzhong model , a 2.7B GPT-2 style Chinese LLM; Damo model https://modelscope.cn/models/damo/nlp_gpt3_text-generation_2.7B/summary, another 2.7B GPT-2 style Chinese LLM; Skytext model https://huggingface.co/SkyWork/SkyTextTiny, a 3B chatbot, and ChatGLM , a 6B ChatGPT-style model trained with instructions.
For the generalization test of black-box models, we use ChatGPT (GPT turbo) and MOSS, a 16B Chinese LLM https://github.com/OpenLMLab/MOSS with instructions that ask LLMs to re-write the given document as black-box origin tests.
We generate Chinese AI-generated contexts with open-source models and black-box models to build the whole Chinese benchmark for origin tracing tests.
As seen in Table 4, the Chinese texts show similar performances with English texts, indicating that the origin tracing strategy can be used in different languages.
We list several case studies that are randomly selected from the testset. We show the texts to be detected, percent-of-perplexity scores, correlation scores calculated by Sniffer, and the tracing results.
As listed below, we show the list of perplexity scores of the given text and the extracted features, then we show the predicted tracing result. As seen, the perplexity list shows a similar trend between different models, but the perplexity value tends to show differences across models. For instance, in the GPT-2 generated text calculated list of perplexity, the GPT-2 perplexity is relatively lower than other perplexities calculated by other models, revealing the feature that helps trace the GPT-2 texts. On the other hand, for human-written texts, and texts generated by strong LLMs such as ChatGPT, the value does not show obvious differences between different models, for instance, some peaks in the perplexity list can be from the LLaMA models and some can be from some other models, which supports the hypothesis made above.