Multiscale Positive-Unlabeled Detection of AI-Generated Texts

Yuchuan Tian, Hanting Chen, Xutao Wang, Zheyuan Bai, Qinghua Zhang, Ruifeng Li, Chao Xu, Yunhe Wang

Introduction

Recent developments in Large Language Models (LLMs) have brought astonishing changes to people’s lives. The GPT-2 model, created in early 2019, is capable of simple question-answering tasks; GPT-3 is a great leap in model size and capability; ChatGPT , announced in late 2022, shows comparable performance to humans as a chatbot; GPT-4 , released this year, has even better generative performance. These advancements are making people’s lives easier with applications like writing aids, search engines, and Office Suites. However, they could be used to generate deceptive fake texts for illegal and unethical purposes.

Previous works have proposed numerous approaches to distinguish fake AI-generated text from genuine human languages. Canonical work used simple machine learning classifiers as baselines; some works [12; 26] proposed zero-shot detection measures based on pretrained models; numerous works [34; 6; 13; 27] perform simple finetuning of pretrained language models on the AI-text classification task.

Despite various methods, few mainstream methods investigated the negative impact of text length: the difficulty to detect significantly increases as texts become shorter. Some latest online ChatGPT detectors have noticed this issue, but they dodge rather than address it by putting up minimum text length requirements [38; 11; 30]. In the era of smartphones where people rely heavily on fragmented mobile media, fake short articles like SMSes, Tweets, and reviews generated by LLMs could pose huge threats to one’s daily life, yet we still lack a comprehensive detector that is capable of detecting both short texts and long-texts.

To improve detectors’ performance on short texts, we rethink the plain "Binary Classification" setting that is intuitively applied. It is seemingly natural to phrase text detection as a binary classification task, as texts have clear origins (from human works or AI outputs) and thus, clear binary labels (real or fake); but interestingly, we observe a handful of machine-generated texts that are overly short and simple, such that these texts are highly similar to human (e.g. Ex. 2 in Table 1). It is not suitable to assign these simple machine texts with either clear human or AI labels; rather, they are in an "Unlabeled" state. Though the case is occasional and most short machine texts (e.g. Ex. 1 in Table 1) are still distinguishable based on manifold features, it prompts us to question the rationality of clear binary labels on general short machine texts. On the contrary, we hold that short machine-generated texts are partially "Unlabeled". As machine-generated texts become shorter and simpler, the "Unlabeled" property could gradually dominate the text.

In this sense, we model the task of AI-generated text detection as a partial Positive-Unlabeled (PU) problem and formulate the Multiscale Positive-Unlabeled (MPU) training framework to address the challenging task of short text detection without sacrificing long texts. PU problems typically address binary classification tasks where positive data and unlabeled data are offered for training. Considering the partially "Unlabeled" property of short machine texts, we rephrase detector training as a partial PU problem and boost detectors’ performance on multiscale texts. In order to improve conventional PU optimization targets for texts of various lengths, a length-aware Multiscale PU (MPU) loss is proposed and applied during the training process. We are aware that the PU prior probability of a text being positive is length-variant. To this end, an abstract recurrent model is designed to adjust the PU prior probability automatically based on corpus length. Further, a Text Multiscaling module is also proposed to exert the effect of Multiscale PU loss by diversifying training corpora in terms of length. Experiments demonstrate that the MPU framework is significantly effective in improving short-text detection performance; meanwhile, detection on long texts is also augmented.

Multiscale Positive-Unlabeled Text Detection

Since the introduction of GPT-2 and its successors, fake texts generated by powerful LLMs are causing ethical and legal issues. Methods are developed to discriminate against these generated texts in various misuse scenarios. Zellers et al. shed light on machine-generated fake news by proposing a GPT-based news generator GROVER, and uses GROVER itself to sort fake news out; Adelani et al. looks at detection of fake online reviews; Fagni et al. focuses on machine-generated fake tweets and proposes the TweepFake dataset. Other proposed detection methods are for general scenarios. Several canonical baselines are mentioned by Solaiman et al. to detect GPT-2 texts, including simple TF-IDF classifiers and finetuned RoBERTa ; GLTR detect generated texts in a zero-shot manner by using token prediction probabilities from available pretrained NLP models like BERT and GPT-2 . After the introduction of ChatGPT , some new detection methods [23; 26; 27; 13] are released.

Despite manifold methods, mainstream detectors seldom take the factor of text length into account, and thus they always fail on short texts. We have tried several existing detection methods for short LLM-generated texts (shown in Table 4), but none of them perform well. As people nowadays are immersed in short, fragmented forms of mobile media, they are vulnerable to LLM attacks with no reliable means to defend themselves. Hence, we are in urgent need of a performant short AI-generated text detector.

Intuitively, past works formulated the task of AI text detection as a binary classification problem, i.e. classifying texts as AI or Human. However, the formulation could be problematic for shorter texts as we found high similarities between extremely simple AI texts and human texts. The phenomenon could be rare in actual applications. But it is fundamentally reasonable, because LLMs learn from human languages; and for sentences whose structures are overly simple, they are seemingly "copied" by LLMs from what they have learned. Therefore, the attribution of these simple machine texts is uncertain: on one hand, they are indeed outputs from Language Models; on the other hand, they are ordinary human languages. Though the completely non-classifiable case mostly happens for extremely short texts or commonly used phrases (that rarely occurs in our benchmarks and detection of which is of no application value), it inspires us to think about the partially "unlabeled" property behind the vast majority of short, distinguishable texts despite their definite labels.

To overcome this issue, we model the task of multiscale text detection as a partial Positive Unlabeled problem (PU). In this problem, corpora from human are regarded as "Positive", but short texts from machines are given an additional "Unlabeled" mark for PU loss calculations (detailed in Sec. 2.3). Then our detector model is optimized within this partial PU context.

2 Preliminaries: PU Classification

Overview of PU methods. Previous works have proposed methods to train a binary classifier with positive and unlabeled data. Many PU methods [2; 9; 20; 35; 14; 5] constructs PU loss based on positive and unlabeled samples, for classifying unlabeled data. Other PU methods include two-step learning and bias learning . The two-step technique first identifies reliable negative examples and then performs learning based on the positives and negatives of the mark [15; 17]; biased learning treats unlabeled data as a negative sample of class-labeled noise [16; 33]. Above all, we refer to applying a PU loss during training to address the task of multiscale AI-generated text detection, because PU losses could be generally applied on powerful finetuning text detectors without much additional computation costs.

However, in the PU problem, only partial positive sample data and unlabeled data are available. Since the negative part of loss (1−π)R^N(g,−1)(1-\pi)\hat{R}_{N}(g,-1) is missing, the unbiased PN loss function cannot be used directly. Hence, an estimation for the negative loss part is calculated via positive and unlabeled samples. The estimation is shown as follows :

The non-negative PU loss. Despite the unbiasedness of the unbiased PU (uPU) loss, Kiryo et al. indicates that uPU loss could cause overfitting problems, limitting the flexibility of models to be trained. As a remedy, Kiryo et al. proposes the non-negative risk estimator for the uPU loss, which is then formulated as the non-negative PU (nnPU) loss. The nnPU loss is obtained as follows:

The nnPU loss is performant and thus widely referred by later PU works and applications [19; 3; 31; 40; 5; 35; 37]. However, to the best of our knowledge, no previous works have applied PU to scenario of length-variant texts, in which simple usage of nnPU might not be effective. We hope to develop an effective PU mechanism in aid of detecting length-variant texts.

3 MPU: A Length-sensitive PU Approach

where ff is some function that merges the classification of all previous tokens Si−1S_{i-1} with the classification of the last token tit_{i}. Next, the abstraction is concretized based on task characteristics of human-generated text discrimination. Since relatively short texts tend to have simple semantic correlations to be captured, human text discrimination is performed via capturing signals from tokens. We hold that each token has a hidden property of origin, and the attribution contributes to the classification of the whole sequence. Tokens, as extreme cases of short texts, could be sorted into two categories: "clear positive", i.e. the token could hardly be generated by AI; or "unlabeled", i.e. the token is mediocre and universally used, giving no signal as "human-spoken". Each token is expected to provide an equal contribution to the overall sequence classification towards the orientation of its own category . In this sense, the merging function ff is formulated as equally-weighted addition:

where δ(ti+1)\delta(t_{i+1}) is defined as the contribution of δ(ti+1)\delta(t_{i+1}). For simplicity, we discretize the transition of classification from i→i+1i\rightarrow i+1 and each token contribution is designated as binary. We also take text length into consideration by normalizing δ(ti+1)\delta(t_{i+1}) with a factor of sequence length ll. Under these assumptions, the transition is formulated as:

Notably, we use a hard clip function to bound the overall classification results in interval [0,1]\left[0,1\right] rather than other non-linear functions, e.g. sigmoid. This is because clear positive tokens could be rare in practice. This assumption is particularly true when we consider recent advancements of generative language models, where human and AI languages are more resembling. In other words, a majority of words are both frequently used by human and AI, while only a few signal words manifest unique human characteristics. This property requires the discriminate model to be highly sensitive to positive token signals. Hence, we set hard boundaries rather than using non-linear standardizing functions to scale the output between $.Further,toencouragepositiveresponses,weinitiallypositiveastheinitialstate. Further, to encourage positive responses, we initially positive as the initial state\Delta(S_{0})$ of the discriminator.

Finally, on top of the canonical non-negative PU loss as defined in Eq. 4, we define the Multiscale PU Loss with text-length-variant priors:

4 Text Multiscaling

The proposed Multiscale PU Loss expects training texts of highly variant lengths, but training sets may contain lengthy paragraphs only. Therefore, we introduce Text Multiscaling Module that generates a variety of short texts to exert the potential of the length-sensitive Multiscale PU loss. We propose random deletion at sentence scale as a solution. Text Multiscaling module is designed as follows: first, a complete training text cc is first tokenized into nn sentences [si]i=1n\left[s_{i}\right]_{i=1}^{n}, denoted as array CC:

where operator ∥\parallel stands for text concatenation. Then the sentences are independently and randomly masked based on a sentence-wise mask probability psentp_{sent}. Namely, each sentence is decided by an independent Bernoulli trial in the sample space X:S→{0,1}X:S\rightarrow\{0,1\}. In the sample space, 0 means the sentence is discarded and 1 stands for the sentence is maintained. The sampling process of mask array M∈{0,1}nM\in\{0,1\}^{n} is denoted as:

Finally, we merge all sentences for a multiscaled training text cmulc_{mul}:

where ⊙\odot stands for the elementwise Hadamard product. Text Multiscaling module is a one-to-one mapping from c→cmulc\rightarrow c_{mul}; we are not generating more training samples, but substituting the original sample for fair comparison in experiments. Notably, it is probable that multiscale could leave the original text intact, or only one sentence is left. The relative sequence of remaining sentences is maintained to avoid breaking excess logical relations between sentences. Multiscaled texts automatically inherit class labels of their original text. The concern for attribution change due to length reduction is to be addressed by the use of Multiscale PU Loss.

Though random deletion is also applied in Easy Data Augmentation (EDA) , our method is different from theirs in two aspects. Firstly, our method is focused on multiscaling, while word-level random deletion proposed by EDA has limited effect in generating texts of various lengths. Secondly, EDA could break semantic meanings in sentences: deletion of keywords could change the class of a sentence; while a more integrated, sentence-level deletion reduces the chance of class property change.

Experiments

Datasets. We choose TweepFake and HC3 as benchmarks for our experiments. TweepFake is a dataset of tweets for AI-generated microblog detection. Since latest LLMs have completely reshaped the task of AI text detection, we also adopt HC3 , which is an up-to-date ChatGPT text detection dataset including both English and Chinese. Additionally, HC3 has short-text benchmarks: HC3-English-Sent and HC3-Chinese-Sent. We use these datasets to demonstrate the effectiveness of our method.

The length statistics in Table 2 show the distribution similarity of English short-text benchmarks, i.e. TweepFake (that consists of tweets) and HC3-En-Sent. We conclude from the statistics that the adopted HC3 short-text benchmark could simulate the fragmented language environment (e.g. Twitter) on mobile apps. Detector evaluation on these short-text benchmarks could reflect their real-world detection capabilities in smartphone-related scenarios.

Detectors. BERT and RoBERTa are adopted to apply our MPU method, due to their popularity and supreme performance in previous AI text detection works [34; 10; 23; 13]. Training-agnostic detection algorithms are excluded from our consideration.

2 TweepFake Detection Results

In TweepFake experiments, we follow Kumarage et al. for our training settings. Kumarage et al. is one of the latest works on AI-generated text detection, and it claims outstanding performance on short-text detection. We strictly follow the original training strategy in Kumarage et al. : the model is trained with the AdamW optimizer at batchsize 16 and learning rate 1e−51e-5.

TweepFake mainly consists of short tweets. we inspect the dataset and find that a vast majority of texts are single or a handful of sentences. Hence, we refrain from using Text Multiscaling that randomly delete sentences for TweepFake datasets; rather, we directly apply Multiscale PU loss during training. As shown in Table 3, the experiment result of the proposed MPU is promising: it greatly improves the performance of finetuned RoBERTa, and its performance outcompetes the latest TweepFake baseline RoBERTa-Stylo that requires an additional module for stylometric feature extraction during finetuning.

3 HC3-English Detection Results

We also experiment our method on ChatGPT corpora that are much harder to detect. In the ChatGPT text detection experiments, we follow the setting of HC3 to test the performance of our method. HC3 is a dataset targeted at ChatGPT text detection. All texts are reduced into shorter texts for a sentence-level variant. We apply the MPU framework on the full-scale dataset of HC3 .

Several baseline detectors are chosen to demonstrate the outstanding detection performance of our MPU method. These baselines are open-source and replicable. Among these baselines, GLTR , PPL , and DetectGPT are zero-shot methods that do not require further training: they rely on the likelihood outputs of a pretrained language model. The OpenAI Detector is a RoBERTa detector finetuned on OpenAI’s GPT-2 corpora. RoBERTa-Stylo is one of the latest detection baseline targeted for short texts. BERT-Finetuned and RoBERTa-Finetuned are language models plainly finetuned on HC3 , following the official setting; while BERT-MPU and RoBERTa-MPU are language models trained on HC3 via the proposed MPU method.

It could be observed from Table 4 that most existing methods perform poorly on short texts. The statistics verify our previous claim that the detection of shorter texts is a difficult problem. Specifically, finetuned BERT and RoBERTa are good at detecting long, full-level texts, but they fail to filter out shorter AI-generated texts. On the contrary, our MPU method could greatly improve short-text performances and boost long AI-generated text detection as well. We will further investigate the effect of solitary MPU components in Sec. 3.5.

4 HC3-Chinese Detection Results

To verify the generality of the proposed MPU method in other languages, we also compare our method with baselines on Chinese AI text detection benchmark HC3-Chinese . Following Guo et al. , we use chinese-roberta-wwm-ext as the pretrained language model. The results are shown in Table 5. Our method could still outcompete other methods by large margins in terms of short-text detection, reaching an F1 score of 89.37 on HC3-Chinese-Sent.

5 Ablations

Framework Components. We perform ablations on the solitary effects of Text Multiscaling and Multiscale PU loss.

From Table 6, it is firm that the addition of Text Multiscaling to training corpus greatly improves performance on sentence-level corpus detection as expected. Unfortunately, the detector’s capability on full corpus decays. This performance drop is attributed to the unreasonable label assignment to short corpus from random sentence deletion: the generated short corpora automatically inherit labels from their full-level predecessors in Text Multiscaling Module, neglecting "unlabeled" properties as introduced in Sec. 2.1. The addition of MPU loss reverses full-level corpus detection performance drop and boosts short-text performance as well. Solitary addition of MPU loss only would have little help for detection performance for lack of short texts.

MPU Loss. We further investigate MPU loss configurations on ChatGPT text detection benchmark HC3-English .

The performance of Multiscale PU loss is evaluated against ordinary PU loss that disregards changes in sentence lengths, as shown in Table 7. Multiscale PU loss is sensitive to training corpora of various lengths and thus is more performant compared with its ordinary counterpart.

Introduced in the abstract recurrent detection model (Sec. 2.3), token-wise prior pp estimates the probability of a token being highly characteristic as human-spoken. As shown in Table 10, we carefully tune pp and found that the best performance is reached at p=0.2p=0.2, which is small as we expect.

We also carefully adjust the affine weight hyperparameter for PU loss γ\gamma, as shown in Table 10. As the affine weight γ\gamma for PU loss gradually increases, the full-level corpus detection performance reaches the peak at γ=0.4\gamma=0.4 and then drops, while the sentence-level performance reaches its peak at γ=0.6\gamma=0.6. From a comprehensive perspective, the best overall performance is reached at γ=0.4\gamma=0.4 where both performances on full and sentence-level corpus are satisfactory. The climb-and-drop trend reveals that short machine-generated sentences are not completely unlabeled; short-text classification should be viewed as a partial PU problem rather than a complete PU problem.

Further, we test the advantage of the non-negative risk estimator in the nnPU loss against uPU loss , as introduced in Sec. 2.2. The results are shown in Table 11.

Text Multiscaling. As introduced in Sec. 2.4, we randomly mask sentences of the training set at probability psentp_{sent} for multiscale text augmentation. We investigate on tuning psentp_{sent} for the optimal value. The statistics are shown in Table 10. When psentp_{sent} is set at 0.250.25, the test performance on both full and sentence level corpus are satisfactory; when psentp_{sent} is set too high, sentence-level detection performance is enhanced, but full-level performance is negatively impacted because the full-scale training texts are overly damaged.

Conclusion

This paper proposes a Multiscale Positve-Unlabeled (MPU) framework for AI-generated text detection. We look at the iffy attribution of short AI-generated corpus, and model AI text detection as a partial PU problem. MPU loss and Text Multiscaling Module are to augment detectors’ discriminative ability on short corpus.

This paper proposes a training method for AI-generated text detectors. Despite outstanding performance on multiscale texts, chances are that the detectors output the wrong attribution of a certain piece of text. This may cause ethical issues when the detector is used for detecting plagarism, fake news, et cetera. Hence, we strongly recommend that results from the detector could only serve as a reference in actual applications.

Experiments are reproducible. We have attached complete training settings in the Appendix; we also fix random seeds in our codes for the ease of replication. All details are in Appendix D.

References

Appendix A Estimation Details of Confidence Expectation

The transition matrix Given positive probability pp of a single token, we express state transition as a band matrix P\mathbf{P}. An example matrix form of P\mathbf{P} is listed as follows:

Interestingly, we could leverage unique features of the sparse band matrix P\mathbf{P}. First, obviously Pl+1[n,:]=[0;Pl[n,:]]\mathbf{P}_{l+1}[n,:]=[0;\mathbf{P}_{l}[n,:]]. Further, if we compare

we would discover that M=[0;K]M=[0;K], namely, array MM is array KK prepended by a zero. (The physical meaning of MM and KK is the last line of matrix Pl+1l\mathbf{P}_{l+1}^{l} and Pll\mathbf{P}_{l}^{l}, respectively.) Based on this discovery, we could simplify Eq. 16:

Then we look at the concrete form of [0;K]Pl+1[0;K]\mathbf{P}_{l+1}. For simplicity, we denote the nthn^{th} element of KK as knk_{n}:

Based on the table above, we could derive the relations between E[Δ(Sl+1)]E\left[\Delta(S_{l+1})\right] and E[Δ(Sl)]E\left[\Delta(S_{l})\right]:

As long as we view {l×E[Δ(Sl)]}\{l\times E\left[\Delta(S_{l})\right]\} as a sequence of corpus length ll starting from 1×E[Δ(S1)]=p1\times E\left[\Delta(S_{1})\right]=p, we could solve E[Δ(Sl)]E\left[\Delta(S_{l})\right] for l>1l>1:

Appendix B Proposal of Imposing Space Cleaning on the HC3-English Benchmark

We use the HC3 benchmark for ChatGPT corpus detection experiments. However, we inspected HC3 corpora and discovered that the corpora are flawed: human corpora have additional spaces before punctuations, while corpora from AI do not have this feature. The extra spacing could directly impact the input to detectors. We list several examples below, demonstrating the obvious difference between Human and ChatGPT corpora in the HC3 benchmark :

In the examples, we show original corpus as well as their token ids after being processed by the RoBERTa-base tokenizer. Most human corpora have an unexpected 479 token (standing for " .", i.e. a space and a period), while ChatGPT corpora does not manifest this feature.

Hence, the detector could judge the attribution of a certain corpus simply by detecting these spacing mistakes. Embarrasingly, if we use the logical judgement of whether token id 479 is contained in the sequence to detect human corpora, the F1 score would reach 82.12%82.12\% on sentence-level test corpora of the HC3 benchmark. The performance of such a simple logic is even better than the officially reported performance (81.89%81.89\%) of finetuned RoBERTa-base . Above all, we strongly recommend later works that involve the HC3 benchmark to remove unnecessary spaces before punctuations. We have open-sourced the code of simple cleaning helper function that removes unnecessary spaces.

Appendix C Baseline Replications

DetectGPT is a latest open-sourced AI corpus detection baseline, but the original paper did not report its performance on latest LLM texts. Hence, we replicate DetectGPT on the HC3-English ChatGPT corpus dataset, and compare it with our MPU method. The experiment results are shown in Table 4, where our MPU method outcompetes DetectGPT by large margins. There is still a visible gap between latest training-agnostic methods (e.g. DetectGPT) and finetuned language models on ChatGPT corpora.

We also provide some detailed procedures to tailor DetectGPT for the HC3 benchmark: 1. Full-scale HC3 corpora are always too long to perturb. Therefore, we truncate corpora as long as they raise perturbation errors, following recommendations from authors of DetectGPT. 2. We use 100 perturbations for full-scale HC3 corpora (following DetectGPT ), but we use 10 perturbations for sentence-level HC3 because there are too many corpora. It also reflects that DetectGPT is not very efficient for large-scale corpora compared to language model detectors, because it requires tens of model runs for a single corpus. 3. DetectGPT uses AUROC as the classification metric; however, this metric is not applicable to finetuned language models that output probabilities for respective classes. Hence, given confidences of all corpora outputted from DetectGPT, we choose 1000 equally-spaced threshold between max and min values, and maintain the threshold with the largest F1 score. Notably, this will provide an upperbound for the performance of DetectGPT, as in real applications the threshold is pre-set; scanning for the best threshold on test sets is strictly prohibited.

C.2 GLTR, PPL, & OpenAI

These methods have already been open-sourced on HuggingFace. We directly input all texts in the testset to these baseline methods and measure their performances.

We have found an inconsistency in comparison to reported values while replicating GLTR and RoBERTa-Finetuned on the HC3-Chinese benchmark, shown in Table 12. This inconsistency is tolerable and won’t affect our final conclusion.

Appendix D Replication Details

Following the training setting of Kumarage et al. , we use batchsize 16, learning rate 1e−51e-5 for TweepFake; following the setting of , we use batchsize 32, learning rate 5e−55e-5 for HC3. AdamW optimizors are adopted. Selected benchmarks are publicly accessible online.

We use a single Nvidia Tesla V100 for experiment devices. A single epoch of training costs around 30 minutes. We replicate all experiments three times to avoid fluctuation, using seed=0,1,2.