A Robust Semantics-based Watermark for Large Language Model against Paraphrasing
Jie Ren, Han Xu, Yiding Liu, Yingqian Cui, Shuaiqiang Wang, Dawei Yin, Jiliang Tang
Introduction
Large language models (LLMs) have shown its great ability in various natural language processing (NLP) tasks like Question Answering (QA) Lu et al. (2022), reasoning tasks Wei et al. (2022); Creswell et al. (2022) and code development Xu et al. (2022). However, tremendous concerns have been raised that LLMs are possible to be used improperly and illegally. For example, fake news that are indistinguishable are easy to be fabricated (Kreps et al., 2022; Zellers et al., 2019), which, when disseminated, could instigate widespread panic. Similarly, in the commercial sphere, convincingly generated reviews can manipulate consumer perceptions, leading to unethical business competition (Salminen et al., 2022). Therefore, detecting LLM-generated text has become crucial in the real-world applications of LLMs.
Among diverse methods to detect LLM-generated texts, the watermark strategies have demonstrate outstanding precision (Kirchenbauer et al., 2023a). It is proposed to encode a secret watermark into the generated texts, such that one can tell whether a text is generated by detecting this watermark. One representative strategy Kirchenbauer et al. (2023a); Yoo et al. (2023) is to encode the watermark based on the “partition of vocabulary”. In detail, given a language model, these methods devise a mapping from precedent tokens to a particular partition of the vocabulary by a partition function for the consequent token. The partition function uses the hashes of the input as the seed of a random generator to split the vocabulary to a green list and a red list. During the text generation phase, the consequent token has an increased probability to be sampled from the green list. In this way, the watermark is encoded through the matching between the precedent tokens and the vocabulary partition for the consequent token. The detection is also facilitated by detecting this matching in generated contents. However, recent works Krishna et al. (2023); Kaddour et al. (2023) reveal that this watermark may be easily eliminated by sentence paraphrasing. Individuals seeking to improperly utilize LLMs without being detected can paraphrase the generated contents, like altering the order and the choices of the words, and only retain the general meaning of the text to achieve their malicious goals like faking news. These paraphrases will change the seed of the partition function, i.e. the token hashes, and consequently disrupt the matching between the precedent tokens and the green list. As a result, the detection effectiveness of the watermarks can be dramatically compromised.
In this paper, we propose to leverage the semantic meaning of precedent token sequences as the seed for partition function, instead of simple hashes of precedent tokens, since the core semantic meaning is expected to be maintained after paraphrasing. To achieve this goal, one key obstacle is how to capture the semantics when applying them for the partition function to watermark generated texts. It is a common practice to quantify the semantics via embeddings Reimers and Gurevych (2019); Gao et al. (2021); Li et al. (2020); Giorgi et al. (2021). Embeddings indeed can represent consistent semantics after paraphrasing. Since the embeddings are high-dimensional vectors in the continuous space, the embedding vectors could present some minor changes after paraphrasing. These minor changes can lead to a substantial difference in the partition of vocabulary as the random generator in the partition function is sensitive to any change of the seed.
To overcome this challenge, i.e., to make the quantified semantics invariant under paraphrase, we propose a new watermark method, SemaMark, which discretizes the continuous embedding space. Intuitively, the discretization can coarsen the representation of the embeddings which could tolerate the potential changes caused by paraphrase. By proper discretization, the paraphrased changes could exceed the same discrete section with a low probability and the quantified semantics will likely remain the same even after paraphrase. Therefore, the partition results will not change. However, directly converting the high-dimensional embedding space into discrete is intricate and challenging. For example, discretizing each dimension will lead to dense discrete values, i.e., the total number of discrete values is too large. Thus, one discrete value is insufficient to accommodate the minor changes by paraphrase and the partitions are still unstable. In other words, the change of only one discrete dimension can have a strong impact on the partition function. To address this problem, SemaMark first uses a Multi-Layer Perception (MLP) to condense the continuous high-dimensional embeddings into normalized vectors in 2D space. The vectors are located on a unit circle named Normalized Embedding Ring (NE-Ring). Then the condensed NE-Ring is equally divided into various sections for distinct values, transforming the continuous space into discrete “semantic values”. Based on the discretization, SemaMark further introduces two strategies to advance the watermark’s concealment and to improve the robustness under paraphrase. First, SemaMark leverages the uniformity Wang and Isola (2020) of Contrastive Learning(CL) Chen et al. (2020) to strength the MLP and mitigate the situation that the semantics are unevenly concentrating on some discrete sections on NE-Ring. The unevenly distribution will cause the resulting discrete semantic values overly monotonous. It raises the concern that the watermark might be cracked by counting token frequency Zhao et al. (2023). Second, SemaMark utilizes an offset detection method to further enhance the robustness at the boundary of different discrete sections whose semantic values are possibly vulnerable to paraphrase. Comprehensive experiments are conducted to demonstrate the effectiveness and robustness of SemaMark under different paraphrases.
Related works
As the development of LLMs, various LLM-generated detection tools have also been proposed. Learning-based methods train a classification model to detect the difference between human-written text and machine-generated text like Guo et al. (2023); Wang et al. (2023); Li et al. (2023). Other works do not rely on the classification model, but try to use the property of the LLM to test whether a given text is generated by LLMs. For example, DetectGPT Mitchell et al. (2023) assumes that the generated text will have high likelihood. GPT-who Venkatraman et al. (2023) uses UID-based features to model the unique statistical signature of each LLM and human author for accurate authorship attribution. These methods do not interact the generation process of LLMs and thus have to explore unknown features of LLMs for detection. Instead, watermarks can change the model with a small but pre-defined rule which accelerates the detection process effectively.
Watermark.
The distinction between watermark and other methods is that watermark can proactively change the generation to insert a concealed watermark into the generated text. This gives clear difference in watermerked and non-watermarked texts. Watermark shifts the text using a small but pre-defined rule to make the detection much more effective. The partition of the vocabulary for each token is a representative watermark method Kirchenbauer et al. (2023a); Yoo et al. (2023); Kirchenbauer et al. (2023b). In each auto-regressive step of generating one token, the method uses the previous tokens’ hashes, to select a part of the vocabulary as “green” at a ratio of . Subsequently, they elevate the likelihood of the tokens by boosting the logits of the softmax by . Through this approach, at each token position, the probability of this matching between the seed and green tokens tends to increase.
For a sentence with tokens, it is viewed as a sample set of size . Each token is one sample from the vocabulary. A non-watermarked sentence is expected to have tokens showing this match. The watermark detection is approached as a -test with null hypothesis that the text is non-watermarke. If the -statistic is large, i.e. it is significant different from the null hypothesis, the null hypothesis can be rejected and the text can be predicted as watermarked:
where is the number of tokens showing the matching between seed and the green list. Yoo et al. (2023) further expand this watermark of green and red list to more lists for multi-bit encoding.
Method
In this section, we introduce the detailed design of SemaMark. We first present how to use the semantic information as the seed for watermark methods that are based on random partition of vocabulary in Section 3.1. Then in Section 3.2 and Section 3.3, we introduce the CL training scheme and the smoothed detection method for further improving the robustness.
As mentioned, the existing watermark methods based on partition of vocabulary are susceptible to paraphrase. Paraphrase can easily change the previous tokens and disrupt the matching between tokens and the partition of vocabulary, without significantly affecting the semantic meaning. Thus, SemaMark uses the invariant semantics for watermarking by discretizing the embedding space to accommodate the minor perturbation of semantics and provide a stable mapping between semantics and vocabulary partition for the consequent token.
However, discretization in a high-dimension space is intricate and non-trivial. Therefore, we first reduce the high-dimensional embedding space onto the 2D NE-Ring and then discretize via NE-Ring. The whole watermarking process is shown in Figure 1. SemaMark first reduces the dimension of the embedding space to obtain the discrete semantic values by two steps, i.e., weighted embedding pooling and discretizing by NE-Ring, and then uses the semantic value to partition the vocabulary. The logits of green list is shifted to increase the probability of matching between semantics and the consequent token for watermarking the LLM, . In the following, we introduce more details about the two steps to obtain a stable semantic value.
It is obvious that more tokens are used for watermarking, more stable it is for paraphrasing. Therefore, to enhance the robustness, we aggregate the semantics of previous tokens by the weighted mean pooling function before dimension reduction, instead of using only one preceding token’s embedding. For the token sequence starting at position , we use their semantics to generate the token in the position, . We denote their embeddings as . can be easily obtained from the LLM that we want to watermark. For pooling, assigns a smaller weight to the token far from and larger weight to the token closer to by:
S2: discretizing by NE-Ring.
After aggregating the embeddings by weighted pooling, SemaMark uses MLP to transform to a normalized vector in 2D embedding space. The normalized vectors locate on a unit circle in the 2D space, which is named as NE-Ring. The discretization function, , discretizes NE-Ring by equally segmenting into different sections. It takes the polar angle of as input and outputs the discretized semantic values , where . is defined as
With the two steps, we can get a stable discrete semantic value as the seed for the partition function to partition the vocabulary for the consequent token. Following Kirchenbauer et al. (2023a), the vocabulary is partitioned into green and red list. We increase the logits of the tokens in green list by and recalculate the probability distribution based on the shifted logits. For each token to generate, we increase the possibility of green list based on its previous tokens’ semantics. Thus, all the generated tokens will be likely to have this matching between the semantics and the consequent green token. By detecting the matching, we can discriminate whether a text is watermarked or not and then detect the LLM-generated contents effectively. Besides, SemaMark proposes two strategies to reduce the risk of being cracked by Contrastive Learning and further increase the robustness by the offset detection in the following sections.
Although the MLP can transform the high-dimensional embeddings onto NE-Ring, we expect that the MLP should obtain a uniform distribution of on NE-Ring. If different semantics unevenly concentrate in some areas on NE-Ring, the resulting discrete semantic values will be overly monotonous and the green list is more changeless. The green list might be revealed by counting the token frequency, which compromises the concealment of watermark and leads to the risk of being cracked. Ideally, SemaMark should generate a wider variety of semantic values for different sentences, while each semantic value is robust and stable if its corresponding sentence is paraphrased. To achieve this goal, we propose to use Contrastive Learning to train MLP since Contrastive Learning has the property of uniformity that the data will be evenly distributed in the whole feature space Wang and Isola (2020). The uniform distribution can help the normalized vectors cover all the semantic values. As a result, NE-Ring can generate a wider variety of semantic values to prevent the watermark from being cracked.
In Contrastive Learning, we first input the sentences into the model to get a batch of sequences of tokens and their pooling embeddings , denoted as , where and is the batch size. To compose a contrastive loss, we construct the positive and negative pairs by a soft augmentation:
where is a Gaussian noise. The soft augmentation can simplify the construction of postive samples. With this soft augmentation, we can assign the the samples sharing similar embeddings from the same sequence as positive pairs and samples from different sequences as negative pairs. This is consistent with our intuition that after paraphrases the semantic embeddings is robust and will not change significantly. Then the contrastive loss is
where is cosine similarity and is the temperature. By Contrastive Learning, the output of reduced semantic embeddings can be evenly distributed in all of the space on NE-Ring, and cover all the discrete sections to improve the robustness of SemaMark.
3 Q𝑄Q-offset detection
To mitigate the influence of this error, we propose -offset detection. As shown in Figure 2(c), we offset the discrete seed by tokens to detect the matching between semantics and the consequent tokens, where and the sign of indicates the direction of the offset. We choose the maximal -statistic in different as the -offset score. However, -offset detection will also increase the -offset score of non-watermark text, which indicates that the detected green word fraction of non-watermark text is higher. The in Eq. (1) is possibly inaccurate. Thus during generation, we set to a fixed value, while in detection process, we treat as a hyper-parameter and use an evaluation set to determine its value in practice.
Experiment
In this section, we provides experiments to demonstrate the robustness of SemaMark in Section 3.1. In Section 4.3, we show that our watermark has almost no influence on the quality of generated text. In Section 4.4, we show the sensitivity of the partition function on continuous embeddings and demonstrate that the discretization can mitigate this problem. In Section 4.5 and Section 4.6, we show the effectiveness of -offset detection and the distribution of NE-Ring, respectively.
Backbone models and datasets. We test our watermark method on two backbone models, OPT-2.7B and OPT-6.7B Zhang et al. (2022). For dataset, we use the news-like subset of C4 Raffel et al. (2020), which covers a variety of topics. From the news-like subset of C4, we extract a training set, a validation set and a test set. For each sample, we use the first half of text as prompt to generate watermark sentences.
Baseline methods. We compare our method with three baselines LeftHash, SelfHash Kirchenbauer et al. (2023b) and EXP-Edit Kuditipudi et al. (2023). LeftHash and SelfHash are two methods based on the partition of vocabulary using the hashes of tokens. EXP-Edit uses a private sequence to encode the watermark by changing the probability distribution of the sequence of tokens.
Paraphrase setups. We use three representative methods to paraphrase the watermarked text, round-trip translation Tiedemann and Thottingal (2020), Dipper Krishna et al. (2023) and GPT-3.5. For round-trip translation, we first translate from English to Chinese and then transform back to English, such that some words and expressions will be changed because the translation is not an one-to-one mapping. For Dipper, we follow the parameter setting in Kirchenbauer et al. (2023b). For GPT-3.5, we use the prompt in Kirchenbauer et al. (2023b) to query GPT-3.5 for paraphrasing.
Evaluation metrics and hyper-parameters. We use F1 score with best threshold and ROC-AUC to measure the performance of the watermark detection. All the metrics are calculated based on at least 500 watermarked samples and 500 non-watermark samples. The length of watermarked samples before paraphrase and non-watermark samples is . In generation, we set for LeftHash, SelfHash and SemaMark. In detection, we set and based on the evaluation set in Section 4.5. In SemaMark, we set , , for OPT-2.7B and for OPT-6.7B.
2 Main Results
In this subsection, we demonstrate the robustness of the proposed SemaMark under paraphrasing by comparing it with three baseline methods on two backbone models. We first generate watermarked texts and use three paraphrase methods to remove the watermarks. The detection performance of both texts with and without paraphrase is reported in Table 1. As we can see, before paraphrase, all the watermarked methods have good detection performance. LeftHash and SemaMark are the best with ROC-AUC even higher than 0.99. After paraphrasing, SemaMark has the best detection performance most of the time across all the backbone models and all the paraphrase methods, which suggests that our method is more robust to paraphrase.
In detail, by round-trip translation, the paraphrase reduces the detection ability of baseline methods effectively, while the watermark of SemaMark is robust. Under round-trip translation, the best ROC-AUC of baselines 0.9091 on OPT-2.7B and 0.8807 on OPT-6.7B. But ROC-AUC of SemaMark is 0.9692 and 0.9308, which is at least 0.05 higher than all the baseline methods. Similarly, under paraphrase of GPT-3.5, SemaMark is better than all the baselines. The best baseline performance under GPT-3.5 is 0.9392 in ROC-AUC on OPT-2.7B and 0.8990 in ROC-AUC on opt-6.7B, but SemaMark has higher AUC-ROC of 0.9406 and 0.9377. For Dipper, We note that all methods are robust to Dipper since it does not significantly reduce the detection performance. However, SemaMark is still one of the most robust. For OPT-2.7B, it performs best in ROC-AUC, while for OPT-6.7B, it has the best F1 score. From Table 1, the results show an obvious improvement of SemaMark in robustness. This implies that using semantics as the seed for the partition function is be effective under paraphrase.
3 Text Quality
Watermark should not compromise the generation quality of LLMs. In this subsection, we compare the text quality by calculating perplexity and demonstrate that our watermark has almost no influence on the generated quality. Perplexity measures the likelihood that a sentence is generated by one model. Lower perplexity means the watermarked text is more predictable. In other words, it is more consistent with the reasoning of the given model. In Figure 3, we use OPT-6.7B with no watermark to get perplexity for all the watermarked methods. All the results in Figure 3 are calculated without paraphrase, because the generation quality of text is not related to paraphrase. From Figure 3(a) for OPT-2.7B, we can see that our watermark, LeftHash and SelfHash have almost no influence on the generation quality. They has perplexity at around 6 which is similar as the generated text without watermark. Instead, EXP-Edit has much higher perplexity, which means that EXP-Edit changes the generated text in an aggressive way and much reduces the generation quality after watermarking. This is probably because EXP-Edit adjusts the logits on the whole vocabulary. From Figure 3(b), we can draw almost the same conclusion for OPT-6.7B. EXP-Exit also increases the perplexity by around 10, while the average perplexity of LeftHash, SelfHash and ours is around 1 higher than the non-watermarked generated text. Although the robustness of EXP-Exit is better than LeftHash and SelfHash especially on GPT-3.5, it remarkably reduces the quality of generated text . Instead, our SemaMark can keep the quality and robustness at the same time. In summary, our watermark has almost no influence on the generation quality of LLMs.
4 Ablation Study
In this subsection, we study the influence of the length of the sequence we used for generating one semantic value and the sensitivity of the partition function.
Length of previous token sequence tokens, . In the first step of SemaMark, i.e., weighted embedding pooling, we use the semantic of the previous tokens to get the more stable embedding. But if the length of the sequence is too long, it will also hurt the robustness. In Figure 4, we test the ROC-AUC using different from 10 to 30. The results show that before , ROC-AUC will increase as the increases. But when , ROC-AUC becomes fluctuating. One possible reason is that the paraphrase will change the length of sentences, and the tokens cannot be exactly mapping before and after paraphrases. Another possible reason is in the beginning of generation for the first tokens, the number of previous tokens is smaller than and NE-Ring can only use the embeddings of limited tokens for prediction, which may be unstable. Thus, too long or too short sequence will both hurt the robustness of SemaMark against paraphrase. In our experiments, we choose for all the settings.
Sensitivity of partition function. As we mentioned, partition function is sensitive to any change of the input as it only uses the input for the seed of the random generator. To validate its sensitivity to continuous embeddings, we adopt the embedding vector as the input to show that with tiny change of the embeddings, the partition of vocabulary can be very different. We propose a hash method based on md5sum Deepakumara et al. (2001) to adopt the partition function by transforming the continuous embeddings to an integral seed. We use 1000 sequences to test the sensitivity. For each sequence embedding, we first get a green list from the partition function. Then we change one dimension of the embedding by only 1e-5 to get a new partition result. The overlapping of the two green lists is 24.99% on the average of 1000 sequences. It is consistent with we use to watermark, because the random partition with the changed embedding is independent from the original one. Instead, after we use NE-Ring to discretize the embeddings, the overlapping of green list after changing embeddings by 1e-5 is 100%, which means the discretization can effectively ignore this change. In practice, SemaMark can provide the tolerance that is much larger than 1e-5, which makes the watermark more robust under paraphrase. With the improvement of -offset, the detection of SemaMark is more robust and effective.
5 Q𝑄Q-offset detection
In this subsection, we show that the effectiveness of the proposed -offset detection. In Figure 5, we demonstrate the change of ROC-AUC of SemaMark with different in offset detection under three different paraphrases. -offset detection searches the highest -statistics from to as the -offset score. From Figure 5, we can see that when increases, ROC-AUC first increases and decreases after is around 15. When , the offset can help correct the errors of semantic values close to the boundary. Compared with detection without offset, i.e. , ROC-AUC of SemaMark is much better, which means the offset can help to solve the errors of semantic values around the boundaries that are more vulnerable to paraphrases. When , the correction of this error is limited, meanwhile, the offset will also increase the -offset score of negative samples since they are also searching the highest -statistics of negative samples. On the other hand, the computation cost will also increase if is too large because it has to search more possible . In practice, we set in all the experiments, which can effective reduce the influence of the errors of semantic values at the boundaries.
Since the -offset detection searches the highest green word fraction, the fraction of green list word of non-watermarked text will be higher than the that we used to randomly select the green list. Thus, it is not accurate to use the original for -statistics. We treat as a hyper-parameter and use a validation set to select its value. As shown in Figure 6, the detection performance of SemaMark under paraphrases of Dipper and GPT-3.5 will reach the highest when is around , while it will continue to increase under round-trip translation. In practice, we set for -offset detection.
6 Distribution on NE-Ring based on CL
In this subsection, we demonstrate that Contrastive Learning can help evenly distribute the semantics on the NE-Ring. The even distribution can help the sequences reach all possible semantic values and provide more diverse semantic values to prevent the watermark from being cracked by counting token frequency. In Figure 7(a), we use Gaussian density estimation Chen (2017) to get the distribution of the semantics on the NE-Ring before discretization. We use different colors to show the density. The NE-Ring in Figure 7(a) shows that, the distribution is uniform. All the density is between 0.052 and 0.054. We further plot the density based on the polar angle in Figure 7(b) where the density has almost no change on all the polar angle from 0 to . This implies that the training based on Contrastive Learning can ensure the semantics will reach all possible discrete values. It can prevent the case where the discrete values will gather in some discrete sections and produce monotonous vocabulary partitions. As a result, it can protect the watermark from being cracked by counting token frequency.
Conclusion
In this paper, we use the semantic information for watermarking to enhance the robustness against paraphrase. The existing watermark methods use the matching between the previous tokens and the partition vocabulary. This matching can be easily broken by paraphrasing. However, we construct the mapping between the semantics and the vocabulary. In this way, the semantics will stay stable under paraphrase and the robustness of watermark will increase. To make use of semantics, we propose SemaMark to discrete the embedding space on NE-Ring and propose a training method based on Contrastive Learning. In addition, we use -offset detection to further increase the robustness by increasing the tolerance of the semantic values close to the discrete boundary. In our experiments, we demonstrate our method can perform much better compared with baseline methods under paraphrase with little influence on the generation quality.