Robust Multi-bit Natural Language Watermarking through Invariant Features

KiYoon Yoo, Wonhyuk Ahn, Jiho Jang, Nojun Kwak

Introduction

Digital watermarking is a technology that enables the embedding of information into multimedia (e.g. image, video, audio) in an unnoticeable way without degrading the original utility of the content. Through embedding information such as owner/purchaser ID, its application includes leakage tracing, ownership identification, meta-data binding, and tamper-proofing. To effectively combat intentional evasion by the adversary or unintentional digital degradation, a watermarking framework should not only be able to embed adequate bits of information but also demonstrate robustness against potential corruption Tao et al. (2014); Zhu et al. (2018). Watermarking in image and video contents has been extensively explored for pre-deep learning methods (Hsu and Wu, 1999; Wolfgang et al., 1999; Wang et al., 2001). With the advent of deep neural networks, deep watermarking has emerged as a new paradigm that improves the three key aspects of watermarking: payload (i.e. the number of bits embedded), robustness (i.e. accuracy of the extracted message), and quality of the embedded media.

Natural language watermarking uses text as the carrier for the watermark by imperceptibly modifying semantics and/or syntactic features. As opposed to altering the visual appearances Rizzo et al. (2019), this type of modification makes natural language watermarking resistant to piracy based on manual transcription. Previous research has focused on techniques such as lexical substitution with predefined rules and dictionaries or structural transformation (Topkara et al., 2006a, b; Atallah et al., 2001). Through utilizing neural networks, recent works have either replaced the predefined set of rules with learning-based methodology (Abdelnabi and Fritz, 2021, AWT), thereby removing heuristics or vastly improved the quality of lexical substitution (Yang et al., 2022, ContextLS). Despite the superiority over traditional methods, however, recent works are not without their limitations: AWT is prone to error during message extraction especially when a higher number of bits are embedded and occasionally generates deteriorated watermarked samples due to its entire reliance on a neural network; ContextLS has a fixed upper-bound on the payload and more importantly, does not consider extracting the bit message under corruption, which leads to low robustness. This work strives to advance both payload and robustness of natural language watermarking.

To build an effective robust watermarking system for natural language, we draw inspiration from a well-known proposition of a classical image watermarking work (Cox et al., 1997): That watermarks should "be placed explicitly in the perceptually most significant components" of an image. If this is achieved, the adversary must corrupt the content’s fundamental structure to destroy the watermark. This degrades the utility of the original content, rendering the purpose of pirating futile.

However, embedding the watermark directly on the "perceptually most significant components" is only possible for images due to the inherent perceptual capacity of images. That is, modification in individual pixels is much more imperceptible than on individual words. Due to this, while we adhere to the gist of the proposition, we do not embed directly on the most significant component. Instead, we identify features that are semantically or syntactically fundamental components of the text and thus, invariant to minor modifications in texts. Then we use them as anchor points to pinpoint the position of watermarks. After formulating a general framework for robust natural watermarking, we empirically study the effectiveness of various potential invariant features derived from the semantic and syntactic components. Through step-by-step analysis of the possible sources of errors during watermark extraction, we further propose a corruption-resistant infill model that is trained explicitly to be robust on possible types of corruption.

Our experimental results encompassing four datasets of various writing styles demonstrate the robustness of (1) relying on invariant features for watermark embedding (2) using a robustly trained infill model. The absolute robustness improvement of our full method compared with the previous work is +16.8% point on average on the four datasets, three corruption types, and two corruption ratios.

Preliminaries

Conversely, the adversary attempts to interfere with the message extraction phase by corrupting the watermarked text, while maintaining the original utility of the text. For instance, an illegal pirating party will want to avoid the watermark being used to trace the leakage point while still wanting to preserve the text for illegal distribution. This constrains the adversary from corrupting the text too much both quantitatively and qualitatively. To this end, we borrow techniques from adversarial attack Jin et al. (2020); Morris et al. (2020a) to alter the text and maintain its original semantics.

We consider word insertion Li et al. (2021), deletion Feng et al. (2018), and substitution Garg and Ramakrishnan (2020) across 2.5% to 5.0% corruption ratios of the number of words in each sentence following Abdelnabi and Fritz (2021). The number of words inserted/substituted/deleted is equal to \textscround(CR×N)\textsc{round}(CR\times N) where CRCR is the corruption ratio and NN is the number of words in the sentence. This ensures shorter sentences containing little to no room for corruption are not severely degraded. To additionally constrain the corrupted text from diverging from the original text, we use the pre-trained sentence transformerhttps://www.sbert.net/ all-MiniLM-L6-v2, which was trained on multiple datasets consisting of 1 billion pairs of sentences, to filter out corrupted texts that have cosine similarity less than 0.98 with the original text.

3 Infill Model

Similar to ContextLS Yang et al. (2022), we use a pre-trained infill model to generate the candidates of watermarked sets. Given a masked sequence X\i={x1,⋯ ,xi−1,MASK,xi+1,⋯ ,xt}X_{\backslash i}=\{x_{1},\cdots,x_{i-1},\text{MASK},x_{i+1},\cdots,x_{t}\}, an infill language model can predict the appropriate words to fill in the mask(s). An infill model parameterized by θ\theta outputs the probability distribution of xix_{i} over the vocabulary (vv):

We denote the set of top-kk token candidates outputted by the infill model as

Framework for Robust Natural Language Watermarking

For the watermarking system to be robust against corruption, SS should be chosen such that it depends on the properties of the text that are relatively invariant to corruption. That is, SS should be a function of the invariant features of the text. More concretely, an ideal invariant feature is characterized by:

A significant portion of the text has to be modified for it to be altered.

Thus, it is invariant to the corruptions that preserve the utility (e.g. semantics, nuance) of the original text.

We sought to discover invariant features in the two easily attainable domains in natural language: semantic and syntactic components. An illustration of these components is shown in Figure 1 Left.

Keyword Component On the semantic level, we first pinpoint keywords that ought to be maintained for the utility of the original text to be maintained. Our intuition is that keywords are semantically fundamental parts of a sentence and thus, are maintained and invariant despite corruption. This includes proper nouns as they are often not replaceable with synonyms without changing the semantics (e.g. name of a movie, person, region), which can be extracted by an off-the-shelf Named Entity Recognition model. In addition, we use an unsupervised method called YAKE (Campos et al., 2018) that outputs semantically essential words. After extracting the keywords, we use them as anchors and can determine the position of the masks by a simple heuristic. For instance, the word adjacent to the keyword can be selected as the mask.

Syntactic Dependency Component On the syntactic level, we construct a dependency parsing tree employing an off-the-shelf parser. A dependency parser describes the syntactic structure of a sentence by constructing a directed edge between a head word and its dependent word(s). Each dependent word is labeled as a specific type of dependency determined by its grammatical role. We hypothesize that the overall grammatical structure outputted by the parsing tree will be relatively robust to minor corruptions in the sentence. To select which type of dependency should be masked, we construct a predefined ordering to maintain the semantics of the watermarked sentences. The ordering is constructed by masking and substituting each type of dependency using an infill model and comparing its entailment score computed by an NLI model(e.g. RoBERTa-Large-NLIhttps://huggingface.co/roberta-large-mnli) on a separate held-out dataset as shown in Alg. 1 (a more detailed procedure and the full list are provided in the Appendix A.4). Using the generated ordering, we mask each dependency until the target number of masks is reached. For both types of components (semantic & syntactic), we ensure that keywords are not masked.

So how well do the aforementioned components fare against corruption? The results in Table 1 bolster our hypothesis that keywords and syntactic components may indeed act as invariant features as both show considerably high robustness across three different types of corruption measured by the ratio of mask matching samples. As opposed to this, ContexLS (Yang et al., 2022), which does not rely on any invariant features has a drastically lower Rg1\mathcal{R}_{g_{1}}. This signifies that a different word is masked out due to the corruption, which hampers the watermark extraction process.

2 Phase 2: Watermark Encoding

In Phase 2, a set of valid watermarked texts is generated by g2(X,S)g_{2}(X,S) to embed or extract the message. For ours, since the state is the set of mask positions, this comprises using an infill model to select top-kk words and alphabetically sort them to generate a valid set of watermarks. Concretely, using the notations from section 2.3, g2(X,S)g_{2}(X,S) can be divided into the following steps:

Ti={t1i,⋯ ,tki}=\textscinfill(X\i;k1),∀i∈S\begin{aligned} \mathcal{T}_{i}&=\{t_{1}^{i},\cdots,t_{k}^{i}\}=\textsc{infill}(X_{\backslash i};k_{1}),\forall i\in S\end{aligned}

Filter Ti\mathcal{T}_{i} to remove any punctuation marks, subwords, stopwords. Update Ti\mathcal{T}_{i} by selecting top-k2k_{2} (≤k1\leq k_{1}) and sort them alphabetically.

However, what happens when there is corruption in the watermarked texts? Even if the exact state is recovered, the same set of watermarked texts may not be recovered as the infill model relies on local contexts to fill in the masks. Noting this in mind, we can also define the robustness of g2g_{2} as

Interestingly, for ContextLS the gap between Rg1\mathcal{R}_{g_{1}} and Rg2\mathcal{R}_{g_{2}} is nearly zero, showing that Phase 1 is already a bottleneck for achieving robustness. The smaller gap can be explained by the use of smaller top-k2k_{2}(=2) and the incremental watermarking scheme, which incrementally increases the sequence to infill. This may reduce the possibility of a corrupted word influencing the infill model.

3 Robust Infill Model

In addition, to improve the training dynamics, we follow the masking strategy proposed in section 3.1 to choose the words to masks, instead of following the random masking strategy used in the original pretraining phase. This aligns distributions of the masked words at train time and test time, which leads to a better performance (robustness) given the same compute time. As opposed to this, since the original masking strategy randomly selects a certain proportion of words to mask out, this will provide a weaker signal for the infill model to follow.

We use the Kullback–Leibler (KL) divergence as our metric. More specifically, we use the ‘reverse KL’ as our loss term in which the predicted distribution (as opposed to the target distribution) is used to weigh the difference of the log distribution as done in Variational Bayes Kingma and Welling (2014). This aids the model from outputting a "zero-forcing" predicted distribution. The consistency loss between the two distributions is defined by

for all ii of the masked tokens. The graph outputting pp is detached to train a model to output a consistent output when given a corrupted input. As we expected, using the robust infill model to the Syntactic component leads to a noticeable improvement in Rg2\mathcal{R}_{g_{2}}, while that of Rg1\mathcal{R}_{g_{1}} is negligible (Table 2).

The corrupted inputs are generated following the same strategy in section 2.2 using a separate train dataset. We ablate our design choices in section 5.3.

allows the embedding and extraction of watermarks faultlessly when there is no corruption.

can incorporate invariant features for watermark embedding, achieving robustness in the presence of corruption.

further enhance robustness in Phase 2 by utilizing a robust infill model.

Experiment

Dataset To evaluate the effectiveness of the proposed method, we use four datasets with various styles. IMDB (Maas et al., 2011) is a movie reviews dataset, making it more colloquial. WikiText-2 Merity et al. (2016), consisting of articles from Wikipedia, has a more informative style. We also experiment with two novels, Dracula and Wuthering Heights (WH), which have a distinct style compared to modern English and are available on Project Gutenberg (Bram, 1897; Emily, 1847).

Metrics For payload, we compute bits per word (BPW). For robustness, we compute the bit error (BER) of the extracted message. We also measure the quality of the watermarked text by comparing it with the original cover text. Following Yang et al. (2022); Abdelnabi and Fritz (2021), we compute the entailment score (ES) using an NLI model (RoBERTa-Large-NLI) and semantic similarity (SS) by comparing the cosine similarity of the representations outputted by a pre-trained sentence transformer (stsb-RoBERTa-base-v2). We also conduct a human evaluation study to assess semantic quality.

Implementation Details For ours and ContextLS (Yang et al., 2022), both of which operate on individual sentences, we use the smallest off-the-shelf model (en-core-web-sm) from Spacy (Honnibal and Montani, 2017) to split the sentences. The same Spacy model is also used for NER (named entity recognizer) and building the dependency parser for ours. Both methods use BERT-base as the infill model and select top-32 (k1k_{1}) tokens. We set our payload to a similar degree with the compared method(s) by controlling the number of masks per sentence (∣S∣|S|) and the top-k2k_{2} tokens (section 3.2); these configurations for each dataset are shown in Appendix Table 12. We watermark the first 5,000 sentences for each dataset and use TextAttack Morris et al. (2020b) to create corrupted samples. For robust infilling, we finetune BERT for 100 epochs on the individual datasets. For more details, refer to the Appendix.

Compared Methods We compare our method with deep learning-based methods (Abdelnabi and Fritz, 2021, AWT)(Yang et al., 2022, ContextLS) for our experiments as pre-deep learning methods (Topkara et al., 2006b; Hao et al., 2018) that are entirely rule-based have low payload and/or low semantic quality (later shown in Table 4). More details about the compared methods are in section 6.

Table 3 shows the watermarking results on all four datasets. Some challenges we faced during training AWT and our approach to overcoming this are detailed in Appendix A.2. Since the loss did not converge on IDMB for AWT as detailed in section A.3, we omit the results for this.

We test the robustness of each method on corruption ratios (CR) of 2.5% and 5%. For ours, we apply robust infilling for the Syntactic Dependency Component, which is indicated in the final column by +RI. AWT suffers less from a larger corruption rate and sometimes outperforms our methods without RI. However, the BER at zero corruption rate is non-negligible, which is crucial for a reliable watermarking system. In addition, we observe qualitatively that AWT often repeats words or replaces pronouns on the watermarked sets, which seems to provide signals for extracting the message – this may provide a distinct signal for message extraction at the cost of severe quality degradation. Some examples are shown in Appendix A.7 and Tab. 17-19.

Our final model largely outperforms ContextLS in all the datasets and corruption rates. Additionally, both semantic and syntactic components are substantially more robust than ContextLS even without robust infilling in all the datasets. The absolute improvements in BER by using Syntactic component across corruption types with respect to ContextLS under CR=2.5% are 13.6%, 8.2%, 14.4%, and 12.9% points for the four datasets respectively when using the Syntactic component; For CR=5%, they are 10.0%, 10.2%, 11.0%, and 11.7% points.

2 Semantic Scores of Watermark

Table 4 shows the results for semantic metrics. While our method falls behind ContextLS, we achieve better semantic scores than all the other methods while achieving robustness. ContextLS is able to maintain a high semantic similarity by explicitly using an NLI model to filter out candidate tokens. However, the accuracy of the extracted message severely deteriorates in the presence of corruption as shown in the previous section. Using ordered dependencies sorted by the entailment score significantly increases the semantic metrics than using a randomly ordered one, denoted by "–NLI Ordering". The results are in Appendix Table 15.

We also conduct human evaluation comparing the fluency of the watermarked text and cover text (FluencyΔ\Delta) and how much semantics is maintained (Semantic Similarity; SS) compared to the original cover text in Tab. 5. The details of the experiment are in section A.6. This is aligned with our findings in automatic metrics, but shows a distinct gap between ours and AWT. Notably, the levels of fluency change of ours and ContextLS compared to the original cover text are nearly the same.

Discussion

Some design choices we differ from ContextLS is top-k2>2k_{2}>2 which determines the number of candidate tokens per mask. We can increase the payload depending on the requirement by choosing a higher k2k_{2}. However, for ContextLS increasing k2k_{2} counter-intuitively leads to a lower payload. This is because ContextLS determines the valid watermark sets (those that can extract the message without error) with much stronger constraints (for details see Eq. 5,6,7 of Yang et al. (2022)). This also requires an exhaustive search over the whole sentence with an incrementally increasing window, which leads to a much longer embedding / extraction time due to the multiple forward passes of the neural network. For instance, the wall clock time of embedding in 1000 sentences on IMDB is more than 20 times on ContextLS (81 vs. 4 minutes). More results are summarized in Table 6. Results for applying our robust infill model to ContextLS are in Appendix A.4.

2 Pitfalls of Automatic Semantic Metrics

Although the automatic semantic metrics do provide a meaningful signal that aids in maintaining the original semantics, they do not show the full picture. First, the scores do not accurately reflect the change in semantics when substituting for the coordination dependency (e.g. and, or, nor, but, yet). As shown in Table 7, both the entailment score and semantic similarity score overlook some semantic changes that are easily perceptible by humans. This is also reflected in the sorted dependency list we constructed in section 3.1 - the average NLI score after infilling a coordination dependency is 0.974, which is ranked second. An easy fix can be made by placing the coordination dependency at the last rank or simply discarding it. We show in Appendix Table 11 that this also provides a comparable BPW and robustness.

Another pathology of the NLI model we observed was when a named entity such as a person or a region is masked out. Table 7 shows an example in ContextLS and how ES is abnormally high. Such watermarks may significantly hurt the utility of novels if the name of a character is modified. This problem is circumvented in ours by disregarding named entities (detected using NER) as possible mask candidates.

3 Ablations and Other Results

Ablations In this section, we ablate some of the design choices. First, we compare the design choices of our masking strategies (random vs. ours) and loss terms (Forward KL and Reverse KL) in Table 8. Our masking strategy improves both BPW and robustness compared to randomly masking out words. Though preliminary experiments showed RKL is more effective for higher payload and robustness, further experiments showed the types of KL do not significantly affect the final robustness when we use our masking strategy. We further present the results under character-based corruption and compare robustness against different corruption types in Appendix A.4.

Stress Testing Syntactic Component We experiment with how our proposed Syntactic component fares in a stronger corruption rate. The results are shown in Appendix Fig. 3. While the robustness is still over 0.9 for both insertion and substitution at CR=0.1, the robustness rapidly drops against deletion. This shows that our syntactic component is most fragile against deletion.

Related Works

Natural language watermarking embeds information via manipulation of semantics or syntactic features rather than altering the visual appearance of words, lines, and documents (Rizzo et al., 2019). This makes natural language watermarking robust to re-formatting of the file or manual transcription of the text (Topkara et al., 2005). Early works in natural language watermarking have relied on synonym substitution (Topkara et al., 2006b), restructuring of syntactic structures (Atallah et al., 2001), or paraphrasing (Atallah et al., 2003). The reliance on a predefined set of rules often leads to a low bit capacity and the lack of contextual consideration during the embedding process may result in a degraded utility of the watermarked text that sounds unnatural or strange.

With the advent of neural networks, some works have done away with the reliance on pre-defined sets of rules as done in previous works. Adversarial Watermarking Transformer (Abdelnabi and Fritz, 2021, AWT) propose an encode-decoder transformer architecture that learns to extract the message from the decoded watermarked text. To maintain the quality of the watermarked text, they use signals from sentence transformers and language models. However, due to entirely relying upon a neural network for message embedding and extraction, the extracted message is prone to error even without corruption, especially when the payload is high and has a noticeable artifact such as repeated tokens in some of the samples. Yang et al. (2022) takes an algorithmic approach for embedding and extraction of messages, making it errorless. Additionally, using a neural infill model along with an NLI model has shown better quality in lexical substitution than more traditional approaches (e.g. WordNet). However, robustness under corruption is not considered.

Image Watermarking Explicitly considering corruption for robustness and using different domains of the multimedia are all highly relevant to blind image watermarking, which has been extensively explored (Mun et al., 2019; Zhu et al., 2018; Zhong et al., 2020; Luo et al., 2020). Like our robust infill training, Zhu et al.; Luo et al. explicitly consider possible image corruptions to improve robustness. Meanwhile, transforming the pixel domain to various frequency domains using transform methods such as Discrete Cosine Transform has shown to be both effective and more robust (Potdar et al., 2005). The use of keywords and dependencies to determine the embedding position in our work can be similarly considered as transforming the raw text into semantic and syntactic domains, respectively.

Other Lines of Work Steganography is a similar line of work concealing secret data into a cover media focusing on covertness rather than robustness. Various methods have been studied in the natural language domain (Tina Fang et al., 2017; Yang et al., 2018; Ziegler et al., 2019; Yang et al., 2020; Ueoka et al., 2021). This line of works differs from watermarking in that the cover text may be arbitrarily generated to conceal the secret message, which eases the constraint of maintaining the original semantics.

Recently, He et al. (2022a) proposed to watermark outputs of language models to prevent model stealing and extraction. While the main objective of these works (He et al., 2022a, b) differs from ours, the methodologies can be adapted to watermark text directly. However, these are only limited to zero-bit watermarking (e.g. whether the text is from a language model or not), while ours allow embedding of any multi-bit information. Similarly, Kirchenbauer et al. (2023) propose to watermark outputs of language models at decoding time in a zero-bit manner to distinguish machine-generated texts from human-written text.

Conclusion

We propose using invariant features of natural language to embed robust watermarks to corruptions. We empirically validate two potential components easily discoverable by off-the-shelf models. The proposed method outperforms recent neural network-based watermarking in robustness and payload while having a comparable semantic quality. We do not claim that the invariant features studied in this work are the optimal approach. Instead, we pave the way for future works to explore other effective domains and solutions following the framework.

Limitations

Despite its robustness, our method has subpar results on the automatic semantic metrics compared to the most recent work. This may be a natural consequence of the perceptibility vs. robustness trade-off Tao et al. (2014); De Vleeschouwer et al. (2002): a stronger watermark tends to interfere with the original content. Nonetheless, by using some technical tricks (e.g. neural infill model, NLI-sorted ordering) our method is able to be superior to all the other methods including two traditional ones and a neural network-based method.

Techniques from adversarial attack were employed to simulate possible corruptions in our work. However, these automatic attacks does not always lead to imperceptible modifications of the original texts (Morris et al., 2020a). Thus, the corruptions used in our work may be a rough estimate of what true adversaries might do to evade watermarking. In addition, our method is not tested against paraphrasing, which may substantially change the syntactic component of the text. One realistic reason that deterred us from experimenting on paraphrase-based attacks was their lack of controllability compared to other attacks that have fine-grained control over the number of corrupted words. Likewise, for text resources like novels that value subtle nuances, the aforementioned property may discourage the adversary from using it to destroy watermarking.

Acknowledgements

This work was supported by Korean Government through the IITP grants 2022-0-00320, 2021-0-01343, NRF grant 2021R1A2C3006659 and by Webtoon AI at NAVER WEBTOON in 2022.

References

Appendix A Appendix

Dataset Split Following ContextLS, we subsampled the first 5000 sentences and used the same subset across all methods. Our preliminary experiments showed subsampling other samples only led to minor variability: standard error of the mean BPW across 3 trials 0.002. We use the same subset for all our experiments to avoid any confounding factors. For the robustness experiment, which had a stochastic element, the standard errors for BER’s for insertion and substitution were also marginal (both 0.004) compared to the performance gap.

To finetune our robust infill model, we required a train set other than the test set that will be watermarked. For IMDB and Wikitext-2, we used the original training split. For the novels datasets, we take the first 40% of the text as the train set and the rest as the test set. The same splits are also used for training AWT as well.

Corruption To test the robustness, we corrupt the first 1000 sentences of the 5000 test sets. Since the watermark embedding processes for ours and ContextLS are deterministic given the message, we run the embedding experiment once for a fixed random seed. Due to the implementation of TextAttack, some corruption modules may be non-deterministic, which will lead to a non-deterministic BER. We find that the deletion module we used is deterministic so we run the robustness experiment once. On the other hand, we create five corrupted samples per sample for insertion and substitution and report the mean for ours and ContextLS.

Computation Time The actual watermarking process does not require gradient computation. The largest bottleneck in the pipeline is the forward passes of the infill model. The actual wall clock time and the number of passes are detailed on section 5.1. Training the infill model requires the most computation time. We finetune all our models in a single GPU environment using either Titan RTX or RTX 3090. Finetuning on Wikitext-2 was the longest among the datasets, which required approximately 22 GPU-hours for 100 epochs.

Training Details of Infill Model We use AdamW with a learning rate of 5e-5 using linear warmup 0.1 of the total training steps. All our models are trained for 100 epochs and we used the last checkpoint. For random masking, we simply mask out 15% of the words using whole word masking strategy.

A.2 AWT Implementation Details

We use the official implementation and mostly adhere to the hyperparameters employed by AWT unless otherwise noted. In the original paper, the experiment was conducted only for a lower payload BPW=0.05 on the Wikitext-2 dataset, so implementation details for a higher payload BPW=0.1 or other datasets needed to be adjusted.

First, we replaced the AWD-LSTM language model with GPT-2, providing a superior language modeling capability. Second, when the payload was increased to BPW=0.1, the weighting term for the reconstruction loss (see Section IV-D) was doubled at the second training stage of AWT to make the model converge. Third, we combined data for Dracula and Wuthering Heights into a single dataset to train and evaluate the AWT model because we were unable to train the model for each dataset separately due to a lack of data.

For a fair comparison in robustness experiments, watermarked segments are concatenated and then split into sentences, to which corruption is applied on a per-sentence basis. Lastly, the corrupted segments are used to report BER against attacks. In addition, AWT constructs a dictionary of tokens using the corpus before watermarking embedding. This may introduce unknown tokens for insertion and substitution, in which case we exclude these tokens.

A.3 AWT on IMDB dataset

The text reconstruction loss did not converge for the IMDB datasets. This led to a severe quality decrease in the watermarked sentence as shown below in Table 13. We nevertheless test the robustness under corruption. The BER@CR=0.05 for the three corruption types were 0.283, 0.278, and 0.299.

A.4 More Results

Ordering of NLI and Discarding Coordination To define the ordering of syntactic dependency, we mask out each of the dependencies on the train set and then infill the masked-out dependencies. The infilled sentences are compared with the original sentence. A Pythonic algorithm for one sample is shown Alg. 1. This is done for 500 samples of IMDB. The resultant ordering is shown in Table 14.

As discussed in section 5.2, substituting the coordination dependency (CC) is often leads to a semantic drift that is undetectable by automatic metrics. We also provide the BPW and robustenss results after discarding CC from the NLI ordering list in Table 11.

Character-based Corruption We also experiment with character-based corruption, which may happen when unintentionally during manual transcription. We simulate this type of corruption by randomly swapping a character with a neighboring character using TextAttack. Similar to our main experiment, we test on CR={2.5%, 5%}. On the IMDB dataset, our Syntactic Dependency Component model has a BER of .079 and .167, respectively. While our RI model did not explicitly train on this type of error, it nevertheless improves robustness to 0.063 and 0.142, respectively.

ContextLS + Robust Infill Using a finetuned infill model gave a meaningful boost in robustness in all datasets for our method. Is this model effective for ContextLS as well? Using an infill model trained using random masks is not always beneficial to the robustness of ContextLS and the improvement is marginal compared to that of ours (Appendix Table 16). This is expected given our analysis in section 3.1 that Phase 1 is a strong bottleneck for ContextLS, yet we believe it can be further improved if a specific masking strategy used in ContextLS is adapted when finetuning the infill model.

A.5 More Discussions

Computing BER For ours and ContextLS, the number of bits varies by sentence. This leads to an issue when computing BER as the predicted message may have less or more bits than the true message. To accurately assess BER, we assume that the true number of bits is unknown during extraction. When the extracted number of bits is less than the ground truth, we consider all unpredicted bits as errors. Conversely, when more bits are extracted, we truncate them and consider all over-extracted bits as errors.

A.6 Human Evaluation

We collected human annotations of the watermarked texts through ClickWorker and disclosed the responses may be used for research purposes. The workers were recruited from United States, United Kingdom, and Ireland at the age of 20-99 who considered themselves with English as their native languages. The survey was designed to take approximately 40-60 minutes and the fee was 20 Euros, which was over the minimum wages of the three countries. We only used the responses that had an adequately high "semantic was completely maintained" answer proportion for those watermarked texts that were not altered from the cover text to ensure the instructions were followed. When thresholding this proportion by 0.5, 2 responses were discarded out of the 7 responses. Screenshots of the survey are in the last page in Figure 4. The survey consisted of 10 random samples each from Dracula and Wuthering Heights. We excluded Wikitext-2 as AWT preprocessed the name of the entities as unknown tokens, which may lead to substantial decrease in fluency for the annotators. IMDB was excluded as the text reconstruction loss did not converge for AWT, which led to incomprehensible sentences. Part 1 consisted of rating the fluency of each sentence including the original cover text. Fluency Δ\Delta was computed by subtracting the fluency of the watermarked sample from the original one. Part 2 consisted of rating how much semantics is maintained given the reference sentence (cover text).

A.7 Watermarked Examples

Examples of watermarked texts are provided in Table 17-20. The watermarked words are marked by color. For ours and ContextLS, some texts may be unaltered from the cover text if the original text is included in the valid watermarked sets. For AWT, this is only possible if the watermark has been embedded at a different section of the segment since it usually takes multiple sentences (40 words) as inputs. Thus, we display only those examples that have been modified for qualitative analysis. (Conversely, for human evaluation, we randomly sample sentences.) For Wikitext-2, which contains considerable amount of entities, many of the entities have been marked as unknown tokens on AWT outputs. We manually substitute these tokens for presentation purposes.