SynGEC: Syntax-Enhanced Grammatical Error Correction with a Tailored GEC-Oriented Parser

Yue Zhang, Bo Zhang, Zhenghua Li, Zuyi Bao, Chen Li, Min Zhang

Introduction

Given an ungrammatical sentence, the grammatical error correction (GEC) task aims to produce a grammatical target sentence with the intended meaning (Grundkiewicz et al., 2020; Wang et al., 2021). Recent mainstream approaches treat GEC as a monolingual machine translation (MT) task (Yuan and Briscoe, 2016; Junczys-Dowmunt et al., 2018). Standard encoder-decoder based MT models, e.g., Transformer (Vaswani et al., 2017), have emerged as a dominant paradigm and achieved state-of-the-art (SOTA) results on various GEC benchmarks (Rothe et al., 2021; Stahlberg and Kumar, 2021; Sun et al., 2022; Zhang et al., 2022). Despite their impressive achievements, most work treats the input sentence as a sequence of tokens, without explicitly exploiting syntactic or semantic information.

Compared with MT, GEC has two peculiarities that directly motivate this work. First, the training data for GEC models is much less abundant, which may be alleviated by incorporating linguistic structure knowledge like syntax. As shown in Table 1 and Table 5, the English and Chinese GEC tasks only have about 126K and 157K high-quality labelled source/target sentence pairs for training, if not considering the highly noisy crowd-annotated Lang8 data (Mita et al., 2020). Second, according to our preliminary observation, many errors in ungrammatical sentences are intrinsically correlated with syntactic information. For example, errors like inconsistency in tense or singular-vs-plural forms can be better detected and corrected with the help of long-range syntactic dependencies.

In this paper, we propose SynGEC, an approach that can effectively inject the syntactic structure of the input sentence into the encoder part of GEC models. The critical challenge here is that off-the-shelf parsers are unreliable when handling ungrammatical sentences. On the one hand, off-the-shelf parsers are trained on clean treebanks that only consist of grammatical sentences. When parsing ungrammatical sentences, their performance may sharply degrade due to the input mismatch. On the other hand, mainstream syntax representation schemes, adopted by existing treebanks, do not cover the non-canonical structures arising from grammatical errors. In consequence, under such schemes, it is sometimes difficult to find a plausible syntactic tree to properly parse an ungrammatical sentence (e.g., the sentence in Figure 1(d)).

Indeed, there have been several prior works that try to improve syntactic parsing for ungrammatical texts by annotating data (Dickinson and Ragheb, 2009; Berzak et al., 2016; Nagata and Sakaguchi, 2016). However, these works do not extend existing syntax representation schemes to accommodate errors, which means that they make little change on the original syntactic label sets. Besides, manual annotation is expensive and time-consuming, so their annotated treebanks for ungrammatical sentences are of a relatively small scale.

To confront the challenge of unreliable performance of off-the-shelf parsers on ungrammatical sentences, we propose to train a tailored GEC-oriented parser (GOPar). The basic idea is to utilize parallel source/target sentence pairs in the GEC training data. First, we parse the target correct sentences using a vanilla off-the-shelf parser. Then, we construct the tree for the source incorrect sentences via tree projection. To accommodate grammatical errors, we propose an extended syntax representation scheme based on several straightforward rules, which allows us to represent both grammatical errors and syntax in a unified tree structure. Finally, we train GOPar directly on the automatically constructed trees of the source incorrect sentences in the GEC training data. During both GEC training and evaluation procedures, GOPar is used to generate syntactic information for the input sentences.

To incorporate syntactic information provided by GOPar, we cascade several label-aware graph convolutional network (GCN) layers (Kipf and Welling, 2017; Zhang et al., 2020a) above the encoder of our baseline Transformer-based GEC model. We conduct experiments on two widely-used English GEC evaluation datasets, i.e., CoNLL-14 (Ng et al., 2014) and BEA-19 (Bryant et al., 2019), and two Chinese GEC evaluation datasets, i.e., NLPCC-18 Zhao et al. (2018) and MuCGEC (Zhang et al., 2022). Extensive experimental results and in-depth analyses show that our SynGEC approach achieves consistent and substantial improvement on all datasets, even when the baseline model is enhanced with large pre-trained language models (PLMs) like BART (Lewis et al., 2020), and outperforming previous SOTA systems under comparable settings.

Our GEC-Oriented Parser

This section describes our tailored GOPar, a dependency parser that is more competent in parsing ungrammatical sentences than off-the-shelf parsers.

The standard scheme for representing dependency syntax is originally designed for grammatical sentences, and thus may not cover many non-canonical structures in grammatically erroneous sentences. Therefore, to obtain a tailored parser, our first task is to extend the syntax representation scheme and, more specifically, to design a complementary set of rules to handle different grammatical mistakes. With this scheme, we can directly use a unified tree structure to represent both grammatical errors and syntactic information.

As shown in Figure 1, we propose a light-weight extended scheme based on several straightforward rules, corresponding to the three types of grammatical errors, i.e., substituted, redundant and missing (Bryant et al., 2017). We treat word-order errors as the combination of redundant and missing errors in this work. In this work, we use the Stanford Dependencies Scheme v3.3.0 Dozat and Manning (2017) as the basic scheme. The rules are designed in such a way that we make as few adjustments as possible to the syntactic tree of the target correct sentence during tree projection. Correspondingly, we add three labels into the original syntactic label set, i.e., “S”, “R” and “M”, to capture three kinds of errors. Since such categorization is also adopted in the grammatical error detection (GED) task (Yuan et al., 2021), we refer to them as GED labels.

Substituted errors (S) include spelling errors, tense errors, singular/plural inconsistency errors, etc. For simplicity, we do not consider such fine-grained categories, and directly use a single “S” label to indicate that the word should be replaced by another one, as shown in Figure 1(b).

Redundant errors (R) mean that some words should be deleted. For each redundant word, we let it depend on its right-side adjacent word, with a label “R”,Even if the right-side adjacent word is redundant. as shown in Figure 1(c). When the redundant word is at the end of the sentence, we instead let it depend on its left-side adjacent word.

Missing errors (M) mean that some words should be inserted. For each missing word, we assign a label “M” to the incoming arc of its right-side adjacent word, as shown in Figure 1(d). When the missing word is at the end of the sentence, we keep the original tree unchanged. If several consecutive words are missing, the structure remains the same as when a single word is missing. Moreover, since a missing word may have children in the tree of the correct sentence, we let them depend on the head word of the missing word, without changing their syntactic labels.

Limitation discussion. Our extended scheme may encounter problems when different types of errors occur consecutively. Taking “But was no buyers” as an example, we need to replace “was” with “were” and then insert “there” before “were” at the same time. Therefore, according to our rules, the label of the incoming arc of “was” can be either “S” or “M”, leading to a label conflict. To decide a unique label, we simply define a priority order: “S”>>“R”>>“M’. Overall, the current version of our scheme is imperfect, and there are still many points that can be improved. For example, when confronting substituted and redundant errors, some original labels will be overwritten by the GED label “S” and “M”, which may cause the loss of some valuable information. One possible solution is to combine GED and syntax labels and use joint labels like “S-Root”, “M-Subj”, etc. We leave such extensions of our scheme as future work.

2 Training GOPar

With the extended syntax representation scheme, we propose to train our tailored GOPar by using the parallel GEC training data D={(xi,yi)}D=\{(x_{i},y_{i})\} as a pivot. The major goal is to automatically generate high-quality parse trees for large-scale sentences with realistic grammatical errors, and use them to train a parser suitable for parsing ungrammatical sentences. Figure 2 illustrates the workflow, consisting of the following four steps.

First, we use an off-the-shelf parser to parse the target correct sentences (i.e., yiy_{i}) of the GEC training data. The off-the-shelf parser can produce reliable parse trees for target-side sentences since they are (ideally) free from grammatical errors.

Second, we employ ERRANT (Bryant et al., 2017)https://github.com/chrisjbryant/errant to extract all grammatical errors in the source incorrect sentence (i.e., xix_{i}) according to the alignments between xix_{i} and yiy_{i}. The errors extracted by ERRANT mainly contain 3 parts: the start and end positions of errors in source sentences, the corresponding corrections, and the error types.

Third, we construct the tree of xix_{i} by projecting the target-side tree of yiy_{i} to the source side. For words that are not related to any errors, dependencies and labels are directly copied; for those related to errors, dependencies and labels are assigned according to the rules introduced in Section 2.1.

Fourth, with constructed parse trees for all source-side sentences in DD, we then use them as a treebank to train our tailored GOPar.

An alternative way to build GOPar is directly utilizing manually labeled treebanks. Since existing treebanks only contain grammatical sentences, we can inject synthetic errors based on rules or back-translation models (Foster et al., 2008; Cahill, 2015). Then, we can produce parse trees for ungrammatical sentences analogously through the above second and third steps. However, our preliminary experiments show that GOPar built in this way is much inferior and can only slightly improves our baseline GEC model. We suspect that the reasons are two-fold. On the one hand, there is a considerable gap between synthetic and real grammatical errors; on the other hand, the generated data is not enough to train GOPar adequately due to the limited scale of existing treebanks, as GOPar needs to learn to accommodate multifarious errors.

The DepGCN-based GEC Model

This section describes our DepGCN-based GEC model, whose architecture is shown in Figure 3. We adopt GCN (Kipf and Welling, 2017) to encode the dependency syntax trees of the source sentence. Then, we feed the encoded syntactic information into a Transformer-based GEC model.

We model GEC as a sequence-to-sequence task and employ the commonly used Transformer model (Vaswani et al., 2017) as the backbone. The Transformer is composed of an encoder and a decoder. The encoder utilizes the multi-head self-attention mechanism to get the contextualized representations of each token in the source sentence. The decoder has a similar architecture while additionally containing a masked multi-head self-attention module to model the generated token information.

During training, the objective function is to minimize the teacher forcing negative log-likelihood loss (Williams and Zipser, 1989), formally:

where θ\theta is trainable model parameters, xx is the source sentence, y={y1,y2,...,yn}y=\{y_{1},y_{2},...,y_{n}\} is the ground-truth target sentence with nn tokens, and y<t={y1,y2,...,yt−1}y_{<t}=\{y_{1},y_{2},...,y_{t-1}\} is the tokens visible in tt-th training time step.

During inference, we utilize beam search decoding (Wiseman and Rush, 2016) to find an optimal sequence y∗y^{*} by maximizing the conditional probability P(y∗∣x;θ)P(y^{*}\mid x;\theta).

Previous work shows that PLMs, e.g., BART (Lewis et al., 2020) and T5 (Raffel et al., 2020), can improve GEC performance over training from scratch by large margins (Rothe et al., 2021; Sun et al., 2022). In this work, we further use BART to build a stronger baseline, since it shares the same model architecture with our Transformer backbone. Specifically, we use the BART parameters to initialize our Transformer backbone and then continue training on GEC training data. More details are discussed in Section 4.1 and Section 5.

2 Dependency GCN (DepGCN)

We employ DepGCN Zhang et al. (2020b) to encode dependency syntax information. The DepGCN module stacks several identical blocks, and each block is composed of a GCN sub-layer and a feed-forward sub-layer.

For the GCN sub-layer, we introduce the information of the dependency arcs and dependency labels simultaneously. We compute the output hi(l)\mathbf{h}_{i}^{(l)} of ll-th GCN at the ii-th token as:

To reduce the error propagation issue, following Zhang et al. (2020a), we use the arc probability matrix obtained from GOPar as the adjacency matrix AA, which may provide richer syntactic structures. For e(i,j)\mathbf{e}^{(i,j)}, we use the 1-best label of the 1-best head word wjw_{j} of wjw_{j}.

We then feed the outputs of the GCN sub-layer to the feed-forward (FF) sub-layer that contains two linear transformations with a ReLU activation function in between, as shown below:

3 Representation Fusion

In order to balance the contribution of the syntax-aware representations from DepGCN (hisyn\mathbf{h}_{i}^{syn}) and the representations from the basic Transformer encoder (hibasic\mathbf{h}_{i}^{basic}), we use their interpolation (i.e., weighted-sum) as the final representations, which are ultimately fed into the Transformer decoder:

where β∈(0,1)\beta\in(0,1) is a hyper-parameter called the fusion factor, and hifinal\mathbf{h}_{i}^{final} represents the final output vector of the Transformer encoder for the ii-th token. As depicted in Figure 3, this operation is analogous to the residual connection.

Experiments on English GEC

Datasets and evaluation. We first pre-train our model on the cleaned version of the Lang8 dataset (CLang8) CLang8 can be downloaded from https://github.com/google-research-datasets/clang8 released by Rothe et al. (2021). Then, we use the FCE dataset (Yannakoudakis et al., 2011), the NUCLE dataset (Dahlmeier et al., 2013) and the W&I+LOCNESS train-set (Bryant et al., 2019) for model fine-tuning following previous studies. Like Omelianchuk et al. (2020), we decompose the fine-tuning procedure into two stages: 1) fine-tuning on FCE+NUCLE+W&I+LOCNESS; 2) further fine-tuning only on the small-scale but high-quality W&I+LOCNESS.

For evaluation, we report average P/R/F0.5 results over three runs with different random seeds on the CoNLL-14 test setWe use the official-2014.combined.m2 (no-alt) version of CoNLL-14, which is adopted by most existing works. (Ng et al., 2014) evaluated by M2Scorer (Dahlmeier and Ng, 2012) and BEA-19 test set (Bryant et al., 2019) evaluated by ERRANT (Bryant et al., 2017). The BEA-19 dev set serves as validation data during the whole training. The statistics of above datasets are shown in Table 1.

Besides, we also experiment on the small-scale JFLEG test-set Napoles et al. (2017) and list the results in Appendix B.

GEC model details. We adopt Fairseqhttps://github.com/pytorch/fairseq (Ott et al., 2019) to build our Transformer baseline and DepGCN-based model. For the DepGCN-based model, we empirically stack N2=3N_{2}=3 DepGCN blocks and set the fusion factor β=0.5\beta=0.5 in Equation 4. We apply BPE (Sennrich et al., 2016) to generate a 32K shared subword vocabulary. We apply the Dropout-Src mechanism Junczys-Dowmunt et al. (2018) to source-side word embeddings to alleviate over-fitting. More model details are discussed in Appendix A.

GOPar details. The training data for GOPar is generated from the CLang8 dataset (Rothe et al., 2021) using the procedure described in Section 2. For all English experiments, GOPar operates at the word level. In contrast, our GEC models perform on the subword level. To fill this gap, we transform word-level syntax trees into subword-level ones by adding arcs. For example, if wiw_{i} is the head word of wjw_{j}, we will add arcs from all subwords of wiw_{i} to all subwords of wjw_{j}. All added arcs copy the arc probability of wi→wjw_{i}\rightarrow w_{j}. We will explore more sophisticated ways to handle this mismatch issue.

For both the off-the-shelf parser and GOPar, we use the biaffine parsing approach (Dozat and Manning, 2017). We directly adopt the implementation of SuParhttps://github.com/yzhangcs/parser (Zhang et al., 2020b) and follow their default hyper-parameter settings. After comparing several popular PLMs, we choose to use ELECTRA (Clark et al., 2020) to provide the contextual token representations for parsers. To obtain word-level representations, we aggregate the subword-level representations from ELECTRA into word-level ones via average pooling. The off-the-shelf parser is trained on PTB Marcinkiewicz (1994). Please kindly notice that all parsers are always enhanced with ELECTRA, even when the GEC model does not use PLM.

Incorporating BART. To explore whether the syntactic knowledge is still useful after introducing powerful PLMs, we use BART (Lewis et al., 2020) to initialize the Transformer backbone of our models. It is noteworthy that we adopt a two-stage training procedure to keep the training stable. Firstly, we fine-tune the BART-initialized Transformer backbone until it converges. Secondly, we add an auxiliary DepGCN module into the converged Transformer backbone and only tune the DepGCN parameters on the same training data. The intuition behind this procedure is that the BART part has been extensively pre-trained whereas the DepGCN part is just randomly initialized. In our preliminary experiments, we frequently encountered training collapse when training two parts simultaneously, and observed that the scale of the gradients of the two parts vary substantially.

2 Main Results

The main results are listed in Table 2. In the top group of results without PLMs, SynGEC achieves 63.5/68.4 \mboxF0.5\mbox{F}_{0.5} scores on CoNLL-14 and BEA-19 test-sets, respectively, outperforming all other systems utilizing syntax. The performance of SynGEC is only lower than Stahlberg and Kumar (2021) on both test-sets, probably because they use an extra huge synthetic corpus with 540M sentence-pairs. The incorporation of syntactic information provided by GOPar leads to 4.4/4.2 \mboxF0.5\mbox{F}_{0.5} improvements over our baseline, which demonstrates that the tailored syntactic knowledge from GOPar is quite helpful for GEC and our DepGCN-based GEC model can effectively capture it.

In the bottom group of results using PLMs, our SynGEC approach augmented with BART achieves 67.6/72.9 \mboxF0.5\mbox{F}_{0.5} scores, which are comparable or even better than other cutting-edge PLM-enhanced models under similar sizes. After removing syntax, the \mboxF0.5\mbox{F}_{0.5} scores decline by 0.9 on both datasets, which reveals that the contribution from adaptive syntax and PLMs does not fully overlap. It is worth noting that Rothe et al. (2021) also build another much larger GEC model based on the T5-11B (Raffel et al., 2020) and achieve 68.9/75.9 \mboxF0.5\mbox{F}_{0.5} scores. For a fair comparison, we do not list this result in Table 2 as this model is about 24×\times larger than ours.

3 Analysis and Discussion

Effectiveness of GOPar. In Table 2, we present the results of using an off-the-shelf parserWe use biaffine-dep-roberta-en model provided by SuPar. to provide syntactic knowledge (GOPar →\rightarrow Off-the-shelf Parser). After changing the parser, the impact of syntax becomes marginal under all settings. This observation implies that the performance gains contributed from syntax are highly contingent on the quality of parses. We look further into the parses and find that GOPar is more robust when facing grammatical errors and can further identify such errors, while the off-the-shelf parser is vulnerable and tends to provide incorrect parses. So we can draw a conclusion that the task adaptation of parsers is essential when applying syntax to the GEC task.

Decomposition of syntactic information. To gain more insights on how adaptive syntactic information works, we decompose it into three parts: 1) the arc information, which means only using the topological structure of the syntax tree; 2) the GED label information, which refers to the special labels “S”, “R” and “M” for marking erroneous tokens; and 3) the syntax label information, such as “subj” and “iobj” for different syntactic relations. We conduct an ablation study to explore the effect of each kind of information for GEC, as shown in Table 3. Concretely, for “w/o GED Labels”, we force the parser to skip GED labels and select the syntax label with the highest probability when predicting. For “w/o Syntax Labels”, we replace all syntax labels with “O” in the results. For “w/o All Labels”, we do not feed the label embeddings into the DepGCN module and use the dependency distance information to re-scale the self-attention weights in the Transformer encoder.

There are several observations. First, removing GED labels or syntax labels reduces the performance of SynGEC to a similar extent, which indicates that they are equally important to GEC. Second, when we only use the arc information, i.e., removing all labels, the recall drops sharply while the precision increases notably compared with the baseline. We speculate that the contribution from arc information is mainly on preventing GEC models from being misled by inappropriate context. Third, the full SynGEC approach utilizing all three kinds of information achieves the best performance, which implies that they have intrinsic complementary strengths.

Influence of self-training. Despite the effectiveness of GOPar compared with off-the-shelf parsers trained on small-scale manually-annotated treebanks of grammatical sentences, it is still not clear whether—or to what extent—the improvement comes from our GEC-oriented adaption of the parser. It is also possible that some or most improvement is due to the larger training set and domain adaptation via self-training McClosky et al. (2006). Self-training is a classical semi-supervised learning method that enhances models with large-scale pseudo-labeled in-domain data. To study this, we directly utilize the pseudo-labeled trees of target-side sentences in GEC data to train a parser (Self-training in Table 4). We observe that SynGEC significantly outperforms only using self-training without the step of tree projection, which demonstrates that the effectiveness of GOPar mainly stems from the GEC-oriented adaptation.

Error type performance. Figure 4 shows more fine-grained evaluation results on different error types on BEA-19-dev. The results support that syntactic information from GOPar is beneficial for most error types. Specifically, syntactic knowledge significantly improves the GEC model’s ability to correct context-sensitive errors, such as DET, PREP, PUCNT, VERB:SVA, and VERB:TENSE. Correcting such errors requires long-distance information, which syntax can effectively provide. The syntax also helps solve word-ordering (WO) errors, which need sentence structure information to correct. Besides, the performance on PRON, OTHER, VERB is also substantially improved. Meanwhile, we note that a small subset of types is negatively affected, like ADJ, MORPH, SPELL, and NOUN. After more careful observation, we find that their corrections mainly depend on local information, where syntactic knowledge may not help much or even introduce noises.

Experiments on Chinese GEC

Datasets and evaluation. For Chinese, we report P/R/\mboxF0.5\mbox{F}_{0.5} values on NLPCC-18-test (Zhao et al., 2018) and MuCGEC-test Zhang et al. (2022) using their official evaluation tools. MuCGEC-dev is used for hyper-parameter tuning and checkpoint selection. For training data, we use the Chinese Lang8 dataset (Zhao et al., 2018) and HSK dataset (Zhang, 2009). The statistics of above-mentioned datasets are shown in Table 5.

Char-based GOPar. Current Chinese GEC models usually treat the input sentence as a character sequence and do not perform word segmentation Zhao and Wang (2020). In contrast, dependency parsers typically treat the input sentence as a word sequence. To handle this mismatch, we follow Yan et al. (2020) and build a char-based GOPar. The basic idea is to convert a word-based tree into a char-based one by letting each character depends on its right-hand one inside multi-character words.

Use of BART. We employ the recently proposed Chinese BART (Shao et al., 2021), which is originally implemented with the HuggingFace Transformers toolkithttps://github.com/huggingface/transformers (Wolf et al., 2020). We manage to wrap their code and use it on our Fairseq implementation. Specifically, we find many common characters are missing in its vocabulary. Therefore, we add 3,866 Chinese characters and punctuation marks from Chinese Gigaword and Wikipedia corpora, leading to a substantial performance boost according to our preliminary experiments. The embeddings of these newly added tokens are randomly initialized and trained on GEC data.

Results are presented in Table 6. When not using BART, our SynGEC outperforms the Transformer baseline by 1.60/2.09 \mboxF0.5\mbox{F}_{0.5} score on NLPCC-18-test and MuCGEC-test, respectively. When using BART, our baseline already outperforms the previous SOTA system (Zhang et al., 2022), thanks to the engineering efforts mentioned in the previous paragraph. Again, SynGEC further improves the \mboxF0.5\mbox{F}_{0.5} score by 0.68/0.58. These results indicate that our proposed SynGEC approach can be effective for different languages. Given that our SynGEC approach is actually language-independent, we plan to test it in more languages in the future.

Related Works

Grammatical error correction. Recent work mainly formulates GEC as a monolingual translation task and handle it with burgeoning encoder-decoder-based MT models (Yuan and Briscoe, 2016; Junczys-Dowmunt et al., 2018), among which Transformer (Vaswani et al., 2017) has become a dominant paradigm. With the help of synthetic training data (Lichtarge et al., 2019; Yasunaga et al., 2021) and large PLMs (Kaneko et al., 2020; Katsumata and Komachi, 2020), Transformer-based GEC models have achieved SOTA performance on various benchmark datasets (Rothe et al., 2021; Stahlberg and Kumar, 2021).

Meanwhile, the sequence-to-edit (Seq2Edit) approach emerges as a competitive alternative, which predicts a sequence of edit operations to achieve correction (Gu et al., 2019; Awasthi et al., 2019; Omelianchuk et al., 2020). Although this work adopts the Transformer-based GEC models as the baseline, our SynGEC approach can also be applied to Seq2Edit models straightforwardly, which we leave to future work.

Parsing ungrammatical sentences. Despite the success of syntactic parsing on clean sentences (Dozat and Manning, 2017; Zhang et al., 2020b), parsing noisy sentences is still under-explored, including but not limited to learner texts (Foster, 2004; Hashemi and Hwa, 2016), speech disfluencies Honnibal and Johnson (2014), and historical texts Pettersson et al. (2012). This work focuses on parsing ungrammatical texts. Previous studies mainly tackle this problem by annotating small-scale trees for ungrammatical sentences and re-training a parser on them (Dickinson and Ragheb, 2009; Petrov and McDonald, 2012; Cahill, 2015; Berzak et al., 2016). Instead, we propose to train a tailored parser on automatically generated syntax trees from parallel GEC data, which avoids the laborious manual annotation. This idea has been mentioned as future work in Wagner (2012).

Syntax-enhanced GEC. Many previous work has demonstrated the effectiveness of utilizing syntactic information for various NLP tasks, such as machine translation (Bastings et al., 2017; Zhang et al., 2019), opinion role labeling (Zhang et al., 2020a), and semantic role labeling (Xia et al., 2019; Sachan et al., 2021). Meanwhile, we have found two recent works on syntax-enhanced GEC (Wan and Wan, 2021; Li et al., 2022). Both works directly produce the dependency tree of the input sentence using an off-the-shelf parser, without tailoring parsers for ungrammatical sentences. They both use graph attention networks (GAT) for tree encoding (Velickovic et al., 2018). Besides the dependency tree, Li et al. (2022) exploits the constituent tree of the input sentence as well.

Compared with the above two works, the major contribution of our work is directly dealing with the severe performance drop issue via our tailored GOPar. We adopt GCN for tree encoding because our preliminary experiments show that it achieves similar performance but is faster. Moreover, our baseline GEC models achieve much higher performance than theirs, as shown in Table 2.

Conclusions

This paper presents a SynGEC approach that incorporates adapted dependency syntax into GEC models. The key idea is adjusting vanilla parsers to accommodate ungrammatical sentences. We first extend the standard syntax representation scheme to use a unified tree structure to encode both grammatical errors and syntactic structure. Then we obtain high-quality parse trees of ungrammatical sentences by projecting target-side trees into source-side ones in parallel GEC training data, which are ultimately used for training a tailored parser named GOPar. We employ GCN to encode syntax produced by GOPar. Experiments on mainstream datasets in two languages show that SynGEC is effective and achieves SOTA results.

Limitations

First, off-the-shelf parsers may still produce noisy parse trees for the target-side correct sentences, which could further lead to noise in our projected trees for the source-side incorrect sentences. Second, we have only employed three coarse-grained labels to distinguish grammatical errors in our syntax representation scheme, while fine-grained categories may further benefit GEC (Yuan et al., 2021). Both limitations may be mitigated by integrating ideas and resources achieved by previous work on manually annotating syntactic trees for ungrammatical sentences (Dickinson and Ragheb, 2009; Berzak et al., 2016). Besides, all of our efforts focus on integrating source-side syntactic information, while there is also some work trying to incorporate target-side syntax and get positive results (Aharoni and Goldberg, 2017; Wang et al., 2018). We will further study how to appropriately utilize such target-side syntax in our future work.

Acknowledgements

We want to thank all the anonymous reviewers for their valuable comments. We also thank Yu Zhang, Houquan Zhou for their great help and insightful suggestions when polishing this paper. This work was partially supported by the National Natural Science Foundation of China (Grant No.62176173 and No.61876116) and by Alibaba Group through Alibaba Innovative Research Program. This work was also partially supported by Projected Funded by the Priority Academic Program Development of Jiangsu Higher Education Institutions.

References

Appendix A Hyper-parameters

The main hyper-parameters adopted by SynGEC are presented in Table 7. When not using PLMs, the total training time is about 3 hours. When using PLMs, the training costs about 7 hours. For fine-tuning BART on GEC data, we directly utilize the same hyper-parameters described in Katsumata and Komachi (2020). When confronting sentences longer than the max input length, we keep them unchanged during predicting.

Appendix B Experiments on JFLEG

JFLEG (Napoles et al., 2017) is an English GEC evaluation dataset which focuses on fluency and uses the GLUE score (Napoles et al., 2015) as the evaluation metric. We evaluate the baseline and the SynGEC approach in Table 2 on JFLEG. Since JFLEG’s scale is relatively small (only 747 sentences), we choose to present the results in the appendix. From Table 8, we can see that the syntactic knowledge still continuously improves the GEC performance over baselines with/without PLMs.

Appendix C The GED ability of GOPar

We evaluate the binary Grammatical Error Detection (GED) performance of GOPar on two mainstream GED dataset, i.e., BEA-19-dev (Bryant et al., 2019) and FCE-test (Yannakoudakis et al., 2011). We follow Rei and Yannakoudakis (2016) and report token-level P/R/F values for detecting incorrect labels. Table 9 shows the performance of GOPar and other leading GED models. When using the same training data, GOPar has a superior ability to detect grammatical errors. This phenomenon is very interesting and worthy of more in-depth study.

Appendix D Download links of PLMs

The download links of PLMs used in our experiments are listed below. We employ ELECTRA (Clark et al., 2020; Cui et al., 2020) to build GOPar and BART (Lewis et al., 2020; Shao et al., 2021) to enhance our GEC model.