Graphix-T5: Mixing Pre-Trained Transformers with Graph-Aware Layers for Text-to-SQL Parsing

Jinyang Li, Binyuan Hui, Reynold Cheng, Bowen Qin, Chenhao Ma, Nan Huo, Fei Huang, Wenyu Du, Luo Si, Yongbin Li

Introduction

Relational database, serving as an important resource for users to make decision in many fields, such as health care, sports, and entertainment, has emerged frequently because of the big data era. It is efficient for data users to access the information from databases via structured query language, e.g., SQL. Despite its effectiveness and efficiency, the complex nature of SQLs leads to extremely expensive learning efforts for non-technical users. Therefore, text-to-SQL (Cai et al. 2018; Zelle and Mooney 1996; Xu, Liu, and Song 2017; Yu et al. 2018a; Yaghmazadeh et al. 2017), aiming to convert natural language instructions or questions into SQL queries, has attracted remarkable attention.

In this work, we explore the challenging cross-domain setting where a text-to-SQL parser needs to achieve domain generalization, i.e. , the ability to generalize to domains that are unseen during training. Achieving this goal would, in principle, contribute to a universal natural language interface that allows users to interact with data in arbitrary domains. The major challenge towards domain generalization (Wang et al. 2020a; Cao et al. 2021; Wang et al. 2022; Cai et al. 2021; Hui et al. 2022) is that generating structure-rich SQLs requires (potentially multi-hop) reasoning, i.e. the ability to properly contextualize a user question against a given database by considering many explicit relations (e.g., table-column relations specified by database schema) and implicit relations (e.g., whether a phrase refers to a column or table). Figure 1 shows an introductory example of multi-hop reasoning in the text-to-SQL parsing and Figure 6 presents two more detailed cases.

From the modeling perspective, there are two critical dimensions along which we can differentiate current text-to-SQL parsers. The first is how to effectively imbue relational structures (both explicit and implicit) in the form of graphs into neural networks, and the second is how to take the most advantage of pre-trained models (e.g.T5 (Raffel et al. 2020)). These two dimensions are inter-connected and form a spectrum of methods. On one end of the spectrum, PICARD (Scholak, Schucher, and Bahdanau 2021) uses the original pre-trained T5 model by linearizing database schemas into sequences, hoping that T5 can successfully capture the underlying relational structures. On the other end of the spectrum, RAT-SQL (Wang et al. 2020a) only utilizes pre-trained encoders (e.g., BERT (Devlin et al. 2019)) and explicitly captures desired relations via specialized relation-aware models. However, more powerful encoder-decoder based pre-trained models are not exploited in this framework, but relational structures are accommodated at most. In this work, we explore the cross zone where the encoder-decoder based pre-trained models (specifically T5) and relation-aware encodings are deeply coupled in favor of better domain generalization. We first observe that naively adding a relational graph-based module in the middle of T5, resulting in a ‘T5-encoder →\rightarrow graph-based module →\rightarrow T5-decoder architecture’ (see also Figure 2(c), namely GNN-T5), does not work very well on standard benchmarks. Presumably, the deficiency comes from the middle graph-based modules breaking the original information flow inside T5.

In order to address this problem, we present a novel architecture called Graphix-T5 that is capable of effectively modelling relational structure information while maintaining the powerful contextual encoding capability of the pretrained T5. First, we design a Graphix layer that simultaneously encodes a mixture of semantic and structural information. Concretely, hidden states of inputs composed by questions and databases are modelled by contextualized semantic encoding, and the structural representation is injected in each transformer layer using a relational GNN block that enhances multi-hop reasoning through message passing (Fang et al. 2020; Velickovic et al. 2018) to capture explicit and implicit relations. Second, we construct a new encoder by stacking the Graphix layers and replacing the original T5 encoder. In each Graphix layer, the parameters of the semantic block are still initialized by T5, in an attempt to maintain the contextualized encoding power of the pre-training. In contrast to the severed GNN-T5 (Figure 2.(c)), the Graphix-T5 (Figure 2.(d)) will allow intensive interaction between semantic and structure from the starting layers.

We empirically show the effectiveness of Graphix-T5 on several cross-domain text-to-SQL benchmarks, i.e. , Spider, Syn, Dk and Realistic. On these datasets, the proposed model achieves new state-of-the-art performance, substantially outperforming all existing models by large margins. Specifically, Graphix-T5-large surprisingly beats the vanilla T5-3B. Furthermore, we verified that Graphix-T5 can also achieve the significant improvement in the low-resource and compositional generalization obviously thanks to the introduction of structural bias. It should be noticed that though we only focus on text-to-SQL parsing in this work, we believe that the general methodology of Graphix-T5 can be extended to structured knowledge grounding tasks, e.g., TableQA (Pasupat and Liang 2015), Data-to-text (Nan et al. 2021) and KBQA (Talmor and Berant 2018).

Task Formulation and Notations

Given a natural language question Q={q1,...,q∣Q∣}\mathcal{Q}=\left\{q_{1},...,q_{\left|\mathcal{Q}\right|}\right\} with its corresponding database schemas D=⟨C,T⟩\mathcal{D}=\left\langle\mathcal{C},\mathcal{T}\right\rangle, where C={c1,...,c∣C∣}\mathcal{C}=\left\{c_{1},...,c_{\left|\mathcal{C}\right|}\right\} and T={t1,...,t∣T∣}\mathcal{T}=\left\{t_{1},...,t_{\left|\mathcal{T}\right|}\right\} represent columns and tables, ∣C∣\left|\mathcal{C}\right| and ∣T∣\left|\mathcal{T}\right| refer to the number of columns and tables in each database respectively. The goal of text-to-SQL is to generate the corresponding SQL query yy.

2 Vanilla T5 Architecture

The most canonical and effective format of inputs to T5 performing text-to-SQL task is PeteShaw (Shaw et al. 2021), which unifies natural language questions Q\mathcal{Q} and database schema D\mathcal{D} as a joint sequence as shown:

where qiq_{i} is ithi^{th} token in the question, tjt_{j} represents jthj^{th} table in the D\mathcal{D}, and cktjc_{k}^{t_{j}} refers to the kthk^{th} column in the jthj^{th} table. ∗* is the special column token in the database. Dname\mathcal{D}_{name} is the name of each database.

Encoder-Decoder Training Mechanism

Following (Shaw et al. 2021), T5 (Raffel et al. 2020) adopt an encoder-decoder mechanism to generate SQLs. First, the bi-directional encoder learns the hidden state hh of input xx, then the decoder generates SQLs based on hh as:

where Θ\Theta and Υ\Upsilon refers to parameters of the encoder and decoder, and hh connects the encoder and decoder. The model is initialized with pretrained T5 parameters and optimized as the following objective.

where xx, yy indicates the input and output tokens respectively and ∣y∣|y| is the max length of generation SQL.

Proposed Approach: Graphix-T5

We continue to take both questions and database schemas as depicted in Eq. (1) to encode the contextual information through the original T5.

Graph Construction

The joint input questions and schemas can be displayed as a heterogeneous graph G=⟨V,R⟩\mathcal{G}=\left\langle\mathcal{V},\mathcal{R}\right\rangle consisting of three types of nodes V=Q∪C∪T\mathcal{V}=\mathcal{Q}\cup\mathcal{C}\cup\mathcal{T} and multiple types of relations \mathcal{R}=\textsl{r_{1},...,, ...,r_{\left|\mathcal{R}\right|}}, where each rir_{i} refers to a one-hop relation between nodes and a multi-hop relation rkr^{k} is defined as a composition of one-hop relations: rk=r1∘r2⋯∘rIr^{k}=r_{1}\circ r_{2}\cdots\circ r_{I} as shown in the Figure 1, where II refers to the length of each rkr^{k}. Inspired by (Wang et al. 2020a; Cao et al. 2021; Qin et al. 2022b; Hui et al. 2022), we enumerated a list of pre-defined relations to connect nodes. The relation sets can be divided into three main categories:

Schema relations: Foreign-Key, Primary-Key, and Same-Table pertain to the particular explicit schema relations that the original T5 cannot obtain from linear inputs.

Schema linking relations: Exact-Match, Partial-Match, and Value-Match are implicit linking relations between question and schema nodes. A new type of relation Bridge is introduced.

Question relations: Modifier and Argument are implicit dependency relations between tokens in a question.

No-Match Mode vs. Bridge Mode

Previous works (Cao et al. 2021; Hui et al. 2022) through adding the dummy edges called No-Match indicate that the there are question tokens and the schema tokens, which should be correlated but cannot be linked due to existing string-matched rules. However, as shown in the Figure 3, No-Match may lead to over-smoothing problem (Chen et al. 2020a) since they bring out too many noisy neighbors to compute the attention score. Suppose there exists A tokens for the question and B schema items that are semantic relevant but not linked by the rule, the number of edges need to be linked as No-Match is A×BA\times B. In contrast, we leverage the special token * as a bridge node, allowing all schema nodes to be reached from the question token nodes by decreasing the number of edges drastically from A×BA\times B to A+BA+B.

2 Graphix-Layer

The Graphix layer is designed to integrate semantic information obtained from each transformer block with structural information of a relational graph neural network (GNN) block.

In each Graphix Layer, structural representations are produced through the relational graph attention network (RGAT) (Wang et al. 2020b) over the pre-defined question-schema heterogeneous graph. Formally, given initial node embeddingVarious initialization strategies could be implemented. In this work, we initialized the node embeddings with their semantic representations. eiinit{e}_{i}^{init} for ithi^{th} node and its jthj^{th} neighbor ejinit{e}_{j}^{init} linked by specific types of relations, it can be computed through:

After computing representations from both semantic and structural space, the lthl^{th} Graphix Layer employs a mixture of semantic and structural information to enable information integration as following:

3 Graphix-T5

Here we present our entire Graphix-T5 model formally. The hidden states of the last layer of Graphix-encoder can be represented as:

where G\mathcal{G} is the question-schema heterogeneous graph, the Ψ\Psi are the additional parameters of the RGAT, which are initialized randomly. In order to preserve the pre-trained semantic knowledge, we migrate parameters Θ\Theta from original T5 encoder as the initial parameters of semantic transformer block of the Graphix layer.

4 Training

Similar to original T5, we also follow a fine-tuning strategy. The whole training framework is to optimize the following log-likelihood.

Experiment

We conduct extensive experiments on four challenging benchmarks for cross-domain text-to-SQLs and two different training settings. (1) Spider (Yu et al. 2018b) is a large-scale cross-domain text-to-SQL benchmark, also including 9 previous classic datasets, e.g., Scholar (Iyer et al. 2017), WikiSQL (Zhong, Xiong, and Socher 2017), GeoQuery (Zelle and Mooney 1996), etc. It contains 8659 training examples and 1034 development examples, which covers 200 complex databases across 138 domains. The testing set is not available for individual review. (2) Syn (Gan et al. 2021a) replaces the simple string-matched question tokens or schema names with their synonyms. (3) Dk (Gan, Chen, and Purver 2021) requires the text-to-SQL parsers to equip with the capability of domain knowledge reasoning. (4) Realistic removes and switches the obvious mentions of schema items in questions, making it closer to the real scenarios. Furthermore, we also test the compositional generalization ability of our model on the Spider-SSP (Shaw et al. 2021) with three splits from Spider: Spider-Length (split dataset based on variant lengths); Spider-TMCD (Target Maximum Compound Divergence) and Spider-Template (split based on different parsing templates). Finally, the performances of Graphix-T5 on Low-Resource setting are evaluated on usage of 10%, 20%, 50% data separately.

Following (Yu et al. 2018b), Exact Match (EM) and Execution Accuracy (EX) are the two standard metrics we use to measure performance of our model. EM can evaluate how much a generated SQL is comparable to the gold SQL. EX can reflect whether a predicted SQL is valid and returns the exact result as desired by users.

We implement our codes https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/graphix mainly based on hugging-face transformers library (Wolf et al. 2020) https://huggingface.co/. We set the max input length as 1024, generation max length as 128, and batch size as 32. We also adopt Adafactor (Shazeer and Stern 2018) as our primary optimizer with a linear decayed learning rate of 5e-5. During the experiment, Graphix layers are mainly injected into the encoder to learn better representations for structural generalization. We evaluate our effectiveness of Graphix-T5 across two main versions: T5-Large with approximately 800M parameters and T5-3B, with more than 3 Billion parameters literally. All experiments are conducted on one NVIDIA Tesla A100, which is available for the most research centers.

Our model are compared mainly to mainstream strong baseline models such as GNNSQL (Bogin, Berant, and Gardner 2019), RATSQL (Wang et al. 2020a), GAZP (Zhong et al. 2020), BRIDEGE (Chen et al. 2020b), SMBOP (Rubin and Berant 2021), NatSQL (Gan et al. 2021b), LGESQL (Cao et al. 2021), S2SQL (Hui et al. 2022) and T5+PICARD (Scholak, Schucher, and Bahdanau 2021) across the disparate datasets and settings.

2 Overall Performance

Table 1 displays the performance of Graphix-T5 and other competitive baseline models on official Spider benchmark. First, we demonstrate that Graphix-T5-3B with a constrained decoding module PICARD (Scholak, Schucher, and Bahdanau 2021) achieves the state-of-the-art on this challenging cross-domain text-to-SQL benchmark. Also, it is evident that Graphix-T5 is vastly superior to the vanilla T5 on large and 3B scale with a significant margin. This indicates that the structural generalization capability of the Graphix layer is crucial for T5 such a text-to-text PLM to perform the text-to-SQL task.

As shown in the Table 2, we further demonstrate the robustness of Graphix-T5 when it confronts with more challenging and closer to realistic evaluations in Syn, Dk, Realistic without any additional training. First of all, the results show that Graphix-T5-3B outperforms other baseline models across all three datasets. Furthermore, we observe that Graphix-T5-large and Graphix-T5-3B surpass the performance of vanilla T5-large and T5-3B with a clear margin, respectively. This demonstrates that vanilla T5 is hungry for structural reasoning when dealing with more flexible and complicated questions for text-to-SQLs from real-world scenarios. And Graphix can mitigate this problem.

As shown in Table 3, on Spider-Ssp, the grammar-based inductive T5 model provided by (Shaw et al. 2021), named NQG-T5, has no obvious advantages over vanilla T5, which indicates that the grammar of natural language is not helpful to enhance T5 for compositional generation. However, Graphix-T5 helps the T5 gain the SQL knowledge and makes it less vulnerable to these modifications through the effective fusion of structural information.

Figure 4 records the performance of Graphix-T5-large and T5-large on different low-resource settings. It displays 1) in each low-resource setting, Graphix-T5-large performs considerably better than vanilla T5-large. It demonstrates that the structural knowledge created by humans can compensate for the inadequate learning due to low-resource data (Ye et al. 2022); 2) notably, Graphix-T5-large can perform obviously better than the vanilla T5-large trained on 100% data even within just usage of 50% data. This further verifies the strengths of Graphix-T5 for training in the low-data resources.

As presented in Table 4, we also compare the more precise performance results of Graphix-T5 to the vanilla T5 in four separate SQL difficulty levels splitted by Spider officially, in order to better comprehend the performance improvements. We observe that Graphix-T5 is more capable of handling harder text-to-SQL cases, as illustrated in the Hard and Extra-hard examples, indicating that structural bias training is beneficial to the text-to-text PLMs to reason over complex scenarios.

3 Ablation Study

As shown in Table 5, to better validate the function of each component of Graphix-T5, ablation studies are performed in large version and expected to answer the following questions.

Graphix-T5-large with Bridge Mode can achieve the better performance than with No-Match Mode. It indicates that No-match mode will greatly increase the number of noisy neighbors, resulting in higher risk of over-smoothing issues (Chen et al. 2020a).

With Double-Graph means that Graphix-T5 incorporate Graphix layer into the both encoder and decoder. The result reveals that adding Graphix layers to the decoder does not lead to any improvements. Since decoder is an auto-regressive model, which only considers the history tokens when generating the current token. However, Graphix-T5, which can forecast the information of future tokens by global linking, may disrupt this characteristic leading to the negative impact on the decoder. Therefore, we propose that the best tactic is to only incorporate Graphix layers into the encoder.

Echoing Figure 2, we access the performance of 4 categories of models using PLMs on Spider. According to Table 5 (c), the performance of GNN-T5 has decreased by roughly 20% when compared to Graphix-T5, proving GNN-T5 training strategy to be ineffective. Moreover, we notice that such severed GNN-T5 encounters a catastrophic forgetting problem (French 1999) during training. Since the accuracy of the GNN-T5 continues to be 0 in the first thousands of steps, as shown in Figure 5, it is evident that all pretrained knowledge from T5 would be forgotten. After convergence, the GNN-T5 performance decreases significantly from the Graphix-T5, indicating that only a small portion of the semantic information from T5 has been utilized. In contrast, Graphix-T5 can achieve almost 50% accuracy inside the first 1000 training steps and more than 20% improvement than GNN-T5 after convergence, which verifies the advantages of Graphix-T5 that can avoid catastrophic forgetting and augment generalization capability.

4 Case Study

To illustrate the effectiveness of Graphix qualitatively, two examples are displayed in Figure 6, which are sampled randomly from Syn. Figure 6 shows the comparison of predicted SQLs by vanilla T5-3B and Graphix-T5-3B. We can observe that Graphix can generate correct SQLs even in the hard scenarios. That is because that, even with a small number of keywords overlapped, Graphix-T5 can accurately identify counterpart column or table objects and generate a high-quality SQL through multi-hop reasoning and structural grounding. For example, in the first case, vanilla T5-3B picks the incorrect columns paper_id, paper_name, and paper_description, which even don’t appear in the table documents. This implies that vanilla T5-3B is unable to reach the target schema elements without the capability of structural grounding when confronting challenging text-to-SQLs. Instead, Graphix-T5-3B can correspond the question entities to the correct column names through multi-hop paths presented in the Figure 6. In the second case, vanilla T5-3B misidentifies the country as their target column, however, "France" only appears in the column countryname of the table countries. This suggests T5-3B is only able to generate semantically valid SQLs, which fails to take into account the real database structure. On contrary, Graphix-T5 can produce truly valid SQLs in terms of both questions and databases via a successful mixture of semantic and structural information during training.

Related Works

The basic principle of a cross-domain text-to-SQL parser is to build an encoder to learn the representations of the questions and schemas, while employing a decoder to generate SQLs with the information learnt in the encoder (Qin et al. 2022a). In particular, IRNET (Guo et al. 2019) proposes to design an encoder to learn the representations of questions and schemas respectively via an attention-based Bi-LSTM and a decoder to predict SQLs via encoded intermediate representations. Later, the graph-based encoders have been successfully proved its effectiveness in text-to-SQL tasks, for example, some works (Bogin, Berant, and Gardner 2019; Chen et al. 2021) construct the schema graph and enhance the representations of inputs. RATSQL (Wang et al. 2020a), SDSQL (Hui et al. 2021b), LGESQL (Cao et al. 2021), S2SQL (Hui et al. 2022) further improve structural reasoning through modelling relations between schema and questions. R2SQL (Hui et al. 2021a), Score (Yu et al. 2021) and Star (Cai et al. 2022) enhance structural reasoning for context-dependent text-to-SQL parsing. These works are performed by the PLM independently building the semantic features, followed by the graph-based module injecting the structural information. However, such training strategy is just effective to encoder-based PLMs (i.e. , BERT (Devlin et al. 2019), ELECTRA (Clark et al. 2020), et al.).

Recently, the text-to-text PLM T5 has been proven effectiveness in text-to-SQL (Shaw et al. 2021; Qin et al. 2022c). Besides, (Scholak, Schucher, and Bahdanau 2021) designs a constrained decoding process, namely PICARD, to detect and refuse erroneous tokens during the beam-search phase. Xie et al. (2022) further injects the knowledge from other structural knowledge grounding tasks into T5 with multi-task to boost performance on text-to-SQL. Despite effectiveness, these methods still struggle to generate SQLs in the more challenging and complex scenarios without explicit and implicit structural information. However, Graphix-T5 can overcome this issue by an argument of graph representation learning in the encoder. Concurrently, RASAT (Qi et al. 2022) also attempts to provide T5 with the structural information by adding edge embedding into the multi-head self-attention, while we keep the pre-trained transformers complete in order to benefit the most from prior semantic knowledge, which leads to better performance.

Conclusion

In this paper, we proposed an effective architecture to boost the capability of structural encoding of T5 cohesively while keeping the pretrained T5’s potent contextual encoding ability. In order to achieve this goal, we designed a Graph-Aware semi-pretrained text-to-text PLM, namely Graphix-T5, to augment the multi-hop reasoning for the challenging text-to-SQL task. The results under the extensive experiments demonstrate the effectiveness of Graphix-T5, proving that structural information is crucial for the current text-to-text PLMs for complicated text-to-SQL cases.

Acknowledgement

We thank Dr. Tao Yu and Tianbao Xie for evaluation of our work on SPIDER leaderboard. We thank Dr. Bailin Wang and Dr. Bowen Li for constructive suggestions. Reynold Cheng, Jinyang Li, Nan Huo, and Wenyu Du were supported by the University of Hong Kong (Project 104006830), the Guangdong–Hong Kong-Macau Joint Laboratory Program 2020 (Project No: 2020B1212030009), and the Innovation Wing Two Research fund. Jinyang Li was also supported by HKU Presidential PhD Scholar Programme and Alibaba Group through Alibaba Research Intern Program. Chenhao Ma was supported in part by Shenzhen Science and Technology Program under grant No.ZDSYS20211021111415025.

References

Appendix A Fine-grained Syntax Relations

The previous work (Cao et al. 2021; Wang et al. 2020a), which employed distances as the only relationship between tokens when constructing a graph, was unable to account for the deterministic relationships between tokens. For example, Given two sentences with the same meanings: "List names of students who are not from France."; "What are the names of students whose nationality is not France?". The relation between not and France should be the same in these two sentences. However, it is represented as two different relations: Distance-2 and Distance-1 respectively according to their defined relations, which will lead PLMs to learn the wrong relation representations. Even though Hui et al. (2022) proposed Forward and Backward as the additional abstract correlations of question tokens, it is still hard to discern the more important relations that is useful to text-to-SQL. In this work, we observe that that nouns and other tokens that potentially indicate characteristics of the nouns can help the model to find the corresponding database items. In order to achieve this goal, we cluster dependency parsing relations manually into two new categories of syntax relations: Modifier and Argment. As shown in Table 6, the Modifier denotes that some properties of the source token node are being modified by the target token node. For example, in the phrase Female Students, the Female is a modifier of the Students; the Production is the modifier of the token Time in the phrase Production Time. All other dependency parsing relations will be marked as Argment.

Appendix B Leadboard Result

After being equipped with PICARD, Graphix-T5 achieves the No.1 on Spider testing leaderboard with the clear margin, as shown in the table 7 and table 8.