TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data
Pengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian Riedel
Introduction
Recent years have witnessed a rapid advance in the ability to understand and answer questions about free-form natural language (NL) text Rajpurkar et al. (2016), largely due to large-scale, pretrained language models (LMs) like BERT Devlin et al. (2019). These models allow us to capture the syntax and semantics of text via representations learned in an unsupervised manner, before fine-tuning the model to downstream tasks Melamud et al. (2016); McCann et al. (2017); Peters et al. (2018); Liu et al. (2019b); Yang et al. (2019); Goldberg (2019). It is also relatively easy to apply such pretrained LMs to comprehension tasks that are modeled as text span selection problems, where the boundary of an answer span can be predicted using a simple classifier on top of the LM Joshi et al. (2019).
However, it is less clear how one could pretrain and fine-tune such models for other QA tasks that involve joint reasoning over both free-form NL text and structured data. One example task is semantic parsing for access to databases (DBs) (Zelle and Mooney, 1996; Berant et al., 2013; Yih et al., 2015), the task of transducing an NL utterance (e.g., “Which country has the largest GDP?”) into a structured query over DB tables (e.g., SQL querying a database of economics). A key challenge in this scenario is understanding the structured schema of DB tables (e.g., the name, data type, and stored values of columns), and more importantly, the alignment between the input text and the schema (e.g., the token “GDP” refers to the Gross Domestic Product column), which is essential for inferring the correct DB query (Berant and Liang, 2014).
Neural semantic parsers tailored to this task therefore attempt to learn joint representations of NL utterances and the (semi-)structured schema of DB tables (e.g., representations of its columns or cell values, as in Krishnamurthy et al. (2017); Bogin et al. (2019b); Wang et al. (2019a), inter alia). However, this unique setting poses several challenges in applying pretrained LMs. First, information stored in DB tables exhibit strong underlying structure, while existing LMs (e.g., BERT) are solely trained for encoding free-form text. Second, a DB table could potentially have a large number of rows, and naively encoding all of them using a resource-heavy LM is computationally intractable. Finally, unlike most text-based QA tasks (e.g., SQuAD, Rajpurkar et al. (2016)) which could be formulated as a generic answer span selection problem and solved by a pretrained model with additional classification layers, semantic parsing is highly domain-specific, and the architecture of a neural parser is strongly coupled with the structure of its underlying DB (e.g., systems for SQL-based and other types of DBs use different encoder models). In fact, existing systems have attempted to leverage BERT, but each with their own domain-specific, in-house strategies to encode the structured information in the DB (Guo et al., 2019; Zhang et al., 2019a; Hwang et al., 2019), and importantly, without pretraining representations on structured data. These challenges call for development of general-purpose pretraining approaches tailored to learning representations for both NL utterances and structured DB tables.
In this paper we present TaBert, a pretraining approach for joint understanding of NL text and (semi-)structured tabular data (§ 3). TaBert is built on top of BERT, and jointly learns contextual representations for utterances and the structured schema of DB tables (e.g., a vector for each utterance token and table column). Specifically, TaBert linearizes the structure of tables to be compatible with a Transformer-based BERT model. To cope with large tables, we propose content snapshots, a method to encode a subset of table content most relevant to the input utterance. This strategy is further combined with a vertical attention mechanism to share information among cell representations in different rows (§ 3.1). To capture the association between tabular data and related NL text, TaBert is pretrained on a parallel corpus of 26 million tables and English paragraphs (§ 3.2).
TaBert can be plugged into a neural semantic parser as a general-purpose encoder to compute representations for utterances and tables. Our key insight is that although semantic parsers are highly domain-specific, most systems rely on representations of input utterances and the table schemas to facilitate subsequent generation of DB queries, and these representations can be provided by TaBert, regardless of the domain of the parsing task.
We apply TaBert to two different semantic parsing paradigms: (1) a classical supervised learning setting on the Spider text-to-SQL dataset (Yu et al., 2018c), where TaBert is fine-tuned together with a task-specific parser using parallel NL utterances and labeled DB queries (§ 4.1); and (2) a challenging weakly-supervised learning benchmark WikiTableQuestions (Pasupat and Liang, 2015), where a system has to infer latent DB queries from its execution results (§ 4.2). We demonstrate TaBert is effective in both scenarios, showing that it is a drop-in replacement of a parser’s original encoder for computing contextual representations of NL utterances and DB tables. Specifically, systems augmented with TaBert outperforms their counterparts using Bert, registering state-of-the-art performance on WikiTableQuestions, while performing competitively on Spider (§ 5).
Background
Semantic parsing tackles the task of translating an NL utterance into a formal meaning representation (MR) . Specifically, we focus on parsing utterances to access database tables, where is a structured query (e.g., an SQL query) executable on a set of relational DB tables . A relational table is a listing of rows of data, with each row consisting of cells , one for each column . Each cell contains a list of tokens.
Depending on the underlying data representation schema used by the DB, a table could either be fully structured with strongly-typed and normalized contents (e.g., a table column named distance has a unit of kilometers, with all of its cell values, like 200, bearing the same unit), as is commonly the case for SQL-based DBs (§ 4.1). Alternatively, it could be semi-structured with unnormalized, textual cell values (e.g., 200 km, § 4.2). The query language could also take a variety of forms, from general-purpose DB access languages like SQL to domain-specific ones tailored to a particular task.
Given an utterance and its associated tables, a neural semantic parser generates a DB query from the vector representations of the utterance tokens and the structured schema of tables. In this paper we refer schema as the set of columns in a table, and its representation as the list of vectors that represent its columnsColumn representations for more complex schemas, e.g., those capturing inter-table dependency via primary and foreign keys, could be derived from these table-wise representations.. We will introduce how TaBert computes these representations in § 3.1.
Masked Language Models
Given a sequence of NL tokens , a masked language model (e.g., BERT) is an LM trained using the masked language modeling objective, which aims to recover the original tokens in from a “corrupted” context created by randomly masking out certain tokens in . Specifically, let be the subset of tokens in selected to be masked out, and denote the masked sequence with tokens in replaced by a [MASK] symbol. A masked LM defines a distribution over the target tokens given the masked context .
BERT parameterizes using a Transformer model. During the pretraining phase, BERT maximizes on large-scale textual corpora. In the fine-tuning phase, the pretrained model is used as an encoder to compute representations of input NL tokens, and its parameters are jointly tuned with other task-specific neural components.
TaBert: Learning Joint Representa- tions over Textual and Tabular Data
We first present how TaBert computes representations for NL utterances and table schemas (§ 3.1), and then describe the pretraining procedure (§ 3.2).
Fig. 1 presents a schematic overview of TaBert. Given an utterance and a table , TaBert first creates a content snapshot of . This snapshot consists of sampled rows that summarize the information in most relevant to the input utterance. The model then linearizes each row in the snapshot, concatenates each linearized row with the utterance, and uses the concatenated string as input to a Transformer (e.g., BERT) model, which outputs row-wise encoding vectors of utterance tokens and cells. The encodings for all the rows in the snapshot are fed into a series of vertical self-attention layers, where a cell representation (or an utterance token representation) is computed by attending to vertically-aligned vectors of the same column (or the same NL token). Finally, representations for each utterance token and column are generated from a pooling layer.
One major feature of TaBert is its use of the table contents, as opposed to just using the column names, in encoding the table schema. This is motivated by the fact that contents provide more detail about the semantics of a column than just the column’s name, which might be ambiguous. For instance, the Venue column in Fig. 1 which is used to answer the example question actually refers to host cities, and encoding the sampled cell values while creating its representation may help match the term “city” in the input utterance to this column.
However, a DB table could potentially have a large number of rows, with only few of them actually relevant to answering the input utterance. Encoding all of the contents using a resource-heavy Transformer is both computationally intractable and likely not necessary. Thus, we instead use a content snapshot consisting of only a few rows that are most relevant to the input utterance, providing an efficient approach to calculate content-sensitive column representations from cell values.
We use a simple strategy to create content snapshots of rows based on the relevance between the utterance and a row. For , we select the top- rows in the input table that have the highest -gram overlap ratio with the utterance.We use in our experiments. Empirically this simple matching heuristic is able to correctly identify the best-matched rows for 40 out of 50 sampled examples on WikiTableQuestions. For , to include in the snapshot as much information relevant to the utterance as possible, we create a synthetic row by selecting the cell values from each column that have the highest -gram overlap with the utterance. Using synthetic rows in this restricted setting is motivated by the fact that cell values most relevant to answer the utterance could come from multiple rows. As an example, consider the utterance “How many more participants were there in 2008 than in the London Olympics?”, and an associating table with columns Year, Host City and Number of Participants, the most relevant cells to the utterance, 2008 (from Year) and London (from Host City), are from different rows, which could be included in a single synthetic row. In the initial experiments we found synthetic rows also help stabilize learning.
Row Linearization
TaBert creates a linearized sequence for each row in the content snapshot as input to the Transformer model. Fig. 1(B) depicts the linearization for , which consists of a concatenation of the utterance, columns, and their cell values. Specifically, each cell is represented by the name and data typeWe use two data types, text, and real for numbers, predicted by majority voting over the NER labels of cell tokens. of the column, together with its actual value, separated by a vertical bar. As an example, the cell valued 2005 in in Fig. 1 is encoded as
The linearization of a row is then formed by concatenating the above string encodings of all the cells, separated by the [SEP] symbol. We then prefix the row linearization with utterance tokens as input sequence to the Transformer.
Existing works have applied different linearization strategies to encode tables with Transformers Hwang et al. (2019); Chen et al. (2019), while our row approach is specifically designed for encoding content snapshots. We present in § 5 results with different linearization choices.
Vertical Self-Attention Mechanism
The base Transformer model in TaBert outputs vector encodings of utterance and cell tokens for each row. These row-level vectors are computed separately and therefore independent of each other. To allow for information flow across cell representations of different rows, we propose vertical self-attention, a self-attention mechanism that operates over vertically aligned vectors from different rows.
As in Fig. 1(C), TaBert has stacked vertical-level self-attention layers. To generate aligned inputs for vertical attention, we first compute a fixed-length initial vector for each cell at position , which is given by mean-pooling over the sequence of the Transformer’s output vectors that correspond to its variable-length linearization as in Eq. (1). Next, the sequence of word vectors for the NL utterance (from the base Transformer model) are concatenated with the cell vectors as initial inputs to the vertical attention layer.
Each vertical attention layer has the same parameterization as the Transformer layer in Vaswani et al. (2017), but operates on vertically aligned elements, i.e., utterance and cell vectors that correspond to the same question token and column, respectively. This vertical self-attention mechanism enables the model to aggregate information from different rows in the content snapshot, allowing TaBert to capture cross-row dependencies on cell values.
Utterance and Column Representations
A representation is computed for each column by mean-pooling over its vertically aligned cell vectors, , from the last vertical layer. A representation for each utterance token, , is computed similarly over the vertically aligned token vectors. These representations will be used by downstream neural semantic parsers. TaBert also outputs an optional fixed-length table representation using the representation of the prefixed [CLS] symbol, which is useful for parsers that operate on multiple DB tables.
2 Pretraining Procedure
Since there is no large-scale, high-quality parallel corpus of NL text and structured tables, we instead use semi-structured tables that commonly exist on the Web as a surrogate data source. As a first step in this line, we focus on collecting parallel data in English, while extending to multilingual scenarios would be an interesting avenue for future work. Specifically, we collect tables and their surrounding NL text from English Wikipedia and the WDC WebTable Corpus Lehmberg et al. (2016), a large-scale table collection from CommonCrawl. The raw data is extremely noisy, and we apply aggressive cleaning heuristics to filter out invalid examples (e.g., examples with HTML snippets or in foreign languages, and non-relational tables without headers). See Appendix § A.1 for details of data pre-processing. The pre-processed corpus contains 26.6 million parallel examples of tables and NL sentences. We perform sub-tokenization using the Wordpiece tokenizer shipped with BERT.
Unsupervised Learning Objectives
We apply different objectives for learning representations of the NL context and structured tables. For NL contexts, we use the standard Masked Language Modeling (MLM) objective Devlin et al. (2019), with a masking rate of 15% sub-tokens in an NL context.
For learning column representations, we design two objectives motivated by the intuition that a column representation should contain both the general information of the column (e.g., its name and data type), and representative cell values relevant to the NL context. First, a Masked Column Prediction (MCP) objective encourages the model to recover the names and data types of masked columns. Specifically, we randomly select 20% of the columns in an input table, masking their names and data types in each row linearization (e.g., if the column Year in Fig. 1 is selected, the tokens Year and real in Eq. (1) will be masked). Given the column representation , TaBert is trained to predict the bag of masked (name and type) tokens from using a multi-label classification objective. Intuitively, MCP encourages the model to recover column information from its contexts.
Next, we use an auxiliary Cell Value Recovery (CVR) objective to ensure information of representative cell values in content snapshots is retained after additional layers of vertical self-attention. Specifically, for each masked column in the above MCP objective, CVR predicts the original tokens of each cell (of ) in the content snapshot conditioned on its cell vector .The cell value tokens are not masked in the input sequence, since predicting masked cell values is challenging even with the presence of its surrounding context. For instance, for the example cell in Eq. (1), we predict its value 2005 from . Since a cell could have multiple value tokens, we apply the span-based prediction objective Joshi et al. (2019). Specifically, to predict a cell token , its positional embedding and the cell representations are fed into a two-layer network with GeLU activations Hendrycks and Gimpel (2016). The output of is then used to predict the original value token from a softmax layer.
Example Application: Semantic Parsing over Tables
We apply TaBert for representation learning on two semantic parsing paradigms, a classical supervised text-to-SQL task over structured DBs (§ 4.1), and a weakly supervised parsing problem on semi-structured Web tables (§ 4.2).
Supervised learning is the typical scenario of learning a parser using parallel data of utterances and queries. We use Spider Yu et al. (2018c), a text-to-SQL dataset with 10,181 examples across 200 DBs. Each example consists of an utterance (e.g., “What is the total number of languages used in Aruba?”), a DB with one or more tables, and an annotated SQL query, which typically involves joining multiple tables to get the answer (e.g., SELECT COUNT(*) FROM Country JOIN Lang ON Country.Code = Lang.CountryCode WHERE Name = ‘Aruba’).
Base Semantic Parser
We aim to show TaBert could help improve upon an already strong parser. Unfortunately, at the time of writing, none of the top systems on Spider were publicly available. To establish a reasonable testbed, we developed our in-house system based on TranX Yin and Neubig (2018), an open-source general-purpose semantic parser. TranX translates an NL utterance into an intermediate meaning representation guided by a user-defined grammar. The generated intermediate MR could then be deterministically converted to domain-specific query languages (e.g., SQL).
We use TaBert as encoder of utterances and table schemas. Specifically, for a given utterance and a DB with a set of tables , we first pair with each table in as inputs to TaBert, which generates sets of table-specific representations of utterances and columns. At each time step, an LSTM decoder performs hierarchical attention Libovický and Helcl (2017) over the list of table-specific representations, constructing an MR based on the predefined grammar. Following the IRNet model Guo et al. (2019) which achieved the best performance on Spider, we use SemQL, a simplified version of the SQL, as the underlying grammar. We refer interested readers to Appendix § B.1 for details of our system.
2 Weakly Supervised Semantic Parsing
Weakly supervised semantic parsing considers the reinforcement learning task of inferring the correct query from its execution results (i.e., whether the answer is correct). Compared to supervised learning, weakly supervised parsing is significantly more challenging, as the parser does not have access to the labeled query, and has to explore the exponentially large search space of possible queries guided by the noisy binary reward signal of execution results.
WikiTableQuestions Pasupat and Liang (2015) is a popular dataset for weakly supervised semantic parsing, which has 22,033 utterances and 2,108 semi-structured Web tables from Wikipedia.While some of the 421 testing Wikipedia tables might be included in our pretraining corpora, they only account for a very tiny fraction. In our pilot study, we also found pretraining only on Wikipedia tables resulted in worse performance. Compared to Spider, examples in this dataset do not involve joining multiple tables, but typically require compositional, multi-hop reasoning over a series of entries in the given table (e.g., to answer the example in Fig. 1 the parser needs to reason over the row set , locating the Venue field with the largest value of Year).
Base Semantic Parser
MAPO Liang et al. (2018) is a strong system for weakly supervised semantic parsing. It improves the sample efficiency of the REINFORCE algorithm by biasing the exploration of queries towards the high-rewarding ones already discovered by the model. MAPO uses a domain-specific query language tailored to answering compositional questions on single tables, and its utterances and column representations are derived from an LSTM encoder, which we replaced with our TaBert model. See Appendix § B.2 for details of MAPO and our adaptation.
Experiments
In this section we evaluate TaBert on downstream tasks of semantic parsing to DB tables.
We train two variants of the model, and , with the underlying Transformer model initialized with the uncased versions of and , respectively.We also attempted to train TaBert on our collected corpus from scratch without initialization from BERT, but with inferior results, potentially due to the average lower quality of web-scraped tables compared to purely textual corpora. We leave improving the quality of training data as future work. During pretraining, for each table and its associated NL context in the corpus, we create a series of training instances of paired NL sentences (as synthetically generated utterances) and tables (as content snapshots) by (1) sliding a (non-overlapping) context window of sentences with a maximum length of tokens, and (2) using the NL tokens in the window as the utterance, and pairing it with randomly sampled rows from the table as content snapshots. TaBert is implemented in PyTorch using distributed training. Refer to Appendix § A.2 for details of pretraining.
Comparing Models
Evaluation Metrics
As standard, we report execution accuracy on WikiTableQuestions and exact-match accuracy of DB queries on Spider.
1 Main Results
Comparing semantic parsers augmented with TaBert and Bert, we found TaBert is more effective across the board. We hypothesize that the performance improvements would be attributed by two factors. First, pre-training on large parallel textual and tabular corpora helps TaBert learn to encode structure-rich tabular inputs in their linearized form (Eq. (1)), whose format is different from the ordinary natural language data that Bert is trained on. Second, pre-training on parallel data could also helps the model produce representations that better capture the alignment between an utterance and the relevant information presented in the structured schema, which is important for semantic parsing.
Overall, the results on the two benchmarks demonstrate that pretraining on aligned textual and tabular data is necessary for joint understanding of NL utterances and tables, and TaBert works well with both structured (Spider) and semi-structured (WikiTableQuestions) DBs, and agnostic of the task-specific structures of semantic parsers.
Effect of Row Linearization
TaBert uses row linearization to represent a table row as sequential input to Transformer. § 5.1 (upper half) presents results using various linearization methods. We find adding type information and content snapshots improves performance, as they provide more hints about the meaning of a column.
We also compare with existing linearization methods in literature using a model, with results shown in § 5.1 (lower half). Hwang et al. (2019) uses BERT to encode concatenated column names to learn column representations. In line with our previous discussion on the effectiveness content snapshots, this simple strategy without encoding cell contents underperforms (although with pretrained on our tabular corpus the results become slightly better). Additionally, we remark that linearizing table contents has also be applied to other BERT-based tabular reasoning tasks. For instance, Chen et al. (2019) propose a “natural” linearization approach for checking if an NL statement entails the factual information listed in a table using a binary classifier with representations from Bert, where a table is linearized by concatenating the semicolon-separated cell linearization for all rows. Each cell is represented by a phrase “column name is cell value”. For completeness, we also tested this cell linearization approach, and find achieved improved results. We leave pretraining TaBert with this linearization strategy as promising future work.
TaBert uses two objectives (§ 3.2), a masked column prediction (MCP) and a cell value recovery (CVR) objective, to learn column representations that could capture both the general information of the column (via MCP) and its representative cell values related to the utterance (via CVR). Tab. 5 shows ablation results of pretraining TaBert with different objectives. We find TaBert trained with both MCP and the auxiliary CVR objectives gets a slight advantage, suggesting CVR could potentially lead to more representative column representations with additional cell information.
Tables are important media of world knowledge. Semantic parsers have been adapted to operate over structured DB tables Wang et al. (2015); Xu et al. (2017); Dong and Lapata (2018); Yu et al. (2018b); Shi et al. (2018); Wang et al. (2018), and open-domain, semi-structured Web tables Pasupat and Liang (2015); Sun et al. (2016); Neelakantan et al. (2016). To improve representations of utterances and tables for neural semantic parsing, existing systems have applied pretrained word embeddings (e.g.., GloVe, as in Zhong et al. (2017); Yu et al. (2018a); Sun et al. (2018); Liang et al. (2018)), and BERT-family models for learning joint contextual representations of utterances and tables, but with domain-specific approaches to encode the structured information in tables Hwang et al. (2019); He et al. (2019); Guo et al. (2019); Zhang et al. (2019a). TaBert advances this line of research by presenting a general-purpose, pretrained encoder over parallel corpora of Web tables and NL context. Another relevant direction is to augment representations of columns from an individual table with global information of its linked tables defined by the DB schema Bogin et al. (2019a); Wang et al. (2019a). TaBert could also potentially improve performance of these systems with improved table-level representations.
Knowledge-enhanced Pretraining
Recent pre-training models have incorporated structured information from knowledge bases (KBs) or other structured semantic annotations into training contextual word representations, either by fusing vector representations of entities and relations on KBs into word representations of LMs Peters et al. (2019); Zhang et al. (2019b, c), or by encouraging the LM to recover KB entities and relations from text Sun et al. (2019); Liu et al. (2019a). TaBert is broadly relevant to this line in that it also exposes an LM with structured data (i.e., tables), while aiming to learn joint representations for both textual and structured tabular data.
We present TaBert, a pretrained encoder for joint understanding of textual and tabular data. We show that semantic parsers using TaBert as a general-purpose feature representation layer achieved strong results on two benchmarks. This work also opens up several avenues for future work. First, we plan to evaluate TaBert on other related tasks involving joint reasoning over textual and tabular data (e.g., table retrieval and table-to-text generation). Second, following the discussions in § 5, we will explore other table linearization strategies with Transformers, improving the quality of pretraining corpora, as well as novel unsupervised objectives. Finally, to extend TaBert to cross-lingual settings with utterances in foreign languages and structured schemas defined in English, we plan to apply more advanced semantic similarity metrics for creating content snapshots.
Appendix A Pretraining Details
We collect parallel examples of tables and their surrounding NL sentences from two sources:
We extract all the tables on English WikipediaWe do not use infoboxes (tables on the top-right of a Wiki page that describe properties of the main topic), as they are not relational tables.. For each table, we use the preceding three paragraphs as the NL context, as we observe that most Wiki tables are located after where they are described in the body text.
WDC WebTable Corpus
Lehmberg et al. (2016) is a large collection of Web tables extracted from the Common Crawl Web scrapehttp://webdatacommons.org/webtables. We use its 2015 English-language relational subset, which consists of million relational tables and their surrounding NL contexts.
Preprocessing
Our dataset is collected from arbitrary Web tables, which are extremely noisy. We develop a set of heuristics to clean the data by: (1) removing columns whose names have more than 10 tokens; (2) filtering cells with more than two non-ASCII characters or 20 tokens; (3) removing empty or repetitive rows and columns; (4) filtering tables with less than three rows and four columns, and (5) running spaCy to identify the data type of columns (text or real value) by majority voting over the NER labels of column tokens, (6) rotating vertically oriented tables. We sub-tokenize the corpus using the Wordpiece tokenizer in Devlin et al. (2019). The pre-processing results in 1.3 million tables from Wikipedia and 25.3 million tables from the WDC corpus.
A.2 Pretraining Setup
Appendix B Semantic Parsers
We develop our text-to-SQL parser based on TranX Yin and Neubig (2018), which translates an NL utterance into a tree-structured abstract meaning representation following user-specified grammar, before deterministically convert the generated abstract MR into an SQL query. TranX models the construction process of an abstract MR (tree-structured representation of an SQL query) using a transition-based system, which decomposes its generation story into a sequence of actions following the user defined grammar.
Formally, given an input NL utterance and a database with a set of tables , the probability of generating of an SQL query (i.e., its semantically equivalent MR) is decomposed as the production of action probabilities:
where is the action applied to the hypothesis at time stamp . denote the previous action history. We refer readers to Yin and Neubig (2018) for details of the transition system and how individual action probabilities are computed. In our adaptation of TranX to text-to-SQL parsing on Spider, we follow Guo et al. (2019) and use SemQL as the underlying grammar, which is a simplification of the SQL language. Fig. 2 lists the SemSQL grammar specified using the abstract syntax description language Wang et al. (1997). Intuitively, the generation starts from a tree-structured derivation with the root production rule select_stmtSelectStatement, which lays out overall the structure of an SQL query. At each time step, the decoder algorithm locates the current opening node on the derivation tree, following a depth-first, left-to-right order. If the opening node is not a leaf node, the decoder invokes an action which expands the opening node using a production rule with appropriate type. If the current opening node is a leaf node (e.g., a node denoting string literal), the decoder fills in the leaf node using actions that emit terminal values.
To use such a transition system to generate SQL queries, we extend its action space with two new types of actions, SelectTable for node of type table_ref in Fig. 2, which selects a table (e.g., for predicting target tables for a FROM clause), and SelectColumn for node of type column_ref, which selects the column from table (e.g., for predicting a result column used in the SELECT clause).
As described in § 4.1, TaBert produces a list of entries, with one entry for each table :
The updated decoder state is then used to compute the probability of carrying out the action defined at time step , . For a SelectTable action, its probability of is defined similarly as Eq. (4). For a SelectColumn action, it is factorized as the probability of selecting the table (given by Eq. (4)), times the probability of selecting the column . The latter is defined as
Configuration
We use the default configuration of TranX. For TaBert parameters, we use an Adam optimizer with a learning rate of and linearly decayed learning rate schedule, and another Adam optimizer with a constant learning rate of for all remaining parameters. During training, we update model parameters for 25000 iterations, and freeze the TaBert parameters at the first 1000 update steps. We use a batch size of 30 and beam size of 3. We use gradient accumulation for large models to fit a batch into GPU memory.
B.2 Weakly-supervised Parsing on WikiTableQuestions
We use MAPO Liang et al. (2018), a strong weakly-supervised semantic parser. The original MAPO models comes with an LSTM encoder, which generates utterance and column representations used by the decoder to predict table queries. We directly substitute the encoder with TaBert, and project the utterance and table representations from TaBert to the original embedding space using a linear transformation. MAPO uses a domain-specific query language tailored to answer compositional questions on a single table. For instance, the example question in Fig. 1 could be answered using the following query
MAPO is written in Tensorflow. In our experiments we use an optimized re-implementation in PyTorch, which yields 4 training speedup.
Configuration
We use the same optimizer and learning rate schedule as in § B.1. We use a batch size of 10, and train the model for 20000 steps, with the TaBert parameters frozen at the first 5000 steps. Other hyper-parameters are kept the same as the original MAPO system.