Learning Joint Semantic Parsers from Disjoint Data
Hao Peng, Sam Thomson, Swabha Swayamdipta, Noah A. Smith
Introduction
Semantic parsing aims to automatically predict formal representations of meaning underlying natural language, and has been useful in question answering Shen and Lapata (2007), text-to-scene generation Coyne et al. (2012), dialog systems Chen et al. (2013) and social-network extraction Agarwal et al. (2014), among others. Various formal meaning representations have been developed corresponding to different semantic theories Fillmore (1982); Palmer et al. (2005); Flickinger et al. (2012); Banarescu et al. (2013). The distributed nature of these efforts results in a set of annotated resources that are similar in spirit, but not strictly compatible. A major axis of structural divergence in semantic formalisms is whether based on spans Baker et al. (1998); Palmer et al. (2005) or dependencies (Surdeanu et al., 2008; Oepen et al., 2014; Banarescu et al., 2013; Copestake et al., 2005, inter alia). Depending on application requirements, either might be most useful in a given situation.
Learning from a union of these resources seems promising, since more data almost always translates into better performance. This is indeed the case for two prior techniques—parameter sharing FitzGerald et al. (2015); Kshirsagar et al. (2015), and joint decoding across multiple formalisms using cross-task factors that score combinations of substructures from each Peng et al. (2017). Parameter sharing can be used in a wide range of multitask scenarios, when there is no data overlap or even any similarity between the tasks Collobert and Weston (2008); Søgaard and Goldberg (2016). But techniques involving joint decoding have so far only been shown to work for parallel annotations of dependency-based formalisms, which are structurally very similar to each other Lluís et al. (2013); Peng et al. (2017). Of particular interest is the approach of Peng et al., where three kinds of semantic graphs are jointly learned on the same input, using parallel annotations. However, as new annotation efforts cannot be expected to use the same original texts as earlier efforts, the utility of this approach is limited.
We propose an extension to Peng et al.’s formulation which addresses this limitation by considering disjoint resources, each containing only a single kind of annotation. Moreover, we consider structurally divergent formalisms, one dealing with semantic spans and the other with semantic dependencies. We experiment on frame-semantic parsing Gildea and Jurafsky (2002); Das et al. (2010), a span-based semantic role labeling (SRL) task (§2.1), and on a dependency-based minimum recursion semantic parsing (DELPH-IN MRS, or DM; Flickinger et al., 2012) task (§2.2). See Figure 1 for an example sentence with gold FrameNet annotations, and author-annotated DM representations.
Our joint inference formulation handles missing annotations by treating the structures that are not present in a given training example as latent variables (§3).Following past work on support vector machines with latent variables (Yu and Joachims, 2009), we use the term “latent variable,” even though the model is not probabilistic. Specifically, semantic dependencies are treated as a collection of latent variables when training on FrameNet examples.
using a latent variable formulation to extend cross-task scoring techniques to scenarios where datasets do not overlap;
learning cross-task parts across structurally divergent formalisms; and
Our approach results in a new state-of-the-art in frame-semantic parsing, improving prior work by 0.8% absolute points (§6), and achieves competitive performance on semantic dependency parsing. Our code is available at https://github.com/Noahs-ARK/NeurboParser.
Tasks and Related Work
We describe the two tasks addressed in this work—frame-semantic parsing (§2.1) and semantic dependency parsing (§2.2)—and discuss how their structures relate to each other (§2.3).
Frame-semantic parsing is a span-based task, under which certain words or phrases in a sentence evoke semantic frames. A frame is a group of events, situations, or relationships that all share the same set of participant and attribute types, called frame elements or roles. Gold supervision for frame-semantic parses comes from the FrameNet lexicon and corpus Baker et al. (1998).
Concretely, for a given sentence, , a frame-semantic parse consists of:
a set of targets, each being a short span (usually a single token96.5% of targets in the training data are single tokens.) that evokes a frame;
for each target , the frame that it evokes; and
for each frame , a set of non-overlapping argument spans in the sentence, each argument having a start token index , end token index and role label .
In this work, we assume gold targets and LUs are given, and parse each target independently, following the literature (Johansson and Nugues, 2007; FitzGerald et al., 2015; Yang and Mitchell, 2017; Swayamdipta et al., 2017, inter alia). Moreover, following Yang and Mitchell (2017), we perform frame and argument identification jointly. Most prior work has enforced the constraint that a role may be filled by at most one argument span, but following Swayamdipta et al. (2017) we do not impose this constraint, requiring only that arguments for the same target do not overlap.
2 Semantic Dependency Parsing
Broad-coverage semantic dependency parsing (SDP; Oepen et al., 2014, 2015, 2016) represents sentential semantics with labeled bilexical dependencies. The SDP task mainly focuses on three semantic formalisms, which have been converted to dependency graphs from their original annotations. In this work we focus on only the DELPH-IN MRS (DM) formalism.
Each semantic dependency corresponds to a labeled, directed edge between two words. A single token is also designated as the top of the parse, usually indicating the main predicate in the sentence. For example in Figure 1, the left-most arc has head “Only”, dependent “few”, and label arg1. In semantic dependencies, the head of an arc is analogous to the target in frame semantics, the destination corresponds to the argument, and the label corresponds to the role. The same set of labels are available for all arcs, in contrast to the frame-specific roles in FrameNet.
3 Spans vs. Dependencies
Early semantic role labeling was span-based (Gildea and Jurafsky, 2002; Toutanova et al., 2008, inter alia), with spans corresponding to syntactic constituents. But, as in syntactic parsing, there are sometimes theoretical or practical reasons to prefer dependency graphs. To this end, Surdeanu et al. (2008) devised heuristics based on syntactic head rules Collins (2003) to transform PropBank Palmer et al. (2005) annotations into dependencies. Hence, for PropBank at least, there is a very direct connection (through syntax) between spans and dependencies.
For many other semantic representations, such a direct relationship might not be present. Some semantic representations are designed as graphs from the start (Hajič et al., 2012; Banarescu et al., 2013), and have no gold alignment to spans. Conversely, some span-based formalisms are not annotated with syntax (Baker et al., 1998; He et al., 2015), In FrameNet, phrase types of arguments and their grammatical function in relation to their target have been annotated. But in order to apply head rules, the internal structure of arguments (or at least their semantic heads) would also require syntactic annotations. and so head rules would require using (noisy and potentially expensive) predicted syntax.
Inspired by the head rules of Surdeanu et al. (2008), we design cross-task parts, without relying on gold or predicted syntax (which may be either unavailable or error-prone) or on heuristics.
Model
The overall score can be decomposed into the sum of frame SRL score , semantic dependency score , and a cross-task score :
and require access to the target and LU, in addition to , but does not. For clarity, we omit the dependence on the input sentence, target, and lexical unit, whenever the context is clear. Below we describe how each of the scores is computed based on the individual parts that make up the candidate parses.
The score of a frame-semantic parse consists of
the score for argument parts, , each associated with a token span and semantic role from .
The computation of is described in §4.2.
Semantic dependency score.
Following Martins and Almeida (2014), we consider three types of parts in a semantic dependency graph: semantic heads, unlabeled semantic arcs, and labeled semantic arcs. Analogous to Equation 3, the score for a dependency graph is the sum of local scores:
The computation of is described in §4.3.
Cross task score.
In addition to task-specific parts, we introduce a set of cross-task parts. Each cross-task part relates an argument part from to an unlabeled dependency arc from . Based on the head-rules described in §2.3, we consider unlabeled arcs from the target to any token inside the span.Most targets are single-words (§2.1). For multi-token targets, we consider only the first token, which is usually content-bearing. Intuitively, an argument in FrameNet would be converted into a dependency from its target to the semantic head of its span. Since we do not know the semantic head of the span, we consider all tokens in the span as potential modifiers of the target. Figure 2 shows examples of cross-task parts. The cross-task score is given by
The computation of is described in §4.4.
In contrast to previous work (Lluís et al., 2013; Peng et al., 2017), where there are parallel annotations for all formalisms, our input sentences contain only one of the two—either the span-based frame SRL annotations, or semantic dependency graphs from DM. To handle missing annotations, we treat semantic dependencies as latent when decoding frame-semantic structures.Semantic dependency parses over a sentence are not constrained to be identical for different frame-semantic targets. Because the DM dataset we use does not have target annotations, we do not use latent variables for frame semantic structures when predicting semantic dependency graphs. The parsing problem here reduces to
Parameterizations of Scores
This section describes the parametrization of the scoring functions from §3. At a very high level: we learn contextualized token and span vectors using a bidirectional LSTM (biLSTM; Graves, 2012) and multilayer perceptrons (MLPs) (§4.1); we learn lookup embeddings for LUs, frames, roles, and arc labels; and to score a part, we combine the relevant representations into a single scalar score using a (learned) low-rank multilinear mapping. Scoring frames and arguments is detailed in §4.2, that of dependency structures in §4.3, and §4.4 shows how to capture interactions between arguments and dependencies. All parameters are learned jointly, through the optimization of a multitask objective (§5).
The order of a tensor is the number of its dimensions—an order-2 tensor is a matrix and an order-1 tensor is a vector. Let denote tensor product; the tensor product of two order-2 tensors and yields an order- tensor where . We use to denote inner products.
1 Token and Span Representations
The representations of tokens and spans are formed using biLSTMs followed by MLPs.
Each token in the input sentence is mapped to an embedding vector. Two LSTMs (Hochreiter and Schmidhuber, 1997) are run in opposite directions over the input vector sequence. We use the concatenation of the two hidden representations at each position as a contextualized word embedding for each token:
Span representations.
Following Lee et al. (2017), span representations are computed based on boundary word representations and discrete length and distance features. Concretely, given a target and its associated argument with boundary indices and , we compute three features based on the length of , and the distances from and to the start of . We concatenate the token representations at ’s boundary with the discrete features . We then use a two-layer -MLP to compute the span representation:
The target representation is similarly computed using a separate , with a length feature but no distance features.
2 Frame and Argument Scoring
As defined in §3, the representation for a predicate part incorporates representations of a target span, the associated LU and the frame evoked by the LU. The score for a predicate part is given by a multilinear mapping:
A candidate argument consists of a span and its role label, which in turn depends on the frame, target and LU. Hence the score for argument part, is given by extending definitions from Equation 9:
where is a low-rank order-2 tensor of learned parameters and is a learned lookup embedding of the role label.
3 Dependency Scoring
Local scores for dependencies are implemented with two-layer -MLPs, followed by a final linear layer reducing the represenation to a single scalar score. For example, let denote an unlabeled arc (ua). Its score is:
where is a vector of learned weights. The scores for other types of parts are computed similarly, but with separate MLPs and weights.
4 Cross-Task Part Scoring
As shown in Figure 2, each cross-task part consists of two first-order parts: a frame argument part , and an unlabeled dependency part, . The score for a cross-task part incorporates both:
where is a low-rank order-2 tensor of parameters. Following previous work (Lei et al., 2014; Peng et al., 2017), we construct the parameter tensors , , and so as to upper-bound their ranks.
Training and Inference
All parameters from the previous sections are trained using a max-margin training objective (§5.1). For inference, we use a linear programming procedure, and a sparsity-promoting penalty term for speeding it up (§5.2).
Let denote the gold frame-semantic parse, and let denote the cost of predicting with respect to . We optimize the latent structured hinge loss (Yu and Joachims, 2009), which gives a subdifferentiable upper-bound on :
Following Martins and Almeida (2014), we use a weighted Hamming distance as the cost function, where, to encourage recall, we use costs 0.6 for false negative predictions and 0.4 for false positives. Equation 13 can be evaluated by applying the same max-decoding algorithm twice—once with cost-augmented inference (Crammer et al., 2006), and once more keeping fixed. Training then aims to minimize the average loss over all training instances.We do not use latent frame structures when decoding semantic dependency graphs (§3). Hence, the loss reduces to structured hinge (Tsochantaridis et al., 2004) when training on semantic dependencies.
Another potential approach to training a model on disjoint data would be to marginalize out the latent structures and optimize the conditional log-likelihood (Naradowsky et al., 2012). Although max-decoding and computing marginals are both NP-hard in general graphical models, there are more efficient off-the-shelf implementations for approximate max-decoding, hence, we adopt a max-margin formulation.
2 Inference
We formulate the maximizations in Equation 13 as 0–1 integer linear programs and use AD3 to solve them (Martins et al., 2011). We only enforce a non-overlapping constraint when decoding FrameNet structures, so that the argument identification subproblem can be efficiently solved by a dynamic program (Kong et al., 2016; Swayamdipta et al., 2017). When decoding semantic dependency graphs, we enforce the determinism constraint (Flanigan et al., 2014), where certain labels may appear on at most one arc outgoing from the same token.
where is a hyperparameter, set to as a practical tradeoff between efficiency and development set performance. Whenever the score for a cross-task part is driven to zero, that part’s score no longer needs to be considered during inference. It is important to note that by promoting sparsity this way, we do not prune out any candidate solutions. We are instead encouraging fewer terms in the scoring function, which leads to smaller, faster inference problems even though the space of feasible parses is unchanged.
Experiments
Our model is evaluated on two different releases of FrameNet: FN 1.5 and FN 1.7,https://FN.icsi.berkeley.edu/fndrupal/ using splits from Swayamdipta et al. (2017). Following Swayamdipta et al. (2017) and Yang and Mitchell (2017), each target annotation is treated as a separate training instance. We also include as training data the exemplar sentences, each annotated for a single target, as they have been reported to improve performance (Kshirsagar et al., 2015; Yang and Mitchell, 2017). For semantic dependencies, we use the English DM dataset from the SemEval 2015 Task 18 closed track (Oepen et al., 2015).http://sdp.delph-in.net/. The closed track does not have access to any syntactic analyses. The impact of syntactic features on SDP performance is extensively studied in Ribeyre et al. (2015). DM contains instances from the WSJ corpus for training and both in-domain (id) and out-of-domain (ood) test sets, the latter from the Brown corpus.Our FN training data does not overlap with the DM test set. We remove the 3 training sentences from DM which appear in FN test data. Table 1 summarizes the sizes of the datasets.
Baselines.
We compare FN performance of our joint learning model (Full) to two baselines:
A single-task frame SRL model, trained using a structured hinge objective.
A joint model without cross-task parts. It demonstrates the effect of sharing parameters in word embeddings and LSTMs (like in Full). It does not use latent semantic dependency structures, and aims to minimize the sum of training losses from both tasks.
We also compare semantic dependency parsing performance against the single task model by Peng et al. (2017), denoted as NeurboParser (Basic). To ensure fair comparison with our Full model, we made several modifications to their implementation (§6.3). We observed performance improvements from our reimplementation, which can be seen in Table 5.
Pruning strategies.
For frame SRL, we discard argument spans longer than 20 tokens (Swayamdipta et al., 2017). We further pretrain an unlabeled model and prune spans with posteriors lower than , with being the input sentence length. For semantic dependencies, we generally follow Martins and Almeida (2014), replacing their feature-rich pruner with neural networks. We observe that spans/arcs remain after pruning, with around 96% FN development recall, and more than 99% for DM.On average, around argument spans, and unlabeled dependency arcs remain after pruning.
1 Empirical Results
Table 2 compares our full frame-semantic parsing results to previous systems. Among them, Täckström et al. (2015) and Roth (2016) implement a two-stage pipeline and use the method from Hermann et al. (2014) to predict frames. FitzGerald et al. (2015) uses the same pipeline formulation, but improves the frame identification of Hermann et al. (2014) with better syntactic features. open-SESAME (Swayamdipta et al., 2017) uses predicted frames from FitzGerald et al. (2015), and improves argument identification using a softmax-margin segmental RNN. They observe further improvements from product of experts ensembles (Hinton, 2002).
The best published FN 1.5 results are due to Yang and Mitchell (2017). Their relational model (Rel) formulates argument identification as a sequence of local classifications. They additionally introduce an ensemble method (denoted as All) to integrate the predictions of a sequential CRF. They use a linear program to jointly predict frames and arguments at test time. As shown in Table 2, our single-model performance outperforms their Rel model, and is on par with their All model. For a fair comparison, we build an ensemble (Full, ) by separately training two models, differing only in random seeds, and averaging their part scores. Our ensembled model outperforms previous best results by 0.8% absolute.
Table 3 compares our frame identification results with previous approaches. Hermann et al. (2014) and Hartmann et al. (2017) use distributed word representations and syntax features. We follow the Full Lexicon setting (Hermann et al., 2014) and extract candidate frames from the official directories. The Ambiguous setting compares lexical units with more than one possible frames. Our approach improves over all previous models under both settings, demonstrating a clear benefit from joint learning.
We observe similar trends on FN 1.7 for both full structure extraction and for frame identification only (Table 4). FN 1.7 extends FN 1.5 with more consistent annotations. Its test set is different from that of FN 1.5, so the results are not directly comparable to Table 2. We are the first to report frame-semantic parsing results on FN 1.7, and we encourage future efforts to do so as well.
Semantic dependency parsing results.
Table 5 compares our semantic dependency parsing performance on DM with the baselines. Our reimplementation of the Basic model slightly improves performance on in-domain test data. The NoCTP model ties parameters from word embeddings and LSTMs when training on FrameNet and DM, but does not use cross-task parts or joint prediction. NoCTP achieves similar in-domain test performance, and improves over Basic on out-of-domain data. By jointly predicting FrameNet structures and semantic dependency graphs, the Full model outperforms the baselines by more than 0.6% absolute scores under both settings.
Previous state-of-the-art results on DM are due to the joint learning model of Peng et al. (2017), denoted as NeurboParser (Freda3). They adopted a multitask learning approach, jointly predicting three different parallel semantic dependency annotations. Our Full model’s in-domain test performance is on par with Freda3, and improves over it by 0.6% absolute on out-of-domain test data. Our ensemble of two Full models achieves a new state-of-the-art in both in-domain and out-of-domain test performance.
2 Analysis
Similarly to He et al. (2017), we categorize prediction errors made by the Basic and Full models in Table 6. Entirely missing an argument accounts for most of the errors for both models, but we observe fewer errors by Full compared to Basic in this category. Full tends to predict more arguments in general, including more incorrect arguments.
Since candidate roles are determined by frames, frame and role errors are highly correlated. Therefore, we also show the role errors when frames are correctly predicted (parenthesized numbers in the second row). When a predicted argument span matches a gold span, predicting the semantic role is less challenging. Role errors account for only around 13% of all errors, and half of them are due to mispredictions of frames.
Performance by argument length.
Figure 3 plots dev. precision and recall of both Basic and Full against binned argument lengths. We observe two trends: (a) Full tends to predict longer arguments (averaging 3.2) compared to Basic (averaging 2.9), while keeping similar precision;Average gold span length is 3.4 after discarding those longer than 20. (b) recall improvement in Full mainly comes from arguments longer than 4.
3 Implementation Details
Our implementation is based on DyNet (Neubig et al., 2017).https://github.com/clab/dynet We use predicted part-of-speech tags and lemmas using NLTK (Bird et al., 2009).http://www.nltk.org/
Modifications to Peng et al. (2017).
To ensure fair comparisons, we note two implementation modifications to Peng et al.’s basic model. We use a more recent version (2.0) of the DyNet toolkit, and we use 50-dimensional lemma embeddings instead of their 25-dimensional randomly-initialized learned word embeddings.
Conclusion
We presented a novel multitask approach to learning semantic parsers from disjoint corpora with structurally divergent formalisms. We showed how joint learning and prediction can be done with scoring functions that explicitly relate spans and dependencies, even when they are never observed together in the data. We handled the resulting inference challenges with a novel adaptation of graphical model structure learning to the deep learning setting. We raised the state-of-the-art on DM and FrameNet parsing by learning from both, despite their structural differences and non-overlapping data. While our selection of factors is specific to spans and dependencies, our general techniques could be adapted to work with more combinations of structured prediction tasks. We have released our implementation at https://github.com/Noahs-ARK/NeurboParser.
Acknowledgments
We thank Kenton Lee, Luheng He, and Rowan Zellers for their helpful comments, and the anonymous reviewers for their valuable feedback. This work was supported in part by NSF grant IIS-1562364.