Assessing The Factual Accuracy of Generated Text

Ben Goodrich, Vinay Rao, Mohammad Saleh, Peter J Liu

Introduction

Recently, there has been wide empirical success in text summarization (Rush et al., 2015; Nallapati et al., 2016; Liu et al., 2018), machine translation (Bahdanau et al., 2015; Wu et al., 2016; Vaswani et al., 2017), dialogue response generation (Li et al., 2017; Serban et al., 2017b; Serban et al., 2017a), and other text generation tasks. For evaluation, these models generally rely on metrics like ROUGE (Recall-Oriented Understudy for Gisting Evaluation) (Lin, 2004), BLEU (Bilingual Evaluation Understudy) (Papineni et al., 2002) and perplexity (Brown et al., 1992) that measure locally constrained n-gram overlap. In this paper, we propose an automatic metric for evaluating the factual accuracy of generated text.

A fact f is defined to be a relation tuple (subject, relation, object), where subject has a binary relation to object and can be assumed to have been inferred from text or a knowledge base, e.g. Barack Hussein Obama II (born August 4, 1961) is an American politician who served as the 44th President of the United States from January 20, 2009 to January 20, 2017 implies a set of facts such as (Barack Obama, president of, United States), (Barack Obama, born on, August 4 1961).

In this paper, we limit our scope to the task of evaluating text summarization. To evaluate a text summarization model, we compare the ground-truth summary text, T and the generated summary, G. Let ft,fg∈Ff_{t},f_{g}\in F, and FT,FG⊂FF_{T},F_{G}\subset F where FF is a set of relation tuples.

The models used in the metric we propose do not make use of world knowledge (e.g. knowledge base) during inference, and to account for that we filter FTF_{T} and FGF_{G} by only considering claims made in GG that can either be verified or refuted by statements in TT. Concretely, if ft=(subjt,relt,objt)∈FTf_{t}=(subj_{t},rel_{t},obj_{t})\in F_{T} and fg=(subjg,relg,objg)∈FGf_{g}=(subj_{g},rel_{g},obj_{g})\in F_{G}

We can then define factual accuracy factaccfact_{acc} as the precisionprecision between FT′F_{T^{\prime}} and FG′F_{G^{\prime}}.

For example, consider ground-truth summary T: Brad Pitt was born in 1963 and generated summary G: Brad Pitt was born in 1961. Then, FTF_{T} = {(Brad Pitt, born-in, 1963)}, FGF_{G} = {(Brad Pitt, born-in, 1961)}. The metric factaccfact_{acc} = 0 indicates there is no factual consistency between the two summaries, whereas another metric like ROUGE-1 (1-gram overlap) measures 0.83. A real example is highlighted in Table 1 where the summarization model commits such a mistake. It is important to be able to measure these mistakes accurately to aid in training factually accurate summarization models.

Extracting fact tuples from text has been previously studied in methods like OpenIE (Open Information Extraction) (Banko et al., 2007). OpenIE extracts triplets with an unspecified schema, and the relation is usually the text linking the two entities. However, it does not leverage information from a knowledge base and leads to outputs that are hard to compare. For example, Person was born in that town ⇒\Rightarrow (Person, born in, town). But That town is the birthplace of Person ⇒\Rightarrow (Town, is the birthplace of, Person).

We standardize comparison by studying structured approaches to relation tuple extraction where the schema is fixed. We compare two approaches for fact extraction. One is a two-step process that first involves recognizing all the named entities in a sentence, and then classifying the relation for every pair of entities in the sentence (Sorokin and Gurevych, 2017; Lin et al., 2016). Our other approach is to use an end-to-end model with a Transformer-based architecture (Vaswani et al., 2017) that is trained to output structured fact tuples. These models are described in Section 4. We create a new dataset for fact extraction using distant supervision (Mintz et al., 2009) on Wikipedia text by cross-referencing facts from the Wikidata knowledge base (Vrandečić and Krötzsch, 2014). To the best of our knowledge, this dataset is bigger and contains more relations and domains than previously used datasets for relation or fact tuple extraction.

We introduce model-based metrics to analyze the factual accuracy of generated text (Sec 4). We compare them against model-free metrics listed in Sec 5.

To train fact tuple extraction models, we release code (as part of the Tensor2Tensorhttps://github.com/tensorflow/tensor2tensor framework along with the model weightshttps://github.com/tensorflow/tensor2tensor/tree/master/tensor2tensor/data_generators/wikifact) and a large dataset (Sec 3) based on Wikidata and Wikipedia at https://github.com/google-research-datasets/wikifact.

We show that a Transformer-based end-to-end fact extraction model is able to perform structured prediction of relation tuples, avoiding the need to split the process into multiple steps (named entity recognition, coreference resolution and relation classification). It is able to extract complete sets of facts from full pages of text in one pass.

We conduct experiments to compare our proposed metric against human evaluation of factual accuracy of generated text (Sec 8.1) and show that model-based metrics are better correlated with human judgment when compared to traditional metrics like ROUGE.

Our models work under some limitations that are discussed in Sec 9.1, and then Sec 9.2 discusses future work and ways to make our models more robust.

Related Work and Motivation

Many evaluation metrics have been proposed for text generation tasks like BLEU (Papineni et al., 2002) and METEOR (Lavie and Agarwal, 2007) for machine translation and ROUGE (Lin, 2004), Basic Elements (Hovy et al., 2006) & Pyramid (Nenkova and Passonneau, 2004) for text summarization. In Steinberger and Jezek (2009), the authors explain the different kinds of evaluation we can perform for summarization. They are broadly classified as extrinsic metrics that are specific to tasks (e.g. in summarizing a person, whether the date of birth has been included) and intrinsic metrics like grammaticality, coherency and non-redundancy that are based on the analysis of the summary. ROUGE, BLEU, sentence level F1 measures, etc are intrinsic content based metrics. Zhang et al. (2018) and other related works study ways to estimate the trustworthiness of answers to a question. With the recent shift towards using neural abstractive methods for text summarization and other text generation tasks, we believe that it is important to assess the factual accuracy of generated text. Wiseman et al. (2017) have also studied some extractive evaluative methods to assess the quality of generated text. This includes a Relation Generator, which predicts the relation between entities to assess the factual correctness of generated records. However, we introduce a much larger dataset and enable training end-to-end models that can extract fact triplets from text. We additionally perform detailed analysis of the fact extraction models.

Typical fact extraction pipelines are a multistage process consisting of part-of-speech tagging, named entity recognition (Finkel et al., 2005; Lample et al., 2016; Chiu and Nichols, 2016) that produces entities {ei}\{e_{i}\} and then relation classification that predicts a relation rkr_{k} for every pair of entities (ei,ej)(e_{i},e_{j}). OpenIE (Banko et al., 2007) predicts a relation by linking the text connecting eie_{i} and eje_{j}. Because it does not have a fixed schema, logical reasoning on its outputs are not possible. Mohamed et al. (2011) extend this to start with a fixed schema that can grow with more training, yet retain a consistent output surface form.

In this paper, we consider fact classification models with fixed schema. This idea has been studied in many previous works including Surdeanu et al. (2012), which considered datasets that have multiple relation labels for an entity pair, which each may have multiple instances in the input text. This was modeled as a graphical model over latent variables. Riedel et al. (2013) treated relation extraction as reasoning with matrix-factorization, and could work with surface-form texts and knowledge-base embeddings simultaneously. However, both of these works had datasets with very few types of relations, and were shown to work over limited domains. Recently, neural networks have been used for classifying relations. Lin et al. (2016) used attention over multiple instances for the same entity pair to predict relations. Sorokin and Gurevych (2017) proposed to predict multiple relations in a sentence by using all the entity pairs and relation labels in the sentence as contextual input. We propose a simpler model where we classify relations between all the entity pairs in a sentence, without any additional context. We also make use of our proposed dataset that is bigger, more diverse and has more relation types. Our dataset also has article-level information that can be used to train models like in Section 4.2. Since using two-step processes may be affected by compounding of errors across the models, some end-to-end approaches (Miwa and Sasaki, 2014; Miwa and Bansal, 2016) have been proposed, where the models extract entities and relations in one pass through the model. However, the method used in Miwa and Sasaki (2014) required designing hand-crafted features and task-specific algorithms. Miwa and Bansal (2016) has a two-phase model that first extracts entity candidates and then predicts relations based on the parsed tree-structure of the sentence. We instead propose a sequence-to-sequence model that is able to output fact tuples directly, and does not require any feature engineering.

We found that the abstractive summarization models such as those described in Liu et al. (2018) may generate sentences with factual inaccuracies (e.g. incorrect month in date of birth, wrong city in the state, etc.). Cao et al. (2017) found that 30% of summaries generated by a state-of-the-art summarization model contained factual inaccuracies. We found by running a large-scale experiment as described in Section 8.1, that the summarization model had factual inaccuracy rate of approximately 17%. We believe that this is because such mistakes are not heavily penalized by cross-entropy or n-gram based model losses and metrics. As further motivation, we synthesized factually inaccurate samples by making simple corruptions to Wikipedia lead sections. We replaced mentions of dates (day and month only), locations or people with other entities of the same type in the text. For example, Barack was born on August 4, 1961 in Honolulu. He married Michelle on October 3, 1992 in Chicago. becomes Barack was born on October 3, 1961 in Chicago. He married Michelle on August 4, 1992 in Honolulu.. Table 2 shows that model-free metrics such as ROUGE and OpenIE-based tuple comparison do not reflect the decline in factual accuracy due to such corruption as much as the model-based metrics do.

Dataset

We create a dataset for fact extraction using distant supervision that is based entirely on the English Wikipedia corpus and the Wikidata knowledge base WKBW_{KB} (Vrandečić and Krötzsch, 2014). Our distant supervisor is very similar to the one proposed by Mintz et al. (2009). Although the inputs and labels for the classifier and end-to-end model are slightly different, we start by running an NER and co-reference resolution system\reffootnote−ner−coref{}^{\ref{footnote-ner-coref}} on each Wikipedia article. The topic of that article is considered as the subject ese_{s}. The other entities eje_{j} found in the article are considered objects. For every pair (es,ej)(e_{s},e_{j}), we say they are related if there is a relation rkr_{k} such that the triplet (es,rk,ej)(e_{s},r_{k},e_{j}) is found in WKBW_{KB}. We add this triplet to a set of positive examples EpE_{p}. If no such relation exists between ese_{s} and eje_{j}, we add the triplet (es,r0,ej)(e_{s},r_{0},e_{j}) (r0r_{0} denotes no-relation) to a set of negative examples EnE_{n}.

Model-based Metrics

In this section we describe models that can extract fact tuples from text and how we use them to define the factual accuracy metric as defined in Eq 1. Given some input text XX, we then extract claims made in XX as fact tuples.

This approach consists of two steps, where we first recognize all the named entities eie_{i} from XX and then classify relations between entity pairs (ei,ej)(e_{i},e_{j}).

Entities are real-world objects like people, locations, organizations etc that can be identified by a proper namehttps://en.wikipedia.org/wiki/Named_entity. Entities can be identified with named-entity recognition (NER) systems like Chiu and Nichols (2016); Lample et al. (2016); Finkel et al. (2005) that take in XX and produce the set {ei}\{e_{i}\}. NER is followed by co-reference resolutionWhile we use an NER and co-reference resolution system that is not available to the public, the dataset we release (Section 3) has the positions of all the recognized and resolved entities that we use for training our classifier. (Clark and Manning, 2016; Recasens et al., 2013; Lee et al., 2011; Raghunathan et al., 2010). Publicly available NER and co-reference systems include Stanford’s CoreNLPhttp://stanfordnlp.github.io/CoreNLP/coref.html and NLTKhttps://www.nltk.org/.

1.2. Relation Classifier

For every pair (ei,ej), ei≠ej(e_{i},e_{j}),\ e_{i}\not=e_{j} we consider all sentences SlS_{l} in XX that contain both entities. The input to the classifier is then each of these sentences SlS_{l}. Because a sentence may contain multiple entities, we also add a prefix SUBJSUBJ for eie_{i} and OBJOBJ to eje_{j} as a hint. For example, XX = Person1 was born in City1 becomes SlS_{l} = SUBJ{ Person1 } was born in OBJ{ City1 }. Unlike Sorokin and Gurevych (2017), our classifier does not require additional context. Let sis_{i} be a token in the input sentence SlS_{l} after NER, and rkr^{k} denote the kkth relation. Our classifier takes in input tokens sis_{i} that are first embedded onto a latent space, and then a stack of Transformer encoder-only layers process the whole sequence. A subsequent max-pooling layer selects one of these outputs that is then converted to a probability estimate of relations by a sigmoid operation. The exact series of operations can be viewed as:

Figure 1(a) also shows the architecture of this model.

1.3. Dataset preparation

For every triplet ff in Ep∪EnE_{p}\cup E_{n}, we have sentence(s)There may be more than one sentence in the article that have mentions of the subject and object entity pair. SlS_{l} in the article that may describe the relation between ese_{s} and eje_{j}. SlS_{l} is processed so that subject and object are prefixed with “SUBJ” and “OBJ” as a hint to the model (Section 4.1). This leads to a dataset with 2.9 million positive examples and 34 million negative examples totaling to 45GiB on disk.

The classifier predicts a relation rkr_{k} for each entity pair (ei,ej)(e_{i},e_{j}). We extract such triplets from the ground-truth TT and generated text GG, and use the definition from eq 1 to calculate the factual accuracy.

2. End-to-End Extraction

We propose an end-to-end fact extraction model to avoid compounding of errors across components in multi-stage approaches like Section 4.1 (Mccallum and Jensen, 2003). This model also does not require any feature engineering or context. The input to the model is text XX of any length (sentence/paragraph/article) and the subjectsubject entity ese_{s} prefixed to XX. All the inputs tokens in [es;X][e_{s};X] are first embedded onto a latent space. A Transformer model consisting of a stack of encoder layers followed by decoder layers produces an output sequence of arbitrary length. A softmax operation is applied to every output token to define a distribution at every timestep. Figure 1(b) shows the architecture of this model. To encourage the model to have structured outputs, we train the model with labels that are a sequence of fact tuples. For example, if XX = “ Person1 was born in Country1. He was a painter”, then the label, YY, for that input is “Person1 ⟨\langlet⟩\rangle born in ⟨\langlet⟩\rangle Country1 ⟨\langlef⟩\rangle Person1 ⟨\langlet⟩\rangle profession ⟨\langlet⟩\rangle painter ⟨\langleend⟩\rangle”, where ⟨\langlet⟩\rangle separates tokens within the fact fif_{i} and ⟨\langlef⟩\rangle separates facts. For prediction, we perform a beam search over all the output timesteps, and continue decoding until ⟨\langleend⟩\rangle is predicted. A length-penalty α\alpha controls the length of this prediction as in (Wu et al., 2016).

If the input article text is XX, every triplet fpf_{p} in EpE_{p} (we ignore the negative examples for end-to-end models because no relations between entity pairs is implied by no output by the model) is appended to the article’s label LL. LL will then contain a series of tokens that describe facts, with seperators between them. For example: es⟨t⟩r1⟨t⟩e1⟨f⟩es⟨t⟩r2⟨t⟩e2⟨f⟩...e_{s}\langle t\rangle r_{1}\langle t\rangle e_{1}\langle f\rangle e_{s}\langle t\rangle r_{2}\langle t\rangle e_{2}\langle f\rangle... We also prepend the input text XX with ese_{s} ([es;X][e_{s};X]) as a hint to the model for generating facts about ese_{s}. This leads to a dataset with 2.5 million examples totaling to 1.5GiB on disk. This dataset is made available at https://github.com/tensorflow/tensor2tensor/tree/master/tensor2tensor/data_generators/wikifact

The End-to-End model is able to produce a sequence of fact tuples in the form, subj1subj_{1} ⟨\langlet⟩\rangle rel1rel_{1} ⟨\langlet⟩\rangle obj1obj_{1} ⟨\langlef⟩\rangle subj1subj_{1} ⟨\langlet⟩\rangle rel2rel_{2} ⟨\langlet⟩\rangle obj2obj_{2}. It is trained to output relations from a fixed schema based on WikiData. Consider an output from this model, Barack Obama⟨t⟩P69⟨t⟩Harvard\textit{Barack Obama}\langle t\rangle P69\langle t\rangle Harvard. P69P69 denotes ‘educated at’https://www.wikidata.org/wiki/Property:P69. These tuples are extracted from TT and GG to fit into the metric defined in eq 1.

3. NER + Binary Relation Classifier

Similar to the typical relation classifier detailed in Sec 4.1, we define a classifier that predicts whether a pair of entities (ei,ej)(e_{i},e_{j}) are related to each other through any relation. This allows for verifying that entities are related in both the ground-truth TT and generated text GG, while being flexible enough to allow for any relation types. We also note that two entities can be related to each other in multiple ways. The inputs to this model are the same as Sec 4.1, but the model is expected to output relrel as

Data for this model is generated with the same procedure detailed in Sec 4.1.3. The only difference is the way we define the label relrel. We consider entities eie_{i} and eje_{j} to be related if there is a relation rkr_{k} such that (ei,rk,ej)(e_{i},r_{k},e_{j}) is found in WKBW_{KB}.

The model predicts relrel for each entity pair (ei,ej)(e_{i},e_{j}), and we are able to extract a set of tuples of the form (ei,rel,ej)(e_{i},rel,e_{j}) from both TT and GG. To use eq 1 to define the factual accuracy, we filter the set by considering only entity pairs (ei,ej)(e_{i},e_{j}) that are found in both TT and GG to then compare the predicted label relrel between them.

Model-free Metrics

We describe model-free automatic metrics in this section. Unlike model-based metrics, they are not susceptible to changes in training data, and might be considered easier to interpret or understand.

ROUGE (Lin, 2004) has been used as an automatic metric to judge the quality of generated text, and has shown to correlate well with human judgment of overall linguistic quality of the text.

2. OpenIE

OpenIE (Banko et al., 2007) is a tool that can extract relation tuples from text, without a specified schema. We use it to extract sets of relation tuples from TT and GG, and then compute the precision like in eq 1.

Model Experiments

In this section, we describe the methods we used to train and evaluate our relation extraction models. All of our proposed classifiers and end-to-end models have 6 Transformer layers and 1 embedding layer, with number of neurons (hidden layer size) set to 512. In the Transformer-based models, we use 8 attention heads. Our models are trained using the AdaFactor (Shazeer and Stern, 2018) optimizer. We use the publicly available Tensor2Tensor (Vaswani et al., 2018)https://github.com/tensorflow/tensor2tensor framework for our experiments and will be releasing our code extensions as part of that framework. On our proposed dataset, the classifiers are trained for 50,000 iterations with batch-size of 1024 and the end-to-end models are trained for 50,000 iterations with batch-size of 256. We evaluate classifiers and end-to-end models on our dataset. These results are presented in Table 3. The end-to-end model is learning to recognize entities, resolving entity co-references, and reason about their relation in one pass through the model. To the best of our knowledge, we are not aware of other end-to-end structured relation extraction models and therefore do not include a comparison against other approaches. Some examples of extracting facts on our dataset are shown in A.2, where we include a comparison to OpenIE’s triplet extraction.

We calculate precision and recall in the above experiments by matching ground-truth fact tuples exactly. This implies that the end-to-end model is not only learning to identify entities and resolve co-references, but also predict structured output, and its outputs can be used for reasoning. Their performance is competitive against relation classifiers while having a simple training and inference routine.

For each model, we sort and select the ten most frequent relation types that appear in our test sets. The F1F1 measure on these relations for classifiers are shown in Table 4, and end-to-end models are shown in Table 5.

Error Analysis of Model Predictions

Distant supervision (Mintz et al., 2009) is a way to create training data by using weak signals. In our dataset, we assign a relation label rkr_{k} for every entity pair (ei,ej)(e_{i},e_{j}) in the input text XX if the relation tuple (ei,rk,ej)(e_{i},r_{k},e_{j}) exists in the Wikidata knowledge base WKBW_{KB}. However, the sentence SlS_{l} containing (ei,ej)(e_{i},e_{j}) may not necessarily entail rkr_{k}. This leads to inaccurate estimates of the true-positive rate for our fact extraction models. We evaluate the effect of this distant supervision by gathering the set of facts extracted from our models that are marked false-positive by the distant supervision scheme. We present a pair of input text (Wikipedia articles) and facts extracted by our models to human evaluators, and ask them to mark a fact to be True only if the relation tuple (subject,relation,object)(subject,relation,object) is implied by the input text. We asked two evaluators to score facts marked false-positive from a random set of 30 Wikipedia articles. We consider the fact to be true if both evaluators agree. We present the results in Table 6, where we can see the rate of false-positive facts that were marked true by the evaluators. This suggests that the end-to-end models could benefit by a better labeling scheme.

In this section, we show the effectiveness of our proposed metric on judging the factual accuracy of generated text. We use the text summarization model proposed in (Liu et al., 2018) to generate lead sections of Wikipedia articles using the dataset and model in that paper, and compare the generated summary against the real lead section. In the following section, we describe the methodology used to compare human judgment of factual accuracy and how we compare our metric against that baseline.

Every claim made in the generated text GG can be considered to belong to one of three categories: supported by a sentence in ground-truth TT, refuted by TT or cannot be verified by TT. The evaluators were asked to only consider claims that are either supported or refuted by TT. This ensures that no external knowledge is used in comparing TT and GG, and ignores all claims that cannot be verified by TT. Four evaluators were asked to rate 30 examples of generated text GG and then give it a score of 1-5 with 5 being highest factual accuracy. A special case is where the generated text has no verifiable claims. In this case, they were asked to give it a score of 1. Figure 2 shows the interface a human evaluator uses in our experiment.

We conduct the same experiment on two sets of data: first is a random sampling from summaries generated for Actors. We consider this an easier subset because we expect our fact extraction models to do well on this subset due to the summaries and Wikipedia lead sections generally containing relationships our models perform well on (see tables 4 and 5). We present these results in Table 7. We analyzed the inter-rater agreement on the scores given to each example, and found that Krippendorff’s alpha (allows for ordinal rankings) was 0.6897. The second is a random sampling from all categories in Wikipedia. The results are presented in Table 8. The inter-rater agreement on this sample was found to be 0.7530.

We see that our end-to-end model (Section 4.2) has the best correlation on both subsets, indicating that it generalizes better to generated text. This may also be because the classifier suffers from a compounding of errors, where it is unable to predict relations if the NER system fails to recognize entities.

Conclusion

The dataset we create only makes use of sentences found in Wikipedia, and facts found in WikiData. This means that our models are biased to sentences structured to the neutral tone set in Wikipedia, and towards popular types of facts expressed in WikiData such as date of birth, profession, etc. Other sources of text may have more complex structures and styles of writing that may make it hard for our models to adapt to easily. An simple example of this is negating a binary relationship with ‘not’, and different ways of expressing the same idea such as ‘wife/husband’ instead of ‘spouse’. WikiData is an incomplete knowledge base, and this also leads to many sentences that in reality imply a fact to be marked containing no facts. This is a very typical problem faced by any work using distant supervision, and is combated with methods like active learning (Shen et al., 2017). It should be noted that ROUGE and to the best of our knowledge, most other automatic metrics, are also susceptible to changes in linguistic style and structure. However, elaborate labeling and bigger datasets will allow for our models to learn to overcome these challenges.

2. Discussion and future work

We have shown that our proposed metric is able to indicate the factual accuracy of generated text, and agrees with human judgment on our datasets. By leveraging a new dataset for both relation classification and end-to-end fact extraction, we also showed that classifiers and end-to-end models with straightforward architectures are able to perform competitive fact extraction. Our end-to-end model avoids compounding of errors over sub-components typically used in other fact-extraction pipelines. We will release the code and datasets used to train this model, so that the proposed metric can be used to standardize comparison. We are in the process of building a bigger dataset that will contain multiple text domains, stronger human supervision and a larger collection of relation tuples that will help overcome many of the limitations discussed in the previous section (9.1). We encourage further development and use of this metric for automating the assessment of factual accuracy of generated text, and the development of better end-to-end models with structured outputs for fact extraction.

References

Appendix A Appendix

We release code to train our fact extraction models as part of the Tensor2Tensor frameworkhttps://github.com/tensorflow/tensor2tensor along with trained model weights at https://github.com/tensorflow/tensor2tensor/tree/master/tensor2tensor/data_generators/wikifact. A large fact extraction dataset (Sec 3) based on Wikidata and Wikipedia is made available https://github.com/tensorflow/tensor2tensor/tree/master/tensor2tensor/data_generators/wikifact. To train our end-to-end and classifier models for fact extraction, we use the hyper-parameter set “transformer_base” defined in the Tensor2Tensor frameworkhttps://github.com/tensorflow/tensor2tensor/blob/master/tensor2tensor/models/transformer.py. We further release code to use our end-to-end models as a fact extractor and calculate the factual accuracy metric at https://github.com/tensorflow/tensor2tensor/tree/master/tensor2tensor/data_generators/wikifact.

A.2. Fact extraction example

We include an example of facts extracted from text using our models where we compare it against OpenIE’s (Banko et al., 2007) triplet extraction in Table 9. This example illustrates the advantage of using structured approaches to fact extraction. OpenIE yields many triplets that mostly cannot be used for reasoning.