A Dataset for Answering Time-Sensitive Questions
Wenhu Chen, Xinyi Wang, William Yang Wang
Introduction
As time evolves, many facts will evolve along with it, such as ‘the U.S. President’, ‘Home Team of Lebron James’, etc. Understanding the scope and interval of knowledge is an essential task studied by previous literature (Allen, 1983). The fact evolution is commonly reflected in our daily text corpora like Wikipedia or Daily News. For example, the Wikipedia of ‘Lebron James’https://en.wikipedia.org/wiki/LeBron_James covers the whole evolution of his home team. The temporal transition of these facts is normally scattered across the long document in very diverse expressions, either represented explicitly or implicitly. Such characteristics pose great challenges to the existing NLP models. For example, a user might pose a question like ‘Which team did Lebron James play for in 2007?’. The only valid evidence in Wikipedia is ‘Cleveland Cavaliers (2003–2010)’. To answer this question, the model is required to perform temporal reasoning. We simulate these time-sensitive trivia questions and present them to the current state-of-art QA models (Zaheer et al., 2020; Izacard and Grave, 2020) trained on large-scale datasets. These models can only achieve a compromised 27% accuracy, much lower than their performance on Natural Questions (Kwiatkowski et al., 2019) with 64% accuracy and SQuAD (Rajpurkar et al., 2016, 2018) with 90% accuracy. This gap indicates the difficulties to handle temporal questions.
In this paper, we are specifically interested in handling these temporal questions. We formally define these questions as time-sensitive questions based on the following criterion: a) the question contains a [time specifier] like ‘in 2007’ or ‘before 2010’, b) modifying the [time specifier] will lead to answer change. c) such questions require temporal reasoning. We found that these time-sensitive questions despite their ubiquity are under-studied in the existing QA datasets. For example, our human study reveals that Natural Questions (Kwiatkowski et al., 2019) only contains less than 5% of questions with [time specifier]. In SQuAD (Rajpurkar et al., 2016, 2018), though there is a larger portion of questions with [time specifier], these questions usually copy the original time-specifying phrases from the passage without requiring any temporal reasoning, thus not meeting condition c). In TriviaQA/WebQuestions/WebComplexQuestions (Berant et al., 2013; Bao et al., 2016; Joshi et al., 2017), our human study reveals that there are more questions involving time specifiers, however, most of these specifiers in these questions are not modifiable. An example is ‘What is the title of the last Harry Potter novel in 2007’, where ‘last Harry Potter’ already implies ‘year=2007’, therefore, the [time specifier] is redundant and not modifiable. Thus, these questions cannot be considered time-sensitive due to condition b).
The closest to ours is Tempquestions (Jia et al., 2018a, b), which investigates the temporal questions with time specifiers. However, their questions are extracted from the above-mentioned datasets, which fails to meet the condition b). Furthermore, Tempquestions studies KG-based QA instead of Text-based QA, which differentiates from our goal of understanding temporal transition in natural text. Therefore, we propose to construct our own dataset called Time-Sensitive Question Answering (TimeQA). We first identify time-evolving facts from WikiData (Vrandečić and Krötzsch, 2014), and then employ human workers to annotate the boundaries of these facts by aligning with Wikipedia passages. We synthesize diverse question-answer pairs based on the annotated time-evolving facts using diverse templates. Finally, we create two datasets (easy and hard) with two levels of difficulty, both containing 20K question-answer pairs regarding 5.5K time-evolving fact and 70 relations. The hard version is more challenging as it requires more temporal reasoning than the easy version. The example in Figure 1 showcases some examples in our TimeQA dataset. The challenges posed by our dataset is in two folds:
Temporal Understanding: understand the time scope (start and end time) of facts in the long text. However, the time information can be expressed implicitly in the text, which requires temporal commonsense to understand, for example, ‘during the second world war’ implies ‘from 1939 to 1945’, ‘one year after 1934’ refers to ‘year 1944’, etc.
Temporal Reasoning: reason over the temporal information in the text conditioned on the query. More formally, the model needs to understand the temporal relationship (‘within’, ‘between’, ‘before’, ‘after’, etc.) between the time presented in the query and document.
We evaluate different state-of-the-art QA models’ performance on both easy and hard versions, the performance drops from 60% to 45% on the hard version, which indicates the model suffers from its incompetence to perform temporal reasoning. When comparing with human performance of 87% on TimeQA-hard, the existing neural models are still significantly lagging behind. Therefore, we believe TimeQA could serve as a valuable benchmark in studying this problem.
Dataset and Problem Definition
Here we demonstrate our dataset construction pipeline, our dataset construction takes two steps: 1) fact annotation, 2) question-answer synthesizing.
The fact annotation takes three steps as depicted in Figure 2, which is comprised of three sub-steps, mining facts, aligning facts, and verify facts.
We first identify which facts are evolving over time and through the existing annotations from WikiData (Vrandečić and Krötzsch, 2014). Therefore, we resort to WikiData to mine these time-evolving facts, please refer to the ‘Lebron James’ example in https://www.wikidata.org/wiki/Q36159. As indicated in the first-row of Figure 2, we traverse the WikiData pages and select the facts with time quantifier P580 (start from), P582 (end in), and P585 (point in time) to mine interesting knowledge triples with their temporal quantifiers. We discard the triples with numeric objects like “(Denver, population of county, 13876)" because these numerical facts are unlikely to appear in Wikipedia text. We succeed to mine roughly 150K time-evolving facts in the form of (, , {: , : , … : }), where denotes the time boundary of -th fact segment.
After mining these time-evolving facts, we need to trace them back to their Wikipedia text. We decide whether the fact segment (, , : ) is mentioned in the Wikipedia page using the following rules: 1) is hyperlinked in the Wikipedia page, 2) ’s name has exact string match, 3) the TF-IDF score between the ’s name and Wikipedia text is above 40%. 4) Otherwise, the segment will be deemed “[unanswerable]". The process is depicted in the second-row of Figure 2. After this automatic tracing procedure, we discard the time-evolving facts having over than 50% “[unanswerable]" segments. On the other hand, we found that some particular relations like ‘play for’ are dominating the dataset. Hence, we further propose to down-sample these over-represented relations and finally identify 5.5K well-balanced facts as our candidate knowledge triples for next-step human verification.
The previous step generates relatively noisy (text, time-evolving fact) pairs. There mainly exist the following sources of errors: 1) the WikiData annotation is erroneous, 2) the object is mentioned in the text, but its time boundary is not mentioned, 3) the object takes a different surface form, therefore the detected ‘[unanswerable]’ is indeed answerable. Therefore, we propose to add human verification to clean these noisy data pairs. We provide the 5K (text, time-evolving fact) pairs as HITs to high-quality Amazon Mechanical TurkersWe select the workers from English-speaking countries, with acceptance rate over 98% and over 1000 annotation HITs being approved.. The workers can take the following actions: a) correcting the erroneous object, b) changing object to [unanswerable], c) changing [unanswerable] to an object in the text. The process is depicted in the third-row of Figure 2. Each HIT is paid with 1.0 dollars with an average finish time of 5 minutes. The average hourly pay is 12.0 dollars, which exceeds the income requirements proposed in human subject research protocolshttps://en.wikipedia.org/wiki/Minimum_wage_in_the_United_States. The annotation interface is demonstrated in the Appendix.
We make following guideline during annotation to help crowd-workers deal with ambivalent cases. We define extracted to be the correct time scope for the given fact (, , ) under the following conditions: a) , are all explicitly mentioned in the passage, b) and is mentioned in the passage and . c) is more coarse grained than (passage mentions 2018, the question mentions 2018 June), d) is mentioned in the passage, by combining with a reasoning function , we are able to derive . For example, ‘X happens in 1987 (), after two years, Y …’ entails ‘Y happening in 1989 ()’, similarly, ‘X went to university in 1987 ()’ entails ‘X graduated in 1991 ()’. If none of the above conditions are met, we will define the time scope to be ‘unknown’, later on, these facts will be used to synthesize unanswerable questions.
In order to harvest a high-quality dataset, we perform very detailed quality control in the collection procedure. In the interface, we will highlight the mentioned objects and time with special fonts to help the annotators identify them. We batch the HITs by their worker id to accelerate our quality assessment procedure. We sample 2 HITs from each batch and send them to our high-quality verifier to evaluate their correctness. If the sampled HITs pass our quality assessment, we will accept the other HITs within the batch. Otherwise, we will reject the whole batch. The overall acceptance rate is maintained between 85-90%.
During the verification step, roughly 41.6% of the segments are revised. The 5.5K worker-annotated examples are further filtered to obtain 5060 ‘golden’ (text, time-evolving fact) pairs as the final release, with each fact having an average of 4 segments (time span). Out of these facts, 12% of the segments are ‘[unanswerable]’, 7% have multiple objects, and 80% have exactly one object. The harvested facts involve over 60 different relations like ‘play for’, ‘employee of’, etc.
2 Synthesizing Question-Answer Pairs
Once we obtain the human-corrected time-evolving facts, the following step is to generate question-answer pairs from these facts.
This is the main dataset we will use throughout our paper, which relies mainly on template generation. The synthesizing procedure is described in Figure 3. For each given relation, we manually write 2-5 different templates.
We propose several common reasoning types (‘in’, ‘between’, ‘before’, ‘after’) to fill in the placeholder. For example, if we want to generate ‘in (implicit)’ type, we will randomly sample a time ‘July 1776’ within the segment ‘1775/6 - 1778/6’ and create ‘in July 1776’ to fill in the placeholder.
As shown in Figure 3, we classify these types into ‘easy’ or ‘hard’ categories depending on whether the [time specifier] exactly matches the boundaries in the time-evolving axis. Since these boundaries are more likely to be mentioned in the passage explicitly, the questions with such [time specifier] on the boundary are easier to be answer based on surface form rather than temporal reasoning. In contrast, [time specifier] falling in the middle of the time span are more likely to necessitate reasoning over the implicit time information.
To better diagnose the model’s capability to perform different levels of temporal reasoning, we generate two versions of TimeQA dataset. In the easy version, we only sample from ‘easy’ reasoning types. In the hard version, we only sample from ‘hard’ reasoning types. In total, we generate 20K questions (average 4 questions per time-evolving fact), for both versions. The easy questions tend to have more explicit mentions in the document, while the hard questions containing more implicit mentions, which is more challenging for the QA models. In the following experiments, we will demonstrate the performance difference between these two versions. The comprehensive statistics of both the easy and hard versions of TimeQA are shown in Table 1. The license and privacy information for the dataset are shown in the Appendix for reference.
To further complement the template-generated questions, we sample a relation-balanced subset of train/test questions for human paraphrasing. In this paraphrasing process, the crowd-workers are required to rewrite the template questions to make them more natural, unambiguous, and diverse. Specifically, the human-written questions will include diverse mentions over the existing relations and entities as shown in Figure 4.
We finally obtain an extra 1171 easy/hard questions regarding 320 time-evolving facts as our training data and 989 easy/hard questions regarding 257 time-evolving facts as our test data. The relations for this human-annotated subset are well balanced to avoid excessive over-fitting.
Models
Here we formally define the problem setup. The model is given the document and question , where and refers to -th token in document and question with a length of N and M. The model needs to predict an answer string . To cope with the existing challenges, especially the long-term dependency, we propose to use two models BigBird (Zaheer et al., 2020) and FiD (Izacard and Grave, 2020), which are known to achieve state-of-the-art performance on the Natural Question (Kwiatkowski et al., 2019) and TriviaQA (Joshi et al., 2017). We briefly describe their design as follows:
This model aims to extract the start and end positions from the given sequence. The input sequence is a concatenated sequence of question and document . Due to the length of the document, the input sequence can easily exceed 4K tokens. Therefore, BigBird (Zaheer et al., 2020) uses a more generalized attention mechanism as:
During inference time, we select as the start and end position of the prediction span, and the answer prediction is .
2 FiD Generative Model
Since the decoder sequence here the attention is much shorter, the attention cost is ignorable. Thus, the total computation complexity is lowered to . With the almost linear approximation, the generative model can handle input sequences with a length of 4K tokens. We demonstrate the model architecture on the right side of Figure 5.
Experiments
We conduct all the experiments based on HugginFace Transformer (Wolf et al., 2020). The BigBird transformer checkpoints fine-tuned on NQ and TriviaQA are downloaded from vasudevgupta/bigbird-roberta-natural-questions and google/bigbird-base-trivia-itc. These two models are based on BigBird-base with 12 layers, 12 attention heads, and a hidden dimension of 768. The maximum position embedding is 4096. Both models use a local block size of 64 and 3 random blocks for global attention. The FiD transformer checkpoints fine-tuned on NQ and TriviaQA are downloaded from https://github.com/facebookresearch/FiD. The FiD-base model also uses 12 layers of encoder/decoder with 12 attention heads. But its maximum position embedding is limited to 512. Thus, Both models are quite comparable in terms of parameter size.
We fine-tune all the models using AdamW (Loshchilov and Hutter, 2018) with a learning rate of 2e-5. We fine-tune all the models for 3 epochs and evaluate the performance after each epoch on the dev set to select the best-performing model. The models are trained on 4 Titan RTX GPU with 24G memory with a per-GPU-batch-size of 1. For the results in Table 2, Easy/Hard-Mode means we use the corresponding dataset for both training and evaluation.
2 Evaluation Metrics
3 Main Results
We demonstrate our main results as Table 2. In the first block, we use BigBird (pre-trained with MLM) and FiD (initialized from T5 checkpoint) and only fine-tune them on our TimeQA training set without relying on external NQ/TriviaQA data. Since our dataset size is rather limited, the achieved performance is lower than 20%.
In the second block, we probe the performance of BigBird and FiD models fine-tuned on large-scale NQ/TriviaQA data. Since the TimeQA questions are linguistically simple and natural, and also coming from the Wikipedia domain, there exists very little distributional shift. However, these fine-tuned models are only achieving 33% under easy mode and 27% under the hard mode, much lower than their performance on NQ/TriviaQA (over 60%). Such a gap reflects concerning incompetence of these models to deal with temporal reasoning in text.
In the third block, we continue to fine-tune these pre-fine-tuned models on our TimeQA training set. This adaptive fine-tuning greatly enhances the models’ capability to perform temporal reasoning. The performance can be significantly boosted to 60% under easy mode and 45% under hard mode. However, the best-performing model is still far behind the human performance, especially under hard mode (87%), which indicates large headroom for future studies.
Throughout the experiments, we observe that the model is consistently getting much lower accuracy under hard mode. The significant 15% performance drop from easy to hard reflects the incompetence of the models to handle robust temporal reasoning. In contrast, humans are suffering only 2% drop under the hard mode, which indicates that humans are more robust in temporal reasoning.
4 Human-Paraphrased Results
We further provide experimental results on human-paraphrased questions in Table 3. We first evaluate the model fine-tuned on NQ to directly answer the human-paraphrased questions, we can observe a 2-4 points drop in EM score. Then we evaluate the model finetuned on NQ+TimeQA (only containing template questions). Surprisingly, the model without being trained on human-paraphrased questions can generalize very well to these human-paraphrases, suffering only 4-5% EM drop. After applying the 1K human-paraphrased training examples for adaptation, the gap is further decreased to only 2% EM score. The narrow performance gap between the realistic human-written questions and artificially synthesized questions indicates that our synthesized dataset is indeed an accurate proxy for estimating the model’s capability to solve real-world time-sensitive questions.
Since our questions are highly decontextualized, i.e. the question is non-ambiguous, leading to a unique answer in the world. We follow DrQA (Chen et al., 2017) to perform open-domain QA, where we first use BM25 retriever to retrieve the most relevant passage from whole Wikipedia dump and then run BigBird to extract the answer. The best retrieval accuracy HITS@1 under easy and hard are 28.8% and 26.8%. The best end-task QA accuracy is rather low, achieving roughly 14% under Test-Easy and 11% under Test-Hard with lot of headroom for future work.
5 Model Analysis
Here we further analyze the model performance from different angles. Specifically, we demonstrate the model’s score concerning different relations, and different document length as follows:
Here we are interested in understanding the models’ capability to handle the long-term dependency in temporal reasoning. We plot the model’s accuracy under different document lengths in Figure 6. As can be seen, the BigBird’s performance degrades rapidly as the length increases to over 5000 tokens, while the FiD’s performance is quite uniformly distributed across different document lengths. This figure demonstrates that the FiD’s better performance is partially attributed to its strong capability to deal with long-term dependency in temporal reasoning.
Here we are also interested in understanding the model’s performance over different relations and demonstrate our findings in Figure 7. We found that the model’s performance is orthogonal to the frequency of relations. For example, the relation ‘P1037 (director of)’ only has 30 instances, however, its performance is much higher than the relation ‘P39 (position held)’, which has over 500 training instances. It’s mainly due to the fact that the time boundary of relation ‘P1037’ is more likely to be explicitly mentioned than ‘P39 (position held)’.
We are interested in whether the best-performing models can make consistent predictions under random perturbation of [time specifier]. For example, if a model perceives that a fact persists between and , then the model should make consistent prediction for any question with [time specifier] falling within the range of and . Therefore, we select the correctly predicted examples and randomly perturb the [time specifier] for 3 times. We observe how many percentages of the model predictions will remain constant/true under these random perturbations. We observe only 66% of model predictions are agnostic to these perturbations. This finding suggests that the existing QA models are not quite consistent with respect to their predictions.
6 Error Analysis
To better investigate the errors made by the QA model, we perform a detailed analysis based on the predictions from the best-performing FiD model. Under the easy-dev set, we categorize questions based on whether the [time specifier] is explicitly mentioned in the given passage. We found that two-thirds of easy-questions have [time specifier] explicitly mentioned with the average performance over 64%, while the rest easy-questions only achieve 49%. Such comparison indicates that implicit time information is a major challenge to our QA model. We further categorize the errors mainly into sources: 1) temporal reasoning: for example, ‘… in May 2010, one year after that’ refers to ‘May 2011’ based on numeric addition, 2) commonsense reasoning: for example, ‘… in 2012 London Olympics, in the next Olympic game’ refers to ‘2016 Olympic’ based on our commonsense. 3) termination reasoning: the termination time of a fact is commonly unmentioned in a text corpus, it needs to be inferred based on the start time of the next event. For example, ‘XXX joined A team in 2017, … in 2019, B team signed a contract with XXX’, we know that B team and A team are mutually exclusive, there for the termination time of A team is in 2019.
These three cases are prevalent in our daily text, which poses great challenges for the existing models. To further boost the performance on our dataset, it is vital to consider better algorithms to inject the temporal commonsense knowledge into the models.
Related Work
Question Answering There have been numerous efforts to tackle the machine reading comprehension problem. Different datasets like DrQA (Chen et al., 2017), TriviaQA Joshi et al. (2017), SearchQA (Dunn et al., 2017) and DROP (Dua et al., 2019) have been proposed. As the SQuAD (Rajpurkar et al., 2016) questions are relatively simple because they usually require no more than one sentence in the paragraph to answer. The following datasets further challenge the QA model’s capability to handle different scenarios like open-domain, long context, multi-hop, discrete operations, etc. However, these QA datasets lack the existence the time-sensitive questions. Thus, we hope our effort could serve as a complement to the existing QA research. Temporal Reasoning over Knowledge Base Understanding time evolution is an important research topic. As our world is constantly changing, it’s vital to understand the time scope of world knowledge. There have been long-standing efforts to inject temporal quantifiers into the knowledge base (KB) using temporal knowledge extraction techniques (Talukdar et al., 2012b; Chang and Manning, 2012; Talukdar et al., 2012a; Wijaya et al., 2014). Adding the temporal information into KB can empower down-stream applications like KBQA (Ahn et al., 2006; Sanampudi and Guda, 2013; Jia et al., 2018a, b; Saxena et al., 2021) to handle time-sensitive queries. Such two-step approach might suffer from cascaded errors and lead to compromised accuracy. In contrast, TimeQA aims to directly answer time-sensitive queries based on unstructured text. The new problem is more realistic yet challenging due to the high variance of human expressions over time information. Temporal Reasoning over Text Recently, a contemporary dataset SituatedQA (Zhang and Choi, 2021) was also released targeting at answering open-domain time-sensitive QA. There are a few major differences: 1) SituatedQA contains more realistic queries selected from NQ dataset (Kwiatkowski et al., 2019) while TimeQA contains mostly synthesized queries (except human-paraphrased subset). 2) TimeQA contains 20K queries, while SitutatedQA only contains 4K temporal queries. 3) The hard version of TimeQA requires reasoning over implicit temporal mentions in passage, which is not emphasized in SituatedQA. Another related dataset was introduced by Dhingra et al. (2021) to diagnose whether the existing LMs are able to sensitive to changes in temporal knowledge. Their questions are mostly cloze-based probing questions under closed-book setting, while both TimeQA and SituatedQA are evaluated under open-book setting. Temporal Reasoning over Events Another popular domain for temporal reasoning is event-centric tasks, which aims at understanding the textual description of real-world events (Chen et al., 2021). There has been studies on extracting time boundaries for events (Ning et al., 2018; Wen et al., 2021; Zhang et al., 2021), reading comprehension over events (Ning et al., 2020). These event-centric NLP tasks focus on understanding the temporal relationship (e.g. ‘before’, ‘after’, ‘include’, etc) between multiple events (e.g. ‘snow melting’, ‘land slide’, etc), which have inherent logical relation like ‘causation’, ‘negation’, ‘inclusion’, etc. Compared to TORQUE (Ning et al., 2020), there are two major differences: 1) TimeQA requires numerical reasoning over time information while TORQUE does not require it, 2) TimeQA requires modeling long-term temporal dependency in the text, while TORQUE’s passages are mostly short text.
Conclusion
Though time-sensitive facts are pervasive in our daily text corpus, there has been little prior work exploring this direction. In this paper, we build the first dataset to investigate whether existing models can understand time-sensitive facts. Our experiments show that the SoTA models are still lagged behind humans in temporal reasoning. In order to empower the future NLP models to understand temporal information, different temporal-aware models need to be proposed. Finally, this paper opens up new research directions for better modeling temporal information in text representations.
References
Appendix A Appendix
We follow datasheets for datasets guideline to document the followings.
For what purpose was the dataset created? Was there a specific task in mind? Was there a specific gap that needed to be filled? TimeQA is created to test current models’ capability to perform diverse temporal reasoning under unstructured text corpus, which can help future NLP models to better capture the time dimension.
Who created the dataset (e.g., which team, research group) and on behalf of which entity (e.g., company, institution, organization)? UCSB NLP team, mostly Wenhu Chen
A.1.2 Composition
What do the instances that comprise the dataset represent (e.g., documents, photos, people, countries)? Are there multiple types of instances (e.g., movies, users, and ratings; people and interactions between them; nodes and edges)? TimeQA only contains documents (text) in the dataset.
How many instances are there in total (of each type, if appropriate)? There are roughly 20K question-answer pairs.
Does the dataset contain all possible instances or is it a sample (not necessarily random) of instances from a larger set? If the dataset is a sample, then what is the larger set? Is the sample representative of the larger set (e.g., geographic coverage)? If so, please describe how this representativeness was validated/verified. If it is not representative of the larger set, please describe why not (e.g., to cover a more diverse range of instances, because instances were withheld or unavailable). It’s sampled from large Wikipedia passages, it’s representative of all the possible temporal-sensitive information.
Are relationships between individual instances made explicitly (e.g., users’ movie ratings, social network links)? If so, please describe how these relationships are made explicit. N/A.
Are there recommended data splits (e.g., training, development/validation, testing)? If so, please provide a description of these splits, explaining the rationale behind them. Yes, we split training, development, and testing set. We split randomly within each data source.
Are there any errors, sources of noise, or redundancies in the dataset? If so, please provide a description. There could have some potential noise of question or answer annotation.
Is the dataset self-contained, or does it link to or otherwise rely on external resources (e.g., websites, tweets, other datasets)? If it links to or relies on external resources, a) are there guarantees that they will exist, and remain constant, over time; b) are there official archival versions of the complete dataset (i.e., including the external resources as they existed at the time the dataset was created); c) are there any restrictions] (e.g., licenses, fees) associated with any of the external resources that might apply to a future user? Please provide descriptions of all external resources and any restrictions associated with them, as well as links or other access points, as appropriate. TimeQA is self-contained.
Does the dataset contain data that might be considered confidential (e.g., data that is protected by legal privilege or by doctorpatient confidentiality, data that includes the content of individuals’ non-public communications)? If so, please provide a description. No, all the samples in TimeQA is public available.
Does the dataset contain data that, if viewed directly, might be offensive, insulting, threatening, or might otherwise cause anxiety? If so, please describe why. No
Does the dataset relate to people? If not, you may skip the remaining questions in this section. No
A.1.3 Uses
Has the dataset been used for any tasks already? If so, please provide a description? It is proposed to use for QA task.
Is there a repository that links to any or all papers or systems that use the dataset? If so, please provide a link or other access point. It is a new dataset. We run existing state-of-the-art models and release the code at https: https://github.com/wenhuchen/Time-Sensitive-QA
What (other) tasks could the dataset be used for? Many other tasks like relation extractions can be also used.
Is there anything about the composition of the dataset or the way it was collected and preprocessed/cleaned/labeled that might impact future uses? For example, is there anything that a future user might need to know to avoid uses that could result in unfair treatment of individuals or groups (e.g., stereotyping, quality of service issues) or other undesirable harms (e.g., financial harms, legal risks) If so, please provide a description. Is there anything a future user could do to mitigate these undesirable harms? N/A
Are there tasks for which the dataset should not be used? If so, please provide a description. N/A
A.1.4 Distribution
Will the dataset be distributed to third parties outside of the entity (e.g., company, institution, organization) on behalf of which the dataset was created? If so, please provide a description. No.
How will the dataset will be distributed (e.g., tarball on website, API, GitHub)? Does the dataset have a digital object identifier (DOI)? Release on Github. No DOI
When will the dataset be distributed? It is released in https://github.com/wenhuchen/Time-Sensitive-QA
Will the dataset be distributed under a copyright or other intellectual property (IP) license, and/or under applicable terms of use (ToU)? If so, please describe this license and/or ToU, and provide a link or other access point to, or otherwise reproduce, any relevant licensing terms or ToU, as well as any fees associated with these restrictions. BSD 3-Clause "New" or "Revised" License. https://github.com/wenhuchen/Time-Sensitive-QA/blob/main/LICENSE
Have any third parties imposed IP-based or other restrictions on the data associated with the instances? If so, please describe these restrictions, and provide a link or other access point to, or otherwise reproduce, any relevant licensing terms, as well as any fees associated with these restrictions. No.
Do any export controls or other regulatory restrictions apply to the dataset or to individual instances? If so, please describe these restrictions, and provide a link or other access point to, or otherwise reproduce, any supporting documentation. No.
A.1.5 Accessibility
Links to access the dataset and its metadata: the github repository https://github.com/wenhuchen/Time-Sensitive-QA.
The data is saved in a json format, where an example is shown in the README.md file.
UCSB NLP group will maintain this dataset on the Github account.
BSD 3-Clause "New" or "Revised" License https://github.com/wenhuchen/Time-Sensitive-QA/blob/main/LICENSE
A.2 Evaluation Metrics
The F1 score is calculated with the following equation to cover the ‘[unanswerable]’ case.
A.3 Annotation Interface
The annotation interface is demonstrated in Figure 8, the original HIT job url is https://s3.amazonaws.com/mturk_bulk/hits/467870457/JV_fEwGzVy_PgHqtWqWXLA.html.
A.4 Relation Distribution over TimeQA
The annotated time-evolving facts include more than 70 relations, which follows a long-tail distribution. The major relations are shown in Figure 9, where the dominant relations are ‘P54’ (play for), ‘P39’ (position held), ‘P108’ (employer of), ‘P69’ (educated at).
A.5 Answerable vs. Unanswerable
We also provide break-down analysis of model performance for answerable and unanswerable questions in Figure 10. As can be seen, the FiD is more aware of the answerability on TimeQA. On the easy and hard mode, FiD’s accuracy on unanswerable is above BigBird’s by 14% and 18%, while the gap in answerable questions are less significant.