What Are People Asking About COVID-19? A Question Classification Dataset
Jerry Wei, Chengyu Huang, Soroush Vosoughi, Jason Wei
Introduction
A major challenge during fast-developing pandemics such as COVID-19 is keeping people updated with the latest and most relevant information. Since the beginning of COVID, several websites have created frequently asked questions (FAQ) pages that they regularly update. But even so, users might struggle to find their questions on FAQ pages, and many questions remain unanswered. In this paper, we ask—what are people really asking about COVID, and how can we use NLP to better understand questions and retrieve relevant content?
We present Covid-Q, a dataset of 1,690 questions about COVID from 13 online sources. We annotate Covid-Q by classifying questions into 15 general question categoriesWe do not count the “other” category. (see Figure 1) and by grouping questions into question clusters, for which all questions in a cluster ask the same thing and can be answered by the same answer, for a total of 207 clusters. Throughout 2, we analyze the distribution of Covid-Q in terms of question category, cluster, and source.
Covid-Q facilitates several question understanding tasks. First, the question categories can be used for a vanilla text classification task to determine the general category of information a question is asking about. Second, the question clusters can be used for retrieval question answering (since the cluster annotations indicate questions of same intent), where given a new question, a system aims to find a question in an existing database that asks the same thing and returns the corresponding answer Romeo et al. (2016); Sakata et al. (2019). We provide baselines for these two tasks in 3.1 and 3.2. In addition to directly aiding the development of potential applied systems, Covid-Q could also serve as a domain-specific resource for evaluating NLP models trained on COVID data.
Dataset Collection and Annotation
Data collection. In May 2020, we scraped questions about COVID from thirteen sources: seven official FAQ websites from recognized organizations such as the Center for Disease Control (CDC) and the Food and Drug Administration (FDA), and six crowd-based sources such as Quora and Yahoo Answers. Table 1 shows the distribution of collected questions from each source. We also post the original scraped websites for each source.
Data cleaning. We performed several pre-processing steps to remove unrelated, low-quality, and nonsensical questions. First, we deleted questions unrelated to COVID and vague questions with too many interpretations (e.g., “Why COVID?”). Second, we removed location-specific and time-specific versions of questions (e.g., “COVID deaths in New York”), since these questions do not contribute linguistic novelty (you could replace “New York” with any state, for example). Questions that only targeted one location or time, however, were not removed—for instance, “Was China responsible for COVID?” was not removed because no questions asked about any other country being responsible for the pandemic.
Finally, to minimize occurrences of questions that trivially differ, we removed all punctuation and replaced synonymous ways of saying COVID, such as “coronavirus,” and “COVID-19” with “covid.” Table 1 also shows the number of removed questions for each source.
Data annotation. We first annotated our dataset by grouping questions that asked the same thing together into question clusters. The first author manually compared each question with existing clusters and questions, using the definition that two questions belong in the same cluster if they have the same answer. In other words, two questions matched to the same question cluster if and only if they could be answered with a common answer. As every new example in our dataset is checked against all existing question clusters, including clusters with only one question, the time complexity for annotating our dataset is , where is the number of questions.
After all questions were grouped into question clusters, the first author gave each question cluster with at least two questions a name summarizing the questions in that cluster, and each question cluster was assigned to one of 15 question categories (as shown in Figure 1), which were conceived during a thorough discussion with the last author. In Table 2, we show the question clusters with the most questions, along with their assigned question categories and some example questions. Figure 2 shows the distribution of question clusters.
Annotation quality. We ran the dataset through multiple annotators to improve the quality of our annotations. First, the last author confirmed all clusters in the dataset, highlighting any questions that might need to be relabeled and discussing them with the first author. Of the 1,245 questions belonging to question clusters with at least two questions, 131 questions were highlighted and 67 labels were modified. For a second pass, an external annotator similarly read through the question cluster labels, for which 31 questions were highlighted and 15 labels were modified. Most modifications involved separating a single question cluster that was too broad into several more specific clusters.
For another round of validation, we showed three questions from each of the 89 question clusters with to three Mechanical Turk workers, who were asked to select the correct question cluster from five choices. The majority vote from the three workers agreed with our ground-truth question-cluster labels 93.3% of the time. The three workers unanimously agreed on 58.1% of the questions, for which 99.4% of these unanimous labels agreed with our ground-truth label. Workers were paid \0.07$ per question.
Finally, it is possible that some questions could fit in several categories—of 207 clusters, 40 arguably mapped to two or more categories, most frequently the transmission and prevention categories. As this annotation involves some degree of subjectivity, we post formal definitions of each question category with our dataset to make these distinctions more transparent.
Single-question clusters. Interestingly, we observe that for the CDC and FDA frequently asked questions websites, a sizable fraction of questions (44.6% for CDC and 42.1% for FDA) did not ask the same thing as questions from any other source (and therefore formed single-question clusters), suggesting that these sources might want adjust the questions on their websites to question clusters that were seen frequently in search engines such as Google or Bing. Moreover, 54.2% of question clusters that had questions from at least two non-official sources went unanswered by an official source. In the Supplementary Materials, Table 7 shows examples of these questions, and conversely, Table 8 shows CDC and FDA questions that did not belong to the same cluster as any other question.
Question Understanding Tasks
We provide baselines for two tasks: question-category classification, where each question belongs to one of 15 categories, and question clustering, where questions asking the same thing belong to the same cluster.
As our dataset is small when split into training and test sets, we manually generate an additional author-generated evaluation set of questions. For these questions, the first author wrote new questions for question clusters with 4 or 5 questions per cluster until those clusters had 6 questions. These questions were checked in the same fashion as the real questions. For clarity, we only refer to them in 3.1 unless explicitly stated.
The question-category classification task assigns each question to one of 15 categories shown in Figure 1. For the train-test split, we randomly choose 20 questions per category for training (as the smallest category has 26 questions), with the remaining questions going into the test set (see Table 3).
We run simple BERT Devlin et al. (2019) feature-extraction baselines with question representations obtained by average-pooling. For this task, we use two models: (1) SVM and (2) cosine-similarity based -nearest neighbor classification (-NN) with . As shown in Table 4, the SVM marginally outperforms -NN on both the real and generated evaluation sets. Since our dataset is small, we also include results from using data augmentation Wei and Zou (2019). Figure 4 (Supplementary Materials) shows the confusion matrix for BERT-feat: SVM + augmentation for this task.
2 Question Clustering
Of a more granular nature, the question clustering task asks, given a database of known questions, whether a new question asks the same thing as an existing question in the database or whether it is a novel question. To simulate a potential applied setting as much as possible, we use all questions clusters in our dataset, including clusters containing only a single question. As shown in Table 5, we make a 70%–30% train–test split by class.For clusters with two questions, one question went into the training set and one into the test set. 70% of single-question clusters went into the training set and 30% into the test set.
In addition to the -NN baseline from 3.1, we also evaluate a simple model that uses a triplet loss function to train a two-layer neural net on BERT features, a method introduced for facial recognition Schroff et al. (2015) and now used in NLP for few-shot learning Yu et al. (2018) and answer selection Kumar et al. (2019).
For evaluation, we compute a single accuracy metric that requires a question to be either correctly matched to a cluster in the database or to be correctly identified as a novel question. Our baseline models use thresholding to determine whether questions were in the database or novel. Table 6 shows the accuracy from the best threshold for both these models, and Supplementary Figure 3 shows their accuracies for different thresholds.
Discussion
Use cases. We imagine several use cases for Covid-q. Our question clusters could help train and evaluate retrieval-QA systems, such as covid.deepset.ai or covid19.dialogue.co, which, given a new question, aim to retrieve the corresponding QA pair in an existing database. Another relevant context is query understanding, as clusters identify queries of the same intent, and categories identify queries asking about the same topic. Finally, Covid-q could be used broadly to evaluate COVID-specific models—our baseline (Huggingface’s bert-base-uncased) does not even have COVID in the vocabulary, and so we suspect that models pre-trained on scientific or COVID-specific data will outperform our baseline. More related areas include COVID-related query expansion, suggestion, and rewriting.
Limitations. Our dataset was collected in May 2020, and we see it as a snapshot in time of questions asked up until then. As the COVID situation further develops, a host of new questions will arise, and the content of these new questions will potentially not be covered by any existing clusters in our dataset. The question categories, on the other hand, are more likely to remain static (i.e., new questions would likely map to an existing category), but the current way that we came up with the categories might be considered subjective—we leave that determination to the reader (refer to Table 9 or the raw dataset on Github). Finally, although the distribution of questions per cluster is highly skewed (Figure 2), we still provide them at least as a reference for applied scenarios where it would be useful to know the number of queries asking the same thing (and perhaps how many answers are needed to answer the majority of questions asked).
References
Supplementary Materials
For the question clustering task, our models used simple thresholding to determine whether a question matched an existing cluster in the database or was novel. That is, if the similarity between a question and its most similar question in the database was lower than some threshold, then the model predicted that it was a novel question. Figure 3 shows the accuracy of the -NN and triplet loss models at different thresholds.
2 Question-Category Classification Error Analysis
Figure 4 shows the confusion matrix for our SVM classifier on the question-category classification task on the test set of real questions. Categories that were challenging to distinguish were Transmission and Having COVID (34% error rate), and Having COVID and Symptoms (33% error rate).
3 Further Dataset Details
Question mismatches. Table 7 shows example questions from at least two non-official sources that went unanswered by an official source. Table 8 shows example questions from the FDA and CDC FAQ websites that did not ask the same thing as any other questions in our dataset.
Example questions. Table 9 shows example questions from each of the 15 question categories.
Corresponding answers. The FAQ websites from reputable sources (denoted with ∗ in Table 1) provide answers to their questions, and so we also provide them as an auxiliary resource. Using these answers, 23.8% of question clusters have at least one corresponding answer. We caution against using these answers in applied settings, however, because information on COVID changes rapidly.
Additional data collection details. In terms of how questions about COVID were determined, for FAQ websites from official organizations, we considered all questions, and for Google, Bing, Yahoo, and Quora, we searched the keywords “COVID” and “coronavirus.”
As for synonymous ways of saying COVID, we considered “SARS-COV-2,” “coronavirus,” “2019-nCOV,” “COVID-19,” and “COVID19.”
Other COVID-19 datasets. We encourage researchers to also explore other COVID-19 datasets: tweets streamed since January 22 Chen et al. (2020), location-tagged tweets in 65 languages Abdul-Mageed et al. (2020), tweets of COVID symptoms Sarker et al. (2020), a multi-lingual Twitter and Weibo dataset Gao et al. (2020), an Instagram dataset Zarei et al. (2020), emotional responses to COVID Kleinberg et al. (2020), and annotated research abstracts Huang et al. (2020).