A First Instagram Dataset on COVID-19

Koosha Zarei, Reza Farahbakhsh, Noel Crespi, Gareth Tyson

I Introduction

The novel coronavirus (COVID-19) was declared a pandemic by the World Health Organisation (WHO) on 11 March 2020.https://tinyurl.com/WHOPandemicAnnouncement Since then the world has experienced almost 3 million cases. To mitigate its spread, many government have therefore imposed unprecedented social distancing measures that have led to millions become housebound. This has resulted in a flurry of research activity surrounding both understanding and countering the outbreak .

As part of this, social media has become a vital tool in disseminating public health information and maintaining connectivity amongst people. Several recent studies have relied on Twitter data to better understand this . These have primarily focused on health related (mis)information, but there have also been studies into online hate . Despite this, there has been only limited exploration of other social modalities, such as image content.

We argue this represents a limitation, particularly considering the importance of image-based content in the dissemination of news (and misinformation) . This paper introduces a COVID-19 Instagram dataset, which we make available for the research community. We have gathered data between January 5 and March 30 2020 (§III). The dataset covers 18.5K comments and 329K likes from 5.3K posts. These posts have been distributed by 2.5K publishers. The data predominantly covers English language posts, and we provide a number of important features covering both the content and the publisher (§IV). We hope that this dataset can help support a number of use cases. Hence, we conclude the paper by highlighting a number of potential uses related to COVID-19 social media analysis (§V). Details of how to access the data is presented in §VI.

II Related Work

Most related to this work is the set of COVID-19 social media datasets recently released. To date, this predominantly covered textual data (e.g. Twitter). To assist in this, Kazemi et al. provide a toolbox for processing textual data related to COVID-19. In terms of data, the first efforts in this direction was from authors in which provide a large Twitter dataset related to Coronavirus (by crawling major hashtags and trusted accounts). Another similar study , provides an arabic Twitter dataset with a similar data collection methodology. Lopez et al. provide another Twitter dataset including the geolocated tweets. There are some further efforts on providing similar datasets from twitter . Sharma et al. also made a public dashboardhttps://usc-melady.github.io/COVID-19-Tweet-Analysis/ available summarising data across more than 5 million real-time tweets.

These Twitter datasets are being used for various use cases. For example, Saire and Navarro use the data to show the epidemiological impact of COVID-19 on press publications. Singh et al. are also monitoring the flow of (mis)information flow across 2.7M tweets, and correlating it with infection rates to find that misinformation and myths are discussed, but at lower volume than other conversations. To the best of our knowledge, the only paper that has covered Instagram is by Cinelli et al. , who analyse Twitter, Instagram, YouTube, Reddit and Gab data about COVID-19. We complement this by making a public Instagram dataset available to the community. We redirect readers to for a comprehensive survey of ongoing data science research related to COVID-19.

III Data Collection

We have collected public posts from Instagram by crawling all posts associated with a set of COVID-19 hashtags presented in Table I.

Methodology. To be able to collect Instagram public content (in the shape of post), we use the official Instagram APIs . In particular, to get posts that are tagged with specific hashtags, the Instagram Hashtag Engine is used . This API returns public photos and videos that have been tagged with particular hashtags. MongoDB is used as the core database and data is stored as JSON records. The crawler is responsible for gathering both posts and reactions. A reaction can be active (comment) or passive (like). As it is infeasible to collect all reactions, in this dataset, we define a limitat of 500 comments and 500 likes per post. Our crawler is running on several virtual machines in parallel 24/7. Note that we do not manually filter any posts and therefore we gather all posts containing the hashtags, regardless of the specific topics discussed within. The complete architecture of our crawler is described in this paper .

Release v1.0 (April 20, 2020). The first version of this data collection process started on January 5, 2020 and continued until March 30, 2020. The data gathering is still running as the lockdown has not been finished in many countries around the world (at the time of writing this paper). During this time 18.5K comments and 329K likes from 5.3K public posts have been collected. These posts are distributed by 2.5K publishers.

Ethics. In line with Instagram policies as well as user privacy, we only gather publicly available data that is obtainable from Instagram.

IV Dataset Description

To provide context for potential users of our dataset, we next brifely summarise the dataset and describe the characteristics of the content.

Hashtags. Recall that we gather the data by querying certain hashtags. Figure 2 presents the top hashtags tagged within the posts. Figure 1 also presents a wordcloud of the hashtags in our dataset. This naturally includes hashtags outside of our seed set used for crawling. There are intuitive examples, such as corona, covid19, covid_19, stayathome, quarantine, love, covid, virus, and instagram. The are therefore the most repeated hashtags that appear with #coronavirus. Note that this means will might miss posts that mention these concepts in other languages.

Post Language. In order to identify the language of the post, we use spaCy library and we apply it on the text of the caption. The language distribution is displayed in Table II. The dataset is dominated by English language content, making up almost 60% of posts. This is driven by our choice of an English-language hashtag seed set used for data collection. That said, we have broad coverage of other widely spoken languages too, e.g. Spanish (9.9%). Notice that there is no official metric to determine the post language. Therefore, we highlight that this analysis could mis-classify certain posts, such as those solely hashtags or emojis.

Features. To keep data organized, the dataset is divided into four parts: (i) post content, (ii) publisher information, (iii) comment metrics, and (iv) like features (Table III). Posts contain key attributes such as a caption, list of hashtags, image/video, number of likes, number of comments, location, date, tagged list, etc. A post is published by a public account (or public Instagram page) and in our dataset, it can be individual, fan page, news agency, influencer, blogger, etc. Each post receives reactions in the form of comment and like that are issued by the audience/followers. The full feature list is presented in Table III. Furthermore, Table IV presents Post and Profile characteristics in detail.

V Potential Research Topics

We hope that the dataset can support diverse research activities. Below we list a subset of potential topics, we believe the dataset could support:

Fake news, misinformation and rumors spreading:. Several researcher have started to inspect COVID-19 misinformation. As an example, an infodemic observatory have analyzed more than 100M public messages to understand the digital response in online social media to COVID-19 outbreak.https://covid19obs.fbk.eu In another study, Sharma et al. made a public dashboard available summarising data from real-time tweets in in https://usc-melady.github.io/COVID-19-Tweet-Analysis/ with a focus to misinformation spread analysis. We believe that our Instagram data could be used to evaluate the flow of misinformation (e.g. memes) on Instagram.

Bot Population and bot generated content: It is well known that bot content plays a prominent role in social media data . These have the capacity of amplify misinformation or even act against public health policies (e.g. encourgaging a breakdown in social distancing). We posit that the data could be used to explore the role of bots in this dissemination.

Behavioral change analysis during the pandemic: The social distancing measures are created an unprecedented change to millions of people’s lives. Understanding the behavioral consequences of this is vital for understanding things like adherence to social distancing policies and mental health consequences.

Information sharing related Covid-19: Information flow is vital during periods of emergency. We posit that the dataset can be used to understand the flow of information, as well as people’s reactions to such information.

VI Dataset Access

The presented dataset is accessible in this address on Github platform: https://github.com/kooshazarei/COVID-19-InstaPostIDs. This is the first version of the dataset and we are still collecting data. Hence, we hope to make further versions available in the coming weeks and months. We publish this dataset in agreement with Instagram’s Terms & Conditions , and as it is not possible to release the post content and reactions, we just distribute the post ID’s. These are known as the shortcodes. Researchers can then simply retrieve post content through IDs by the help of some open-source projects such as Instaloader that have been developed for such purposes. For any further question, please contact Koosha Zarei at koosha.zarei@telecom-sudparis.eu.

References