Content Moderation on Social Media in the EU: Insights From the DSA Transparency Database

Chiara Drolsbach, Nicolas Pröllochs

Introduction

Social media platforms (e. g., Facebook, X, YouTube) have become primary gateways to information and other content in the digital age (Bakshy et al., 2015). While these platforms offer various social and economic benefits (Utz, 2016; Nisar et al., 2019), they also facilitate the dissemination of harmful (or even illegal) content, including, but not limited to, hate speech (Solovev and Pröllochs, 2022, 2023; Mathew et al., 2019), disinformation (Feuerriegel et al., 2023; Bennett and Livingston, 2020; Geissler et al., 2023), and calls for violence (Bär et al., 2023; Jakubik et al., 2023; Bär et al., 2023). Concerns about harmful content on social media have been rising in recent years, particularly given its potential impact on elections (Allcott and Gentzkow, 2017; Aral and Eckles, 2019; Bakshy et al., 2015; Grinberg et al., 2019; Guess et al., 2020; Moore et al., 2023), public health (Broniatowski et al., 2018; Rocha et al., 2021; Gallotti et al., 2020; Roozenbeek et al., 2020), and public safety (Müller and Schwarz, 2021; Bär et al., 2023; Oh et al., 2013; Geissler et al., 2023). In response to such threats, many platforms have developed more or less sophisticated systems for content moderation, i. e., mechanisms that aim to prevent harm by removing or reducing the visibility of rule-breaking content (Grimmelmann, 2015). However, their degree of strictness varies greatly from platform to platform, and each platform keeps the specifics of how it enacts its moderation decisions largely opaque (Jhaver et al., 2019). Additionally, questions of who decides what is allowed on social media platforms and how providers decide to publish or remove third-party content have become pressing issues in public debate (Goldman, 2021). Over the last years, there have been increasing calls for legislation to revise the present model of self‐regulation, where primarily social media platforms define the rules and procedures of online content moderation (Schlag, 2023; Turillazzi et al., 2023).

As a remedy, the European Union (EU) and its member states have recently introduced and ratified an increasing number of laws and policies aimed at governing online content. The legislation should facilitate greater democratic control and oversight over systemic platforms, including social media (Schlag, 2023; Turillazzi et al., 2023). A key component of the European Commission’s strategy is the Digital Services Act (DSA) (The European Parliament and the Council of the European Union, 2022), which has established a new set of obligations for social media providers to build a safer and more trustworthy digital space (Schlag, 2023; Cauffman and Goanta, 2021; Leerssen, 2023; Turillazzi et al., 2023). Their primary focus is on strengthening the accountability of the platforms when it comes to illegal content uploaded by users. Specifically, major social media providers such as Facebook and X/Twitter can now be held responsible for the risks illegal content on their platforms poses to society. Furthermore, the DSA should establish a harmonized legal framework that avoids inconsistencies and uncoordinated procedures in content moderation amongst platforms. For this purpose, providers of large social media platforms are required to file detailed Statements of Reasons (SoRs) explaining why content was moderated, by reference to the specific legal provision infringed. To ensure scrutiny of content moderation decisions and transparency for both platforms and users, all SoRs are made publicly available by the EU via the DSA Transparency Database (European Commission, 2023).

Research goal: In this work, we provide a holistic early look at the DSA Transparency database. Due to this unique data source, we are, for the first time, able to empirically analyze real-world content moderation decisions of major social media platforms in the EU. Specifically, we address the following research questions:

(RQ1) How frequently is social media content subject to content moderation in the EU?

(RQ2) How often are different types of content (text, images, video, etc.) moderated on social media?

(RQ3) What are specific reasons (i. e., legal grounds) due to which social media providers moderate content?

(RQ4) To what extent are content moderation decisions on social media automated?

(RQ5) What types of content moderation actions do social media platforms implement?

Data & Methods: To address our research questions, we collected all Statements of Reasons (SoRs) that were submitted by major social media platforms within the first two months of the DSA Transparency database in September 25, 2023, from the EU’s official website. During our observation period, more than 156 million SoRs were transmitted by social media platforms, each representing one content moderation action on one of the platforms. We then extract a wide variety of variables (e. g., content types, legal grounds, etc.) from the SoRs in order to empirically analyze content moderation decisions on social media in the EU. Additionally, we implement regression analysis to estimate how the likelihood of automation of content moderation decisions varies across different content types and legal grounds. This allows us to characterize content moderation decisions that are more likely to be performed without human intervention by social media providers.

Contributions: To the best of our knowledge, this study is the first to present an empirical analysis of the Statements of Reasons in the EU’s DSA Transparency Database. Our empirical analysis yields the following main findings: (i) There are vast differences in the frequency of content moderation across platforms. For instance, TikTok performs more than 350 times more content moderation decisions per user than X/Twitter. (ii) Content moderation is most commonly applied for text and videos, whereas images and other content formats undergo moderation less frequently. (ii) The primary reasons for moderation include content falling outside the platform’s scope of service, illegal/harmful speech, and pornography/sexualized content, with moderation of misinformation being relatively uncommon. (iii) The majority of rule-breaking content is detected and decided upon via automated means rather than manual intervention. However, X/Twitter reports that it relies solely on non-automated methods. (iv) There is significant variation in the content moderation actions taken across platforms. While most platforms commonly remove rule-breaking posts, others are more likely to opt for reducing their visibility. Altogether, our study implies inconsistencies in how social media platforms implement their obligations under the DSA – resulting in a fragmented outcome that the DSA is meant to avoid. Our findings have important implications for regulators, who might be inclined to lay out more specific rules that ensure common standards on how social media providers handle rule-breaking content on their platforms.

Background

Content moderation describes mechanisms that are designed to prevent the dissemination of illegal and undesirable content in online communities (Grimmelmann, 2015; Roberts, 2020). In the context of social media, providers can choose from a wide catalogue of possible measures to prevent harm resulting from rule-breaking content on their platforms, such as, for example, content removal, visibility reduction (demotion), labeling, or account suspensions/terminations (Jiang et al., 2023). Effective content moderation mechanisms are often considered to be essential to the functioning of social networks (Gillespie, 2018). For instance, moderation interventions may increase compliance with community guidelines (Horta Ribeiro et al., 2023) and reduce uncivil behaviour (Katsaros et al., 2022). However, recent studies also indicate that content moderation efforts can backfire and increase rule breaking (e. g., because sanctioned users perceive the decision as unfair (Chang and Danescu-Niculescu-Mizil, 2019)), or lead to increased production of harmful content on other platforms (Mitts et al., 2022; Ali et al., 2021; Russo et al., 2023).

Over the last couple of years, content moderation systems on mainstream platforms have become increasingly sophisticated. Historically, the moderation of content on social media was overseen by relatively small review teams and platform rules were limited in their scope (Klonick, 2017). However, over time, as the demand from the public to address and remove harmful content grew, these practices evolved and platforms have established increasingly sophisticated systems to aid their content moderation efforts (Meta, 2023a; Youtube, 2023; X, 2023a; TikTok, 2023a). Many social media providers now enforce platform guidelines using automated content moderation systems that detect and intervene when rule-breaking occurs. Several platforms employ automated filters aimed at removing blatantly rule-breaking content (e. g., child sexual abuse material) (Gillespie, 2018, 2020). Moreover, platforms deploy moderators, who are either compensated (Roberts, 2020) or volunteer (Matias, 2016), to regulate the content that remains after automated filtering. Notably, each social media platform has developed various systems to implement these processes (Gillespie, 2018), yet each platform keeps the specifics of how it enacts its moderation decisions opaque (Jhaver et al., 2019).

2. Digital Services Act

The Digital Services Act (DSA) represents a significant legislative framework developed by the European Union (EU) with the primary goal of modernizing the digital space, ensuring safer and more open digital platforms for all users (Schlag, 2023; Cauffman and Goanta, 2021; Leerssen, 2023; Turillazzi et al., 2023). At its core, the DSA aims to address the challenges brought by the rapid evolution and influence of digital services, particularly large online platforms. The DSA was adopted by the European Parliament in July 2022 and entered into force on November 16, 2022 (The European Parliament and the Council of the European Union, 2022). It includes various requirements that providers of online platforms must fulfill. Each provider had to report their user numbers by February 17, 2023. These figures were used to identify all Very Large Online Platforms (VLOPs) and Search Engines (VLOEs) with more than 45 million users in the EU (corresponding to 10% of the population). Starting on August 25, 2023, all VLOEs and VLOPs must fulfill all obligations contained in the DSA. For all other (smaller) platforms, the rules apply from February 17, 2024. This includes publishing a transparency report every six months in which platforms must describe their own content moderation measures, as well as other relevant information such as, for example, the number of reports they receive from users, the error rate of automated content moderation systems and the composition of their content moderation teams (qualifications and linguistic expertise).

Furthermore, as mandated by Article 17 of the DSA, all VLOPs and VLOEs are required to provide detailed Statements of Reasons (SoRs) for any content moderation activity, including content removal, reach restriction, or account suspension/termination. The intention is to inform users about content moderation decisions and explain the reasons for the respective decisions. In accordance with Article 24(5) of the DSA, all submitted statements are collected and made publicly available on the DSA Transparency Database, which is managed by the Directorate-General for Communications Networks, Content and Technology of the European Commission. SoRs shall be clear and specific and easily comprehensible and as precise and specific as reasonably possible under the given circumstances. To ensure this, they must contain information about the content and the implemented type of moderation, whether automated systems were used, and whether the subject contains illegal content or is incompatible with the terms and conditions of the provider.

Data

We downloaded all Statements of Reasons (SoRs) that were submitted between the introduction of the database on September 25, 2023, and November 25, 2023, from the website of the DSA Transparency Database, i. e. within an observation period of two months. In total, more than 550 million SoRs were transmitted during the observation period. SoRs are submitted by all VLOPs and VLOEs. As we focus on content moderation on social media, we only included SoRs submitted by large social media platforms, namely Facebook, Instagram, YouTube, TikTok, Snapchat, X, LinkedIn and Pinterest. Furthermore, we excluded SoRs for content published before August 25, 2023, as the obligations associated with the DSA framework became effective from this date. The resulting dataset contains more than 156 Mio. SoRs, each including information on one content moderation action.

2. Key Variables

We now present the key variables that we extracted from the SoRs provided in the DSA Transparency Database. All of these variables will be empirically analyzed in the next sections:

Platform Name: The platform that submitted the SoR (Facebook, Instagram, Youtube, TikTok, Snapchat, X, or Pinterest).

Content Type: The type of content that is moderated (Audio, Product, Synthetic Media, Image, Video, Text, or Other).

Category: A categorical variable to specify the type of illegality/incompatibility because of which the content/account was moderated (Violence, Unsafe & Illegal Products, Self Harm, Scope of Platform Service, Scams & Fraud, Risk for Public Security, Protection of Minors, Pornography/Sexualized Content, Non Consensual Behavior, Mis-/Disinformation, Intellectual Property Infringements, Illegal/ Harmful Speech, Data Protection/Privacy Violations or Animal Welfare).

Decision Ground: Whether the content is classified as Illegal or Incompatible.

Automated Decision: Whether the content moderation decision was performed Not Automated, Partially Automated, or Fully Automated.

Automated Detection: Whether automated means were used to identify the content addressed by the decision (Yes or No).

Decision Type: The content moderation action implemented by the platform (Account Suspended, Account Terminated, Content Age Restricted, Content Demoted, Content Disabled, Content Labelled, Content Removed, Content Demoted/Removed or Other).

Empirical Analysis

We start by analyzing how many content moderation actions were submitted by each social media platform within the EU (see Fig. 1). The largest number of SoRs (#SoR) was submitted by TikTok (100.15 Mio; 64.09%), followed by Facebook (33.70 Mio; 21.56%), Pinterest (12.45 Mio; 7.97%), YouTube (5.12 Mio; 3.28%), and Instagram (3.95 Mio; 2.53%). Snapchat (0.61 Mio; 0.39%), X/Twitter (0.27 Mio; 0.17%), and LinkedIn (0.03 Mio; 0.02%) submitted less than one million SoRs during our observation period. It is striking that TikTok moderates substantially more content than all other platforms, some of which are much more relevant in the European market in terms of user numbers (e. g. YouTube, Facebook, and Instagram).

Fig. 1 shows how many content moderation decisions were made by platforms in relation to their size. For this, we calculated the ratio of content moderation actions (SoRs) relative to the monthly active users (MAU) per platform within the EU (#SoRs per MAU).We collected information on the number of monthly active users (MAU) within the EU from the platforms’ Transparency Reports (TikTok, 2023b; Meta, 2023b, c; LinkedIn, 2023; Snapchat, 2023; Youtube, 2023; X, 2023b). According to these numbers, YouTube is the largest platform with more than 400 million MAUs, for Instagram and Facebook Meta reported 259 million MAUs each, TikTok has 150 million MAUs, Pinterest around 150 million MAUs, X/Twitter around 100 million, Snapchat just under 97 million MAUs, and LinkedIn reported around 45 million MAUs. It is evident that TikTok is by far the most active in terms of content moderation, even when considering differences in platform sizes. TikTok submitted over 0.67 SoRs per MAU, surpassing Facebook (0.13) and Pinterest (0.10). The remaining platforms had comparatively few SoRs per MAU (each less than 0.02) For example, the rate at which TikTok carried out content moderation decisions per MAU was more than 350 times that of X/Twitter.

2. Content Types (RQ2)

Next, we study how often different types of content (i. e., text, images, videos, etc.) were moderated on social media platforms in the EU. To assign the SoRs to individual content types, we used the categories specified by the platforms when submitting the SoRs (Content Type). As shown in Fig. 2, the most frequently moderated content is Text (38.74%), followed by Videos (24.42%), Images (6.57%), and Synthetic Media (38.74%). Approximately 27% of all content moderation actions were assigned to the category Other, i. e., content types that are not predefined by the DSA. The majority of this content concerns violations that affect entire accounts/profiles, platform-specific content types (e. g. pins or boards on Pinterest), and (job) advertisements (in particular on Youtube, Pinterest, and LinkedIn).

To analyze whether the distribution of moderated content types is platform-specific, Fig. 3 visualizes the distribution of the seven predefined content categories per platform. We observe that TikTok primarily focuses on Text (57.14%) and Videos (33.97%), whereas moderation of Images (7.78%) is relatively rare. In contrast, Snapchat rarely moderates Text (4.21%) but is relatively more likely to moderate Videos (63.92%) and Images (16.96%). The platform X/Twitter seems to limit its content moderation activities mainly to content categorized as Synthetic Media (99.83%). Note, however, that content assigned to this category may include artificially created content in the form of, for example, text, images, video and/or audio content. The remaining platforms each categorized more than 50% of all content moderation actions as content type Other.

Altogether, we find that the type of content subject to moderation is, in many cases, strongly related to the content that is predominantly published on the respective platform (e. g., Video and Image on Snapchat, Video and Advertisement on YouTube, Pins on Pinterest, job-related content on LinkedIn). It is also worth noting that the content type is specified by the platform (i. e., self-reported), and each content piece is assigned to a single type. However, on social media, posts can consist of a mixture of different content types (e. g., video/image and text). In a similar vein, AI-generated content may be classified as image/text/video or as synthetic media. Overall, it seems likely that the content type reported by platforms often describes only one dimension of a social media post.

3. Reasons for Moderation (RQ3)

We now analyze specific reasons (i. e., legal grounds) due to which social media providers moderate content. When submitting a SoR, the platforms must assign the content to one of 14 given categories to specify the type of illegality/incompatibility because of which the content/account was moderated. Assigning a second category is possible, but not frequently used by the platforms (less than 0.001%). Overall, the most frequent reasons for moderation is content falling outside of the Scope of Platform Service (49.06%), Illegal/Harmful Speech (28.50%), Pornography/Sexualized Content (8.77%), and Violence (6.70%). Content that is classified in the remaining categories (Data Protection/Privacy Violations, Protection of Minors, Scams & Fraud, Intellectual Property Infringements, Self Harm, Mis-/Disinformation, Unsafe & Illegal Products, Non Consensual Behavior, Animal Welfare, and Risk for Public Security) is only relatively rarely subject to moderation (6.97% in total).

Fig. 4 illustrates the distribution of the different categories per platform. We find that there is a relatively highly similarity across most platforms. With the exceptions of X/Twitter and Pinterest, all platforms moderated a large proportion of content that does not correspond to the Scope of Platform Service, Illegal/Harmful Speech, and Violence. In contrast, Pinterest focused primarily on Pornography/Sexualized Content (80.95%). In the case of X/Twitter, the vast majority of content moderation actions were attributed to Violence (41.18%) or Pornography/Sexualized Content (44.37%). Snapchat is the platform that moderated content in the widest variety of categories (e. g. 16.69% in Scams & Fraud and 10.77% Unsafe & Illegal Products).

In addition to the category, platforms must state whether the moderation decision was taken in line with article 17(3)(d) DSA, meaning the content is considered as illegal, or in line with article 17(3)(e), meaning the content is considered incompatible with the service’s terms and conditions. Across all platforms, most content moderation actions were performed because of content being incompatible (99.80%) rather than illegal (0.20%). When comparing the individual platforms, however, it is striking that all platforms except X/Twitter moderated more than 99% of incompatible content. In contrast, X/Twitter reported that 100% of their content moderation actions were taken due to illegal content (see Fig. 5). We thus again notice a significant difference in how X/Twitter vs. other platforms interpret their obligations under the DSA.

4. Automation of Content Moderation (RQ4)

Platforms are also asked to indicate whether the content moderation decision was taken automatically (Automated Decision), and whether the decision was taken on content that has been detected or identified using automated means (Automated Detection). Fig. 6 shows that more than 90 million (60.67%) of all submitted content moderation decisions were performed fully automated (i. e., without human intervention), 48 million (31.48%) partially automated, and 11.8 million (7.74%) not automated. Furthermore, the vast majority of violations have been detected using some sort of automation (89.45%).

Interestingly, there is a strong link between the methods used for identification and decision making. Within the fully automated content moderation activities, more than 99% of all content was also first identified automatically. Within the other decision-making categories, this proportion is significantly lower at 76.06% (partially automated), and 71.75% (not automated), respectively. In other words, automated identification was usually followed by automated decision-making, while a larger proportion of content that was moderated using manual intervention was also previously identified or reported by humans (e. g., users or content moderators).

Analyzing automated decision-making across social media platforms reveals further differences (see Fig. 7). The vast majority of TikTok’s content moderation actions was performed fully automated, while Facebook, Pinterest and Instagram tend to combine automated means and human intervention (i. e., partially automated). In contrast, YouTube, Snapchat, X/Twitter and LinkedIn performed the majority of their content moderation not automated. It is striking that X/Twitter again stands out with a completely manual content moderation (i. e., not automated).

In order to better understand situations in which content moderation is more likely to be performed automatically, we implement a logistic regression model estimating how the likelihood of automated content moderation decision varies across different content types and decision grounds. The dependent variable is Automated Decision, a binary variable (=1=1 if yes; otherwise =0=0) that describes whether a moderation decision ii was performed fully automated or not (i. e., not/partially automated vs. fully automated). The key explanatory variables are AutomatedDetectioni\mathit{AutomatedDetection}_{i} (reference type: No Automated Detection), ContentTypei\mathit{ContentType}_{i} (reference type: Other) and Categoryi\mathit{Category}_{i} (reference type: Scope of Platform Service). Additionally, we control for the time elapsed (in days) between the publication of the content on the platform and the application of the content moderation Delayi\mathit{Delay}_{i} (zz-standardized). The resulting model is

with intercept β0\beta_{0}, monthly fixed effects uiu_{i} to adjust for differences in the date the content was published, and platform fixed effects λi\lambda_{i}. Note that since we apply a logistic regression, the odds ratio (i. e., the exponentiated coefficients) must be calculated to determine the effect sizes (see, eg, (Ai and Norton, 2003)).

The coefficient estimates and their 99% confidence intervals are visualized in Fig. 8. We find that the odds of a content moderation decision being performed automatically are e3.470≈32.143e^{3.470}\approx 32.143 times higher if the content was detected automatically (OR: 32.14332.143, coef: 3.4703.470, p<0.01p<0.01). Across the different content types, we find a strong positive association for Text (OR: 1158.8041158.804, coef: 7.0557.055, p<0.01p<0.01). Specifically, the odds of a content moderation decision performed automatically are significantly higher for text content compared to other content (i. e., Product, Synthetic Media, Audio, Other). At the same time, we find a smaller positive association for Images (OR: 3.2433.243, coef: 1.1771.177, p<0.01p<0.01), while Videos are less likely to be moderated automatically (OR: 0.6620.662, coef: −0.413-0.413,p<0.01p<0.01). With regard to legal grounds, we find that all categories are statistically significantly less likely to be automatically moderated compared to content that does correspond to the Scope of Platform Service (each p<0.01p<0.01). The strongest negative association can be observed for the (relatively rare) type of Non Consensual Behaviour (OR: 0.00010.0001, coef: −8.753-8.753, p<0.01p<0.01). The effect sizes for the more frequently moderated types Illegal/Harmful Speech (OR: 0.0130.013, coef: −4.342-4.342, p<0.01p<0.01), Pornography/Sexualized Content (OR: 0.0120.012, coef: −4.357-4.357, p<0.01p<0.01) and Violence (OR: 0.0180.018, coef: −3.995-3.995, p<0.01p<0.01) are all very similar to each other. We also observe a small negative association between the time elapsed between the date the content was published on the platform and the moderation date (OR: 0.7610.761, coef: −0.273-0.273, p<0.01p<0.01).

5. Content Moderation Actions (RQ5)

While the focus in the previous sections was on which and why content is moderated, we now take a closer look at the specific content moderation actions. Content moderation actions performed under the DSA can be grouped into two categories: (i) actions that affect a specific content piece (i. e., removal, labelling, disabling, demotion and age restriction of content); (ii) actions that affect an account (i. e., suspension or termination of an account). The platforms describe in the SoRs how the content was moderated, either by selecting one of these predefined actions or by selecting Other. In the case of the latter, platforms can to provide a short description of the action that was taken. However, in many cases, the content of the text descriptions is very close to the the predefined action categories (e. g., “Limited distribution” →\rightarrow Content Demoted). To accommodate such cases, we employed string matching to assign the text descriptions to the predefined action categories.The descriptions were assigned to the predefined action categories as follows: “Limited distribution” →\rightarrow Content Demoted, “not eligible for recommendation” →\rightarrow Content Demoted, “mute audio” →\rightarrow Content Disabled, “Bounce” →\rightarrow Content Demoted, “Ban” →\rightarrow Content Disabled, “AddTweetAnnotation” →\rightarrow Content Labelled (ordered by descending frequency). X/Twitter describes a large part of the moderated content as “not suitable for work” (NSFW) (71.20%) whereby, according to the platform’s own guidelines, either a reduction in visibility (i. e., demotion) or removal of the content is implemented (see striped area in Fig. 9) (X, 2023c). In cases where a description is absent or cannot be unambiguously assigned, the category Other has been retained (0.36%).

Fig. 9 shows the distribution of implemented content moderation across platforms. Overall, we observe that the most frequent types of content moderation are the removal of content (55.15%), the demotion of content (25.15%), and the suspension of accounts (14.96%). The tendency to focus on content removal and account suspension/termination as the primary choice of content moderation is prevalent across most platforms. However, there are also differences. For instance, Pinterest primarily employed content demotion. Conversely, Snapchat implemented content disabling in over 50% of its content moderation actions. X/Twitter does not distinguish between removal and demotion in its reporting to the DSA Transparency Database.

Discussion

Relevance: Effective mechanisms for content moderation are important tools to prevent the dissemination of illegal and undesirable content in social networks (Gillespie, 2018). However, content moderation decisions by social media providers are oftentimes perceived as intransparent and publicly inaccessible. Previous work was mostly limited to investigating moderation interventions in artificial environments (e. g., interviews, surveys) (e. g., Mena, 2020; Moravec et al., 2020; Pennycook et al., 2020; Bode and Vraga, 2015; Clayton et al., 2020) or based on observational datasets restricted in scope (e. g., Zannettou, 2021; Pröllochs, 2022; Pröllochs and Feuerriegel, 2023; Ling et al., 2023; Drolsbach and Pröllochs, 2023). The reason is that it was extremely challenging, if not impossible, for research to systematically collect and holistically analyze content moderation decisions. In an attempt to foster transparency and accountability, the DSA has committed major social media platforms in the EU to make key information on their content moderation decisions accessible to the public. Here, we leverage the EU’s Transparency Database – a unique and previously unavailable data source – to shed light on how major social media providers moderate user-generated content on their platforms.

Implications: Our study implies differences and inconsistencies in the content moderation practices of major social media platforms. There are vast differences in the volume of content pieces that platforms moderate. For instance, TikTok performs more than 350 times more content moderation decisions per user than X/Twitter. It remains speculative whether platforms with higher content moderation volumes tend to encounter rule-breaking content more frequently or if they simply handle such content differently. However, our findings do indicate that different platforms tend to focus their content moderation efforts on different types of rule-breaking content. While all platforms consistently take action on violence and pornography/sexualized content, their approach to moderating other pertinent online harms, such as illegal/harmful speech and misinformation, varies considerably. For instance, on X/Twitter, illegal/harmful speech and misinformation are rarely moderated. Furthermore, the type of actions social media providers take against rule-breaking content varies across platforms. While most platforms frequently remove rule-breaking content, others are more inclined to reduce its visibility. Altogether, our study suggests that social media platforms interpret their obligations under the DSA differently – resulting in a fragmented outcome that the DSA is meant to avoid. These findings have important implications for regulators, which may be inclined to clarify existing guidelines and/or lay out more specific rules that ensure common standards on how social media providers handle rule-breaking content on their platforms.

Furthermore, our findings contribute fresh insights to the ongoing debate (see (Gillespie, 2020)) regarding whether automation should have a role in social media content moderation or if this responsibility should predominantly rest with humans. Our results suggest that the majority of rule-breaking content is already identified and processed through automated methods rather than manual intervention. While many platforms (e. g., Facebook, Pinterest, and Instagram) employ a combination of automation and human moderation, others (e. g., TikTok) predominantly rely on fully automated content moderation. X/Twitter stands out as an exception, as it reports that it exclusively relies on non-automated means in its content moderation efforts. However, X/Twitter is also a platform that moderates a drastically lower volume of rule-breaking content compared to other platforms. Considering the immense scale of content moderation in the EU (more than 156 million instances within two months), this indicates there could be a necessity for some level of automation to meet the regulatory obligations imposed on social media providers by the DSA.

Limitations and future work: Our work has several limitations, which provide promising opportunities for future research. First, our inferences are limited to the first two months after the introduction of the DSA Transparency Database in the EU. For this observation period, however, we analyze all SoRs that were submitted by social media platforms. Second, content moderation efforts by social media platforms may evolve to a different steady-state due to growing experience, changes in functionality of the database, and clarification of rules by the EU. Future work may analyze how the patterns observed in this paper change over time. Third, more research is necessary to understand how content removal or visibility reduction affect on-platform user behavior. Fourth, it would be interesting to additionally analyze the source content that has been moderated (e. g., removed social media posts). However, this data is not publicly available from the EU. Notwithstanding these limitations, we believe that observing and understanding how social media providers moderate content on their platforms is the first step towards improving future policies, guidelines, and regulations. We hope that our early work inspires more research into improving transparency on content moderation on social media.

Conclusion

The Digital Services Act (DSA) represents a major legislative framework in the EU that obligates large social media providers to publish clear and specific information whenever they remove or restrict access to certain content on their platforms. Due to this new level of transparency, we are, for the first time, able to empirically analyze real-world content moderation decisions of major social media platforms in the EU. Our empirical analysis based on more than 156 million SoRs suggests that there are inconsistencies in how content moderation is carried out and how large social media platforms implement their obligations under the DSA. These findings hold important implications for regulators, who might be motivated to elucidate current guidelines or establish more specific regulations to ensure consistent standards for how social media platforms handle rule-breaking content on their platforms.

References