MentalBERT: Publicly Available Pretrained Language Models for Mental Healthcare

Shaoxiong Ji, Tianlin Zhang, Luna Ansari, Jie Fu, Prayag Tiwari, Erik Cambria

Introduction

Mental health is a global issue, especially severe in most developed countries and many emerging markets. According to the mental health action plan (2013 - 2020) from the World Health Organization, 1 in 4 people worldwide suffer from mental disorders to some extent. Moreover, 3 out of 4 people with severe mental disorders do not receive treatment, worsening the problem. During some periods like the pandemic, people struggle with mental health issues, and many may not get mental health practitioners’ help. Previous studies reveal that suicide risk usually has a connection to mental disorders (Windfuhr and Kapur, 2011). Partly due to severe mental disorders, 900,000 people commit suicide each year worldwide, making suicide the second most common cause of death among the young. Suicide attempters have been reported as suffering from mental disorders, with an investigation on a shift from mental health to suicidal ideation conducted by language and interactional measures (De Choudhury et al., 2016).

Early identification is a practical approach to mental illness and suicidal ideation prevention. Except for traditional proactive screening, social media is a good channel for mental health care. Social media platforms such as Reddit and Twitter provide anonymous space for users to discuss stigmatic topics and self-report personal issues. Social content from users who wrote about mental health issues and posted suicidal ideation has been widely used to study mental health issues (e.g., Ji et al., 2018; Tadesse et al., 2019). Machine learning-based detection techniques can empower healthcare workers in early detection and assessment to take an action of proactive prevention.

Recent advances in deep learning facilitate the development of effective early detection methods (Ji et al., 2021a). A new trend in natural language processing (NLP), contextualized pretrained language models, has attracted much attention for various text processing tasks. The seminal work on a pretrained language model called BERT (Devlin et al., 2019) utilizes bidirectional transformer-based text encoders and trains the model on a large-scale corpus. With the success of BERT, several domain-specific pretrained language models for learning text representations have also been developed and released, such as biomedical BERT (Lee et al., 2020) and clinical BERT (Alsentzer et al., 2019; Huang et al., 2019) for the biomedical and clinical domain, respectively.

However, there are no pretrained language models customized for the domain of mental healthcare. Our paper trains and releases two representative bidirectional masked language models, i.e., BERT and RoBERTa (Liu et al., 2019), with corpus collected from social forums for mental health discussion. The pretrained models in the mental health domain are dubbed MentalBERT and MentalRoBERTa. To our best knowledge, this work is the first to pre-train language models for mental healthcare. Besides, we conduct a comprehensive evaluation on several mental health detection datasets with pretrained language models in different domains. We release the pretrained MentalBERTs with Huggingface’s model repository, available at https://huggingface.co/mental.

Methods and Setup

This section introduces the language model pretraining technique and the pretraining corpus we collected. We then present the downstream tasks that we aim to solve by fine-tuning the pretrained models, and describing the setup of language model fine-tuning. Note that we aim to provide publicly available pretrained text embeddings as language resources and evaluate the usability in downstream tasks rather than propose novel pretraining techniques.

We follow the standard pretraining protocols of BERT and RoBERTa with Huggingface’s Transformers framework (Wolf et al., 2020). These two models work in a bidirectional manner, and we follow their mechanism and adopt the same loss of masked language modeling during pretraining. We use the base network architecture for both models. The BERT model we use is base uncased, which is 12-layer, 768-hidden, and 12-heads and has 110M parameters. For the pretraining of RoBERTa-based MentalBERT, we apply the dynamic masking mechanism that converges slightly slower than the static masking. Instead of the domain-specific pretraining (Gu et al., 2020) that trains language models from scratch, we adopt the training scheme similar to the domain-adaptive pretraining (Gururangan et al., 2020) that continues the pretraining in specific downstream domains. Specifically, we start the training of language models from the checkpoint of original BERT and RoBERTa. In this way, we can utilize the learned knowledge from the general domain and save computing resources, and continued pretraining makes the model adaptive to the target domain of mental health.

We use four Nvidia Tesla v100 GPUs to train the two language models. The computing resources are also one of the main assets of this paper. We set the batch size to 16 per GPU, evaluate every 1,000 steps, and train for 624,000 iterations. Training with four GPUs takes around eight days, i.e., around 32 days with only one GPU.

2 Pretraining Corpus

We collect our pretraining corpus from Reddit, an anonymous network of communities for discussion among people of similar interests. Focusing on the mental health domain, we select several relevant subreddits (i.e., Reddit communities that have a specific topic of interest) and crawl the users’ posts. We do not collect user profiles when collecting the pretraining corpus, even though those profiles are publicly available. The selected mental health-related subreddits include “r/depression”, “r/SuicideWatch”, “r/Anxiety”, “r/offmychest”, “r/bipolar”, “r/mentalillness/”, and “r/mentalhealth”. Eventually, we make the training corpus with a total of 13,671,785 sentences.

3 Downstream Task Fine-tuning

We apply the pretrained MentalBERT and MentalRoBERTa in binary mental disorder detection and multi-class mental disorder classification of various mental disorders such as stress, anxiety, and depression. We fine-tune the language models in downstream tasks. Specifically, we use the embedding of the special token [CLS] of the last hidden layer as the final feature of the input text. We adopt the multilayer perceptron (MLP) with the hyperbolic tangent activation function for the classification model. We set the learning rate of the transformer text encoder to be 1e-05 and the learning rate of classification layers to be 3e-05. The optimizer is Adam (Kingma and Ba, 2014).

Results

We evaluate and compare mental disorder detection methods on different datasets with various mental disorders (e.g., depression, anxiety, and suicidal ideation) collected from popular social platforms (e.g., Reddit and Twitter). Table 1 summarizes the datasets used in this paper. We carefully choose those benchmarks to cover a relatively wide range of mental health categories and social platforms. Some datasets do not provide a validation set. Thus, we partition a small set from the original training set to make the validation set.

Depression is one of the most common mental disorders discussed on many social platforms. We take it as a representative to evaluate the performance of different pretrained models. The first dataset for depression comes from the CLPsych 2015 Shared Task (Coppersmith et al., 2015)http://www.cs.jhu.edu/~mdredze/datasets/clpsych_shared_task_2015/. The first task of CLPsych 2015 contains user-generated posts from users with depression on Twitter. The train partition consists of 327 depression users, and the test data contains 150 depression users. Note that there is an unknown data missing issue in the dataset of the CLPsych 2015 shared task. The second dataset used is from eRisk shared task 1 (Losada and Crestani, 2016), which is a public competition for early risk detection in health-related areas. The eRisk dataset contains posts from 2,810 users, where 1,370 users express depression in their posts and 1,440 act as the control group without depression.

Pirina and Çöltekin (2018) collected additional social data form Reddit and combined them with previously collected data to identify depressionhttps://github.com/Inusette/Identifying-depression. We term this dataset as Depression_Reddit in this paper.

Suicidal Ideation

We use data collected from Reddit and Twitter to test the performance. Firstly, we use the UMD Reddit Suicidality Dataset (Shing et al., 2018) that has a total of 865 users in the subreddit of “SuicideWatch” in Reddithttp://users.umiacs.umd.edu/~resnik/umd_reddit_suicidality_dataset.html. The raw data annotation labels the user posts with four levels of risks. We include the control users and transform the label space into three classes according to the level of risks. In addition to data from Reddit, we also evaluate the performance of data collected from Twitter. We use the Twitter dataset with tweets expressing suicidal ideation and normal posts as the control group, which is collected by Ji et al. (2018, 2021a). We term this dataset as T-SID.

Other Mental Disorders

We also evaluate the performance of classifying other mental disorders such as stress, anxiety, and bipolar. Dreaddit (Turcan and McKeown, 2019) is a dataset for stress analysis with posts collected from five different forums of Reddithttp://www.cs.columbia.edu/~eturcan/data/dreaddit.zip. Specifically, it considers three major stressful topics, i.e., interpersonal conflict, mental illness, and financial need, and collects posts from ten related subreddits, including some mental health domains such as anxiety and PTSD. This dataset consists of a total of 3,553 posts split into train and test sets. Many factors may cause stress. We then use another dataset for recognizing everyday stressors called SAD, which contains 6,850 SMS-like sentences (Mauriello et al., 2021). The SAD dataset derives nine stress factors from stress management articles, chatbot-based conversation systems, crowdsourcing, and web crawling. The specific stressor categories include work, health, fatigue, or physical pain, financial problem, emotional turmoil, school, everyday decision making, family issues, social relationships, and other unspecified stressors. Lastly, we use a dataset called SWMH (Ji et al., 2021a) that contains Reddit posts with various mental disorders, including stress, anxiety, bipolar, depression, and suicidal ideation. Note that this dataset uses weak labels during the annotation process.

2 Baselines

We compare our pretrained language models for mental health with various existing pretrained models in different domains. They are BERT and RoBERTa pretrained with general corpus, BioBERT pretrained in the biomedical domain, and ClinicalBERT pretrained with clinical notes. Note that the aim of this paper is not to achieve the state-of-the-art performance but to demonstrate the usability and evaluate the performance of our pretrained models, though we have achieved competitive performance in some datasets when compared with the state of the art.

3 Results and Discussion

We evaluate the model performance by comparing the recall and F1 scores. Mental disorder detection is usually a task with unbalanced classes, leading to using the F1 score metric. It is also essential to reduce the false negatives, i.e., to ensure as few cases as possible that the detection model misses people with mental disorders. Thus, we also report recall scores.

We first compare the performance of depression detection. Table 2 reports the results on three depression dataset collected from Reddit. MentalRoBERTa archives the best performance on eRisk and CLPsych datasets, and MentalBERT is the second best model on the Depression_\_Reddit dataset.

Results of Classifying Other Mental Disorders

We then compare the performance of classifying other mental disorders and suicidal ideation. Table 3 shows the performance on various datasets with different mental disorder classification tasks. In T-SID, SAD, and Dreaddit, MentalRoBERTa is the best model with the highest recall and F1 scores. The MentalBERT has the highest F1 score in the UMD dataset, while its F1 score is not competitive to other models. While for the SWMH dataset with several mental disorders, the MentalRoBERT obtained the best F1 score.

Discussion

When comparing the domain-specific pretrained models for mental health with models pretrained with general corpora, MentalBERT and MentalRoBERTa gain better performance in most cases. Domain-specific pretraining in the biomedical or clinical domain turns out to be less helpful than pretraining on the target domain of mental health. Those results show that continued pretraining on the mental health domain improves prediction performance in downstream tasks of mental health classification.

Related Work

Contextualized embeddings have been intensively studied in NLP. Self-supervised large-scale pretraining facilitates the learning of semantic and contextual information and benefits various downstream applications such as text classification (Sun et al., 2019), sentiment analysis (Tang et al., 2020; Song et al., 2020) and relation extraction (Alt et al., 2019). There are also many domain-specific variants of pretrained contextualized text embeddings. Embeddings in specific domains aim to encode domain-specific information to boost the performance of a specific domain. For example, BioBERT (Lee et al., 2020) pretrained the BERT model in the biomedical domain using research articles from PubMed, which was applied to many biomedical tasks such as biomedical relation extraction and named entity recognition. ClinicalBERT (Alsentzer et al., 2019) used clinical notes as the pretraining corpus to continue the pretraining of the BERT model. Those domain-specific variants also foster variable downstream applications by fine-tuning pretrained embeddings such as Lin et al. (2019) and Ji et al. (2020).

NLP for Mental Healthcare

Mental healthcare research in social media is increasingly applying NLP techniques to capture users’ behavioral tendencies. Various methods are implemented for labeling, i.e., identifying emotions, mood, and profiles that might indicate mental health problems Calvo et al. (2017). One of the most representative tasks is mental health detection that categorizes given social posts into different classes of mental disorders such as depression (Tadesse et al., 2019). Mental state understanding requires effective feature representation learning and complex emotive processes. Resnik et al. (2013) applied topic modeling, an unsupervised approach that reduces the input of textual data feature space to a fixed number of topics to feature engineering in depression detection. Feature engineering-based machine learning method designs manual features and builds classifiers for mental health detection (Shatte et al., 2019; Abd Rahman et al., 2020). Various features such as sensor signals from personal devices (Mohr et al., 2017) and EEG signals (Gore and Rathi, 2019) have been applied. For detection from textual data in particular, text features include word counts, TF-IDF (Campillo-Ageitos et al., 2021), topic features (Shickel et al., 2020) and sentiment traits (Yoo et al., 2019). Severe mental disorders but without intervention may lead to suicidal ideation (Windfuhr and Kapur, 2011). Many machine learning-based methods have been applied for suicidal ideation detection (Ji et al., 2021b).

Recent work applies deep representation learning methods, which enable automatic feature learning to solve the early mental disorder identification task. Those methods typically build text embeddings and feed the embeddings into neural architectures such as convolutional neural networks (Rao et al., 2020), recurrent networks (Bouarara, 2021), self attention-based Transformers, hybrid architectures like CNN-LSTM (Kang et al., 2021) and more other deep learning architectures (Su et al., 2020). Recent works, e.g., Jiang et al. (2020), Martínez-Castaño et al. (2021) and Bucur et al. (2021), use pretrained language models and fine-tune the model for mental health tasks. However, there are no existing pretrained language models trained with mental health-related text to benefit the domain application directly.

Conclusion and Future Work

This paper trains and releases two masked language models, i.e., MentalBERT and MentalRoBERTa, on the domain data of mental health collected from the Reddit social platform. This paper is the first work that trains domain-specific language models for mental healthcare. Our pretrained models are publicly available and can be reused by the research community. Besides, we conduct a comprehensive evaluation on the performance for downstream mental health detection tasks, including depression, stress, and suicidal ideation detection. Our empirical results show that continued pretraining with mental health-related corpus can improve classification performance.

Our paper is a positive attempt to benefit the research community by releasing the pretrained models for other practical studies and with the hope to facilitate some possible real-world applications to relieve people’s mental health issues. However, we only focus on the English language in this study since English corpora are relatively easy to obtain. In the future work, we plan to collect multilingual mental health-related posts, especially those less studied by the research community, and train a multilingual language model to benefit more people speaking languages other than English.

Social Impact

The paper trains and releases masked language models for mental health to facilitate the automatic detection of mental disorders in online social content for non-clinical use. The models may help social workers find potential individuals in need of early prevention. However, the model predictions are not psychiatric diagnoses. We recommend anyone who suffers from mental health issues to call the local mental health helpline and seek professional help if possible.

Data privacy is an important issue, and we try to minimize the privacy impact when using social posts for model training. During the data collection process, we only use anonymous posts that are manifestly available to the public. We do not collect user profiles even though they are also manifestly public online. We have not attempted to identify the anonymous users or interact with any anonymous users. The collected data are stored securely with password protection even though they are collected from the open web. There might also be some bias, fairness, uncertainty, and interpretability issues during the data collection and model training. Evaluation of those issues is essential in future research.

Acknowledgments

The authors would like to thank Philip Resnik for providing the UMD Reddit Suicidality Dataset, Mark Dredze for providing the dataset in the CLPsych 2015 shared task, and other researchers who make their datasets publicly available. We acknowledge the computational resources provided by the Aalto Science-IT project. The authors wish to acknowledge CSC - IT Center for Science, Finland, for computational resources.

References