Mental Health Disorder Detection Beyond Social Media:
A Systematic Review of Available Datasets


1 Introduction↩︎

The prevalence of mental health disorders is a global concern. In the USA, for example, one in every four adults experiences a diagnosable mental health disorder each year1. Furthermore, research shows that the majority of individuals who die by suicide have an identifiable mental health condition such as depression [1] or substance use disorder [2].

The limited access to mental health care services has become an urgent societal issue [3], [4]. This motivated research on applying NLP and ML models to social media data for early identification of mental disorders and supporting individuals at risk. Social media platforms such as Twitter ([5]; [6]), Reddit ([7]; [8]) and Facebook ([9]; [10]) offer a rich, unobtrusive stream of user-generated content that captures real-life expressions with reduced reporting bias, aiding in the detection and understanding of mental disorders [11][13].

Using social media data for mental health screening raises ethical and privacy concerns regarding user consent, construct validity, and the potential for algorithmic misuse [14][16]. Collecting data from social media can lead to unintended biases as the data may only reflect the experiences of individuals who are willing to openly discuss their mental health online, which primarily includes those who are active on social media platforms [17]. Finally, demographics vary across platforms and they may not be representative of the general population. For instance, X users are primarily male, while TikTok and Instagram users are mostly female; likewise, many platforms tend to be used by teens and young adults [18], [19]. Pragmatically, many social media sites have limited access to APIs for research.

The widespread use of social media data in this domain is partly due to the significant shortage of non-social media free-text based datasets related to mental health disorders derived from clinical and other reliable sources. Clinical data, such as electronic medical records (EMRs) or electronic health records (EHRs) with discharge summaries or clinical notes, psychiatric interview transcripts, and responses to standardized open-ended questionnaires, offer rich, detailed insights into patients’ mental health. Unlike social media data, these sources contain carefully documented information from healthcare professionals, including diagnostic details, symptom descriptions, and treatment histories, often supported by validated clinical scales.

Despite their great potential, non-social media free-text based datasets remain underexplored primarily due to privacy concerns, data access challenges, and annotation complexities. This scarcity presents a barrier to advancing robust, generalizable NLP models that can be effectively integrated into clinical practice. We aim to address this gap by systematically reviewing non-social media free-text based datasets for mental health research. We explore their diversity in terms of source, structure, clinical annotation, types, and research adoption to guide future efforts in dataset development and application. Previous related surveys have primarily focused on datasets from social media platforms [20][25]. To the best of our knowledge, this is the first systematic review2 [26] of free-text based mental health data sources beyond social media.

We address the following research questions:

1.0.0.1 RQ1:

What non-social media free-text based datasets are available for mental health research, and how do they vary by source, structure, and population?

1.0.0.2 RQ2:

How are mental health conditions defined and labeled in these datasets, and what are the clinical implications of these labeling methods?

1.0.0.3 RQ3:

What factors contribute to the popularity and adoption of non-social media free-text based mental health datasets in research?

2 Methods↩︎

This systematic review adopts the PRISMA3 reporting guidelines to comprehensively map and systematically analyze the landscape of the free-text based datasets in mental health research beyond social media, building on prior frameworks (e.g. [27][29]). PRISMA is a standardized guideline aimed at ensuring clear and thorough reporting of systematic reviews and meta-analyses. It features a 27-item checklist that helps authors include all essential components of their review, from initial identification to final conclusions. A key part of PRISMA is its flow diagram (Figure 1), which visually maps the process of selecting studies, making the review process more transparent and easy to follow [30].

Figure 1: PRISMA flowchart illustrating the systematic literature search and screening process.

To follow the checklists of PRISMA, we began with a systematic literature search that continued through June 2025. We used Publish or Perish4 to query multiple academic databases, including Google Scholar, and PubMed. The searches were carried out separately for key mental health related terms: ‘depression’, ‘anxiety’, ‘suicidal ideation’, ‘suicide’, ‘mental disorder’, ‘mental health crisis’, and ‘mental health disorder’, combined with targeted keywords ‘identification’, ‘detection’, ‘prediction’, and ‘analysis’ to capture each relevant disorder type of mental health and associated dataset corpus. We then manually screened the abstracts and included studies that explicitly referenced the use of any datasets or the application of any computational models or techniques after de-duplication, resulting in 543 papers within our initial scope. The rationale for selecting depression and anxiety as key terms is that they are among the most prevalent mental health disorders worldwide [31]. Additionally, given that suicide remains a leading cause of death globally among adults [32], we prioritize the inclusion of these mental health conditions in our study.

To ensure focus on non-social media sources, we excluded papers whose titles or abstracts included keywords related to social media platforms or mentions of datasets scraped from them, such as ‘social media’, ‘Reddit’, ‘Facebook’, or ‘Twitter’. This filtering was essential to isolate papers based on clinical notes, semi-structured interviews, and discharge summaries from other non-social media platforms, and we ended up with 259 total papers after this filtering. Subsequently, we conducted corpus novelty screening to remove studies based on pre-existing datasets, followed by modality screening to exclude papers that focused solely on audio or visual data, and data type screening to filter out datasets containing only numerical or categorical values, such as scale-based survey responses. Through this multi-phase screening, we identified 57 relevant papers, of which 40 proposed or analyzed newly collected free-text based mental health datasets.

In addition to keyword-based search, we followed a backtracking strategy; whenever a paper used an existing dataset, we traced it back to the original dataset publication, even if it did not contain our predefined keywords. We also examined all papers cited in the literature review or background sections of included studies and explored any dataset-related papers mentioned in relevant survey papers found from the initial search. As a result, our final collection may include important dataset papers whose titles or abstracts do not directly match the initial search terms, but were directly aligned with our objectives. This process yielded 5 additional papers, for a total of 45.

Several datasets were initially annotated for mental disorders, but their final task focused on emotion classification, including loneliness, fear, anger, hopelessness, and self-identification. Since emotions like loneliness and hopelessness can serve as indicators of depression [33] or suicidal ideation [34], we decided to include these studies in our survey.

3 Distribution Analysis↩︎

In this section, we provide information on the 45 free-text based datasets included in this systematic review. Table 1 provides a comprehensive summary of the final set of datasets. The datasets were published between 2004 and 2025 in multiple languages, including English, Chinese, Polish, Korean, Japanese, Arabic, and some code-mixed. These datasets cover a broad spectrum of mental health conditions, such as depression, postnatal depression, anxiety, schizophrenia, suicidal ideation, Post-Traumatic Stress Disorder, and bipolar disorder, collected from platforms like clinics, colleges, mobile apps, and therapy sessions. Data types include interview transcripts, essays, discharge summaries, clinical records, forum posts, and suicide notes. While some datasets are publicly available, others are restricted or require agreements, with a few lacking availability details. We show their distribution below across disorders, languages, platforms, data types, demographics, and availability types.

Table 1: Overview of non-social media, text-based mental health datasets. The table summarizes each dataset’s language, mental disorder(s), platform, data form, annotation method, labeling instrument, availability, and citation count (till June 2025).Abbreviations: SD - Suicidal, DP - Depression, MinDP - Minor Depression, PD - Postnatal Depression, PTSD - Post Traumatic Stress Disorder, OCD - Obsessive Compulsive Disorder, SMI - Severe Mental Illness, MDD - Major Depressive Disorder, MHD - Mental Health Distress, AjD - Adjustment Disorder, BD - Bipolar Disorder, SZ - Schizophrenia, AX - Anxiety, SDI - Suicidal Ideation, FEP - First Episode Psychosis, DEM - Dementia, PA - Panic Attack, PUB - Public, DUA- Data Use Agreement, RSTR - Restricted, UNK - Unknown, DS - Discharge Summaries
Dataset Language
Disorder Platform
Type
Procedure
Instrument Label Size Availability Citation
[35] EN SD CLINIC SD Notes Manual Krippendorff’s \(\alpha\) 15 1004 DUA 236
[36] EN DP, PTSD, OCD, BD CLINIC Clinical Notes Manual - 4 816 DUA 45
[37] EN SMI CLINIC EHR(DS) Manual Text Hunter 50 37,211 RSTR 233
[38] EN MDD CLINIC Interview Manual LIFE 2 139 RSTR 313
[39] EN SD CLINIC SD Notes Manual Ontology 2 66 RSTR 322
[40] EN MDD & DD CLINIC EMR Manual DSM-IV 2 861 RSTR 88
[41] EN DP CLINIC EHR (DS) Manual - 3 1200 UNK 74
[42] EN DP CLINIC EHR Manual ICD-9 & AD 2 10,148 UNK 53
[43] EN DP & AX CLINIC Interview Manual DSM-IV - 308 UNK 51
[44] EN SZ CLINIC Essay Manual - 2 56 UNK 71
[45] EN SD CLINIC Clinical Notes Manual - 3 210 UNK 233
[46] EN
SD & Other CLINIC
Clinical Notes Manual - 2 198 UNK 26
[47] EN MHD FORUM Post Manual
Cohen’s Kappa 4 1227 PUB 142
[48] EN PTSD FORUM Essay Manual DSM-IV & CAP Scale 2 300 UNK 95
[49] EN DP & AX FORUM Post Self-disclosure - 2 16,975 UNK 55
[50] EN DP, BD & SD FORUM Post Self-disclosure - 2 267,964 UNK 303
[51] EN SZ & BD COLLEGE Interview Manual DSM-V & DSM-IV 3 644 DUA 11
[52] EN DP COLLEGE Essay Manual BDI & IDD-L 3 124 UNK 1543
[53] EN AX DEBATE Political Speech Manual - 2 4000 PUB 17
[54] EN DP & AX Telemedicine Platform Message Manual
GAD-7
GAD-7 10,718 DUA 43
[55] EN PD APP Essay Manual EPDS 2 1,091 UNK 14
[56] EN DP & SD Phone SMS Manual Self-Disclosure 2 94 UNK 127
[57] EN DP & AX Online Therapy Chat Dialogue Mixed MALLET & LIWC 2 882 UNK 78
[58] EN DP & PTSD Virtual Agent Interview Manual PHQ-8 5 275 DUA 402
[59] EN DP & SD MIXED SD Notes & Articles Manual - 12 426 UNK 64
[60] EN DP, AX & PTSD MIXED Interview Manual - - 621 DUA 759
[61] EN DP & SDI MIXED SD Notes & Book Manual Cohen’s Kappa 15 2393 PUB 45
[62] EN DP & AX mTurk Interview Manual
GAD-7
GAD-7 2674 UNK 29
[63] EN-ES DP & SDI Social Network & Forum Post Manual Cohen’s Kappa 4 102 PUB 24
[64] EN-ZH
BD, AjD, DEM & Other CLINIC EHR(DS) Manual
Sheehan Disability Scale 9 4,836 RSTR 122
[65] ZH MDD CLINIC Interview Manual HAMD & PHQ-9 2 78 PUB 48
[66] ZH DP & AX CLINIC Interview Manual HAMD & HAMA 3 1025 PUB 9
[67] ZH DP & AX CLINIC EMR Manual DSM-V & ICD-10 4 1,160 PUB 1
[68] ZH DP CLINIC Interview Manual MADRS 2 113 DUA 5
[69] ZH AX CLINIC EHR Manual ICD-9 & ICD-10 2 84,426 RSTR 0
[70] ZH DP & SD CLINIC Interview Manual HAMD 3 305 UNK 10
[71] ZH DP APP Interview Manual SDS 2 162 PUB 146
[72] ZH AX Phn. Recording Essay Manual GAD-7 3 227 UNK 11
[73] KO DP, AX & SD CLINIC Interview Manual
BAI & BSS 2 166 DUA 22
[74] KO SZ & FEP CLINIC Interview Manual DSM-V, PANSS 3 133 RSTR 33
[75] PL SZ CLINIC Interview Manual ICD-10 2 94 UNK 20
[76] ES-CL DP CLINIC Interview Manual DSM-V 8 451 UNK 2
[77] TH DP FORUM Post Manual Key-word Search 2 944 PUB 18
[78] JA DP FORUM Post Self-disclosure - 2 108 UNK 29
[79] AR DP FORUM Post
Manual
QIDS-SR 2 20,000 UNK 75
[80] EN SD Online DB Song Lyrics Manual - 2 810 UNK 37
[81] EL SDI - POEM Manual - 2 90 UNK 12
[82] CF DP, AX & PA Tablets - - - - - UNK 50
Figure 2: Distribution of the datasets by year

Temporal Distribution Figure 2 illustrates the number of free-text based datasets proposed and analyzed in published studies each year between 2004 and 2025. The introduction of such datasets remained minimal and steady up to 2012. In 2014 and 2017, there was a noticeable rise in the number of proposed datasets, indicating intensified research efforts in this domain during that period. The proposal of such datasets peaked in 2022, which can be attributed to greater attention to mental health concerns after the COVID-19 pandemic [25]. The lower count in 2025 is probably a result of the systematic search being conducted until June 2025. Despite this overall growth, the process of collecting such datasets often involves complex procedures, including ethics reviews, annotator training, and repeated permission requests, which may explain the limited number of datasets over the years.

Figure 3: Distribution of datasets into by language groups.

3.0.0.1 Language Distribution

Figure 3 illustrates the distribution of the datasets by language group. English (EN) accounts for the majority, comprising 62.2% of the datasets, highlighting its predominant role in this research area. Chinese (ZH) represents 17.8% of the total. The ‘Other’ group includes a variety of less-represented languages, such as Korean (KO), Polish (PL), Greek (EL), Japanese (JA), Thai (TH), and Arabic (AR); collectively representing another 13.3%. Additionally, 6.4% of the datasets fall under the “Code-mixed” category, which includes multilingual combinations like EN-ES (English-Spanish), EN-ZH (English-Chinese), and ES-CL (Spanish-Chilean). While there is some linguistic variety, English datasets dominate in mental health research, indicating a need for more inclusive and multilingual dataset development.

3.0.0.2 Mental Health Disorder Distribution

We consider the distribution of datasets for different mental health disorders and language diversity. In Figure 4, we present a heatmap with the distribution of mental health datasets across different disorders and languages. Since some datasets cover more than one disorder, the total count of disorders is higher than the number of datasets. To make things clearer, we group similar disorders; depression related conditions, grouped under ‘DP’ (including Depression (DP), Major Depressive Disorder (MDD), Minor Depression (MinDP), and Postnatal Depression (PD)) are the most studied, appearing in 33 datasets. Anxiety (AX) and suicidal ideation/suicide(SD) follow with 12 and 11 datasets, respectively. Other less common disorders, such as Adjustment Disorder (AjD), Obsessive–compulsive disorder (OCD), Dysthymic Disorder (DD), and Dementia (DEM) are grouped under ‘Others’; each accounts for 8 datasets. Schizophrenia (SZ) and First-episode Psychosis (FEP), grouped together as ‘SZ’, are analyzed in 5 datasets, while Post-Traumatic Stress Disorder (PTSD), and Bipolar Disorder (BD) are the least represented ones.

English datasets dominate the landscape, especially for depression (DP) and suicidal ideation (SD). Code-mixed and non-English data remain underrepresented across most disorders. This highlights a concentration on depression and a linguistic imbalance, emphasizing the need for more diverse mental health datasets.

Figure 4: Heatmap showing the distribution of datasets across language groups and disorders.

3.0.0.3 Platform & Data Type Distribution

Figure 5 shows the distribution of data types across four platform types: CLINIC, FORUM, MIXED (from multiple sources), and OTHER (including apps, virtual agents, online therapy chat, telemedicine platform and suicide notes). For simplicity, we group similar data types, like Electronic Health Records (EHR), Electronic Medical Records (EMR), clinical records (CR), and Discharge Summaries (DS) are grouped under ‘EHR’; essays and questionnaires fall under ‘questionnaires’ category; while types like interviews and posts are kept distinct.

Figure 5: Distribution of data types across platforms.

As shown in Figure 5, clinical datasets mainly stem from interviews (48%) and EHRs (39%), reflecting structured clinical data sources. Forum datasets are heavily composed of user-generated posts (86%). Mixed platform datasets are more balanced, with one dataset each from interviews, posts, questionnaires, and other forms (each 25%). In the “Other” category, data forms are diverse: interviews and unconventional sources (like phone recordings, SMS, app data) both account for above 80%, and questionnaires make up the remaining.

3.0.0.4 Availability & Impact

Figure 6 displays the distribution of citation counts (as per Google Scholar5) for mental health datasets categorized by their availability: datasets that are publicly available (PUB), datasets that require a Data Use Agreement (DUA), datasets labeled as Restricted (RSTR), and datasets with unspecified availability (UNK). Citation counts are normalized by year to account for differences in publication age. Datasets categorized as DUA and RSTR generally have higher citation counts compared to those that are publicly accessible or have unclear availability. Notably, DUA datasets show a wider range and a higher median citation rate, while RSTR datasets demonstrate a more consistent citation pattern. In contrast, PUB and UNK datasets have lower medians and are more tightly clustered around fewer citations, though some outliers exist.

The citation trends can also be partly explained by language distribution (Figure 7).

Figure 6: Distribution of citation counts by dataset availability type.
Figure 7: Availability of datasets across language groups.

Many English datasets are in the UNK category, while Chinese datasets are more publicly available (PUB). However, English-language PUB datasets rarely include depression-related data, and none originate from clinical settings. In contrast, English-language DUA datasets predominantly focus on depression and are collected in clinical contexts – this contributes to their higher credibility and research utility. Then again, most academic research is conducted in English; the higher citation rates of DUA and RSTR datasets reflect language dominance. Meanwhile, the lower citation rates for PUB datasets may be influenced by their association with less widely used languages like Chinese.

3.0.0.5 Demographic Representation

Figure 8 shows how frequently different demographic attributes appear in the datasets shown in this survey.

Figure 8: Frequency of demographic attributes across datasets.

Demographic attributes like gender and age dominate, appearing in over 24 datasets, while others like education, race, and ethnicity are moderately included. Fewer datasets report attributes like income, relationship status, or health indicators (e.g., height, weight, blood pressure). Some authors [38], [78] use demographic information to select negative samples that match the same profile type as the positive ones, which is rarely done in social media-based data, resulting in poor performance on the minority class [83], [84]. Some datasets focus on specific populations such as students [52], [56], [71], veterans [45], [60], seafarers [72], new mothers [55], or politicians [53]. Some have used demographic features in classification tasks, often reporting improved model performance [42], [62], [85].

4 Tools & Techniques for Data Annotation & Labeling↩︎

Most datasets in our review are manually annotated by annotators or expert clinicians, following widely used tools and techniques. This section describes these datasets and the tools employed. Some datasets, such as those by [49], [50], [78], and [79], are based on self-reported user content, typically collected from forums where individuals voluntarily share their thoughts under relevant discussion threads.

4.1 Clinical Diagnostic Instruments↩︎

Clinical diagnostic tools are assessments carried out by qualified healthcare professionals using standardized frameworks. They typically involve face-to-face evaluations, structured interviews, and expert judgment to ensure reliable and consistent mental health diagnoses. This systematic review highlights multiple studies that utilized clinical diagnostic tools for labeling and assessment.

The DSM-IV (Diagnostic and Statistical Manual of Mental Disorders, Fourth Edition) [86] and its updated version, the DSM-V [87], offer standardized criteria for diagnosing a wide range of mental health conditions. Diagnoses based on DSM are often recorded during clinical visits and stored in EHRs, making them a key source of labeled data. Several works in this survey rely on DSM-IV for annotation, including [40], [48], [43] and [64], while studies such as [67], [79], and [76] utilize the updated DSM-V guidelines. Notably, [51] incorporated both frameworks to account for changes in diagnostic criteria across versions.

The ICD (International Classification of Diseases), particularly ICD-9 and ICD-10 [88], is used for coding clinical diagnoses, including mental and behavioral disorders. While the DSM focuses on mental health, the ICD covers all diseases. For instance, the code F32.9 in ICD-10 represents an unspecified single episode of major depressive disorder. In this survey, ICD-9 is used by [42], while [67] and [75] utilize ICD-10 for data annotation, and [69] uses both editions to label anxiety.

The LIFE (Longitudinal Interval Follow-Up Evaluation) is a structured method used to monitor the long-term progression of psychiatric conditions, typically every six months, particularly within longitudinal studies [89]. [90] underlines that LIFE is used in both clinical practice and research to assess the long-term impact of psychiatric disorders. In this survey, only [38] uses LIFE to classify the interviews.

The CAPS (Clinician-Administered PTSD Scale) is a structured interview for diagnosing PTSD. It assesses 20 core PTSD symptoms, onset, duration, impairment, and dissociative features related to a specific traumatic event and is considered the gold standard for PTSD evaluation [91]. Clinicians in [48] use both the DSM-IV PTSD module and the CAPS scale to annotate the dataset with binary labels for PTSD presence or absence.

The HAMD (Hamilton Depression Rating Scale) [92] is a clinician-administered scale used to assess depression severity through patient interviews, including 17 core items that rate depression from mild to severe. Similarly, HAMA (Hamilton Anxiety Rating Scale) [93] evaluates anxiety symptoms. In this survey, [65], [70], [73] use HAMD to annotate the dataset for MDD or suicidal ideation, while [66] use both HAMD and HAMA for depression and anxiety labeling.

The MADRS (Montgomery-Åsberg Depression Rating Scale) [94] is a clinician-administered tool for assessing depression severity, emphasizing core symptoms. It is often used in clinical trials and is sensitive to changes in symptoms. [68] uses MADRS to identify depression from transcribed interviews.

4.2 Screening Questionnaires and Self-Report Scales↩︎

Screening questionnaires and self-report scales are tools where individuals assess their own mental health by answering standardized questions. These are often based on validated clinical criteria for quickly screening for symptoms. Several studies included in this review employ screening questionnaires and self-report scales for mental health disorder identification.

Some studies in this review, [62], [65], [79], [54] and [73] use the Patient Health Questionnaire (PHQ-9) [95] for dataset annotations and defining gold standard labels for depression. [58] uses the older version of PHQ-9, PHQ-8, which consists of the same questions, excluding the suicidal ideation item [96]. [52] employs the Inventory to Diagnose Depression–Lifetime (IDD-L) [97] to analyze college students’ essays, identifying depression through their reflections on past experiences. The Beck Depression Inventory (BDI-II) [98], a self-reported tool used to assess the intensity of depressive symptoms, [52] employs both BDI-II and IDD-L for dataset annotation. In addition, the Beck Anxiety Inventory (BAI) [99] and the Beck Scale for Suicide Ideation (BSS) [100] are used to measure anxiety and suicidal thoughts, respectively, with [73] applying both to label transcribed interviews. In a study by [55], the Edinburgh Postnatal Depression Scale (EPDS) [101] is used to identify depression in essays collected from pregnant women through an app. [64] combines Sheehan Disability Scale [102] with DSM-IV guidelines to identify major depressive disorder (MDD), schizophrenia (SZ), bipolar disorder (BD), and other psychiatric conditions from discharge summaries, extracted from patients’ EHRs; whereas [71] uses Zung Self-Rating Depression Scale (SDS) [103] to label depression from transcribed interviews collected through an app. [62], [54] and [72] use the GAD-7 (Generalized Anxiety Disorder-7) [104] to classify anxiety from their respective datasets. [79] uses the Quick Inventory of Depressive Symptomatology- Self-Report (QIDS-SR) [105] tool to label depression in posts from a psychological forum in Arabic.

4.3 Manual or Contextual Labeling↩︎

Several datasets in this review do not explicitly mention the annotation tools used; however, they do indicate that trained annotators or clinicians conducted the labeling [36]. In some instances, studies report inter-rater agreement measures like Cohen’s Kappa (e.g. [47], [61], [63]), Fleiss’s Kappa [47], or Krippendorff’s \(\alpha\) [35] to justify the labeling decisions, and occasionally a third annotator was used to resolve conflicts. A few datasets generate gold-standard labels using NLP-based methods such as Text Hunter [37], custom ontologies [39], or frameworks like MALLET and LIWC [57]. [77] uses keyword search (i.e., “depression",”anxiety") to label the datasets.

Few miscellaneous datasets contain song lyrics [80], poems [81] from artists who have committed suicide, even ancient texts scraped from Babylonian tablets [82]; the authors do not mention using any labeling techniques.

4.4 Clinical Implications↩︎

Different labeling strategies in mental health datasets carry distinct clinical implications. Clinical diagnostic tools are widely considered the most reliable or gold standard [106], offering standardized and consistent evaluations that support accuracy across clinicians6. However, these tools have limitations, including time constraints (e.g., typical sessions lasting 15–20 minutes), limited accessibility (shortage of clinicians), and missed opportunities for monitoring between appointments (often referred to as “clinical whitespace") [107].

To address these gaps, self-report tools have become increasingly valuable. They capture patients’ first-hand perspectives and are ideal for digital use [106], especially when clinicians are unavailable, such as during clinical whitespace periods [108] or while patients await intake appointments [109]. While self-reports offer efficiency and scalability, they should not replace in-person evaluations, as they are susceptible to social desirability effects, recall bias [110], and trust issues. Instead, they should serve as supplementary tools [106]. In fact, research supports combining clinician ratings and self-reports for a more comprehensive understanding of patient conditions [111].

When clinical labels are unavailable, manual annotation, together with annotator agreement techniques, is often used to scale data labeling. While useful, these methods can suffer from reduced accuracy if not guided by trained professionals [112]. Overall, each labeling method presents trade-offs between scalability, clinical rigor, and data quality.

5 Computational Modeling↩︎

The studies reviewed in this systematic review have been used to develop mental health classification and prediction systems using a broad spectrum of approaches [20], [21], [25]. While a thorough review of computational approaches is outside the scope of this survey, in this section, we provide the reader with a brief summary of the computational models applied to the task. Exploring computational models allows us to better understand prevailing methodological trends, evaluate benchmarks, and identify potential limitations or biases inherent in different approaches. This insight is vital for ensuring replicability and shaping future research directions.

Early works use manual or qualitative techniques while later studies applied traditional machine learning (ML) models like support vector machine (SVM) [35][37], [39], [44], [48], [49], [56], Logistic Regression [35], [42], [51], [57], [62], [70], [72], [78], Decision Trees [41], [48], [71], Adaboost [36], [39], [66], [72], XGboost [68], [75], [76], and Naive Bayes [48], [59], [61], [78]. Traditional ML classifiers are often supported by feature engineering techniques such as Linguistic Inquiry and Word Count (LIWC) [113] analysis [52], [56], [57], [70], Term Frequency-Inverse Document Frequency (tf-idf) [40], [56], [69], or keyword extraction. Deep learning models, particularly Convolutional Neural Networks (CNN) [36], [49], [61], Long Short-Term Memory (LSTM) [36], [55], [68], [77], and Gated Recurrent Unit (GRU) [37], [114] are also used to capture nuanced patterns in unstructured text.

Some datasets have been used to develop interpretable models or hybrid systems that combine rule-based methods with ML (e.g., Conditional Random Field (CRF) [35], [64], TextHunter + ConText [37]). More recent work adopts transformer-based models like Clinical-BigBird [69], MentalBERT [51], MentalRoBERTa [51], and large language models (LLMs) like Qwen2-72B [67], highlighting a shift from fine-tuning pre-trained language models to using instruction-tuned LLMs for domain-specific tasks with little or no fine-tuning.

6 Identified Trends & Research Gaps↩︎

This section summarizes the main findings of this review along with suggestions for future directions. Unsurprisingly, most available datasets are in English. Chinese follows in prevalence, while there are very few datasets in other languages, such as Korean, Arabic, or Polish. There are no datasets in low-resource languages. This gap may be partly due to the complexity of creating these datasets, particularly in clinical contexts. Collecting data often requires patient consent, approval from ethics boards, and substantial manual effort to annotate datasets.

Another important finding is that depression is the most studied mental health disorder within these datasets. Studies either focus exclusively on depression or include related conditions such as anxiety and bipolar disorder. Other conditions, such as PTSD, schizophrenia, and eating disorders, are not well represented. This highlights the need to develop datasets that encompass a broader range of mental health issues. On the other hand, compared to the number of social media datasets [115], [116], clinical datasets are scarce due to their sensitive nature. Releasing more public or even agreement-governed clinical datasets could expand research opportunities and improve reproducibility.

Moreover, the annotation procedures and labeling techniques vary across the datasets. Some use clinical diagnostic tools such as DSM and ICD, while others rely on self-reported questionnaires like PHQ and BDI. However, most papers do not explain why a specific tool was chosen, indicating a lack of reporting standards [16]. Additionally, some datasets are labeled manually or through keyword searches, and a few do not clarify their labeling methods at all. This lack of transparency can make it difficult to trust or replicate the research findings.

Advancements in NLP offer promising solutions to the challenges posed by inconsistent and opaque labeling practices in mental health datasets. With instruction-tuned language models and structured prompting techniques, NLP can support (semi-)automation and standardization of labeling based on formal diagnostic criteria such as DSM or ICD. Approaches like chain-of-thought (CoT) prompting [117] can emulate clinical reasoning and even improve labeling accuracy [118], while chain-of-empathy (CoE) [119] frameworks are well-suited for understanding emotionally nuanced texts, such as therapy transcripts or suicide notes. These strategies can enhance both the consistency and interpretability of labels, enabling the creation of scalable and clinically relevant datasets, even when direct clinician involvement is limited.

7 Conclusion & Future Directions↩︎

In this paper, we presented the first systematic review of mental health free-text datasets beyond social media. We analyze these datasets with respect to different dimensions such as their distribution in terms of languages and mental health disorders, the data types included in them, and their availability.

We revisit the research questions (RQ) posed in the introduction (Section 1) and present the main findings of our review below:

7.0.0.1 RQ1:

What non-social media free-text based datasets are available for mental health research, and how do they vary by source, structure, and population?

RQ1 Findings: Most datasets are in English and primarily focus on depression, while other languages and mental health conditions remain underrepresented. More Chinese datasets are publicly available than English ones. In terms of data collection, some actively balance samples by selecting equal numbers from the control group matching the same profile type. The nature of text-based data also varies widely, ranging from phone recordings and messages to essays, interview transcripts, and suicide notes, mostly collected in clinical settings.

7.0.0.2 RQ2:

How are mental health conditions defined and labeled in these datasets, and what are the clinical implications of these labeling methods?

RQ2 Findings: Clinical diagnostic tools, screening questionnaires, self-report scales, and manual contextual labeling have been used to annotate mental health disorders. These labeling techniques can complement one another and collectively serve as gold standards.

7.0.0.3 RQ3:

What factors contribute to the popularity and adoption of non-social media free-text based mental health datasets in research?

RQ3 Findings: Along with dataset availability, factors such as the data collection setting (clinical vs. informal), language, data type (interviews, EHRs, essays), mental health disorders covered, and the credibility and reliability of the labeling and annotation methods all contribute to the popularity and adoption of non-social media free-text based mental health datasets in research.

There is significant room for improvement in the development and use of these types of datasets. Expanding dataset creation to include more languages and geographic regions can help address current imbalances. Additionally, increasing the representation of a broader range of mental health disorders, not just depression, would make research findings more comprehensive. Establishing clear and consistent labeling methods is essential for ensuring reproducibility and building trust in research outcomes, and advances in NLP, particularly in explainable AI, can help bridge these gaps. Lastly, improving access to high-quality datasets, while maintaining ethical and privacy standards, can facilitate collaborative research.

Limitations↩︎

In this systematic review, we used the PRISMA methodology and conducted a comprehensive literature search using the Publish or Perish software. This approach ensured thorough coverage of datasets on mental health disorders across both the NLP and clinical domains. Although we used mental health-related terms for our keyword searches, it is possible that studies may use alternative terms or less common phrases to describe mental health disorders, which may not have been included in our search strategy. As a result, some relevant works might have been unintentionally overlooked. Additionally, the search queries were in English, which may have excluded relevant non-English publications. Finally, Publish or Perish does not index certain databases, such as CINAHL, potentially limiting coverage of some clinically oriented studies.

Ethics Statement↩︎

This systematic review draws from previously published studies to guide future research in identifying mental health disorders beyond social media. While we present the available datasets used for this purpose in our review, we did not make any attempts to build prediction models using this data. We acknowledge that detecting early signs of mental health disorders requires adherence to ethical protocols. Moreover, misuse of these sensitive data or of the models trained on the data can lead to stigmatization or harm to individuals with mental health disorders [17].

Acknowledgements↩︎

We would like to thank the anonymous reviewers for their constructive feedback and thoughtful suggestions, which helped improve the clarity and quality of this work. We are also grateful to our collaborators for their valuable contributions and insightful discussions throughout the project. Finally, we acknowledge the creators and maintainers of the datasets studied in our review for developing and curating these resources, thus enabling and advancing research in this domain.

References↩︎

References↩︎

[1]
R. K. Bailey et al., “Suicide: Current trends,” Journal of the National Medical Association, vol. 103, no. 7, pp. 614–617, 2011.
[2]
F. L. Lynch et al., “Substance use disorders and risk of suicide in a general US population: A case control study,” Addiction science & clinical practice, vol. 15, no. 1, p. 14, 2020.
[3]
J. R. Cummings, H. Wen, and B. G. Druss, “Improving access to mental health services for youth in the united states,” Jama, vol. 309, no. 6, pp. 553–554, 2013.
[4]
R. D. Hester, “Lack of access to mental health services contributing to the high suicide rates among veterans,” International journal of mental health systems, vol. 11, no. 1, p. 47, 2017.
[5]
M. Chatterjee, P. Samanta, P. Kumar, and D. Sarkar, “Suicide ideation detection using multiple feature analysis from twitter data,” in IEEE DELCON, 2022.
[6]
D. S. Khafaga, M. Auvdaiappan, K. Deepa, M. Abouhawwash, and F. K. Karim, “Deep learning for depression detection using twitter data,” Intelligent Automation & Soft Computing, vol. 36, no. 2, pp. 1301–1313, 2023.
[7]
N. Boettcher, “Studies of depression and anxiety using reddit as a data source: Scoping review,” JMIR mental health, vol. 8, no. 11, p. e29487, 2021.
[8]
E. Adil Jaafar and H. Abdul-Salam Jasim, “A corpus-based stylistic analysis of online suicide notes retrieved from reddit,” Cogent Arts & Humanities, vol. 9, no. 1, p. 2047434, 2022.
[9]
R. A. Calvo, D. N. Milne, M. S. Hussain, and H. Christensen, “Natural language processing in mental health applications using non-clinical texts,” Natural Language Engineering, vol. 23, no. 5, pp. 649–685, 2017.
[10]
M. R. Islam, M. A. Kabir, A. Ahmed, A. R. M. Kamal, H. Wang, and A. Ulhaq, “Depression detection from social network data using machine learning techniques,” Health information science and systems, vol. 6, pp. 1–12, 2018.
[11]
S. Sametoğlu, D. Pelt, J. C. Eichstaedt, L. H. Ungar, and M. Bartels, “The value of social media language for the assessment of wellbeing: A systematic review and meta-analysis,” The Journal of Positive Psychology, vol. 19, no. 3, pp. 471–489, 2024.
[12]
N. Raihan, S. S. C. Puspo, S. Farabi, A.-M. Bucur, T. Ranasinghe, and M. Zampieri, “Mentalhelp: A multi-task dataset for mental health in social media,” in Proceedings of LREC-COLING, 2024.
[13]
N. Raihan, S. S. C. Puspo, A.-M. Bucur, S. Chancellor, and M. Zampieri, “Large language models for mental health: A multilingual evaluation,” in Proceedings of LowResLM, 2026.
[14]
S. Chancellor, E. P. Baumer, and M. De Choudhury, “Who is the" human" in human-centered machine learning: The case of predicting mental health from social media,” Proceedings of the ACM on Human-Computer Interaction, vol. 3, no. CSCW, pp. 1–32, 2019.
[15]
J. Nicholas, S. Onie, and M. E. Larsen, “Ethics and privacy in social media research for mental health,” Current psychiatry reports, vol. 22, pp. 1–7, 2020.
[16]
S. Chancellor and M. De Choudhury, “Methods in predictive techniques for mental health status on social media: A critical review,” NPJ digital medicine, vol. 3, no. 1, p. 43, 2020.
[17]
S. Chancellor, M. L. Birnbaum, E. D. Caine, V. M. Silenzio, and M. De Choudhury, “A taxonomy of ethical tensions in inferring mental health states from social media,” in Proceedings of FACCT, 2019.
[18]
A. Olteanu, C. Castillo, F. Diaz, and E. Kıcıman, “Social data: Biases, methodological pitfalls, and ethical boundaries,” Frontiers in big data, vol. 2, p. 13, 2019.
[19]
Y. Zhao et al., “Biases in using social media data for public health surveillance: A scoping review,” International Journal of Medical Informatics, vol. 164, p. 104804, 2022.
[20]
K. Harrigian, C. Aguirre, and M. Dredze, “On the state of social media data for mental health research,” in Proceedings of CLPsych, 2021.
[21]
E. A. Ríssola, M. Aliannejadi, and F. Crestani, “Beyond modelling: Understanding mental disorders in online social media,” in Proc. Of ECIR, 2020, pp. 296–310.
[22]
R. Skaik and D. Inkpen, “Using social media for mental health surveillance: A review,” ACM Computing Surveys, vol. 53, no. 6, pp. 1–31, 2020.
[23]
A. Abdulsalam and A. Alhothali, “Suicidal ideation detection on social media: A review of machine learning methods,” Social Network Analysis and Mining, vol. 14, no. 1, p. 188, 2024.
[24]
A.-M. Bucur, A. Moldovan, K. Parvatikar, M. Zampieri, A. Khudabukhsh, and L. P. Dinu, “Datasets for depression modeling in social media: An overview,” in Proceedings of CLPsych, 2025.
[25]
A.-M. Bucur, A.-C. Moldovan, K. Parvatikar, M. Zampieri, A. R. KhudaBukhsh, and L. P. Dinu, “On the state of NLP approaches to modeling depression in social media: A post-COVID-19 outlook,” IEEE Journal of Biomedical and Health Informatics, vol. 29, no. 6, pp. 4439–4451, 2025.
[26]
M. J. Grant and A. Booth, “A typology of reviews: An analysis of 14 review types and associated methodologies,” Health information & libraries journal, vol. 26, no. 2, pp. 91–108, 2009.
[27]
M. Templier and G. Paré, “A framework for guiding and evaluating literature reviews,” Communications of the Association for Information Systems, vol. 37, no. 1, p. 6, 2015.
[28]
D. Zarate, V. Stavropoulos, M. Ball, G. de Sena Collier, and N. C. Jacobson, “Exploring the digital footprint of depression: A PRISMA systematic literature review of the empirical evidence,” BMC psychiatry, vol. 22, no. 1, p. 421, 2022.
[29]
M. Pazdur, D. Tutus, and A.-C. Haag, “Risk factors for problematic social media use in youth: A systematic review of longitudinal studies,” Adolescent Research Review, pp. 1–17, 2025.
[30]
M. J. Page et al., “The PRISMA 2020 statement: An updated guideline for reporting systematic reviews,” bmj, vol. 372, 2021.
[31]
Institute for Health Metrics and Evaluation, Online database. Seattle, WA. Accessed 13 August 2025“Global burden of disease study 2021 (GBD 2021) results.” https://vizhub.healthdata.org/gbd-results/, 2024.
[32]
Centers for Disease Control and Prevention, Accessed 22 February 2026“Suicidal thoughts and behavior.” https://www.cdc.gov/mental-health/about-data/suicidal-thoughts-and-behavior.html, 2025.
[33]
W. S. Rholes, J. H. Riskind, and B. Neville, “The relationship of cognitions and hopelessness to depression and anxiety,” Social Cognition, vol. 3, no. 1, pp. 36–50, 1985.
[34]
I. Baryshnikov et al., “Role of hopelessness in suicidal ideation among patients with depressive disorders,” The Journal of clinical psychiatry, vol. 81, no. 2, p. 8339, 2020.
[35]
J. P. Pestian et al., “Sentiment analysis of suicide notes: A shared task,” Biomedical informatics insights, vol. 5, pp. BII–S9042, 2012.
[36]
O. Uzuner, A. Stubbs, and M. Filannino, “A natural language processing challenge for clinical records: Research domains criteria (RDoC) for psychiatry,” Journal of biomedical informatics, vol. 75, p. S1, 2017.
[37]
R. G. Jackson et al., “Natural language processing to extract symptoms of severe mental illness from clinical text: The clinical record interactive search comprehensive data extraction (CRIS-CODE) project,” BMJ open, vol. 7, no. 1, p. e012012, 2017.
[38]
L.-S. A. Low, N. C. Maddage, M. Lech, L. B. Sheeber, and N. B. Allen, “Detection of clinical depression in adolescents’ speech during family interactions,” IEEE transactions on biomedical engineering, vol. 58, no. 3, pp. 574–586, 2010.
[39]
J. Pestian, H. Nasrallah, P. Matykiewicz, A. Bennett, and A. Leenaars, “Suicide note classification using natural language processing: A content analysis,” Biomedical informatics insights, vol. 3, pp. BII–S4706, 2010.
[40]
J. Geraci, P. Wilansky, V. de Luca, A. Roy, J. L. Kennedy, and J. Strauss, “Applying deep neural networks to unstructured text notes in electronic medical records for phenotyping youth depression,” BMJ Ment Health, vol. 20, no. 3, pp. 83–87, 2017.
[41]
L. Zhou et al., “Identifying patients with depression using free-text clinical documents,” in MEDINFO, 2015.
[42]
Y. Meng, W. Speier, M. Ong, and C. W. Arnold, “HCET: Hierarchical clinical embedding with topic modeling on electronic health records for predicting future depression,” IEEE Journal of Biomedical and Health Informatics, vol. 25, no. 4, pp. 1265–1272, 2021.
[43]
R. A. Marrie et al., “Effects of psychiatric comorbidity in immune-mediated inflammatory disease: Protocol for a prospective study,” JMIR Research Protocols, vol. 7, no. 1, p. e8794, 2018.
[44]
J. Diederich, A. Al-Ajmi, and P. Yellowlees, “Ex-ray: Data mining and mental health,” Applied Soft Computing, vol. 7, no. 3, pp. 923–928, 2007.
[45]
C. Poulin et al., “Predicting the risk of suicide by analyzing the text of clinical notes,” PloS one, vol. 9, no. 1, p. e85733, 2014.
[46]
P. Saini, D. While, K. Chantler, K. Windfuhr, and N. Kapur, “Assessment and management of suicide risk in primary care,” Crisis, 2014.
[47]
D. N. Milne, G. Pink, B. Hachey, and R. A. Calvo, “Clpsych 2016 shared task: Triaging content in online peer-support forums,” in Proceedings of the third workshop on computational linguistics and clinical psychology, 2016, pp. 118–127.
[48]
Q. He, B. P. Veldkamp, C. A. Glas, and T. de Vries, “Automated assessment of patients’ self-narratives for posttraumatic stress disorder screening using natural language processing and text mining,” Assessment, vol. 24, no. 2, pp. 157–172, 2017.
[49]
Y. Tyshchenko, “Depression and anxiety detection from blog posts data,” Nature Precis. Sci., Inst. Comput. Sci., Univ. Tartu, Tartu, Estonia, pp. 6–46, 2018.
[50]
T. Nguyen, D. Phung, B. Dao, S. Venkatesh, and M. Berk, “Affective and content analysis of online depression communities,” IEEE Transactions on Affective Computing, vol. 5, no. 3, pp. 217–226, 2014.
[51]
A. Aich et al., “Towards intelligent clinically-informed language analyses of people with bipolar disorder and schizophrenia,” in Findings of EMNLP, 2022.
[52]
S. Rude, E.-M. Gortner, and J. Pennebaker, “Language use of depressed and depression-vulnerable college students,” Cognition & Emotion, vol. 18, no. 8, pp. 1121–1133, 2004.
[53]
L. Rheault, “Expressions of anxiety in political texts,” in Proceedings of NLP-CSS, 2016.
[54]
T. D. Hull, M. Malgaroli, P. S. Connolly, S. Feuerstein, and N. M. Simon, “Two-way messaging therapy for depression and anxiety: Longitudinal response trajectories,” BMC psychiatry, vol. 20, no. 1, p. 297, 2020.
[55]
T. Krishnamurti, K. Allen, L. Hayani, S. Rodriguez, and A. L. Davis, “Identification of maternal depression risk from natural language collected in a mobile health app,” Procedia computer science, vol. 206, pp. 132–140, 2022.
[56]
A. L. Nobles, J. J. Glenn, K. Kowsari, B. A. Teachman, and L. E. Barnes, “Identification of imminent suicide risk among young adults using text messages,” in Proceedings of CHI, 2018.
[57]
C. Howes, M. Purver, and R. McCabe, “Linguistic indicators of severity and progress in online text-based therapy for depression,” in Proceedings of CLPsych, 2014.
[58]
F. Ringeval et al., “AVEC 2019 workshop and challenge: State-of-mind, detecting depression with AI, and cross-cultural affect recognition,” in Proceedings of AVEC, 2019.
[59]
A. M. Schoene and N. Dethlefs, “Automatic identification of suicide notes from linguistic and sentiment features,” in Proceedings of SIGHUM, LaTeCH, 2016.
[60]
J. Gratch et al., “The distress analysis interview corpus of human and computer interviews,” in Proceedings of LREC, 2014.
[61]
S. Ghosh, A. Ekbal, and P. Bhattacharyya, “Cease, a corpus of emotion annotated suicide notes in english,” in Proceedings of LREC, 2020.
[62]
M. Tasnim, M. Ehghaghi, B. Diep, and J. Novikova, DEPAC: A corpus for depression and anxiety detection from speech,” in Proceedings of CLPsych, 2022.
[63]
R. W. A. Caicedo, J. M. G. Soriano, and H. A. M. Sasieta, “Assessment of supervised classifiers for the task of detecting messages with suicidal ideation,” Heliyon, vol. 6, no. 8, 2020.
[64]
C.-S. Wu, C.-J. Kuo, C.-H. Su, S.-H. Wang, and H.-J. Dai, “Using text mining to extract depressive symptoms and to validate the diagnosis of major depressive disorder from electronic health records,” Journal of affective disorders, vol. 260, pp. 617–623, 2020.
[65]
B. Zou et al., “Semi-structural interview-based chinese multimodal depression corpus towards automatic preliminary screening of depressive disorders,” IEEE Transactions on Affective Computing, vol. 14, no. 4, pp. 2823–2838, 2023.
[66]
Y. Jiang, Z. Zhang, and X. Sun, “MMDA: A multimodal dataset for depression and anxiety detection,” in Proceedings of ICPR, 2022.
[67]
S. Xu et al., “Identifying psychiatric manifestations in outpatients with depression and anxiety: A large language model-based approach,” medRxiv, pp. 2025–01, 2025.
[68]
K. Mao et al., “Analysis of automated clinical depression diagnosis in a chinese corpus,” IEEE Transactions on Biomedical Circuits and Systems, vol. 17, no. 5, pp. 1135–1152, 2023.
[69]
J. Ive et al., “A data-centric approach to detecting and mitigating demographic bias in pediatric mental health text: A case study in anxiety detection,” arXiv preprint arXiv:2501.00129, 2024.
[70]
T. M. Li et al., “Detection of suicidal ideation in clinical interviews for depression using natural language processing and machine learning: Cross-sectional study,” JMIR medical informatics, vol. 11, no. 1, p. e50221, 2023.
[71]
Y. Shen, H. Yang, and L. Lin, “Automatic depression detection: An emotional audio-textual corpus and a gru/bilstm-based model,” in IEEE ICASSP, 2022.
[72]
H. Mo, S. C. Hui, X. Liao, Y. Li, W. Zhang, and S. Ding, “A multimodal data-driven framework for anxiety screening,” IEEE Transactions on Instrumentation and Measurement, vol. 73, pp. 1–13, 2024.
[73]
D. Shin et al., “Detection of depression and suicide risk based on text from clinical interviews using machine learning: Possibility of a new objective diagnostic marker,” Frontiers in psychiatry, vol. 13, p. 801301, 2022.
[74]
A. Figueroa-Barra et al., “Automatic language analysis identifies and predicts schizophrenia in first-episode of psychosis,” Schizophrenia, vol. 8, no. 1, p. 53, 2022.
[75]
A. Wawer, I. Chojnicka, L. Okruszek, and J. Sarzynska-Wawer, “Single and cross-disorder detection for autism and schizophrenia,” Cognitive Computation, vol. 14, no. 1, pp. 461–473, 2022.
[76]
J. Oh et al., “Development of depression detection algorithm using text scripts of routine psychiatric interview,” Frontiers in psychiatry, vol. 14, p. 1256571, 2024.
[77]
M. Hämäläinen, P. Patpong, K. Alnajjar, N. Partanen, and J. Rueter, “Detecting depression in Thai blog posts: A dataset and a baseline,” in Proceedings of w-NUT, 2021.
[78]
M. Hiraga, “Predicting depression for japanese blog text,” in Proceedings of ACL, 2017.
[79]
N. S. Alghamdi, H. A. H. Mahmoud, A. Abraham, S. A. Alanazi, and L. Garcı́a-Hernández, “Predicting depression symptoms in an arabic psychological forum,” IEEE access, vol. 8, pp. 57317–57334, 2020.
[80]
M. Mulholland and J. Quinn, “Suicidal tendencies: The automatic classification of suicidal and non-suicidal lyricists using nlp,” in Proceedings of IJCNLP, 2013.
[81]
A. D. Zervopoulos et al., “Language processing for predicting suicidal tendencies: A case study in greek poetry,” in IFIP AIAI, 2019.
[82]
E. H. Reynolds and J. V. K. Wilson, “Depression and anxiety in babylon,” Journal of the Royal Society of Medicine, vol. 106, no. 12, pp. 478–481, 2013.
[83]
S. Rai et al., “Key language markers of depression on social media depend on race,” Proceedings of the National Academy of Sciences, vol. 121, no. 14, p. e2319837121, 2024.
[84]
Y. Cao et al., “Machine learning approaches for mental illness detection on social media: A systematic review of biases and methodological challenges,” arXiv preprint arXiv:2410.16204, 2024.
[85]
J. Shin, H. Yu, H. Moon, A. Madotto, and J. Park, “Dialogue summaries as dialogue states (DS2), template-guided summarization for few-shot dialogue state tracking,” in Findings of ACL, 2022.
[86]
APA, Diagnostic and statistical manual of mental disorders, fourth edition (DSM-IV). American Psychiatric Association, 1994.
[87]
APA, Diagnostic and statistical manual of mental disorders, 5th ed. American Psychiatric Publishing, 2013.
[88]
World Health Organization, The ICD-10 classification of mental and behavioural disorders: Clinical descriptions and diagnostic guidelines, vol. 1. World Health Organization, 1992.
[89]
M. B. Keller et al., “The longitudinal interval follow-up evaluation: A comprehensive method for assessing outcome in prospective longitudinal studies,” Archives of general psychiatry, vol. 44, no. 6, pp. 540–548, 1987.
[90]
R. J. Porter et al., “Validation of the longitudinal interval follow-up evaluation for the long-term measurement of mood symptoms in bipolar disorder,” Brain Sciences, vol. 12, no. 12, 2022.
[91]
D. D. Blake et al., “The development of a clinician-administered PTSD scale,” Journal of traumatic stress, vol. 8, pp. 75–90, 1995.
[92]
L. Renemane and J. Vrublevska, “Chapter 17 - hamilton depression rating scale: Uses and applications,” in The neuroscience of depression, C. R. Martin, L.-A. Hunter, V. B. Patel, V. R. Preedy, and R. Rajendram, Eds. Academic Press, 2021, pp. 175–183.
[93]
L. S. Matza, R. Morlock, C. Sexton, K. Malley, and D. Feltner, “Identifying HAM-a cutoffs for mild, moderate, and severe generalized anxiety disorder,” International journal of methods in psychiatric research, vol. 19, no. 4, pp. 223–232, 2010.
[94]
M. J. Müller, H. Himmerich, B. Kienzle, and A. Szegedi, “Differentiating moderate and severe depression using the montgomery–åsberg depression rating scale (MADRS),” Journal of Affective Disorders, vol. 77, no. 3, pp. 255–260, 2003.
[95]
K. Kroenke, R. L. Spitzer, and J. B. Williams, “The PHQ-9: Validity of a brief depression severity measure,” Journal of general internal medicine, vol. 16, no. 9, pp. 606–613, 2001.
[96]
K. Kroenke, T. W. Strine, R. L. Spitzer, J. B. Williams, J. T. Berry, and A. H. Mokdad, “The PHQ-8 as a measure of current depression in the general population,” Journal of Affective Disorders, vol. 114, no. 1–3, pp. 163–173, 2009.
[97]
M. Zimmerman and W. Coryell, “The inventory to diagnose depression, lifetime version,” Acta Psychiatrica Scandinavica, vol. 75, no. 5, pp. 495–499, 1987.
[98]
A. T. Beck, R. A. Steer, and G. Brown, “Beck depression inventory–II,” Psychological assessment, 1996.
[99]
A. T. Beck, N. Epstein, G. Brown, and R. Steer, “Beck anxiety inventory,” Journal of consulting and clinical psychology, 1993.
[100]
A. T. Beck and R. A. Steer, “Manual for the beck scale for suicide ideation,” San Antonio, TX: Psychological Corporation, vol. 63, 1991.
[101]
J. L. Cox, J. M. Holden, and R. Sagovsky, “Detection of postnatal depression: Development of the 10-item edinburgh postnatal depression scale,” The British journal of psychiatry, vol. 150, no. 6, pp. 782–786, 1987.
[102]
D. V. Sheehan, K. Harnett-Sheehan, and B. Raj, “The measurement of disability,” International clinical psychopharmacology, vol. 11, pp. 89–95, 1996.
[103]
M. Tenti, W. Raffaeli, and P. Gremigni, “A narrative review of the assessment of depression in chronic pain,” Pain Management Nursing, vol. 23, no. 2, pp. 158–167, 2022.
[104]
R. L. Spitzer, K. Kroenke, J. B. W. Williams, and B. Löwe, “A brief measure for assessing generalized anxiety disorder: The GAD-7,” Archives of Internal Medicine, vol. 166, no. 10, pp. 1092–1097, 2006.
[105]
A. J. Rush et al., “The 16-item quick inventory of depressive symptomatology (QIDS), clinician rating (QIDS-c), and self-report (QIDS-SR): A psychometric evaluation in patients with chronic major depression,” Biological Psychiatry, vol. 54, no. 5, pp. 573–583, 2003.
[106]
K. Arrow et al., “Evaluating the use of online self-report questionnaires as clinically valid mental health monitoring tools in the clinical whitespace,” Psychiatric Quarterly, vol. 94, no. 2, pp. 221–231, 2023.
[107]
K. Stene-Larsen and A. Reneflot, “Contact with primary and mental health care prior to suicide: A systematic review of the literature from 2000 to 2017,” Scandinavian journal of public health, vol. 47, no. 1, pp. 9–17, 2019.
[108]
G. Coppersmith, C. Hilland, O. Frieder, and R. Leary, “Scalable mental health analysis in the clinical whitespace via natural language processing,” in IEEE BHI, 2017.
[109]
M. Cruz et al., “Appointment length, psychiatrists’ communication behaviors, and medication management appointment adherence,” Psychiatric Services, vol. 64, no. 9, pp. 886–892, 2013.
[110]
A. Althubaiti, “Information bias in health research: Definition, pitfalls, and adjustment methods,” Journal of multidisciplinary healthcare, pp. 211–217, 2016.
[111]
R. Uher et al., “Self-report and clinician-rated measures of depression severity: Can one replace the other?” Depression and anxiety, vol. 29, no. 12, pp. 1043–1049, 2012.
[112]
A. Sylolypavan, D. Sleeman, H. Wu, and M. Sim, “The impact of inconsistent human annotations on AI driven clinical decision making,” NPJ Digital Medicine, vol. 6, no. 1, p. 26, 2023.
[113]
R. L. Boyd, A. Ashokkumar, S. Seraj, and J. W. Pennebaker, “The development and psychometric properties of LIWC-22,” Austin, TX: University of Texas at Austin, vol. 10, no. 1–47, p. 6, 2022.
[114]
S. Ghosh, D. Baker, D. Jurgens, and V. Prabhakaran, “Detecting cross-geographic biases in toxicity modeling on social media,” in Proceedings of the seventh workshop on noisy user-generated text (w-NUT 2021), Nov. 2021, pp. 313–328, doi: 10.18653/v1/2021.wnut-1.35.
[115]
M. Garg, “Mental health analysis in social media posts: A survey,” Archives of Computational Methods in Engineering, vol. 30, no. 3, p. 1819, 2023.
[116]
A.-M. Bucur, M. Zampieri, T. Ranasinghe, and F. Crestani, “A survey on multilingual mental disorders detection from social media data,” in Proceedings of EACL, 2026.
[117]
J. Wei et al., “Chain-of-thought prompting elicits reasoning in large language models,” Advances in neural information processing systems, vol. 35, pp. 24824–24837, 2022.
[118]
E. Shi, A. Manda, L. Chowdhury, R. Arun, K. Zhu, and M. Lam, “Enhancing depression diagnosis with chain-of-thought prompting,” arXiv preprint arXiv:2408.14053, 2024.
[119]
Y. K. Lee, I. Lee, M. Shin, S. Bae, and S. Hahn, “Chain of empathy: Enhancing empathetic response of large language models based on psychotherapy models,” arXiv preprint arXiv:2311.04915, 2023.

  1. https://www.hopkinsmedicine.org/health/wellness-and-prevention/mental-health-disorder-statistics↩︎

  2. https://github.com/SadiyaPuspo/MHD-Beyond-Social-Media-Datasets-Review↩︎

  3. https://www.prisma-statement.org/prisma-2020↩︎

  4. https://harzing.com/resources/publish-or-perish↩︎

  5. https://scholar.google.com/↩︎

  6. https://www.verywellmind.com/what-is-reliability-2795786↩︎