A computational linguistic study of personal recovery in bipolar disorder

Glorianna Jagfeld
Spectrum Centre for Mental Health Research
Lancaster University
United Kingdom
g.jagfeld@lancaster.ac.uk


Abstract

Mental health research can benefit increasingly fruitfully from computational linguistics methods, given the abundant availability of language data in the internet and advances of computational tools. This interdisciplinary project will collect and analyse social media data of individuals diagnosed with bipolar disorder with regard to their recovery experiences. Personal recovery - living a satisfying and contributing life along symptoms of severe mental health issues - so far has only been investigated qualitatively with structured interviews and quantitatively with standardised questionnaires with mainly English-speaking participants in Western countries. Complementary to this evidence, computational linguistic methods allow us to analyse first-person accounts shared online in large quantities, representing unstructured settings and a more heterogeneous, multilingual population, to draw a more complete picture of the aspects and mechanisms of personal recovery in bipolar disorder.

1 Introduction and background↩︎

Recent years have witnessed increased performance in many computational linguistics tasks such as syntactic and semantic parsing [1], [2], emotion classification [3], and sentiment analysis [4][6], especially concerning the applicability of such tools to noisy online data. Moreover, the field has made substantial progress in developing multilingual models and extending semantic annotation resources to languages beyond English [7][10].

Concurrently, it has been argued for mental health research that it would constitute a ‘valuable critical step’ [11] to analyse first-hand accounts by individuals with lived experience of severe mental health issues in blog posts, tweets, and discussion forums. Several severe mental health difficulties, e.g., bipolar disorder (BD) and schizophrenia are considered as chronic and clinical recovery, defined as being relapse and symptom free for a sustained period of time [12], is considered difficult to achieve [13][15]. Moreover, clinically recovered individuals often do not regain full social and educational/vocational functioning [16], [17]. Therefore, research originating from initiatives by people with lived experience of mental health issues has been advocating emphasis on the individual’s goals in recovery  [18], [19]. This movement gave rise to the concept of personal recovery [20], [21], loosely defined as a ‘way of living a satisfying, hopeful, and contributing life even with limitations caused by illness’ [19]. The aspects of personal recovery have been conceptualised in various ways [22][24]. According to the frequently used CHIME model [25], its main components are Connectedness, Hope and optimism, Identity, Meaning and purpose, and Empowerment. Here, we focus on BD, which is characterised by recurring episodes of depressed and elated (hypomanic or manic) mood [13], [26]. Bipolar spectrum disorders were estimated to affect approximately 2% of the UK population [14] with rates ranging from 0.1%-4.4% across 11 other European, American and Asian countries [27]. Moreover, BD is associated with a high risk of suicide [28], making its prevention and treatment important tasks for society. BD-specific personal recovery research is motivated by mainly two facts: First, the pole of positive/elevated mood and ongoing mood instability constitute core features of BD and pose special challenges compared to other mental health issues, such as unipolar depression [26]. Second, unlike for some other severe mental health difficulties, return to normal functioning is achievable given appropriate treatment [17], [29], [30].

A substantial body of qualitative and quantitative research has shown the importance of personal recovery for individuals diagnosed with BD [23], [24], [26], [31], [32]. Qualitative evidence mainly comes from (semi-)structured interviews and focus groups and has been criticised for small numbers of participants [11], lacking complementary quantitative evidence from larger samples [33]. Some quantitative evidence stems from the standardised bipolar recovery questionnaire [31] and a randomised control trial for recovery-focused cognitive-behavioural therapy [32]. Critically, previous research has taken place only in structured settings. What is more, the recovery concept emerged from research primarily conducted in English-speaking countries, mainly involving researchers and participants of Western ethnicity. This might have led to a lack of non-Western notions of wellbeing in the concept, such as those found in indigenous peoples [33], limiting its the applicability to a general population. Indeed, the variation in BD prevalence rates from 0.1% in India to 4.4% in the US is striking. It has been shown that culture is an important factor in the diagnosis of BD [34], as well as on the causes attributed to mental health difficulties in general and treatments considered appropriate [35], [36]. While approaches to mental health classification from texts have long ignored the cultural dimension [37], first studies show that online language of individuals affected by depression or related mental health difficulties differs significantly across cultures  [37], [38].

Hence, it seems timely to take into account the wealth of accounts of mental health difficulties and recovery stories from individuals of diverse ethnic and cultural backgrounds that are available in a multitude of languages on the internet. Corpus and computational linguistic methods are explicitly designed for processing large amounts of linguistic data [39][42], and as discussed above, recent advances have made it feasible to apply them to noisy user-generated texts from diverse domains, including mental health [43], [44]. Computer-aided analysis of public social media data enables us to address several shortcomings in the scientific underpinning of personal recovery in BD by overcoming the small sample sizes of lab-collected data and including accounts from a more heterogeneous population.

In sum, our research questions are as follows: (1) How is personal recovery discussed online by individuals meeting criteria for BD? (2) What new insights do we get about personal recovery and factors that facilitate or hinder it? We will investigate these questions in two parts, looking at English-language data by westerners and at multilingual data by individuals of diverse ethnicities.

2 Data↩︎

Previous work in computational linguistics and clinical psychology has tended to focus on the detection of mental health issues as classification tasks [45]. Datasets have been collected for various conditions including BD using publicly available social-media data from Twitter [46] and Reddit [47], [48]. Unfortunately, the Twitter dataset is unavailable for further research.1 In both Reddit datasets, mental health-related content was deliberately removed. This allows the training of classifiers that try to predict the mental health of authors from excerpts that do not explicitly address mental health, yet it renders the data useless for analyses on how mental health is talked about online. Due to this lack of appropriate existing publicly accessible datasets, we will create such resources and make them available to subsequent researchers.

We plan to collect data relevant for BD in general as well as for personal recovery in BD from three sources varying in their available amount versus depth of the accounts we expect to find: 1) Twitter, 2) Reddit (focusing on mental health-related content unlike previous work), 3) blogs authored by affected individuals. Twitter and Reddit users with a BD diagnosis will be identified automatically via self-reported diagnosis statements, such as ‘I was diagnosed with BD-I last week’. To do so, we will extend on the diagnosis patterns and terms for BD provided by [48]2. Implicit consent is assumed from users on these platforms to use their public tweets and posts.3 Relevant blogs will be manually identified, and their authors will be contacted to obtain informed consent for using their texts.

Since language and culture are important factors in our research questions, we need information on the language of the texts and the country of residence of their authors3, which is not provided in a structured format in the three data sources. For language identification, Twitter employs an automatic tool [49], which can be used to filter tweets according to 60 language codes, and there are free, fairly accurate tools such as the Google Compact Language Detector4, which can be applied to Reddit and blog posts. The location of Twitter users can be automatically inferred from their tweets [50] or the (albeit noisy) location field in their user profiles [51]. Only one attempt to classify the location of Reddit users has been published so far [52] showing meagre results, indicating that the development of robust location classification approaches on this platform would constitute a valuable contribution. Some companies collect mental health-related online data and make them available to researchers subject to approval of their internal review boards, e.g., OurDataHelps5 by Qntfy or the peer-support forum provider 7 Cups6. Unlike ‘raw’ social media data, these datasets have richer user-provided metadata and explicit consent for research usage. On the other hand, less data is available, the process to obtain access might be tedious within the short timeline of a PhD project and it might be impossible to share the used portions of the data with other researchers. Therefore, we will follow up the possibilities of obtaining access to these datasets, but in parallel also collect our own datasets to avoid dependence on external data providers.

3 Methodology and Resources↩︎

As explained in the introduction, the overarching aim of this project is to investigate in how far information conveyed in social media posts can complement more traditional research methods in clinical psychology to get insights into the recovery experience of individuals with a BD diagnosis. Therefore, we will first conduct a systematic literature review of qualitative evidence to establish a solid base of what is already known about personal recovery experiences in BD for the subsequent social media studies.

Our research questions, which regard the experiences of different populations, lend themselves to several subprojects. First, we will collect and analyse English-language data from westerners. Then, we will address ethnically diverse English-speaking populations and finally multilingual accounts. This has the advantage that we can build data processing and methodological workflows along an increase in complexity of the data collection and analysis throughout the project.

In each project phase, we will employ a mixed-methods approach to combine the advantages of quantitative and qualitative methods [53], [54], which is established in mental health research [55][58] and specifically recommended to investigate personal recovery [59]. Quantitative methods are suitable to study observable behaviour such as language and yield more generalisable results by taking into account large samples. However, they fall short of capturing the subjective, idiosyncratic meaning of socially constructed reality, which is important when studying individuals’ recovery experience [23], [24], [60], [61]. Therefore, we will apply an explanatory sequential research design [54], starting with statistical analysis of the full dataset followed by a manual investigation of fewer examples, similar to ‘distant reading’ [62] in digital humanities.

Since previous research mainly employed (semi-)structured interviews and we do not expect to necessarily find the same aspects emphasised in unstructured settings, even less so when looking at a more diverse and non-English speaking population, we will not derive hypotheses from existing recovery models for testing on the online data. Instead, we will start off with exploratory quantitative research using comparative analysis tools such as Wmatrix [63] to uncover important linguistic features, e.g., on keywords and key concepts that occur with unexpected frequency in our collected datasets relative to reference corpora. The underlying assumption is that keywords and key concepts are indicative of certain aspects of personal recovery, such as those specified in the CHIME model [25], other previous research [23], [24], [61], or novel ones. Comparing online sources with transcripts of structured interviews or subcorpora originating from different cultural backgrounds might uncover aspects that were not prominently represented in the accounts studied in prior research.

A specific challenge will be to narrow down the data to parts relevant for personal recovery, since there is no control over the discussed topics compared to structured interviews. To investigate how individuals discuss personal recovery online and what (potentially unrecorded) aspects they associate with it, without a priori narrowing down the search-space to specific known keywords seems like a chicken-and-egg problem. We propose to address this challenge by an iterative approach similar to the one taken in a corpus linguistic study of cancer metaphors [64]. Drawing on results from previous qualitative research [24], [25], we will compile an initial dictionary of recovery-related terms. Next, we will examine a small portion of the dataset manually, which will be partly randomly sampled and partly selected to contain recovery-related terms. Based on this, we will be able to expand the dictionary and additionally automatically annotate semantic concepts of the identified relevant text passages using a semantic tagging approach such as the UCREL Semantic Analysis System (USAS) [65]. Crucially for the multilingual aspect of the project, USAS can tag semantic categories in eight languages [9]. Then, semantic tagging will be applied to the full corpus to retrieve all text passages mentioning relevant concepts. Furthermore, distributional semantics methods [66], [67] can be used to find terms that frequently co-occur with words from our keyword dictionary. Occurrences of the identified keywords or concepts can be quantified in the full corpus to identify the importance of the related personal recovery aspects.

Linguistic Inquiry and Word Count (LIWC) [68] is a frequently used tool in social-science text analysis to analyse emotional and cognitive components of texts and derive features for classification models [47], [48], [69], [70]. LIWC counts target words organised in a manually constructed hierarchical dictionary without contextual disambiguation in the texts under analysis and has been psychometrically validated and developed for English exclusively. While translations for several languages exist, e.g., Dutch [10], and it is questionable to what extent LIWC concepts can be transferred to other languages and cultures by mere translation. We therefore aim to apply and develop methods that require less manual labour and are applicable to many languages and cultures. One option constitute unsupervised methods, such as topic modelling, which has been applied to explore cultural differences in mental-health related online data already [37], [38]. The Differential Language Analysis ToolKit (DLATK) [71] facilitates social-scientific language analyses, including tools for preprocessing, such as emoticon-aware tokenisers, filtering according to meta data, and analysis, e.g. via robust topic modelling methods.

Furthermore, emotion and sentiment analysis constitute useful tools to investigate the emotions involved in talking about recovery and identify factors that facilitate or hinder it. There are many annotated datasets to train supervised classifiers [4], [72] for these actively researched NLP tasks. Machine learning methods were found to usually outperform rule-based approaches based on look-ups in dictionaries such as LIWC. Again, most annotated resources are English, but state of the art approaches based on multilingual embeddings allow transferring models between languages [5].

4 Ethical considerations↩︎

Ethical considerations are established as essential part in planning mental health research and most research projects undergo approval by an ethics committee. On the contrary, the computational linguistics community has started only recently to consider ethical questions [73], [74]. Likely, this is because computational linguistics was traditionally concerned with publicly available, impersonal texts such as newspapers or texts published with some temporal distance, which left a distance between the text and author. Conversely, recent social media research often deals with highly personal information of living individuals, who can be directly affected by the outcomes [73].

[73] discuss issues that can arise when constructing datasets from social media and conducting analyses or developing predictive models based on these data, which we review here in relation to our project: Demographic bias in sampling the data can lead to exclusion of minority groups, resulting in overgeneralisation of models based on these data. As discussed in the introduction, personal recovery research suffers from a bias towards English-speaking Western individuals of white ethnicity. By studying multilingual accounts of ethnically diverse populations we explicitly address the demographic bias of previous research. Topic overexposure is tricky to address, where certain groups are perceived as abnormal when research repeatedly finds that their language is different or more difficult to process. Unlike previous research [46][48] our goal is not to reveal particularities in the language of individuals affected by mental health problems. Instead, we will compare accounts of individuals with BD from different settings (structured interviews versus informal online discourse) and of different backgrounds. While the latter bears the risk to overexpose certain minority groups, we will pay special attention to this in the dissemination of our results.

Lastly, most research, even when conducted with the best intentions, suffers from the dual-use problem [75], in that it can be misused or have consequences that affect people’s life negatively. For this reason, we refrain from publishing mental health classification methods, which could be used, for example, by health insurance companies for the risk assessment of applicants based on their social media profiles.

If and how informed consent needs to be obtained for research on social media data is a debated issue [76][78], mainly because it is not straightforward to determine if posts are made in a public or private context. From a legal point of view, the privacy policies of Twitter7 and Reddit8, explicitly allow analysis of the user contents by third party, but it is unclear to what extent users are aware of this when posting to these platforms [79]. However, in practice it is often infeasible to seek retrospective consent from hundreds or thousands of social media users. According to current ethical guidelines for social media research [80], [81] and practice in comparable research projects [79], [82], it is regarded as acceptable to waive explicit consent if the anonymity of the users is preserved. Therefore, we will not ask the account holders of Twitter and Reddit posts included in our datasets for their consent.

[80] formulate guidelines for ethical social media health research that pertain especially to data collection and sharing. In line with these, we will only share anonymised and paraphrased excerpts from the texts, as it is often possible to recover a user name via a web search for the verbatim text of a post. However, we will make the original texts available as datasets to subsequent research under a data usage agreement. Since the (automatic) annotation of demographic variables in parts of our dataset constitutes especially sensitive information on minority status in conjunction with mental health, we will only share these annotations with researchers that demonstrate a genuine need for them, i.e. to verify our results or to investigate certain research questions.

Another important question is in which situations of encountering content indicative of a risk of self-harm or harm to others it would be appropriate or even required by duty of care for the research team to pass on information to authorities. Surprisingly, we could only find two mentions of this issue in social media research [82], [83]. Acknowledging that suicidal ideation fluctuates [84], we accord with the ethical review board’s requirement in [82] to only analyse content posted at least three months ago. If the research team, which includes clinical psychologists, still perceives users at risk we will make use of the reporting facilities of Twitter and Reddit.

As a central component we consider the involvement of individuals with lived experience in our project, an aspect which is missing in the discussion of ethical social media health research so far. The proposal has been presented to an advisory board of individuals with a BD diagnosis and was received positively. The advisory board will be consulted at several stages of the project to inform the research design, analysis, and publication of results. We believe that board members can help to address several of the raised ethical problems, e.g., shaping the research questions to avoid feeding into existing biases or overexposing certain groups and highlighting potentially harmful interpretations and uses of our results.

5 Impact and conclusion↩︎

The importance of the recovery concept in the design of mental health services has recently been prominently reinforced, suggesting ‘recovery-oriented social enterprises as key component of the integrated service’ [21]. We think that a recovery approach as leading principle for national or global health service strategies, should be informed by voices of individuals as diverse as those it is supposed to serve. Therefore, we expect the proposed investigations of views on recovery by previously under-researched ethnic, language, and cultural groups to yield valuable insights on the appropriateness of the recovery approach for a wider population. The datasets collected in this project can serve as useful resources for future research. More generally, our social-media data-driven approach could be applied to investigate other areas of mental health if it proves successful in leading to relevant new insights.

Finally, this project is an interdisciplinary endeavour, combining clinical psychology, input from individuals with lived experience of BD, and computational linguistics. While this comes with the challenges of cross-disciplinary research, it has the potential to apply and develop state-of-the-art NLP methods in a way that is psychologically and ethically sound as well as informed and approved by affected people to increase our knowledge of severe mental illnesses such as BD.

Acknowledgments↩︎

I would like to thank my supervisors Steven Jones, Fiona Lobban, and Paul Rayson for their guidance in this project. My heartfelt thanks go also to Chris Lodge, service user researcher at the Spectrum Centre, and the members of the advisory panel he coordinates that offer feedback on this project based on their lived experience of BD. Further, I would like to thank Masoud Rouhizadeh for his helpful comments during pre-submission mentoring and the anonymous reviewers. This project is funded by the Faculty of Health and Medicine at Lancaster University as part of a doctoral scholarship.

References↩︎

[1]
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa. 2011. https://doi.org/10.1.1.231.4614. Journal of Machine Learning Research, 12:2493–2537.
[2]
Daniel Zeman, Jan Hajic, Martin Popel, Milan Straka, Joakim Nivre, Filip Ginter, Slav Petrov, and Stephan Oepen. 2018. . In The SIGNLL Conference on Computational Natural Language Learning.
[3]
Karin Becker, Viviane P. Moreira, and Aline G. L. dos Santos. 2017. https://doi.org/10.1016/j.ipm.2016.12.008. Information Processing and Management, 53(3):684–704.
[4]
Jeremy Barnes, Roman Klinger, and Sabine Schulte im Walde. 2017. https://doi.org/10.18653/V1/W17-5202. In Proceedings of the 8th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, pages 2–12.
[5]
Jeremy Barnes, Roman Klinger, and Sabine Schulte im Walde. 2018. http://arxiv.org/abs/1805.09016 http://aclweb.org/anthology/C18-1070. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL), pages 2483–2493, Melbourne.
[6]
Jeremy Barnes, Roman Klinger, and Sabine Schulte im Walde. 2018. http://arxiv.org/abs/1806.04381. In Proceedings of the 27th International Conference on Computational Linguistics, pages 818–830.
[7]
Emanuele Pianta, Luisa Bentivogli, and Christian Girardi. 2002. http://multiwordnet.fbk.eu/paper/MWN-India-published.pdfMultiWordNet: developing an aligned multilingual database. In Proceedings of the 1st International WordNet Conference, pages 293–302.
[8]
Hans C. Boas, editor. 2009. https://doi.org/10.1093/ijl/ecp034 Mouton de Gruyter, Berlin.
[9]
Scott Piao, Paul Rayson, Dawn Archer, Francesca Bianchi, Carmen Dayrell, Ricardo-maría Jiménez, Dawn Knight, Michal Křen, Laura Löfberg, Muhammad Adeel Nawab, Jawad Shafi, Phoey Lee Teh, and Olga Mudraya. 2016. . Tenth International Conference on Language Resources and Evaluation, pages 2614–2619.
[10]
Peter Boot, Hanna Zijlstra, and Rinie Geenen. 2017. https://doi.org/10.1075/dujal.6.1.04boo. Dutch Journal of Applied Linguistics, 6(1):65 – 76.
[11]
Simon Robertson Stuart, Louise Tansey, and Ethel Quayle. 2017. https://doi.org/10.1080/09638237.2016.1222056. Journal of Mental Health, 26(3):291–304.
[12]
K. N. Roy Chengappa, John Hennen, Ross J. Baldessarini, David J. Kupfer, Lakshmi N. Yatham, Samuel Gershon, Robert W. Baker, and Mauricio Tohen. 2005. https://doi.org/10.1111/j.1399-5618.2004.00171.x. Bipolar Disorders, 7(1):68–76.
[13]
Peter Forster. 2014. https://doi.org/https://doi.org/10.1016/B978-0-12-385157-4.01077-0Bipolar Disorder. Encyclopedia of the Neurological Sciences, pages 420–424.
[14]
Ann Heylighen, Herman Neuckermans, Yoko Akazawa-Ogawa, Mototada Shichiri, Keiko Nishio, Yasukazu Yoshida, Etsuo Niki, Yoshihisa Hagihara, R. C. Dempsey, P. A. Gooding, Steven Huntley Jones, Nadia Akers, Jayne Eaton, Elizabeth Tyler, Amanda Gatherer, Alison Brabban, Rita Marie Long, Anne Fiona Lobban, Raya A. Jones, Kallia Apazoglou, Anne-Lise Küng, Paolo Cordera, Jean-Michel Aubry, Alexandre Dayer, Patrik Vuilleumier, Camille Piguet, Russell S.J., Prof Steven Jones, Anne Cooke, Anne Cooke, Karin Falk, I. Marshal, Steven Huntley Jones, Gina Smith, Lee D Mulligan, Fiona Lobban, Heather Law, Graham Dunn, Mary Welford, James Kelly, John Mulligan, Anthony P Morrison, Elizabeth Tyler, Anne Fiona Lobban, Chris Sutton, Colin Depp, Sheri L Johnson, Ken Laidlaw, Steven Huntley Jones, Greg Murray, Nuwan D Leitan, Neil Thomas, Erin E Michalak, Sheri L Johnson, Steven Huntley Jones, Tania Perich, Lesley Berk, Michael Berk, Timothy H. Monk, Joseph F. Flaherty, Ellen Frank, Kathleen Hoskinson, David J. Kupfer, Ailbhe Spillane, Karen Matvienko-Sikar, Celine Larkin, Paul Corcoran, and Ella Arensman. 2014. https://doi.org/10.1016/j.cpr.2017.01.002, volume 7. National Institute for Health and Care Excellence.
[15]
U.S. Department of Health and Human Services: The National Institute of Mental Health. 2016. https://www.nimh.nih.gov/health/topics/schizophrenia/index.shtmlSchizophrenia.
[16]
Stephen M. Strakowski, Paul E. Keck, Susan L. McElroy, Scott A. West, Kenji W. Sax, John M. Hawkins, Geri F. Kmetz, Vidya H. Upadhyaya, Karen C. Tugrul, and Michelle L. Bourne. 1998. https://doi.org/10.1001/archpsyc.55.1.49. Archives of General Psychiatry, 55(1):49–55.
[17]
Mauricio Tohen, Carlos A. Zarate, John Hennen, Hari Mandir Kaur Khalsa, Stephen M. Strakowski, Priscilla Gebre-Medhin, Paola Salvatore, and Ross J. Baldessarini. 2003. https://doi.org/10.1176/appi.ajp.160.12.2099. American Journal of Psychiatry, 160(12):2099–2107.
[18]
Patricia E. Deegan. 1988. https://doi.org/10.1037/h0099565Psychosocial Rehabilitation Journal, 11(4):11–19.
[19]
William A. Anthony. 1993. . Psychosocial Rehabilitation Journal, 16(4):11–23.
[20]
Retta Andresen, Peter Caputi, and Lindsay G Oades. 2011. https://doi.org/10.1002/9781119975182. John Wiley & Sons, Ltd, Chichester, West Sussex.
[21]
Jim van Os, Sinan Guloksuz, Thomas Willem Vijn, Anton Hafkenscheid, and Philippe Delespaul. 2019. https://doi.org/10.1002/wps.20609World Psychiatry, 18(1):88–96.
[22]
Sharon L. Young and David S. Ensing. 1999. http://ezproxy.deakin.edu.au/login?url=http://search.ebscohost.com/login.aspx?direct=true&db=cinref&AN=PRJ.BB.BAI.YOUNG.ERFPPP&site=ehost-liveExploring recovery from the perspective of people with psychiatric disabilities.Psychiatric Rehabilitation Journal, 22(3):219–231.
[23]
Warren Mansell, Seth Powell, Rebecca Pedley, Nia Thomas, and Sarah Amelia Jones. 2010. https://doi.org/10.1348/014466509X451447. British Journal of Clinical Psychology, 49(2):193–215.
[24]
Anthony P. Morrison, Heather Law, Christine Barrowclough, Richard P. Bentall, Gillian Haddock, Steven Huntley Jones, Martina Kilbride, Elizabeth Pitt, Nicholas Shryane, Nicholas Tarrier, Mary Welford, and Graham Dunn. 2016. https://doi.org/10.3310/pgfar04050. Programme Grants for Applied Research, 4(5):1–272.
[25]
Mary Leamy, Victoria Bird, Clair Le Boutillier, Julie Williams, and Mike Slade. 2011. https://doi.org/10.1192/bjp.bp.110.083733. British Journal of Psychiatry, 199(6):445–452.
[26]
Steven Jones, Fiona Lobban, and Anne Cook. 2010. https://www1.bps.org.uk/system/files/Public files/cat-653.pdfUnderstanding Bipolar Disorder - Why some people experience extreme mood states and what can help. British Psychological Society.
[27]
Kathleen R. Merikangas, Robert Jin, Jian-ping He, Ronald C. Kessler, Sing Lee, Nancy A. Sampson, Maria Carmen Viana, Laura Helena Andrade, Chiyi Hu, Elie G. Karam, Maria Ladea, Maria Elena Medina Mora, Mark Oakley Browne, Yutaka Ono, Jose Posada-Villa, Rajesh Sagar, and Zahari Zarkov. 2011. https://doi.org/10.1001/archgenpsychiatry.2011.12Prevalence and correlates of bipolar spectrum disorder in the world mental health survey initiative. Archives of general psychiatry, 68(3):241–251.
[28]
Danielle M. Novick, Holly A. Swartz, and Ellen Frank. 2010. https://doi.org/10.1111/j.1399-5618.2009.00786.x. Bipolar disorders, 12(1):1–9.
[29]
William Coryell, Carolyn Turvey, Jean Endicott, Andrew C. Leon, Timothy Mueller, David Solomon, and Martin Keller. 1998. https://doi.org/10.1016/S0165-0327(98)00043-3. Journal of Affective Disorders, 50(2-3):109–116.
[30]
Joseph F. Goldberg and Martin Harrow. 2004. https://doi.org/10.1016/S0165-0327(03)00161-7. Journal of Affective Disorders, 81(2):123–131.
[31]
Steven Jones, Lee D. Mulligan, Sally Higginson, Graham Dunn, and Anthony P Morrison. 2012. https://doi.org/10.1016/j.jad.2012.10.003. Journal of Affective Disorders, 147(1-3):34–43.
[32]
Steven H. Jones, Gina Smith, Lee D. Mulligan, Fiona Lobban, Heather Law, Graham Dunn, Mary Welford, James Kelly, John Mulligan, and Anthony P. Morrison. 2015. https://doi.org/10.1192/bjp.bp.113.141259. British Journal of Psychiatry, 206(1):58–66.
[33]
M. Slade, M. Leamy, F. Bacon, M. Janosik, C. Le Boutillier, J. Williams, and V. Bird. 2012. https://doi.org/10.1017/S2045796012000133. Epidemiology and Psychiatric Sciences, 21(4):353–364.
[34]
Paul Mackin, Steven D. Targum, Amir Kalali, Dror Rom, and Allan H. Young. 2006. https://doi.org/10.1192/bjp.bp.105.013920. British Journal of Psychiatry, 189(04):379–380.
[35]
Marsal Sanches and Miguel Roberto Jorge. 2004. . Brazilian Journal of Psychiatry, 26(3):54–56.
[36]
Yulia E. Chentsova-Dutton, Andrew G. Ryder, and Jeanne Tsai. 2014. https://doi.org/10.1017/CBO9781107415324.004. In Ian H. Gotlib and Constance L. Hammen, editors, Handbook of Depression, pages 337–354. Guilford Press.
[37]
Kate Loveys, Jonathan Torrez, Alex Fine, Glen Moriarty, and Glen Coppersmith. 2018. https://doi.org/10.1002/jcu.10150. In Proceedings of the Fifth Workshop on Computational Linguistics and Clinical Psychology: From Keyboard to Clinic, pages 78–87.
[38]
Munmun De Choudhury, Tomaz Logar, Sanket S. Sharma, Wouter Eekhout, and René Clausen Nielsen. 2017. https://doi.org/10.1145/2998181.2998220. In Proceedings of the 2017 ACM Conference on Computer Supported Cooperative Work and Social Computing, pages 353–369.
[39]
Daniel Jurafsky and James H. Martin. 2009. Speech and Language Processing (2nd Edition). Prentice-Hall, Inc., Upper Saddle River, USA.
[40]
Anne O’Keeffe and Michael McCarthy. 2010. https://doi.org/10.1109/IEMBS.2010.5626267. Routledge Handbooks in Applied Linguistics. Routledge.
[41]
Tony McEnery and Andrew Hardie. 2011. https://books.google.co.uk/books?id=3j3Wn_ZT1qwCCorpus Linguistics: Method, Theory and Practice. Cambridge Textbooks in Linguistics. Cambridge University Press.
[42]
Paul Rayson. 2015. . In Douglas Biber and Randi Reppen, editors, The Cambridge Handbook of English corpus linguistics, pages 32–49. Cambridge University Press.
[43]
Philip Resnik, Rebeca Resnik, and Margaret Mitchell. 2014. https://doi.org/10.1083/jcb.106.5.1795. Association for Computational Linguistics.
[44]
Adrian Benton, Margaret Mitchell, and Dirk Hovy. 2017. https://doi.org/10.1890/06-0645.1. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (EACL), volume 1, pages 152–162.
[45]
Alina Arseniev-Koehler, Sharon Mozgai, and Stefan Scherer. 2018. https://doi.org/10.18653/v1/W18-0601. In Proceedings of the Fifth Workshop on Computational Linguistics and Clinical Psychology: From Keyboard to Clinic, pages 1–12.
[46]
Glen Coppersmith, Mark Dredze, Craig Harman, and Kristy Hollingshead. 2015. https://doi.org/10.1890/04-0298. In Conference of the North American Chapter of the Association for Computational Linguistics – Human Language Technologies (NAACL), pages 1–10.
[47]
Ivan Sekulić, Matej Gjurković, and Jan Šnajder. 2018. https://doi.org/10.18653/v1/P17. In WASSA@EMNLP, 2001, pages 72–78, Brussels. Association for Computational Linguistics.
[48]
Arman Cohan, Bart Desmet, Sean Macavaney, Andrew Yates, Luca Soldaini, Sean Macavaney, and Nazli Goharian. 2018. https://www.aclweb.org/anthology/C18-1126. In Proceedings of the 27th International Conference on Computational Linguistics (COLING), pages 1485––1497, Santa Fe. Association for Computational Linguistics.
[49]
Mitja Trampus. 2015. https://blog.twitter.com/2015/evaluating-language-identification-performanceEvaluating language identification performance.
[50]
Zhiyuan Cheng, James Caverlee, and Kyumin Lee. 2010. https://doi.org/10.1145/1871437.1871535. Proceedings of the 19th ACM International Conference on Information and Knowledge Management, pages 759–768.
[51]
Brent Hecht, Lichan Hong, Bongwon Suh, and Ed H. Chi. 2011. https://doi.org/10.1145/1978942.1978976. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, pages 237–246.
[52]
Keith Harrigian. 2018. https://praw.readthedocs.io/en/latest/. In Proceedings of the 2018 EMNLP Workshop W-NUT: The 4th Workshop on Noisy User-generated Text, pages 17–27.
[53]
Abbas Tashakkori and Charles Teddlie. 1998. Mixed methodology: Combining qualitative and quantitative approaches, volume 46. Sage.
[54]
John W. Creswell and Vicki L. Plano Clark. 2011. https://books.google.ch/books?id=YcdlPWPJRBcCDesigning and Conducting Mixed Methods Research. SAGE Publications.
[55]
Allan Steckler, Kenneth R. McLeroy, Robert M. Goodman, Sheryl T. Bird, and Lauri McCormick. 1992. https://doi.org/10.1177/109019819201900101. Health Education & Behavior, 19(1):1–8.
[56]
Frances Baum. 1995. https://doi.org/10.1016/0277-9536(94)E0103-Y. Social Science and Medicine, 40(4):459–468.
[57]
Joanna E. M. Sale, Lynne H. Lohfeld, and Kevin Brazil. 2002. https://doi.org/10.1023/A:1014301607592. Quality & Quantity, 36:43–53.
[58]
Thorleif Lund. 2012. https://doi.org/10.1080/00313831.2011.568674. Scandinavian Journal of Educational Research, 56(2):155–165.
[59]
Bethany L. Leonhardt, Kelsey Huling, Jay A. Hamm, David Roe, Ilanit Hasson-Ohayon, Hamish J. McLeod, and Paul H. Lysaker. 2017. https://doi.org/10.1080/14737175.2017.1378099. Expert Review of Neurotherapeutics, 17(11):1117–1130.
[60]
Sarah J. Russell and Jan L. Browne. 2005. . Australian & New Zealand Journal of Psychiatry, 39(3):187–193.
[61]
Marie Crowe and Maree Inder. 2018. https://doi.org/10.1111/jpm.12455. Journal of Psychiatric and Mental Health Nursing, 25(4):236–244.
[62]
Franco Moretti. 2013. Distant reading. Verso, London.
[63]
Paul Rayson. 2008. https://doi.org/10.1075/ijcl.13.4.06ray. International Journal of Corpus Linguistics, 13(4):519–549.
[64]
Elena Semino, Zsófia Demjén, Andrew Hardie, Sheila Payne, and Paul Rayson. 2017. https://doi.org/10.4324/9781315629834.
[65]
Paul Rayson, Dawn Archer, Scott Piao, and Tony McEnery. 2004. http://eprints.lancs.ac.uk/1783/Proceedings of the beyond named entity recognition semantic labelling for NLP tasks workshop, (February 2017):7–12.
[66]
Alessandro Lenci. 2008. http://linguistica.sns.it/RdL/20.1/ALenci.pdfDistributional semantics in linguistic and cognitive research. Italian Journal of Linguistics, 20(1):1–31.
[67]
Peter D. Turney and Patrick Pantel. 2010. . Journal of Artificial Intelligence Research, 37:141–188.
[68]
James W. Pennebaker, Ryan L. Boyd, Kayla Jordan, and Kate Blackburn. 2015. https://doi.org/10.15781/T29G6Z. Technical report, University of Texas at Austin, Austin.
[69]
Allison M. Tackman, David A. Sbarra, Angela L. Carey, M. Brent Donnellan, Andrea B. Horn, Nicholas S. Holtzman, To’Meisha S. Edwards, James W. Pennebaker, and Matthias R. Mehl. 2018. https://doi.org/10.1037/pspp0000187. Journal of Personality and Social Psychology, (March).
[70]
Zijian Wang and David Jurgens. 2018. http://www.aclweb.org/anthology/D18-1004. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 33–45.
[71]
H. Andrew Schwartz, Salvatore Giorgi, Maarten Sap, Patrick Crutchley, Johannes C. Eichstaedt, and Lyle Ungar. 2017. https://doi.org/10.18653/v1/d17-2010. In Proceedings of the 2017 EMNLP System Demonstrations, pages 55–60.
[72]
Laura-Ana-Maria Ana Maria Bostan and Roman Klinger. 2018. https://doi.org/10.17226/11340. In Proceedings of the 27th International Conference on Computational Linguistics, pages 2104–2119. Association for Computational Linguistics.
[73]
Dirk Hovy and Shannon L. Spruit. 2016. https://doi.org/10.18653/v1/P16-2096o. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (ACL), pages 591–598.
[74]
Dirk Hovy, Shannon Spruit, Margaret Mitchell, Emily M Bender, Michael Strube, and Hanna Wallach. 2017. http://aclweb.org/anthology/W17-1600. Association for Computational Linguistics.
[75]
Hans Jonas. 1984. The Imperative of Responsibility: Foundations of an Ethics for the Technological Age. University of Chicago Press, Chicago.
[76]
Gunther Eysenbach and James E. Till. 2001. https://doi.org/10.1136/bmj.313.7055.438. BMJ, 323(7055):1103–1105.
[77]
Kelsey Beninger, Alexandra Fry, Natalie Jago, Hayley Lepps, Laura Nass, and Hannah Silvester. 2014. http://www.natcen.ac.uk/media/282288/p0639-research-using-social-media-report-final-190214.pdfResearch using Social Media; Users’ Views.
[78]
Michael J. Paul and Mark Dredze. 2017. https://doi.org/10.2200/S00791ED1V01Y201707ICR060. Synthesis Lectures on Information Concepts, Retrieval, and Services, 9(5):1–183.
[79]
Wasim Ahmed, Peter A. Bath, and Gianluca Demartini. 2017. https://doi.org/10.1016/j.jocn.2005.03.017. In Kandy Woodfield, editor, The Ethics of Online Research, pages 79–107. Emerald Books.
[80]
Adrian Benton, Glen Coppersmith, and Mark Dredze. 2017. https://doi.org/10.18653/v1/W17-1612. Proceedings of the First Workshop on Ethics in Natural Language Processing, page 94–102.
[81]
Matthew L. Williams, Pete Burnap, and Luke Sloan. 2017. https://doi.org/10.1177/0038038517708140. Sociology, 51(6):1149–1168.
[82]
Bridianne O’Dea, Stephen Wan, Philip J. Batterham, Alison L. Calear, Cecile Paris, and Helen Christensen. 2015. https://doi.org/10.1016/j.invent.2015.03.005. Internet Interventions, 2(2):183–188.
[83]
Sean D. Young and Renee Garett. 2018. https://doi.org/10.2196/mental.8971. Journal of Medical Internet Research, 20(5):1–5.
[84]
Mitchell J. Prinstein, Matthew K. Nock, Valerie Simon, Julie Wargo Aikins, Charissa S. L. Cheah, and Anthony Spirito. 2008. https://doi.org/10.1037/0022-006X.76.1.92.LongitudinalLongitudinal Trajectories and Predictors of Adolescent Suicidal Ideation and Attempts Following Inpatient Hospitalization. Journal of Consulting and Clinical Psychology, 76(1):92–103.

  1. Email communication with the first author of  [46].↩︎

  2. http://ir.cs.georgetown.edu/data/smhd/↩︎

  3. See Section 4 for ethical considerations on this.↩︎

  4. https://github.com/CLD2Owners/cld2↩︎

  5. https://ourdatahelps.org/↩︎

  6. https://7cups.com/↩︎

  7. https://cdn.cms-twdigitalassets.com/content/dam/legal-twitter/site-assets/privacy-policy-new/Privacy-Policy-Terms-of-Service_EN.pdf↩︎

  8. www.redditinc.com/policies/privacy-policy↩︎