Identifying Chinese Microblog Users With High Suicide Probability Using Internet-Based Profile and Linguistic Features: Classification Model
Open Access
- 12 May 2015
- journal article
- Published by JMIR Publications Inc. in JMIR Mental Health
- Vol. 2 (2), e17
- https://doi.org/10.2196/mental.4227
Abstract
Background: Traditional offline assessment of suicide probability is time consuming and difficult in convincing at-risk individuals to participate. Identifying individuals with high suicide probability through online social media has an advantage in its efficiency and potential to reach out to hidden individuals, yet little research has been focused on this specific field.Objective: The objective of this study was to apply two classification models, Simple Logistic Regression (SLR) and Random Forest (RF), to examine the feasibility and effectiveness of identifying high suicide possibility microblog users in China through profile and linguistic features extracted from Internet-based data.Methods: There were nine hundred and nine Chinese microblog users that completed an Internet survey, and those scoring one SD above the mean of the total Suicide Probability Scale (SPS) score, as well as one SD above the mean in each of the four subscale scores in the participant sample were labeled as high-risk individuals, respectively. Profile and linguistic features were fed into two machine learning algorithms (SLR and RF) to train the model that aims to identify high-risk individuals in general suicide probability and in its four dimensions. Models were trained and then tested by 5-fold cross validation; in which both training set and test set were generated under the stratified random sampling rule from the whole sample. There were three classic performance metrics (Precision, Recall, F1 measure) and a specifically defined metric “Screening Efficiency” that were adopted to evaluate model effectiveness.Results: Classification performance was generally matched between SLR and RF. Given the best performance of the classification models, we were able to retrieve over 70% of the labeled high-risk individuals in overall suicide probability as well as in the four dimensions. Screening Efficiency of most models varied from 1/4 to 1/2. Precision of the models was generally below 30%.Conclusions: Individuals in China with high suicide probability are recognizable by profile and text-based information from microblogs. Although there is still much space to improve the performance of classification models in the future, this study may shed light on preliminary screening of risky individuals via machine learning algorithms, which can work side-by-side with expert scrutiny to increase efficiency in large-scale-surveillance of suicide probability from online social media.Keywords
This publication has 38 references indexed in Scilit:
- The Representation of Suicide on the Internet: Implications for CliniciansJournal of Medical Internet Research, 2012
- Overcoming the Clinical–MR Imaging Paradox of Multiple Sclerosis: MR Imaging Data Assessed with a Random Forest ApproachAmerican Journal of Neuroradiology, 2011
- Hyperlinked SuicideCrisis, 2011
- Internet monitoring of suicide risk in the populationJournal of Affective Disorders, 2010
- Connecting the invisible dots: Reaching lesbian, gay, and bisexual adolescents and young adults at risk for suicide through online social networksSocial Science & Medicine (1982), 2009
- Risk factors for the incidence and persistence of suicide-related outcomes: A 10-year follow-up study using the National Comorbidity SurveysJournal of Affective Disorders, 2008
- Associated Factors of Suicide Among University Students: Importance of Family EnvironmentContemporary Family Therapy, 2006
- Life problems and physical illness as risk factors for suicide in older people: a descriptive and case-control studyPsychological Medicine, 2006
- Factors associated with suicidal ideation in an elderly urban Japanese population: A community‐based, cross‐sectional studyPsychiatry and Clinical Neurosciences, 2005
- The Association of Irritability and Impulsivity with Suicidal Ideation Among 15‐ to 20‐year‐old MalesSuicide and Life-Threatening Behavior, 2004