Compared to physical health, population mental health measurement in the U.S. is very coarse-grained. Currently, in the largest population surveys, such as those carried out by the Centers for Disease Control or Gallup, mental health is only broadly captured through "mentally unhealthy days" or "sadness", and limited to relatively infrequent state or metropolitan estimates. Through the large scale analysis of social media data, robust estimation of population mental health is feasible at much higher resolutions, up to weekly estimates for counties. In the present work, we validate a pipeline that uses a sample of 1.2 billion Tweets from 2 million geo-located users to estimate mental health changes for the two leading mental health conditions, depression and anxiety. We find moderate to large associations between the language-based mental health assessments and survey scores from Gallup for multiple levels of granularity, down to the county-week (fixed effects $\beta = .25$ to $1.58$; $p<.001$). Language-based assessment allows for the cost-effective and scalable monitoring of population mental health at weekly time scales. Such spatially fine-grained time series are well suited to monitor effects of societal events and policies as well as enable quasi-experimental study designs in population health and other disciplines. Beyond mental health in the U.S., this method generalizes to a broad set of psychological outcomes and allows for community measurement in under-resourced settings where no traditional survey measures - but social media data - are available.
翻译:与身体健康相比,美国人口心理健康的测量粒度非常粗糙。目前,在疾病控制中心或盖洛普等机构开展的最大规模人口调查中,心理健康仅通过"心理不健康天数"或"悲伤"等指标进行粗略捕捉,且局限于相对低频的州或都市区估计。通过对社交媒体数据的大规模分析,我们能够在更高分辨率(直至县级周度估计)上对人口心理健康进行稳健估计。本研究验证了一套流程,该流程使用来自200万地理定位用户的12亿条推文样本,评估两种主要心理健康状况(抑郁和焦虑)的变化。我们发现,基于语言的心理健康评估与盖洛普调查得分在多个粒度层级(直至县级-周度)上存在中等到较强的关联(固定效应$\beta = .25$至$1.58$;$p<.001$)。基于语言的评估方法能够以周为时间尺度,对人口心理健康进行经济高效且可扩展的监测。这种空间上精细的时间序列非常适合监测社会事件和政策的影响,并为人口健康及其他学科领域的准实验研究设计提供支持。除美国心理健康外,该方法还可推广至广泛的心理健康结果,并适用于缺乏传统调查手段但可获得社交媒体数据的资源匮乏地区的社区测量。