Data scarcity is a common obstacle in medical research due to the high costs associated with data collection and the complexity of gaining access to and utilizing data. Synthesizing health data may provide an efficient and cost-effective solution to this shortage, enabling researchers to explore distributions and populations that are not represented in existing observations or difficult to access due to privacy considerations. To that end, we have developed a multi-task self-attention model that produces realistic wearable activity data. We examine the characteristics of the generated data and quantify its similarity to genuine samples with both quantitative and qualitative approaches.
翻译:数据稀缺是医学研究中的常见障碍,其原因是数据采集成本高昂,且获取和利用数据的流程复杂。合成健康数据可能为此短缺提供高效且低成本的解决方案,使研究人员能够探索现有观测数据中未被涵盖或出于隐私考虑难以获取的分布与群体。为此,我们开发了一种多任务自注意力模型,用于生成逼真的可穿戴活动数据。我们通过定量与定性相结合的方法,考察了生成数据的特征,并量化了其与真实样本的相似性。