Diversity-aware data are essential for a robust modeling of human behavior in context. In addition, being the human behavior of interest for numerous applications, data must also be reusable across domain, to ensure diversity of interpretations. Current data collection techniques allow only a partial representation of the diversity of people and often generate data that is difficult to reuse. To fill this gap, we propose a data collection methodology, within a hybrid machine-artificial intelligence approach, and its related dataset, based on a comprehensive ontological notion of context which enables data reusability. The dataset has a sample of 158 participants and is collected via the iLog smartphone application. It contains more than 170 GB of subjective and objective data, which comes from 27 smartphone sensors that are associated with 168,095 self-reported annotations on the participants context. The dataset is highly reusable, as demonstrated by its diverse applications.
翻译:多样性感知数据对于在上下文中对人类行为进行稳健建模至关重要。此外,由于人类行为在众多应用中具有研究价值,数据必须跨领域可重用,以确保解释的多样性。当前的数据收集技术仅能部分表征人群的多样性,且往往产生难以复用的数据。为填补这一空白,我们提出了一种混合机器与人工智能方法的数据收集方法论,及其相关数据集,该方法基于一种全面的本体论上下文概念,能够实现数据的可重用性。该数据集包含158名参与者的样本,并通过iLog智能手机应用程序收集。它包含超过170GB的主观与客观数据,这些数据来自27个智能手机传感器,并与168095个关于参与者上下文的自我报告注释相关联。该数据集具有高度的可重用性,其多样化的应用场景已证明了这一点。