Being predominant in digital entertainment for decades, video games have been recognized as valuable software artifacts by the software engineering (SE) community just recently. Such an acknowledgment has unveiled several research opportunities, spanning from empirical studies to the application of AI techniques for classification tasks. In this respect, several curated game datasets have been disclosed for research purposes even though the collected data are insufficient to support the application of advanced models or to enable interdisciplinary studies. Moreover, the majority of those are limited to PC games, thus excluding notorious gaming platforms, e.g., PlayStation, Xbox, and Nintendo. In this paper, we propose PlayMyData, a curated dataset composed of 99,864 multi-platform games gathered by IGDB website. By exploiting a dedicated API, we collect relevant metadata for each game, e.g., description, genre, rating, gameplay video URLs, and screenshots. Furthermore, we enrich PlayMyData with the timing needed to complete each game by mining the HLTB website. To the best of our knowledge, this is the most comprehensive dataset in the domain that can be used to support different automated tasks in SE. More importantly, PlayMyData can be used to foster cross-domain investigations built on top of the provided multimedia data.
翻译:电子游戏作为数字娱乐的主导形式已持续数十年,但直到近期才被软件工程领域视作有价值的软件制品。这一认知催生了诸多研究契机,涵盖从实证研究到面向分类任务的人工智能技术应用。为此,研究者已公开多个精选游戏数据集,然而现有采集数据仍不足以支撑先进模型的应用或跨学科研究的开展。此外,多数数据集局限于PC平台,忽略了PlayStation、Xbox、任天堂等主流游戏平台。本文提出PlayMyData——一个基于IGDB网站构建的包含99,864款多平台游戏的精选数据集。通过专用API接口,我们为每款游戏采集了描述、类型、评分、游戏视频链接及截图等相关元数据。进一步地,通过挖掘HLTB网站数据,我们为PlayMyData补充了各游戏的通关耗时信息。据我们所知,这是该领域内最全面的数据集,可支持软件工程中多种自动化任务的研究。更重要的是,PlayMyData依托所提供的多媒体数据,有望推动跨领域交叉研究。