This study develops and evaluates a systematic methodology for constructing news datasets from Google News, combining automated web scraping, large language model (LLM)-based metadata extraction, and SCImago Media Rankings enrichment. Using the IFMIF-DONES fusion energy project as a case study, we implemented a five-stage data collection pipeline across 81 region-language combinations, yielding 1,482 validated records after a 56% noise reduction. Results are compared against two licensed press databases: MyNews (2,280 records) and ProQuest Newsstream Collection (148 records). Overlap analysis reveals high complementarity, with 76% of Google News records exclusive to this platform. The dataset captures content types absent from proprietary databases, including specialized outlets, institutional communications, and social media posts. However, significant methodological challenges emerge: temporal instability requiring synchronic collection, a 100-result cap per query demanding multi-stage strategies, and unexpected noise including academic PDFs, false positives, and pornographic content infiltrating results through black hat SEO techniques. LLM-assisted extraction proved effective for structured articles but exhibited systematic hallucination patterns requiring validation protocols. We conclude that Google News offers valuable complementary coverage for communication research but demands substantial methodological investment, multi-source triangulation, and robust filtering mechanisms to ensure dataset integrity.
翻译:本研究开发并评估了一种从Google News构建新闻数据集的系统性方法,该方法融合了自动化网络爬虫、基于大语言模型的元数据提取以及SCImago媒体排名增强技术。以IFMIF-DONES聚变能项目为案例,我们构建了一个涵盖81种区域-语言组合的五阶段数据采集流水线,经56%的噪声削减后,最终获得1,482条有效记录。将结果与两个授权新闻数据库(MyNews的2,280条记录和ProQuest Newsstream Collection的148条记录)进行对比。重叠分析显示高度互补性,其中76%的Google News记录为该平台独有。数据集捕获了专有数据库未涵盖的内容类型,包括专业媒体、机构通讯和社交媒体帖子。然而,方法学面临显著挑战:需同步采集的时间不稳定性、每次查询100条结果上限要求多阶段策略,以及通过黑帽SEO技术混入结果的意外噪声(包括学术PDF、误报及色情内容)。大语言模型辅助提取对结构化文章效果显著,但出现系统性幻觉模式,需建立验证协议。我们得出结论:Google News为传播研究提供了有价值的补充覆盖,但需投入大量方法论工作、采用多源三角验证及稳健的过滤机制以确保数据集完整性。