Large language models (LLMs) are increasingly used for analytical tasks, yet their effectiveness in real-world applications remains underexamined, partly due to the opacity of proprietary models. We evaluate ChatGPT (GPT-3.5 and GPT-4) on the practical task of extracting research challenges from a large scholarly corpus in Human-Computer Interaction (HCI). Using a two-step approach, we first apply GPT-3.5 to extract candidate challenges from the 879 papers in the 2023 ACM CHI Conference proceedings, then use GPT-4 to select the most relevant challenges per paper. This process yielded 4,392 research challenges across 113 topics, which we organized through topic modeling and present in an interactive visualization. We compare the identified challenges with previously established HCI grand challenges and the United Nations Sustainable Development Goals, finding both strong alignment in areas such as ethics and accessibility, and gaps in areas such as human-AI collaboration. A task-specific evaluation with human raters confirmed near-perfect agreement that the extracted statements represent plausible research challenges (\k{appa} = 0.97). The two-step approach proved cost-effective at approximately US$50 for the full corpus, suggesting that LLMs offer a practical means for qualitative text analysis at scale, particularly for prototyping research ideas and examining corpora from multiple analytical perspectives.
翻译:大型语言模型(LLMs)越来越多地被用于分析任务,但其在真实应用中的有效性仍缺乏充分检验,部分原因在于专有模型的不透明性。我们评估了ChatGPT(GPT-3.5与GPT-4)在从人机交互(HCI)领域的大规模学术语料库中提取研究挑战这一实际任务中的表现。采用两步法,我们首先应用GPT-3.5从2023年ACM CHI会议论文集的879篇论文中提取候选挑战,随后使用GPT-4为每篇论文筛选最相关的挑战。该流程共获得涵盖113个主题的4,392项研究挑战,我们通过主题建模进行组织,并以交互式可视化形式呈现。我们将识别出的挑战与先前确立的HCI重大挑战及联合国可持续发展目标进行对比,发现两者在伦理、可访问性等领域高度一致,但在人机协作等方面存在缺失。一项针对特定任务、由人类评审员参与的评估证实,提取出的陈述具有近乎完美的研究挑战合理性(\kappa = 0.97)。该两步法在完整语料库上成本约为50美元,显示出高性价比,表明LLMs为大规模定性文本分析提供了实用途径,尤其适用于研究思路原型设计及从多分析视角审视语料库。