We introduce REALTIME QA, a dynamic question answering (QA) platform that announces questions and evaluates systems on a regular basis (weekly in this version). REALTIME QA inquires about the current world, and QA systems need to answer questions about novel events or information. It therefore challenges static, conventional assumptions in open-domain QA datasets and pursues instantaneous applications. We build strong baseline models upon large pretrained language models, including GPT-3 and T5. Our benchmark is an ongoing effort, and this paper presents real-time evaluation results over the past year. Our experimental results show that GPT-3 can often properly update its generation results, based on newly-retrieved documents, highlighting the importance of up-to-date information retrieval. Nonetheless, we find that GPT-3 tends to return outdated answers when retrieved documents do not provide sufficient information to find an answer. This suggests an important avenue for future research: can an open-domain QA system identify such unanswerable cases and communicate with the user or even the retrieval module to modify the retrieval results? We hope that REALTIME QA will spur progress in instantaneous applications of question answering and beyond.
翻译:我们介绍了实时问答(REALTIME QA),这是一个动态问答平台,定期(本版本为每周)发布问题并评估系统。实时问答关注当下世界,问答系统需要回答关于新事件或信息的问题。因此,它挑战了开放域问答数据集中静态、传统的假设,并追求即时应用。我们基于大型预训练语言模型(包括GPT-3和T5)构建了强大的基线模型。我们的基准是一项持续进行的工作,本文展示了过去一年的实时评估结果。实验结果表明,基于新检索的文档,GPT-3通常能适当地更新其生成结果,突显了最新信息检索的重要性。然而,我们发现当检索到的文档无法提供足够信息来找到答案时,GPT-3倾向于返回过时的答案。这为未来研究指明了一个重要方向:开放域问答系统能否识别此类不可回答的情况,并与用户甚至检索模块进行交互以修改检索结果?我们希望实时问答能推动问答即时应用及更广泛领域的发展。