From grading papers to summarizing medical documents, large language models (LLMs) are evermore used for evaluation of text generated by humans and AI alike. However, despite their extensive utility, LLMs exhibit distinct failure modes, necessitating a thorough audit and improvement of their text evaluation capabilities. Here we introduce ALLURE, a systematic approach to Auditing Large Language Models Understanding and Reasoning Errors. ALLURE involves comparing LLM-generated evaluations with annotated data, and iteratively incorporating instances of significant deviation into the evaluator, which leverages in-context learning (ICL) to enhance and improve robust evaluation of text by LLMs. Through this iterative process, we refine the performance of the evaluator LLM, ultimately reducing reliance on human annotators in the evaluation process. We anticipate ALLURE to serve diverse applications of LLMs in various domains related to evaluation of textual data, such as medical summarization, education, and and productivity.
翻译:从批改论文到总结医疗文档,大语言模型(LLM)正越来越多地被用于评估人类和AI生成的文本。然而,尽管具有广泛的实用性,LLM仍表现出明显的失败模式,亟需对其文本评估能力进行系统审计和改进。本文提出ALLURE——一种审计大语言模型理解与推理错误的系统性方法。ALLURE通过比较LLM生成的评估结果与标注数据,将存在显著偏差的实例迭代融入评估器,利用上下文学习(ICL)增强并改进LLM对文本的稳健评估能力。通过这一迭代过程,我们优化了评估器LLM的性能,最终降低了评估流程中对人工标注者的依赖。我们预期ALLURE可服务于LLM在文本数据评估相关领域的多样化应用,如医疗摘要、教育及生产力领域。