The FETCH classifier generates follow-up questions to help refine the best match for the applicant's legal problem, using a low-cost ensemble of LLMs. In this paper, we describe an expert attorney and LLM-assisted evaluation of the follow-up question approach in FETCH and show that while low-cost LLMs perform well at classification tasks, generating high-quality plain-language questions in this setting appears to require a more sophisticated and higher-cost model. Through discussion with legal intake workers, we propose a rubric for the evaluation of legal intake classification questions, and we find that prompt engineering alone is not enough to improve question quality for intake purposes. We also find that LLM-as-judge and human ratings diverge. We demonstrate that with the addition of a single high-cost model, GPT-5, the classifier can elicit relevant information from applicants for legal help, and that the questions lead to more accurate performance at classification tasks. We also find uneven fact elicitation across different categories, including domestic violence, at odds with family law screening protocols, suggesting the value of including dedicated screening panels for certain areas of law.
翻译:FETCH分类器利用低成本大型语言模型集成,生成追问问题以辅助优化申请人法律问题的最佳匹配。本文通过专家律师和大型语言模型辅助评估,描述了FETCH中追问问题的实现方法。研究表明,尽管低成本大型语言模型在分类任务中表现良好,但在该场景下生成高质量通俗语言问题仍需更复杂且成本更高的模型。通过与法律接待人员的讨论,我们提出了法律接待分类问题的评估标准,并发现仅凭提示工程不足以提升接待场景中的问题质量。此外,我们注意到基于大型语言模型的评分与人工评分存在偏差。实验证明,引入单一高成本模型GPT-5后,分类器能有效从法律求助申请人处获取相关信息,且这些问题有助于提升分类任务的准确率。我们还发现不同法律类别(包括家庭暴力案件)的事实获取存在不均衡性,这与家庭法律筛查规程存在冲突,表明针对特定法律领域设置专项筛查机制具有重要价值。