Large Language Models excel in generative tasks but exhibit inefficiencies in structured text selection, particularly in extractive question answering. This challenge is magnified in resource-constrained environments, where deploying multiple specialized models for different tasks is impractical. We propose a Learning-to-Defer framework that allocates queries to specialized experts, ensuring high-confidence predictions while optimizing computational efficiency. Our approach integrates a principled allocation strategy with theoretical guarantees on optimal deferral that balances performance and cost. Empirical evaluations on SQuADv1, SQuADv2, and TriviaQA demonstrate that our method enhances answer reliability while significantly reducing computational overhead, making it well-suited for scalable and efficient EQA deployment.
翻译:大语言模型在生成任务中表现出色,但在结构化文本选择方面效率低下,尤其是在抽取式问答任务中。这一问题在资源受限环境中更为突出,因为部署多个专用模型处理不同任务并不现实。我们提出一种延迟学习框架,将查询分配给专用专家模型,在确保高置信度预测的同时优化计算效率。该方法融合了原则性的分配策略与基于最优延迟理论保障的性能-成本平衡机制。在SQuADv1、SQuADv2及TriviaQA上的实验评估表明,本方法在显著降低计算开销的同时提升了答案可靠性,适用于可扩展且高效的抽取式问答部署场景。