Email importance labeling has long been a critical yet challenging problem for businesses and individuals. Traditional approaches; such as keyword matching, user-defined rules, and sender-based heuristics; demand extensive manual feature engineering and fail to scale effectively or generalize. Recent advances in large language models (LLMs) demonstrate strong potential and a natural fit for this task, offering deep contextual understanding and superior labeling quality. However, using LLM models like GPT-4.1 at enterprise email volumes incurs prohibitive computational costs and hinders real-world deployment. We explore the trade-off space of using alternative labeling schemes as opposed to GPT4.1 scale LLMs, with the goal of achieving near GPT level labeling quality with significantly lower cost. We develop Argo, an enterprise email labeling framework, where we construct a profiler to efficiently search the cost quality trade-off space of labeling and identify cost-efficient alternatives to labeling emails. Additionally, we design an on-demand provisioning scheme to intelligently scale Argo with real time load, to minimize cost increases during peak load inference. Over 3 open-source email datasets, Argo achieves 148-167X inference cost reduction with negligible quality degradation and 20-640000X lower profiling costs, making large-scale, context-aware email labeling practical for enterprises.
翻译:邮件重要性标注长期以来一直是企业和个人面临的关键且具有挑战性的问题。传统方法(如关键词匹配、用户定义规则和基于发件人的启发式方法)需要大量人工特征工程,难以有效扩展或泛化。近期大语言模型(LLM)的进展展示了其在此任务上的强大潜力和天然适配性,提供了深度上下文理解和优越的标注质量。然而,在企业级邮件数据量下使用GPT-4.1等LLM模型会带来高昂的计算成本,阻碍了实际部署。我们探索了使用替代标注方案与GPT-4.1规模LLM之间的成本-质量权衡空间,旨在以显著更低的成本实现接近GPT级别的标注质量。我们开发了Argo——一个企业邮件标注框架,其中构建了剖析器以高效搜索标注中的成本-质量权衡空间,并识别出标注邮件的高性价比替代方案。此外,我们设计了一种按需供应方案,使Argo能够智能地随实时负载进行扩展,以最小化峰值负载推理期间的成本增加。在三个开源邮件数据集上,Argo实现了148-167倍的推理成本降低,标注质量下降可忽略不计,且剖析成本降低20-640,000倍,使得大规模、上下文感知的邮件标注在企业中变得可行。