CIRCLE：基于现实视角的人工智能评估框架 (CIRCLE: A Framework for Evaluating AI from a Real-World Lens)

Reva Schwartz,Carina Westling,Morgan Briggs,Marzieh Fadaee,Isar Nejadgholi,Matthew Holmes,Fariza Rashid,Maya Carlyle,Afaf Taïk,Kyra Wilson,Peter Douglas,Theodora Skeadas,Gabriella Waters,Rumman Chowdhury,Thiago Lacerda

from arxiv, Accepted at Intelligent Systems Conference (IntelliSys) 2026

This paper proposes CIRCLE, a six-stage, lifecycle-based framework to bridge the reality gap between model-centric performance metrics and AI's materialized outcomes in deployment. While existing frameworks like MLOps focus on system stability and benchmarks measure abstract capabilities, decision-makers outside the AI stack lack systematic evidence about the behavior of AI technologies under real-world user variability and constraints. CIRCLE operationalizes the Validation phase of TEVV (Test, Evaluation, Verification, and Validation) by formalizing the translation of stakeholder concerns outside the stack into measurable signals. Unlike participatory design, which often remains localized, or algorithmic audits, which are often retrospective, CIRCLE provides a structured, prospective protocol for linking context-sensitive qualitative insights to scalable quantitative metrics. By integrating methods such as field testing, red teaming, and longitudinal studies into a coordinated pipeline, CIRCLE produces systematic knowledge: evidence that is comparable across sites yet sensitive to local context. This can enable governance based on materialized downstream effects rather than theoretical capabilities.

翻译：本文提出CIRCLE——一个基于生命周期的六阶段框架，旨在弥合模型中心性能指标与AI在部署中实际成效之间的现实鸿沟。尽管现有框架（如MLOps）关注系统稳定性，基准测试衡量抽象能力，但AI技术栈外的决策者仍缺乏关于AI技术在真实世界用户多样性和约束条件下行为的系统性证据。CIRCLE通过将技术栈外部利益相关者的关切形式化为可测量信号，实现了TEVV（测试、评估、验证与确认）中验证阶段的可操作化。与通常局限于局部范围的参与式设计或常属事后追溯的算法审计不同，CIRCLE提供了一种结构化的前瞻性方案，能够将情境敏感的定性洞察与可扩展的定量指标相连接。通过将现场测试、红队演练和纵向研究等方法整合至协调流程中，CIRCLE生成系统性知识：这些证据既具备跨场景可比性，又能保持对局部情境的敏感性。该框架可使治理机制基于实际产生的下游效应而非理论能力建立。

相关内容

关注 7103

人工智能杂志AI(Artificial Intelligence)是目前公认的发表该领域最新研究成果的主要国际论坛。该期刊欢迎有关AI广泛方面的论文，这些论文构成了整个领域的进步，也欢迎介绍人工智能应用的论文，但重点应该放在新的和新颖的人工智能方法如何提高应用领域的性能，而不是介绍传统人工智能方法的另一个应用。关于应用的论文应该描述一个原则性的解决方案，强调其新颖性，并对正在开发的人工智能技术进行深入的评估。官网地址：http://dblp.uni-trier.de/db/journals/ai/

构建面向终端的 AI 编程智能体：脚手架、测试环境、上下文工程及实践经验

专知会员服务

20+阅读 · 3月8日

AI 智能体系统：体系架构、应用场景及评估范式

专知会员服务

66+阅读 · 1月6日

《人工智能暗战：SaaS与边缘计算架构之争》

专知会员服务

14+阅读 · 2025年7月23日

《人工智能系统测试与评估框架》美国防部联合人工智能中心

专知会员服务

82+阅读 · 2024年1月4日