Artificial intelligence (AI) is increasingly used in clinical settings, yet limited oversight and domain expertise can allow algorithmic bias and safety risks to persist. This study evaluates whether an agentic AI system can support auditing biomedical machine learning models for fairness in early-onset colorectal cancer (EO-CRC), a condition with documented demographic disparities. We implemented a two-agent architecture consisting of a Domain Expert Agent that synthesizes literature on EO-CRC disparities and a Fairness Consultant Agent that recommends sensitive attributes and fairness metrics for model evaluation. An ablation study compared three Ollama large language models (8B, 20B, and 120B parameters) across three configurations: pretrained LLM-only, Agent without Retrieval-Augmented Generation (RAG), and Agent with RAG. Across models, the Agent with RAG achieved the highest semantic similarity to expert-derived reference statements, particularly for disparity identification, suggesting agentic systems with retrieval may help scale fairness auditing in clinical AI.
翻译:人工智能(AI)在临床环境中的应用日益广泛,但有限的监管和领域专业知识可能导致算法偏差和安全风险持续存在。本研究评估了智能体AI系统是否能够支持对生物医学机器学习模型进行公平性审计,重点关注早发性结直肠癌(EO-CRC)这一已被证实存在人口统计学差异的疾病。我们实现了一个双智能体架构,包括一个综合EO-CRC差异相关文献的领域专家智能体,以及一个推荐用于模型评估的敏感属性和公平性指标的公平顾问智能体。一项消融研究比较了三种Ollama大语言模型(参数规模分别为8B、20B和120B)在三种配置下的表现:仅使用预训练LLM、无检索增强生成(RAG)的智能体、以及带RAG的智能体。在所有模型中,带RAG的智能体在与专家推导的参考陈述的语义相似度上表现最佳,尤其在差异识别方面,这表明具备检索能力的智能体系统可能有助于扩展临床AI中的公平性审计规模。