集成方法能否提升证据召回率？一项案例研究 (Can ensembles improve evidence recall? A case study) - 专知论文

会员服务 ·

0

集成 · 集成方法 · 模型评估 · 分析 · 识别 ·

2025 年 12 月 30 日

Can ensembles improve evidence recall? A case study

翻译：集成方法能否提升证据召回率？一项案例研究

Katharina Beckh,Sven Heuser,Stefan Rüping

from arxiv, Submitted to ESANN 2026

Feature attribution methods typically provide minimal sufficient evidence justifying a model decision. However, in many applications, such as compliance and cataloging, the full set of contributing features must be identified: complete evidence. We present a case study using existing language models and a medical dataset which contains human-annotated complete evidence. Our findings show that an ensemble approach, aggregating evidence from several models, improves evidence recall over individual models. We examine different ensemble sizes, the effect of evidence-guided training, and provide qualitative insights.

翻译：特征归因方法通常仅提供证明模型决策所需的最小充分证据。然而，在许多应用场景（如合规审查与信息编目）中，必须识别出所有贡献特征：即完整证据。本文通过一项案例研究，利用现有语言模型和包含人工标注完整证据的医学数据集展开分析。研究发现，集成方法通过聚合多个模型的证据，相较于单一模型能够提升证据召回率。我们考察了不同集成规模的影响，探讨了证据引导训练的效果，并提供了定性分析见解。

0

相关内容

何恺明NeurIPS 2024论文《无条件生成的回归：一种自监督表征生成方法》

何恺明NeurIPS 2024论文《无条件生成的回归：一种自监督表征生成方法》

专知会员服务

21+阅读 · 2024年11月4日

【剑桥大学博士论文】使用机器学习的因果推断中的两个问题的半参数方法

【剑桥大学博士论文】使用机器学习的因果推断中的两个问题的半参数方法

专知会员服务

26+阅读 · 2024年5月25日

【MIT博士论文】基于数据的模型可靠性视角，322页pdf

【MIT博士论文】基于数据的模型可靠性视角，322页pdf

专知会员服务

39+阅读 · 2024年3月25日

集成学习研究现状及展望

集成学习研究现状及展望

专知会员服务

58+阅读 · 2023年7月20日

【AI+军事】附论文《混合决策的证据跟踪》美国海军信息战中心

【AI+军事】附论文《混合决策的证据跟踪》美国海军信息战中心

专知会员服务

71+阅读 · 2022年4月28日

【ICLR 2022】MIT论文解读：谈到人工智能，我们可以抛弃数据集吗？基于ML创建合成数据，Generative Models As A Data Source For Multiview Representation Learning

【ICLR 2022】MIT论文解读：谈到人工智能，我们可以抛弃数据集吗？基于ML创建合成数据，Generative Models As A Data Source For Multiview Representation Learning

专知会员服务

41+阅读 · 2022年3月15日

【O’Reilly讲座】基于深度学习的异常检测方法用于检测大型数据集的质量：Anomaly detection using deep learning to measure the quality of large datasets

【O’Reilly讲座】基于深度学习的异常检测方法用于检测大型数据集的质量：Anomaly detection using deep learning to measure the quality of large datasets

专知会员服务

31+阅读 · 2020年1月11日

【自监督学习新成果】基于对比预测编码的数据高效图像识别（Data-Efficient Image Recognition with Contrastive Predictive Coding）

【自监督学习新成果】基于对比预测编码的数据高效图像识别（Data-Efficient Image Recognition with Contrastive Predictive Coding）

专知会员服务

16+阅读 · 2019年12月10日

【元学习 | 论文】元学习聚类，Meta-Learning to Cluster，哥伦比亚大学

【元学习 | 论文】元学习聚类，Meta-Learning to Cluster，哥伦比亚大学

专知会员服务

42+阅读 · 2019年11月21日

From Data to Model Programming: Injecting Structured Priors for Knowledge Extraction，南加州大学计算机科学系任翔助理教授，CIPS ATT 16（2019）

From Data to Model Programming: Injecting Structured Priors for Knowledge Extraction，南加州大学计算机科学系任翔助理教授，CIPS ATT 16（2019）

专知会员服务

14+阅读 · 2019年10月25日

【2023新书】机器学习集成方法，354页pdf

【2023新书】机器学习集成方法，354页pdf

专知

40+阅读 · 2023年4月11日

基于深度学习的数据融合方法研究综述

基于深度学习的数据融合方法研究综述

专知

37+阅读 · 2020年12月10日

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

专知

11+阅读 · 2020年8月28日

基于深度元学习的因果推断新方法

基于深度元学习的因果推断新方法

图与推荐

12+阅读 · 2020年7月21日

【资源推荐】公开数据集收集汇总

【资源推荐】公开数据集收集汇总

专知

19+阅读 · 2019年6月5日

一文教你如何处理不平衡数据集（附代码）

一文教你如何处理不平衡数据集（附代码）

大数据文摘

11+阅读 · 2019年6月2日

常用的模型集成方法介绍：bagging、boosting 、stacking

常用的模型集成方法介绍：bagging、boosting 、stacking

机器之心

14+阅读 · 2019年5月15日

《小样本学习(Few-shot learning)》最新41页综述论文，来自港科大和第四范式

《小样本学习(Few-shot learning)》最新41页综述论文，来自港科大和第四范式

专知

363+阅读 · 2019年4月12日

机器学习数据集哪里找：优秀数据集来源盘点

机器学习数据集哪里找：优秀数据集来源盘点

云栖社区

12+阅读 · 2019年1月30日

学界 | FAIR提出用聚类方法结合卷积网络，实现无监督端到端图像分类

学界 | FAIR提出用聚类方法结合卷积网络，实现无监督端到端图像分类

机器之心

11+阅读 · 2018年8月6日

面向跨领域异构数据的患者相似性学习方法及应用

国家自然科学基金

23+阅读 · 2016年12月31日

有效融合多源异构数据的集成分类器研究

国家自然科学基金

5+阅读 · 2015年12月31日

生物医疗大数据集成分析的统计与计算方法研究

国家自然科学基金

4+阅读 · 2015年12月31日

基于成像环境约束的低质量图像篡改取证研究

国家自然科学基金

1+阅读 · 2015年12月31日

高维不平衡数据的集成学习算法研究

国家自然科学基金

16+阅读 · 2015年12月31日

关联规则集上的知识发现

国家自然科学基金

9+阅读 · 2015年12月31日

面向DS证据理论的关联信息融合研究

国家自然科学基金

4+阅读 · 2015年12月31日

面向异分布数据的主动学习方法

国家自然科学基金

12+阅读 · 2015年12月31日

多源基因表达数据横向整合的统计方法比较

国家自然科学基金

0+阅读 · 2015年12月31日

面向大规模数据流的集成学习模型与方法研究

国家自然科学基金

5+阅读 · 2014年12月31日

Enabling the Reuse of Personal Data in Research: A Classification Model for Legal Compliance

Arxiv

0+阅读 · 1月29日

Higher-Order Feature Attribution: Bridging Statistics, Explainable AI, and Topological Signal Processing

Arxiv

0+阅读 · 1月28日

Scaling Medical Reasoning Verification via Tool-Integrated Reinforcement Learning

Arxiv

0+阅读 · 1月28日

Multimodal data integration and cross-modal querying via orchestrated approximate message passing

Arxiv

0+阅读 · 1月21日

When Are Two Scores Better Than One? Investigating Ensembles of Diffusion Models

Arxiv

0+阅读 · 1月21日

When Are Two Scores Better Than One? Investigating Ensembles of Diffusion Models

Arxiv

0+阅读 · 1月16日

Show me the evidence: Evaluating the role of evidence and natural language explanations in AI-supported fact-checking

Arxiv

0+阅读 · 1月16日

Difficulty Controlled Diffusion Model for Synthesizing Effective Training Data

Arxiv

0+阅读 · 1月7日

Subsampled Ensemble Can Improve Generalization Tail Exponentially

Arxiv

0+阅读 · 1月3日

Are Ensembles Getting Better all the Time?

Arxiv

0+阅读 · 2025年12月30日

VIP会员

文章信息

相关主题

相关VIP内容

何恺明NeurIPS 2024论文《无条件生成的回归：一种自监督表征生成方法》

何恺明NeurIPS 2024论文《无条件生成的回归：一种自监督表征生成方法》

专知会员服务

21+阅读 · 2024年11月4日

【剑桥大学博士论文】使用机器学习的因果推断中的两个问题的半参数方法

【剑桥大学博士论文】使用机器学习的因果推断中的两个问题的半参数方法

专知会员服务

26+阅读 · 2024年5月25日

【MIT博士论文】基于数据的模型可靠性视角，322页pdf

【MIT博士论文】基于数据的模型可靠性视角，322页pdf

专知会员服务

39+阅读 · 2024年3月25日

集成学习研究现状及展望

集成学习研究现状及展望

专知会员服务

58+阅读 · 2023年7月20日

【AI+军事】附论文《混合决策的证据跟踪》美国海军信息战中心

【AI+军事】附论文《混合决策的证据跟踪》美国海军信息战中心

专知会员服务

71+阅读 · 2022年4月28日

【ICLR 2022】MIT论文解读：谈到人工智能，我们可以抛弃数据集吗？基于ML创建合成数据，Generative Models As A Data Source For Multiview Representation Learning

【ICLR 2022】MIT论文解读：谈到人工智能，我们可以抛弃数据集吗？基于ML创建合成数据，Generative Models As A Data Source For Multiview Representation Learning

专知会员服务

41+阅读 · 2022年3月15日

【O’Reilly讲座】基于深度学习的异常检测方法用于检测大型数据集的质量：Anomaly detection using deep learning to measure the quality of large datasets

【O’Reilly讲座】基于深度学习的异常检测方法用于检测大型数据集的质量：Anomaly detection using deep learning to measure the quality of large datasets

专知会员服务

31+阅读 · 2020年1月11日

【自监督学习新成果】基于对比预测编码的数据高效图像识别（Data-Efficient Image Recognition with Contrastive Predictive Coding）

【自监督学习新成果】基于对比预测编码的数据高效图像识别（Data-Efficient Image Recognition with Contrastive Predictive Coding）

专知会员服务

16+阅读 · 2019年12月10日

【元学习 | 论文】元学习聚类，Meta-Learning to Cluster，哥伦比亚大学

【元学习 | 论文】元学习聚类，Meta-Learning to Cluster，哥伦比亚大学

专知会员服务

42+阅读 · 2019年11月21日

From Data to Model Programming: Injecting Structured Priors for Knowledge Extraction，南加州大学计算机科学系任翔助理教授，CIPS ATT 16（2019）

From Data to Model Programming: Injecting Structured Priors for Knowledge Extraction，南加州大学计算机科学系任翔助理教授，CIPS ATT 16（2019）

专知会员服务

14+阅读 · 2019年10月25日

热门VIP内容

开通专知VIP会员享更多权益服务

通用智能体评估的逻辑架构

《无人机与战争：被忽视的环境影响及无人机保护潜力》

论学习、公平性与复杂度

《整合杀伤链：一个用于边缘目标验证与战术推理的零样本框架》最新资料

相关资讯

【2023新书】机器学习集成方法，354页pdf

【2023新书】机器学习集成方法，354页pdf

专知

40+阅读 · 2023年4月11日

基于深度学习的数据融合方法研究综述

基于深度学习的数据融合方法研究综述

专知

37+阅读 · 2020年12月10日

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

专知

11+阅读 · 2020年8月28日

基于深度元学习的因果推断新方法

基于深度元学习的因果推断新方法

图与推荐

12+阅读 · 2020年7月21日

【资源推荐】公开数据集收集汇总

【资源推荐】公开数据集收集汇总

专知

19+阅读 · 2019年6月5日

一文教你如何处理不平衡数据集（附代码）

一文教你如何处理不平衡数据集（附代码）

大数据文摘

11+阅读 · 2019年6月2日

常用的模型集成方法介绍：bagging、boosting 、stacking

常用的模型集成方法介绍：bagging、boosting 、stacking

机器之心

14+阅读 · 2019年5月15日

《小样本学习(Few-shot learning)》最新41页综述论文，来自港科大和第四范式

《小样本学习(Few-shot learning)》最新41页综述论文，来自港科大和第四范式

专知

363+阅读 · 2019年4月12日

机器学习数据集哪里找：优秀数据集来源盘点

机器学习数据集哪里找：优秀数据集来源盘点

云栖社区

12+阅读 · 2019年1月30日

学界 | FAIR提出用聚类方法结合卷积网络，实现无监督端到端图像分类

学界 | FAIR提出用聚类方法结合卷积网络，实现无监督端到端图像分类

机器之心

11+阅读 · 2018年8月6日

相关论文

Enabling the Reuse of Personal Data in Research: A Classification Model for Legal Compliance

Arxiv

0+阅读 · 1月29日

Higher-Order Feature Attribution: Bridging Statistics, Explainable AI, and Topological Signal Processing

Arxiv

0+阅读 · 1月28日

Scaling Medical Reasoning Verification via Tool-Integrated Reinforcement Learning

Arxiv

0+阅读 · 1月28日

Multimodal data integration and cross-modal querying via orchestrated approximate message passing

Arxiv

0+阅读 · 1月21日

When Are Two Scores Better Than One? Investigating Ensembles of Diffusion Models

Arxiv

0+阅读 · 1月21日

When Are Two Scores Better Than One? Investigating Ensembles of Diffusion Models

Arxiv

0+阅读 · 1月16日

Show me the evidence: Evaluating the role of evidence and natural language explanations in AI-supported fact-checking

Arxiv

0+阅读 · 1月16日

Difficulty Controlled Diffusion Model for Synthesizing Effective Training Data

Arxiv

0+阅读 · 1月7日

Subsampled Ensemble Can Improve Generalization Tail Exponentially

Arxiv

0+阅读 · 1月3日

Are Ensembles Getting Better all the Time?

Arxiv

0+阅读 · 2025年12月30日

相关基金

面向跨领域异构数据的患者相似性学习方法及应用

国家自然科学基金

23+阅读 · 2016年12月31日

有效融合多源异构数据的集成分类器研究

国家自然科学基金

5+阅读 · 2015年12月31日

生物医疗大数据集成分析的统计与计算方法研究

国家自然科学基金

4+阅读 · 2015年12月31日

基于成像环境约束的低质量图像篡改取证研究

国家自然科学基金

1+阅读 · 2015年12月31日

高维不平衡数据的集成学习算法研究

国家自然科学基金

16+阅读 · 2015年12月31日

关联规则集上的知识发现

国家自然科学基金

9+阅读 · 2015年12月31日

面向DS证据理论的关联信息融合研究

国家自然科学基金

4+阅读 · 2015年12月31日

面向异分布数据的主动学习方法

国家自然科学基金

12+阅读 · 2015年12月31日

多源基因表达数据横向整合的统计方法比较

国家自然科学基金

0+阅读 · 2015年12月31日

面向大规模数据流的集成学习模型与方法研究

国家自然科学基金

5+阅读 · 2014年12月31日

微信扫码咨询专知VIP会员