Generative Foundation Models (GenFMs) have emerged as transformative tools. However, their widespread adoption raises critical concerns regarding trustworthiness across dimensions. This paper presents a comprehensive framework to address these challenges through three key contributions. First, we systematically review global AI governance laws and policies from governments and regulatory bodies, as well as industry practices and standards. Based on this analysis, we propose a set of guiding principles for GenFMs, developed through extensive multidisciplinary collaboration that integrates technical, ethical, legal, and societal perspectives. Second, we introduce TrustGen, the first dynamic benchmarking platform designed to evaluate trustworthiness across multiple dimensions and model types, including text-to-image, large language, and vision-language models. TrustGen leverages modular components--metadata curation, test case generation, and contextual variation--to enable adaptive and iterative assessments, overcoming the limitations of static evaluation methods. Using TrustGen, we reveal significant progress in trustworthiness while identifying persistent challenges. Finally, we provide an in-depth discussion of the challenges and future directions for trustworthy GenFMs, which reveals the complex, evolving nature of trustworthiness, highlighting the nuanced trade-offs between utility and trustworthiness, and consideration for various downstream applications, identifying persistent challenges and providing a strategic roadmap for future research. This work establishes a holistic framework for advancing trustworthiness in GenAI, paving the way for safer and more responsible integration of GenFMs into critical applications. To facilitate advancement in the community, we release the toolkit for dynamic evaluation.


翻译:生成式基础模型(GenFMs)已成为变革性工具。然而,其广泛应用引发了跨维度的可信性关键问题。本文通过三项核心贡献,提出一个综合性框架以应对这些挑战。首先,我们系统性地回顾了全球政府及监管机构的人工智能治理法律与政策,以及行业实践与标准。基于这一分析,我们通过整合技术、伦理、法律及社会视角的广泛多学科协作,提出了一套面向GenFMs的指导原则。其次,我们引入了TrustGen——首个旨在跨多维度及模型类型(包括文本到图像、大语言模型及视觉语言模型)评估可信性的动态基准平台。TrustGen利用模块化组件——元数据整理、测试用例生成及上下文变异——实现自适应与迭代评估,克服了静态评估方法的局限。借助TrustGen,我们揭示了可信性方面的显著进展,同时识别了持续存在的挑战。最后,我们深入探讨了可信GenFMs面临的挑战与未来方向,揭示了可信性复杂且不断演化的本质,强调了效用与可信性之间微妙的权衡,以及针对各类下游应用的考量,识别了持续存在的挑战,并为未来研究提供了战略路线图。本工作为推进GenAI的可信性建立了整体框架,为GenFMs更安全、更负责任地融入关键应用铺平了道路。为促进社区发展,我们开源了动态评估工具包。

0
下载
关闭预览

相关内容

ACM/IEEE第23届模型驱动工程语言和系统国际会议,是模型驱动软件和系统工程的首要会议系列,由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来,模型涵盖了建模的各个方面,从语言和方法到工具和应用程序。模特的参加者来自不同的背景,包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛,参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会,并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。 官网链接:http://www.modelsconference.org/
可靠且负责任的基础模型:全面综述
专知会员服务
20+阅读 · 2月10日
以人为中心的基础模型:感知、生成与代理建模
专知会员服务
24+阅读 · 2025年2月13日
生成式检索基础
专知会员服务
18+阅读 · 2025年1月28日
生成式建模:综述
专知会员服务
33+阅读 · 2025年1月13日
深度学习模型可解释性的研究进展
专知
26+阅读 · 2020年8月1日
AAAI 2020 | 多模态基准指导的生成式多模态自动文摘
AI科技评论
16+阅读 · 2020年1月5日
斯坦福CS236-深度生成模型2019-全套课程资料分享
深度学习与NLP
20+阅读 · 2019年8月20日
展望:模型驱动的深度学习
人工智能学家
12+阅读 · 2018年1月23日
VAE、GAN、Info-GAN:全解深度学习三大生成模型
数据派THU
20+阅读 · 2017年9月23日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
VIP会员
最新内容
非对称防御中的自组织临界性:俄乌战争
专知会员服务
1+阅读 · 13分钟前
《战争中的大语言模型监管》
专知会员服务
1+阅读 · 23分钟前
边缘计算的军事应用
专知会员服务
7+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
9+阅读 · 8月8日
相关基金
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员