Generative Foundation Models (GenFMs) have emerged as transformative tools. However, their widespread adoption raises critical concerns regarding trustworthiness across dimensions. This paper presents a comprehensive framework to address these challenges through three key contributions. First, we systematically review global AI governance laws and policies from governments and regulatory bodies, as well as industry practices and standards. Based on this analysis, we propose a set of guiding principles for GenFMs, developed through extensive multidisciplinary collaboration that integrates technical, ethical, legal, and societal perspectives. Second, we introduce TrustGen, the first dynamic benchmarking platform designed to evaluate trustworthiness across multiple dimensions and model types, including text-to-image, large language, and vision-language models. TrustGen leverages modular components--metadata curation, test case generation, and contextual variation--to enable adaptive and iterative assessments, overcoming the limitations of static evaluation methods. Using TrustGen, we reveal significant progress in trustworthiness while identifying persistent challenges. Finally, we provide an in-depth discussion of the challenges and future directions for trustworthy GenFMs, which reveals the complex, evolving nature of trustworthiness, highlighting the nuanced trade-offs between utility and trustworthiness, and consideration for various downstream applications, identifying persistent challenges and providing a strategic roadmap for future research. This work establishes a holistic framework for advancing trustworthiness in GenAI, paving the way for safer and more responsible integration of GenFMs into critical applications. To facilitate advancement in the community, we release the toolkit for dynamic evaluation.
翻译:生成式基础模型(GenFMs)已成为变革性工具。然而,其广泛应用引发了跨维度的可信性关键问题。本文通过三项核心贡献,提出一个综合性框架以应对这些挑战。首先,我们系统性地回顾了全球政府及监管机构的人工智能治理法律与政策,以及行业实践与标准。基于这一分析,我们通过整合技术、伦理、法律及社会视角的广泛多学科协作,提出了一套面向GenFMs的指导原则。其次,我们引入了TrustGen——首个旨在跨多维度及模型类型(包括文本到图像、大语言模型及视觉语言模型)评估可信性的动态基准平台。TrustGen利用模块化组件——元数据整理、测试用例生成及上下文变异——实现自适应与迭代评估,克服了静态评估方法的局限。借助TrustGen,我们揭示了可信性方面的显著进展,同时识别了持续存在的挑战。最后,我们深入探讨了可信GenFMs面临的挑战与未来方向,揭示了可信性复杂且不断演化的本质,强调了效用与可信性之间微妙的权衡,以及针对各类下游应用的考量,识别了持续存在的挑战,并为未来研究提供了战略路线图。本工作为推进GenAI的可信性建立了整体框架,为GenFMs更安全、更负责任地融入关键应用铺平了道路。为促进社区发展,我们开源了动态评估工具包。