Neural network verifiers aim to provide formal guarantees on model behavior, but existing verification benchmarks are fundamentally limited by their lack of ground-truth labels. As a result, verifier evaluation relies on indirect heuristics, which prevents exact scoring and systematic study of verifier failure modes. We address this gap by introducing a reusable framework for generating verification instances whose ground-truth robustness labels are known a priori through analytic construction. Our framework led to the discovery of multiple numeric tolerance concerns and an implementation bug in popular verifiers, highlighting the need for ground-truth labels. Additionally, to systematically study verifier failure modes, we introduce the verification Difficulty Profile, a collection of estimable quantities capturing distinct sources of instance hardness. Using our framework and these profiles, we evaluate five state-of-the-art verifiers and show that different instances stress distinct aspects of the verification pipeline. We show that these results can aid the future development of verifiers as they provide actionable targets for improving numerical reliability, relaxation quality, and search behavior. Our code is publicly available: https://github.com/dtroxell19/VeriStressGT.git.


翻译:神经网络验证器旨在为模型行为提供形式化保证,但现有验证基准因缺乏真值标签而存在根本局限性。因此,验证器评估依赖于间接启发式方法,这阻碍了精确评分和验证器失效模式的系统性研究。我们通过引入可复用框架来弥补这一空白,该框架通过解析构造生成已知真值稳健性标签的验证实例。该框架使我们发现多个主流验证器存在数值容差问题及实现缺陷,凸显了真值标签的必要性。此外,为系统性研究验证器失效模式,我们提出验证难度剖面——一组可估计的量化指标,用于捕捉实例难度的不同来源。基于该框架与剖面,我们评估了五种最先进的验证器,表明不同实例会对验证管线的不同方面施加压力。研究结果表明,这些成果通过提供数值可靠性、松弛质量及搜索行为优化的可行目标,可助力验证器的未来发展。我们的代码已公开:https://github.com/dtroxell19/VeriStressGT.git。

0
下载
关闭预览

相关内容

【ETHZ博士论文】神经网络训练与认证,101页pdf
专知会员服务
20+阅读 · 2024年7月28日
【ETHZ博士论文】认证神经网络的表达能力,86页pdf
专知会员服务
20+阅读 · 2024年6月16日
卷积神经网络的可解释性研究综述
专知会员服务
91+阅读 · 2023年6月5日
【AAAI2023】深度神经网络的可解释性验证
专知会员服务
49+阅读 · 2022年12月6日
专知会员服务
92+阅读 · 2021年7月9日
专知会员服务
29+阅读 · 2020年8月8日
机器学习的可解释性:因果推理和稳定学习
DataFunTalk
13+阅读 · 2020年3月3日
你的算法可靠吗? 神经网络不确定性度量
专知
40+阅读 · 2019年4月27日
NetworkMiner - 网络取证分析工具
黑白之道
16+阅读 · 2018年6月29日
神经网络可解释性最新进展
专知
18+阅读 · 2018年3月10日
【AAAI专题】论文分享:以生物可塑性为核心的类脑脉冲神经网络
中国科学院自动化研究所
15+阅读 · 2018年1月23日
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
7+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
9+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
14+阅读 · 8月7日
相关VIP内容
【ETHZ博士论文】神经网络训练与认证,101页pdf
专知会员服务
20+阅读 · 2024年7月28日
【ETHZ博士论文】认证神经网络的表达能力,86页pdf
专知会员服务
20+阅读 · 2024年6月16日
卷积神经网络的可解释性研究综述
专知会员服务
91+阅读 · 2023年6月5日
【AAAI2023】深度神经网络的可解释性验证
专知会员服务
49+阅读 · 2022年12月6日
专知会员服务
92+阅读 · 2021年7月9日
专知会员服务
29+阅读 · 2020年8月8日
相关基金
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
Top
微信扫码咨询专知VIP会员