Neural networks are used as generative surrogate models for scientific discovery, which are trainable approximations of scientific simulations. These models enable users to replace time-consuming numerical simulations with learned alternatives, providing quick solutions. However, high-fidelity generative surrogate models require massive training datasets, which can create storage and I/O challenges. Lossy compression is a promising way to reduce this burden, but compression errors may affect the model quality in subtle ways, making it challenging to quantify their impact. In this work, we examine how lossy compression of training data impacts the quality of generative surrogate models. We begin by characterizing the uncertainty inherent in training neural networks, showing that identical training configurations can produce different models. By exploiting this variability, we propose a method to estimate how much compression-induced error a surrogate model can tolerate without affecting its accuracy. Evaluation of two application simulations demonstrates that our approach significantly reduces memory/storage requirements and speeds up training while producing high-quality surrogate models. These results show that lossy compression saves data storage up to 23.7x and 39x with negligible impact on the quality of the surrogate model. Meanwhile, reducing the size of the training data set also enhances the data loading speed and reduces the training time by up to 3x.


翻译:神经网络被用作科学发现的生成式替代模型,这些模型是对科学模拟的可训练近似。它们使用户能够用学习到的替代方案取代耗时的数值模拟,从而提供快速解决方案。然而,高保真生成式替代模型需要大规模训练数据集,这可能导致存储和I/O方面的挑战。有损压缩是减轻这一负担的一种有前景的方法,但压缩误差可能以微妙的方式影响模型质量,使得量化其影响充满挑战。本研究探讨训练数据的有损压缩如何影响生成式替代模型的质量。我们首先刻画神经网络训练中固有的不确定性,表明相同的训练配置可能产生不同的模型。利用这种变异性,我们提出了一种方法,用于估计替代模型在不影响其精度的情况下所能容忍的压缩引入误差。对两个应用模拟的评估表明,我们的方法在生成高质量替代模型的同时,显著降低了内存/存储需求并加速了训练过程。这些结果显示,有损压缩可将数据存储节省高达23.7倍和39倍,而对替代模型质量的影响可忽略不计。同时,训练数据集规模的减小也提升了数据加载速度,并将训练时间缩短了高达3倍。

0
下载
关闭预览

相关内容

能耗优化的神经网络轻量化方法研究进展
专知会员服务
27+阅读 · 2023年1月29日
最新《神经数据压缩导论》综述
专知会员服务
40+阅读 · 2022年7月19日
专知会员服务
118+阅读 · 2020年8月22日
专知会员服务
74+阅读 · 2020年5月21日
模型压缩究竟在做什么?我们真的需要模型压缩么?
专知会员服务
28+阅读 · 2020年1月16日
深度神经网络模型压缩与加速综述
专知会员服务
130+阅读 · 2019年10月12日
深度学习模型可解释性的研究进展
专知
26+阅读 · 2020年8月1日
基于关系网络的视觉建模:有望替代卷积神经网络
微软研究院AI头条
10+阅读 · 2019年7月12日
2019年新书推荐-《神经网络与深度学习》-Michael Nielsen
深度学习与NLP
14+阅读 · 2019年2月21日
【优青论文】深度神经网络压缩与加速综述
计算机研究与发展
17+阅读 · 2018年9月20日
概览CVPR 2018神经网络图像压缩领域进展
论智
13+阅读 · 2018年6月13日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Arxiv
0+阅读 · 5月14日
Arxiv
0+阅读 · 5月4日
VIP会员
最新内容
从采集到决策:美军视角下的战术情报范式重构
专知会员服务
0+阅读 · 今天2:42
《履带式无人地面战车技术发展现状》
专知会员服务
2+阅读 · 今天1:46
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
相关VIP内容
能耗优化的神经网络轻量化方法研究进展
专知会员服务
27+阅读 · 2023年1月29日
最新《神经数据压缩导论》综述
专知会员服务
40+阅读 · 2022年7月19日
专知会员服务
118+阅读 · 2020年8月22日
专知会员服务
74+阅读 · 2020年5月21日
模型压缩究竟在做什么?我们真的需要模型压缩么?
专知会员服务
28+阅读 · 2020年1月16日
深度神经网络模型压缩与加速综述
专知会员服务
130+阅读 · 2019年10月12日
相关基金
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员