In robotics systems, vast amounts of visual data are easily captured at high resolution using low-cost, low-power hardware. Yet, limited bandwidth and on-device compute resources prevent full utilization when transmitted via conventional codecs like JPEG/MPEG. Newer codecs, like AV1/AVIF, improve the rate-distortion trade-off, but demand far more resources for encoding, impractical without custom ASICs. Recent asymmetric autoencoders deliver high quality under extreme power and bandwidth constraints, but add prohibitive decoding cost and use bespoke formats that ignore decades of infrastructure built around standards like JPEG. To address these limitations, we introduce a compression framework for cloud robotics based on a Sensor Embedded Autoencoder paired with a One-Time Transcode for Efficient Reconstruction (SEAOTTER). Because the sensor, cloud, and consumer stages face very different power and bandwidth budgets, SEAOTTER combines the compactness of a learned latent with the broad usability of a standard JPEG file. Since naive transcoding degrades performance, we propose a learnable JPEG color and quantization transform that enables increased accuracy for global, dense, and vision-language-based perception. Using SEAOTTER, we train both general-purpose and task-aware transcoding pipelines for a pre-trained, frozen encoder. At a compression ratio of 200:1 and compared to AVIF, we observe 7 times faster encoding, 3.5 times faster decoding, and +8% ImageNet top-1 accuracy, while retaining compatibility with JPEG infrastructure. Our code is available at https://github.com/UT-SysML/seaotter .


翻译:在机器人系统中,可通过低功耗低成本硬件轻松采集高分辨率海量视觉数据。然而,当通过JPEG/MPEG等传统编解码器传输时,有限的带宽和边缘计算资源阻碍了数据的充分利用。AV1/AVIF等新型编解码器虽改善了率失真权衡,但编码过程需消耗更多资源,无定制ASIC时难以实际应用。近年来非对称自编码器在极端功耗和带宽约束下实现了高质量重建,但解码成本过高,且采用非标准格式导致无法利用基于JPEG等标准构建的现有基础设施。为解决上述局限,我们提出一种面向云机器人的压缩框架,其核心为传感器嵌入式自编码器与一次性转码实现高效重建(SEAOTTER)。由于传感器、云端与消费终端面临迥异的功耗和带宽预算,SEAOTTER将学习型潜变量的紧凑性与标准JPEG文件的广泛兼容性有机结合。针对简单转码会降低性能的问题,我们提出可学习的JPEG颜色与量化变换方法,可显著提升全局、密集及视觉-语言感知任务的精度。基于SEAOTTER,我们为预训练的冻结编码器训练了通用型与任务感知型转码流水线。在200:1压缩比下,与AVIF相比,编码速度提升7倍,解码速度提升3.5倍,ImageNet Top-1准确率提高8%,同时保持与JPEG基础设施的兼容性。代码已开源:https://github.com/UT-SysML/seaotter。

0
下载
关闭预览

相关内容

传感器(英文名称:transducer/sensor)是一种检测装置,能感受到被测量的信息,并能将感受到的信息,按一定规律变换成为电信号或其他所需形式的信息输出,以满足信息的传输、处理、存储、显示、记录和控制等要求。
《空基机器人系统的传感器融合技术》美陆军最新58页
专知会员服务
31+阅读 · 2025年4月20日
【NeurIPS2021】ResT:一个有效的视觉识别转换器
专知会员服务
23+阅读 · 2021年10月25日
专知会员服务
37+阅读 · 2021年10月16日
深度学习之图像超分辨重建技术
机器学习研究会
12+阅读 · 2018年3月24日
一文概览基于深度学习的超分辨率重建架构
【干货】深入理解变分自编码器
专知
21+阅读 · 2018年3月22日
【干货】深入理解自编码器(附代码实现)
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
《履带式无人地面战车技术发展现状》
专知会员服务
4+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
4+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
13+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
12+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
7+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
10+阅读 · 7月31日
相关VIP内容
《空基机器人系统的传感器融合技术》美陆军最新58页
专知会员服务
31+阅读 · 2025年4月20日
【NeurIPS2021】ResT:一个有效的视觉识别转换器
专知会员服务
23+阅读 · 2021年10月25日
专知会员服务
37+阅读 · 2021年10月16日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员