As transformer architectures and dataset sizes continue to scale, the need to understand the specific dataset factors affecting model performance becomes increasingly urgent. This paper investigates how object physics attributes (color, friction coefficient, shape) and background characteristics (static, dynamic, background complexity) influence the performance of Video Transformers in trajectory prediction tasks under occlusion. Beyond mere occlusion challenges, this study aims to investigate three questions: How do object physics attributes and background characteristics influence the model performance? What kinds of attributes are most influential to the model generalization? Is there a data saturation point for large transformer model performance within a single task? To facilitate this research, we present OccluManip, a real-world video-based robot pushing dataset comprising 460,000 consistent recordings of objects with different physics and varying backgrounds. 1.4 TB and in total 1278 hours of high-quality videos of flexible temporal length along with target object trajectories are collected, accommodating tasks with different temporal requirements. Additionally, we propose Video Occlusion Transformer (VOT), a generic video-transformer-based network achieving an average 96% accuracy across all 18 sub-datasets provided in OccluManip. OccluManip and VOT will be released at: https://github.com/ShutongJIN/OccluManip.git
翻译:随着Transformer架构与数据集规模持续扩大,理解影响模型性能的具体数据集因素变得愈发紧迫。本文研究了物体物理属性(颜色、摩擦系数、形状)与背景特征(静态/动态特性、背景复杂度)如何影响遮挡条件下视频Transformer在轨迹预测任务中的表现。除遮挡挑战本身外,本研究旨在探究三个问题:物体物理属性与背景特征如何影响模型性能?哪些属性对模型泛化能力最具影响力?在单一任务中,大型Transformer模型的性能是否存在数据饱和点?为支撑研究,我们提出了OccluManip——一个基于真实世界视频的机器人推拽数据集,包含46万段关于不同物理属性物体与多样化背景的连续记录。该数据集包含1.4TB、总计1278小时的高质量视频,视频长度可灵活适配不同时间尺度的任务需求,并附带目标物体轨迹信息。此外,我们提出了视频遮挡Transformer(VOT),一种基于通用视频Transformer的网络架构,在OccluManip提供的全部18个子数据集上实现了平均96%的准确率。OccluManip与VOT将发布于:https://github.com/ShutongJIN/OccluManip.git