Current methods in training and benchmarking vision models exhibit an over-reliance on passive, curated datasets. Although models trained on these datasets have shown strong performance in a wide variety of tasks such as classification, detection, and segmentation, they fundamentally are unable to generalize to an ever-evolving world due to constant out-of-distribution shifts of input data. Therefore, instead of training on fixed datasets, can we approach learning in a more human-centric and adaptive manner? In this paper, we introduce Action-Aware Embodied Learning for Perception (ALP), an embodied learning framework that incorporates action information into representation learning through a combination of optimizing a reinforcement learning policy and an inverse dynamics prediction objective. Our method actively explores in complex 3D environments to both learn generalizable task-agnostic visual representations as well as collect downstream training data. We show that ALP outperforms existing baselines in several downstream perception tasks. In addition, we show that by training on actively collected data more relevant to the environment and task, our method generalizes more robustly to downstream tasks compared to models pre-trained on fixed datasets such as ImageNet.
翻译:当前视觉模型的训练和基准测试方法过度依赖被动、精心策划的数据集。尽管基于这些数据集训练的模型在分类、检测和分割等广泛任务中表现优异,但由于输入数据持续存在的分布偏移,它们从根本上无法泛化至不断演化的世界。因此,我们能否摒弃固定数据集的训练模式,以更贴近人类认知的自适应方式进行学习?本文提出了面向感知的动作感知具身学习框架(ALP),该框架通过联合优化强化学习策略与逆动力学预测目标,将动作信息融入表征学习。我们的方法能在复杂3D环境中主动探索,既学习可泛化的任务无关视觉表征,又收集下游训练数据。实验表明,ALP在多项下游感知任务中优于现有基线方法。此外,相较于在ImageNet等固定数据集上预训练的模型,由于在更贴近环境与任务的主动采集数据上训练,我们方法在下游任务中展现了更强的泛化鲁棒性。