We present an end-to-end procedure for embodied exploration inspired by two biological computations: predictive coding and uncertainty minimization. The procedure can be applied to exploration settings in a task-independent and intrinsically driven manner. We first demonstrate our approach in a maze navigation task and show that it can discover the underlying transition distributions and spatial features of the environment. Second, we apply our model to a more complex active vision task, where an agent actively samples its visual environment to gather information. We show that our model builds unsupervised representations through exploration that allow it to efficiently categorize visual scenes. We further show that using these representations for downstream classification leads to superior data efficiency and learning speed compared to other baselines while maintaining lower parameter complexity. Finally, the modularity of our model allows us to probe its internal mechanisms and analyze the interaction between perception and action during exploration.
翻译:我们提出了一种受两种生物计算机制启发——预测编码与不确定性最小化——的具身探索端到端流程。该流程以任务无关且内在驱动的方式适用于探索场景。我们首先在迷宫导航任务中验证该方法,证明其能发现环境中的潜在转移分布与空间特征。其次,我们将模型应用于更复杂的主动视觉任务:智能体主动采样视觉环境以收集信息。我们展示模型通过探索构建的免监督表征可高效分类视觉场景。进一步表明,将此类表征用于下游分类任务时,相较于其他基线方法,能在保持更低参数复杂度的情况下实现更优的数据效率与学习速度。最后,模型的模块化特性使我们能探知其内部机制,并分析探索过程中感知与动作的交互关系。