In agricultural environments, viewpoint planning can be a critical functionality for a robot with visual sensors to obtain informative observations of objects of interest (e.g., fruits) from complex structures of plant with random occlusions. Although recent studies on active vision have shown some potential for agricultural tasks, each model has been designed and validated on a unique environment that would not easily be replicated for benchmarking novel methods being developed later. In this paper, hence, we introduce a dataset for more extensive research on Domain-inspired Active VISion in Agriculture (DAVIS-Ag). To be specific, we utilized our open-source "AgML" framework and the 3D plant simulator of "Helios" to produce 502K RGB images from 30K dense spatial locations in 632 realistically synthesized orchards of strawberries, tomatoes, and grapes. In addition, useful labels are provided for each image, including (1) bounding boxes and (2) pixel-wise instance segmentations for all identifiable fruits, and also (3) pointers to other images that are reachable by an execution of action so as to simulate the active selection of viewpoint at each time step. Using DAVIS-Ag, we show the motivating examples in which performance of fruit detection for the same plant can significantly vary depending on the position and orientation of camera view primarily due to occlusions by other components such as leaves. Furthermore, we develop several baseline models to showcase the "usage" of data with one of agricultural active vision tasks--fruit search optimization--providing evaluation results against which future studies could benchmark their methodologies. For encouraging relevant research, our dataset is released online to be freely available at: https://github.com/ctyeong/DAVIS-Ag
翻译:在农业环境中,视点规划是配备视觉传感器的机器人从具有随机遮挡的复杂植物结构中获取感兴趣对象(如水果)信息性观测的关键功能。尽管近期关于主动视觉的研究已展现出在农业任务中的潜力,但每个模型均在独特环境下设计并验证,导致难以复现其环境以对后续新方法进行基准测试。为此,本文引入一个用于农业领域启发式主动视觉研究(DAVIS-Ag)的更广泛数据集。具体而言,我们利用开源框架"AgML"与三维植物模拟器"Helios",在632个逼真合成的草莓、番茄和葡萄果园中,从3万个密集空间位置生成50.2万张RGB图像。此外,每张图像还提供了实用标注,包括:(1)所有可识别水果的边界框;(2)像素级实例分割;以及(3)指向可通过动作执行到达的其他图像的指针,以模拟每步主动视点选择过程。通过DAVIS-Ag,我们展示了示例性案例:由于叶片等其他组件的遮挡,同一植物的水果检测性能会因相机视角的位置和朝向而显著变化。此外,我们开发了若干基线模型,展示数据在农业主动视觉任务之一——水果搜索优化中的"使用方法",并提供评估结果供后续研究作为基准比较。为鼓励相关研究,本数据集已在网上免费发布:https://github.com/ctyeong/DAVIS-Ag