Understanding geometric concepts, such as distance and shape, is essential for understanding the real world and also for many vision tasks. To incorporate such information into a visual representation of a scene, we propose learning to represent the scene by sketching, inspired by human behavior. Our method, coined Learning by Sketching (LBS), learns to convert an image into a set of colored strokes that explicitly incorporate the geometric information of the scene in a single inference step without requiring a sketch dataset. A sketch is then generated from the strokes where CLIP-based perceptual loss maintains a semantic similarity between the sketch and the image. We show theoretically that sketching is equivariant with respect to arbitrary affine transformations and thus provably preserves geometric information. Experimental results show that LBS substantially improves the performance of object attribute classification on the unlabeled CLEVR dataset, domain transfer between CLEVR and STL-10 datasets, and for diverse downstream tasks, confirming that LBS provides rich geometric information.
翻译:理解距离和形状等几何概念,对于理解真实世界以及完成许多视觉任务至关重要。为了将此类信息融入场景的视觉表征中,我们受人类行为启发,提出通过素描来学习场景表征。我们提出的方法名为“通过素描学习”(LBS),该方法能在单次推理步骤中将图像转换为一组彩色笔画,这些笔画显式地融入场景的几何信息,且无需使用素描数据集。随后根据这些笔画生成素描,其中基于CLIP的感知损失保持了素描与图像之间的语义相似性。我们从理论上证明,素描在任意仿射变换下具有等变性,因此可证明其能保留几何信息。实验结果表明,LBS显著提升了无标注CLEVR数据集上的物体属性分类性能、CLEVR与STL-10数据集之间的域迁移能力,以及多种下游任务的表现,证实LBS能够提供丰富的几何信息。