We introduce PhotoBot, a framework for fully automated photo acquisition based on an interplay between high-level human language guidance and a robot photographer. We propose to communicate photography suggestions to the user via reference images that are selected from a curated gallery. We leverage a visual language model (VLM) and an object detector to characterize the reference images via textual descriptions and then use a large language model (LLM) to retrieve relevant reference images based on a user's language query through text-based reasoning. To correspond the reference image and the observed scene, we exploit pre-trained features from a vision transformer capable of capturing semantic similarity across marked appearance variations. Using these features, we compute pose adjustments for an RGB-D camera by solving a perspective-n-point (PnP) problem. We demonstrate our approach using a manipulator equipped with a wrist camera. Our user studies show that photos taken by PhotoBot are often more aesthetically pleasing than those taken by users themselves, as measured by human feedback. We also show that PhotoBot can generalize to other reference sources such as paintings.
翻译:摘要:我们提出PhotoBot框架,该框架通过高层级人类语言指令与机器人摄影师之间的交互实现全自动图像采集。我们提出通过从精选图库中选取参考图像来向用户传达摄影建议。我们利用视觉语言模型(VLM)和目标检测器,通过文本描述对参考图像进行特征表征,再借助大型语言模型(LLM)基于用户的自然语言查询,通过文本推理检索相关参考图像。为实现参考图像与观测场景的对应,我们利用视觉Transformer中预训练的特征,该特征能捕捉显著外观变化下的语义相似性。基于这些特征,我们通过求解透视n点(PnP)问题计算RGB-D相机的位姿调整。我们使用配备腕部相机的机械臂验证该方法。用户研究表明,根据人类反馈评估,PhotoBot拍摄的照片通常比用户自拍更具美学价值。我们还证明PhotoBot可泛化至其他参考源,例如绘画作品。