We present the All-Seeing (AS) project: a large-scale data and model for recognizing and understanding everything in the open world. Using a scalable data engine that incorporates human feedback and efficient models in the loop, we create a new dataset (AS-1B) with over 1 billion regions annotated with semantic tags, question-answering pairs, and detailed captions. It covers a wide range of 3.5 million common and rare concepts in the real world, and has 132.2 billion tokens that describe the concepts and their attributes. Leveraging this new dataset, we develop the All-Seeing model (ASM), a unified framework for panoptic visual recognition and understanding. The model is trained with open-ended language prompts and locations, which allows it to generalize to various vision and language tasks with remarkable zero-shot performance, including region-text retrieval, region recognition, captioning, and question-answering. We hope that this project can serve as a foundation for vision-language artificial general intelligence research. Models and the dataset shall be released at https://github.com/OpenGVLab/All-Seeing, and demo can be seen at https://huggingface.co/spaces/OpenGVLab/all-seeing.
翻译:我们提出了全视(AS)项目:一个用于识别和理解开放世界中万事万物的大规模数据与模型。通过采用融合人类反馈的可扩展数据引擎以及高效模型进行循环迭代,我们创建了一个新数据集(AS-1B),其中包含超过10亿个区域,并标注了语义标签、问答对以及详细描述。该数据集涵盖了现实世界中广泛存在的350万个常见与罕见概念,并拥有1322亿个用于描述这些概念及其属性的词元。借助这一新数据集,我们开发了全视模型(ASM),这是一个用于全景视觉识别与理解的统一框架。该模型通过开放式语言提示和位置信息进行训练,使其能够泛化至各种视觉与语言任务,并展现出卓越的零样本性能,包括区域-文本检索、区域识别、描述生成以及问答。我们希望该项目能作为视觉语言通用人工智能研究的基石。模型与数据集将在 https://github.com/OpenGVLab/All-Seeing 发布,演示可在 https://huggingface.co/spaces/OpenGVLab/all-seeing 查看。