A Comprehensive Survey on Segment Anything Model for Vision and Beyond

Artificial intelligence (AI) is evolving towards artificial general intelligence, which refers to the ability of an AI system to perform a wide range of tasks and exhibit a level of intelligence similar to that of a human being. This is in contrast to narrow or specialized AI, which is designed to perform specific tasks with a high degree of efficiency. Therefore, it is urgent to design a general class of models, which we term foundation models, trained on broad data that can be adapted to various downstream tasks. The recently proposed segment anything model (SAM) has made significant progress in breaking the boundaries of segmentation, greatly promoting the development of foundation models for computer vision. To fully comprehend SAM, we conduct a survey study. As the first to comprehensively review the progress of segmenting anything task for vision and beyond based on the foundation model of SAM, this work focuses on its applications to various tasks and data types by discussing its historical development, recent progress, and profound impact on broad applications. We first introduce the background and terminology for foundation models including SAM, as well as state-of-the-art methods contemporaneous with SAM that are significant for segmenting anything task. Then, we analyze and summarize the advantages and limitations of SAM across various image processing applications, including software scenes, real-world scenes, and complex scenes. Importantly, some insights are drawn to guide future research to develop more versatile foundation models and improve the architecture of SAM. We also summarize massive other amazing applications of SAM in vision and beyond.

翻译：人工智能正朝着通用人工智能方向发展，即AI系统具备执行广泛任务并展现类人智能水平的能力，这与旨在高效完成特定任务的狭义或专用AI形成对比。因此，亟待设计一类通用模型（我们称之为基础模型），这些模型基于广泛数据训练，并能适配多种下游任务。近期提出的“分割一切”模型（SAM）在打破分割任务边界方面取得了显著进展，极大地推动了计算机视觉基础模型的发展。为全面理解SAM，我们开展了一项综述研究。作为首篇基于SAM基础模型全面回顾视觉及更广领域“分割一切”任务进展的工作，本文聚焦于SAM在多种任务与数据类型中的应用，通过探讨其历史发展、最新进展及对广泛应用的深远影响展开论述。我们首先介绍包括SAM在内的基础模型的背景与术语，以及与SAM同期提出的对“分割一切”任务具有重要意义的先进方法。随后，我们分析并总结了SAM在各类图像处理应用中的优势与局限，涵盖软件场景、真实场景及复杂场景。重要的是，本文提炼出若干见解，以指导未来研究开发更通用的基础模型并改进SAM架构。此外，我们还汇总了SAM在视觉及更广领域中大量令人惊奇的其他应用。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日