Assessing built-environment interaction, such as wheelchair accessibility, is difficult because real-world mobility is shaped by distributed, context-dependent, and temporary barriers that are hard to capture at scale. To support scalable assessment, this paper examines whether vision-language models (VLMs) can identify accessibility barriers from Google Street View (GSV) imagery. We propose an expert-guided retrieval-augmented framework that combines GSV images, ADA-informed guidance, and expert-derived rubrics to evaluate accessibility dimensions. We collect a campus-scale dataset at the University of Florida, linking 407 unique GSV locations with GPS-derived wheelchair dwell behavior as a mobility-friction signal. Results show that VLM ratings are both negatively correlated and distributionally similar with dwell time, indicating partial but consistent alignment with a behavioral proxy for mobility friction. Visual cue analysis shows that certain environmental objects, such as curb ramps and crosswalks, are associated with higher VLM accessibility scores, while alignment remains limited for subtle surface conditions, transient obstructions, and viewpoint-dependent barriers. Overall, our findings show the potential of expert-guided VLMs for scalable accessibility assessment aligning with sensor-derived indicators of real-world wheelchair navigation.
翻译:摘要:评估建筑环境交互(如轮椅无障碍性)具有挑战性,因为现实世界的移动受到分布式、情境依赖且临时性的障碍物影响,这些障碍物难以大规模捕捉。为实现可扩展评估,本文探究视觉-语言模型(VLMs)能否从谷歌街景(GSV)影像中识别无障碍障碍。我们提出一种专家引导的检索增强框架,结合GSV图像、《美国残疾人法案》(ADA)指导原则以及专家制定的评估准则,对无障碍维度进行评价。我们在佛罗里达大学收集了校园级数据集,将407个独特GSV位置与GPS导出的轮椅停留行为(作为移动摩擦信号)相关联。结果表明,VLM评分与停留时间呈负相关且分布相似,表明其与移动摩擦的行为代理指标存在部分但一致的对应关系。视觉线索分析显示,特定环境对象(如路缘坡道和人行横道)与较高VLM无障碍评分相关,但模型对细微路面状况、临时性障碍物及视角依赖型障碍的识别能力仍有限。总体而言,我们的发现揭示了专家引导型VLM在结合传感器衍生指标进行真实世界轮椅导航可扩展无障碍评估中的潜力。