On-device Vision-Language Models (VLMs) promise data privacy via local execution. However, we show that the architectural shift toward Dynamic High-Resolution preprocessing (e.g., AnyRes) introduces an inherent algorithmic side-channel. Unlike static models, dynamic preprocessing decomposes images into a variable number of patches based on their aspect ratio, creating workload-dependent inputs. We demonstrate a dual-layer attack framework against local VLMs. In Tier 1, an unprivileged attacker can exploit significant execution-time variations using standard unprivileged OS metrics to reliably fingerprint the input's geometry. In Tier 2, by profiling Last-Level Cache (LLC) contention, the attacker can resolve semantic ambiguity within identical geometries, distinguishing between visually dense (e.g., medical X-rays) and sparse (e.g., text documents) content. By evaluating state-of-the-art models such as LLaVA-NeXT and Qwen2-VL, we show that combining these signals enables reliable inference of privacy-sensitive contexts. Finally, we analyze the security engineering trade-offs of mitigating this vulnerability, reveal substantial performance overhead with constant-work padding, and propose practical design recommendations for secure Edge AI deployments.
翻译:设备端视觉语言模型(VLM)通过本地执行承诺数据隐私。然而,我们揭示了向动态高分辨率预处理(如AnyRes)的架构转变引入了固有的算法侧信道。与静态模型不同,动态预处理根据图像宽高比将其分解为可变数量的图像块,产生依赖工作负载的输入。我们提出了一种针对本地VLM的双层攻击框架。在第一层,未授权攻击者可以利用标准非特权操作系统指标,通过显著执行时间差异可靠地识别输入几何形状。在第二层,通过分析末级缓存(LLC)争用,攻击者可以消除相同几何形状内的语义歧义,区分视觉密集(如医学X光片)与稀疏(如文本文档)内容。通过评估LLaVA-NeXT和Qwen2-VL等最先进模型,我们展示了结合这些信号能够可靠推断隐私敏感上下文。最后,我们分析了缓解该漏洞的安全工程权衡,揭示了恒定工作负载填充的显著性能开销,并提出了面向安全边缘AI部署的实用设计建议。