On-device Vision-Language Models (VLMs) promise data privacy via local execution. However, we show that the architectural shift toward Dynamic High-Resolution preprocessing (e.g., AnyRes) introduces an inherent algorithmic side-channel. Unlike static models, dynamic preprocessing decomposes images into a variable number of patches based on their aspect ratio, creating workload-dependent inputs. We demonstrate a dual-layer attack framework against local VLMs. In Tier 1, an unprivileged attacker can exploit significant execution-time variations using standard unprivileged OS metrics to reliably fingerprint the input's geometry. In Tier 2, by profiling Last-Level Cache (LLC) contention, the attacker can resolve semantic ambiguity within identical geometries, distinguishing between visually dense (e.g., medical X-rays) and sparse (e.g., text documents) content. By evaluating state-of-the-art models such as LLaVA-NeXT and Qwen2-VL, we show that combining these signals enables reliable inference of privacy-sensitive contexts. Finally, we analyze the security engineering trade-offs of mitigating this vulnerability, reveal substantial performance overhead with constant-work padding, and propose practical design recommendations for secure Edge AI deployments.
翻译:设备端视觉语言模型(VLM)通过本地执行承诺数据隐私。然而,我们证明了向动态高分辨率预处理(例如AnyRes)的架构转变引入了一种固有的算法侧信道。与静态模型不同,动态预处理根据图像宽高比将其分解为可变数量的图块,从而产生依赖于工作负载的输入。我们展示了一个针对本地VLM的双层攻击框架。在第一层中,无特权攻击者可以利用标准无特权操作系统指标下的显著执行时间变化,可靠地识别输入几何形状。在第二层中,通过分析最后一级缓存(LLC)竞争,攻击者可以解析相同几何形状内的语义歧义,区分视觉密集(例如医学X射线)和稀疏(例如文本文档)内容。通过评估LLaVA-NeXT和Qwen2-VL等先进模型,我们展示了结合这些信号能够可靠地推断隐私敏感上下文。最后,我们分析了缓解此漏洞的安全工程权衡,揭示了恒定工作填充会带来显著性能开销,并提出了针对安全边缘AI部署的实用设计建议。