A key challenge in multimodal reasoning is determining which visual dependencies become relevant under a specific task, rather than merely recognizing visible content. We study this through edit-induced constraint discovery in text-in-image editing, a controlled diagnostic setting where a local text change can activate secondary consistency constraints: given a valid editing instruction and an image, can a model identify the secondary regions that must also change? Across 461 diagnostic cases, four MLLMs, and 19 constraint subtypes, models recover only 46% case-level macro recall under unguided prompting versus 94% when constraints are explicitly provided, suggesting that a substantial portion of the failure arises when models must decide which unstated dependencies to surface. Oracle-field decomposition shows that case-specific causal explanations are the most effective partial guidance (0.782 recall), above region names (0.610) or type labels (0.646), suggesting that edit-specific causal cues account for much of the oracle gain. A downstream experiment further shows that higher self-discovery recall does not necessarily improve task performance: unverified self-discovery introduces false positives that offset recall gains, motivating precision-aware constraint elicitation.
翻译:多模态推理的关键挑战在于确定在特定任务下哪些视觉依赖关系变得相关,而不仅仅是识别可见内容。我们通过文本图像编辑中的编辑诱导约束发现来研究这一问题,这是一种受控的诊断设置,其中局部文本变化可能激活二级一致性约束:给定有效的编辑指令和图像,模型能否识别也必须发生变化的其他区域?在461个诊断案例、四种多模态大语言模型(MLLM)和19种约束子类型中,模型在无引导提示下仅恢复46%的案例级宏观召回率,而当明确提供约束时则为94%,这表明当模型必须决定浮现哪些未声明的依赖关系时,大部分失败发生。神谕字段分解显示,案例特定的因果解释是最有效的部分引导(召回率0.782),高于区域名称(0.610)或类型标签(0.646),这表明特定编辑的因果线索解释了神谕增益的大部分。后续实验进一步表明,较高的自我发现召回率并不必然改善任务性能:未经验证的自我发现引入了误报,抵消了召回率的提升,这促使需要关注精度的约束诱导。