We introduce V-tableR1, a process-supervised reinforcement learning framework that elicits rigorous, verifiable reasoning from multimodal large language models (MLLMs). Current MLLMs trained solely on final outcomes often treat visual reasoning as a black box, relying on superficial pattern matching rather than performing rigorous multi-step inference. While Reinforcement Learning with Verifiable Rewards could enforce transparent reasoning trajectories, extending it to visual domains remains severely hindered by the ambiguity of grounding abstract logic into continuous pixel space. We solve this by leveraging the deterministic grid structure of tables as an ideal visual testbed. V-tableR1 employs a specialized critic VLM to provide dense, step-level feedback on the explicit visual chain-of-thought generated by a policy VLM. To optimize this system, we propose Process-Guided Direct Alignment Policy Optimization (PGPO), a novel RL algorithm integrating process rewards, decoupled policy constraints, and length-aware dynamic sampling. Extensive evaluations demonstrate that V-tableR1 explicitly penalizes visual hallucinations and shortcut guessing. By fundamentally shifting multimodal inference from black-box pattern matching to verifiable logical derivation, V-tableR1 4B establishes state-of-the-art accuracy among open-source models on complex tabular benchmarks, outperforming models up to 18x its size and improving over its SFT baseline
翻译:我们提出V-tableR1,一种基于过程监督的强化学习框架,能够激发多模态大语言模型进行严谨、可验证的推理。当前仅在最终结果上训练的多模态大语言模型通常将视觉推理视为黑箱,依赖浅层模式匹配而非执行严格的多步推理。尽管基于可验证奖励的强化学习能够强制执行透明的推理轨迹,但由于将抽象逻辑锚定到连续像素空间存在模糊性,将其扩展至视觉领域仍面临严重阻碍。我们通过利用表格的确定性网格结构作为理想视觉测试平台来解决这一问题。V-tableR1采用专门设计的评论者视觉语言模型,为策略视觉语言模型生成的显式视觉思维链提供密集的逐步级反馈。为优化该系统,我们提出过程引导直接对齐策略优化(PGPO),一种融合过程奖励、解耦策略约束与长度感知动态采样的新型强化学习算法。大量评估表明,V-tableR1能够显式惩罚视觉幻觉与捷径式猜测。通过从根本上将多模态推理从黑箱模式匹配转变为可验证的逻辑推导,V-tableR1 4B在复杂表格基准测试中建立了开源模型的最优准确率,性能超越参数规模大18倍的模型,并显著优于其SFT基线。