Reinforcement learning has recently improved the reasoning ability of Large Language Models and Multimodal LLMs, yet prevailing reward designs emphasise final-answer correctness and consequently tolerate process hallucinations--cases where models reach the right answer while misperceiving visual evidence. We address this process-level misalignment with PaLMR, a framework that aligns not only outcomes but also the reasoning process itself. PaLMR comprises two complementary components: a perception-aligned data layer that constructs process-aware reasoning data with structured pseudo-ground-truths and verifiable visual facts, and a process-aligned optimisation layer that constructs a hierarchical reward fusion scheme with a process-aware scoring function to encourage visually faithful chains-of-thought and improve training stability. Experiments on Qwen2.5-VL-7B show that our approach substantially reduces reasoning hallucinations and improves visual reasoning fidelity, achieving state-of-the-art results on HallusionBench while maintaining strong performance on MMMU, MathVista, and MathVerse. These findings indicate that PaLMR offers a principled and practical route to process-aligned multimodal reasoning, advancing the reliability and interpretability of MLLMs.
翻译:强化学习近期提升了大语言模型与多模态大语言模型的推理能力,但现有奖励机制侧重于最终答案正确性,因而容忍了过程幻觉——即模型在得出正确答案的同时对视觉证据产生错误感知的现象。我们针对这一过程级错位提出PaLMR框架,该框架不仅对齐结果,更对齐推理过程本身。PaLMR由两大互补模块构成:感知对齐数据层,构建包含结构化伪真实标注与可验证视觉事实的过程感知推理数据;过程对齐优化层,设计分层奖励融合机制并引入过程感知评分函数,以鼓励生成忠实于视觉证据的思维链并提升训练稳定性。在Qwen2.5-VL-7B上的实验表明,本方法显著降低了推理幻觉,提升了视觉推理保真度:在HallusionBench上取得最先进结果,同时在MMMU、MathVista及MathVerse上保持强劲性能。这些发现表明,PaLMR为过程对齐多模态推理提供了原则性与实用性兼备的路径,推动了多模态大语言模型的可信度与可解释性发展。