Spoken dialogue models typically start from text LLM backbones, yet reasoning often degrades when conditioning on speech instead of text. We attribute part of this modality gap to a temporal-granularity mismatch: speech tokens are temporally redundant and far longer than text under matched semantics, diluting per-token semantic density and weakening text-native reasoning dynamics. We study speech token design as a representation selection problem and sweep frame rates under a frozen LLM backbone with a fixed information rate. To make low frame rates feasible, we introduce factorized FSQ and a lightweight non-autoregressive audio LM head, scaling capacity to nearly 300\,bits/frame without sacrificing efficient prediction. With the bottleneck removed, we sweep frame rates (50$\rightarrow$2.08\,Hz) and alignment depth, and observe a consistent best regime for speech QA at 4.17\,Hz with intermediate-layer representation alignment.
翻译:口语对话模型通常以文本大语言模型骨干为基础,但在以语音而非文本为条件时,推理能力往往会下降。我们将这种模态差异部分归因于时间粒度的不匹配:语义匹配时,语音令牌在时间上冗余且长度远超文本,这会稀释每个令牌的语义密度,削弱文本原生推理的动态特性。我们将语音令牌设计视为表示选择问题,在固定信息率下对冻结的LLM骨干进行帧率扫描。为使低帧率可行,我们引入分解式FSQ和轻量级非自回归音频LM头,在保持高效预测的同时将容量扩展至近300比特/帧。移除瓶颈后,我们扫描帧率(50→2.08Hz)和对齐深度,观察到语音QA在4.17Hz帧率下且采用中间层表示对齐时存在一致的最佳区域。