Current advancements in Audio Reasoning rely on massive Large Audio-Language Models (LALMs), hindering deployment in resource-constrained environments. We introduce TinyGiantALM, a compact 1.5B efficiency-oriented alternative. Instead of brute-force scaling, we propose an Instruction-Aware Feature Refinement framework using a Query-guided Projector and Semantic Gating to filter acoustic signals based on user intent. On the MMAR benchmark, TinyGiantALM achieves 46.4% zero-shot accuracy, significantly outperforming 7B-13B baselines. While a reasoning gap in logical narrative remains versus 30B+ models and certain trade-offs exist in overly dense or spatial scenes, our approach notably surpasses models up to 8x larger in disentangling mixed-modality environments. These findings demonstrate that architectural precision offers a tangible pathway to secure robust perception capabilities on edge-friendly scales.
翻译:当前音频推理领域的进展依赖大规模音频-语言模型(LALMs),这阻碍了其在资源受限环境中的部署。我们提出TinyGiantALM,一种面向效率的紧凑型1.5B参数替代方案。不同于暴力扩展模型规模,我们提出基于指令感知的特征细化框架,通过查询引导投影器和语义门控机制,根据用户意图筛选声学信号。在MMAR基准测试中,TinyGiantALM以46.4%的零样本准确率显著超越7B-13B基线模型。尽管在与30B+参数模型的逻辑叙事推理存在差距,且在过密或空间场景中存在特定权衡,但本方法在解耦混合模态环境方面显著优于规模大至8倍的模型。这些发现表明,架构精度为在边缘友好规模上获取鲁棒感知能力提供了切实可行的路径。