Rescaling MLM-Head for Neural Sparse Retrieval - 专知论文

会员服务 ·

0

稀疏 · MoDELS · 再缩放 · Backbone · 缩放 ·

Rescaling MLM-Head for Neural Sparse Retrieval

翻译：暂无翻译

Youngjoon Jang,Seongtae Hong,Jonah Turner,Heuiseok Lim

Learned sparse retrieval (LSR) models such as SPLADE have traditionally used BERT-style masked language models as backbone encoders. A natural expectation is that replacing BERT with stronger pretrained encoders should improve retrieval effectiveness. However, we find that under standard SPLADE training recipes, backbones with large MLM-head L2 norms can suffer performance degradation and even training collapse under standard SPLADE training recipes. We identify this failure as a scale mismatch in the MLM head: SPLADE directly uses MLM-head outputs to construct sparse lexical representations, and query-document relevance is computed by an unnormalized dot product over these representations. As a result, an inflated MLM-head scale can amplify sparse activations, distort matching scores, and destabilize contrastive training under common training settings. To address this issue, we introduce a simple initialization-time correction that rescales the MLM-head projection by a constant factor before SPLADE training. This zero-cost adjustment improves training stability without modifying the model architecture or training objective. Across both in-domain and out-of-domain retrieval benchmarks, this simple correction substantially improves large-norm backbones such as ModernBERT and Ettin, turning unstable training runs into competitive sparse retrievers. In several settings, the corrected models further match or surpass the classic BERT-SPLADE baseline. These findings suggest that the bottleneck in adapting pretrained encoders to LSR is not encoder capacity alone, but the calibration of the MLM-head scale used to construct sparse lexical representations.

翻译：暂无翻译

0

相关内容

EMNLP 2024 | 基于知识编辑的大模型敏感知识擦除

EMNLP 2024 | 基于知识编辑的大模型敏感知识擦除

专知会员服务

22+阅读 · 2024年11月19日

ICLR2024｜Mol-Instructions: 面向大模型的大规模生物分子指令数据集

ICLR2024｜Mol-Instructions: 面向大模型的大规模生物分子指令数据集

专知会员服务

12+阅读 · 2024年2月10日

EMNLP2023：预训练模型的知识反刍

EMNLP2023：预训练模型的知识反刍

专知会员服务

32+阅读 · 2023年11月20日

LLM in Medical Domain: 大语言模型在医学领域的应用

LLM in Medical Domain: 大语言模型在医学领域的应用

专知会员服务

103+阅读 · 2023年6月17日

ICLR 2022 | BEIT论文解读：将MLM无监督预训练应用到CV领域

ICLR 2022 | BEIT论文解读：将MLM无监督预训练应用到CV领域

专知会员服务

33+阅读 · 2022年3月24日

【CVPR 2022】跨模态检索的协同双流视觉-语言前训练模型，COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal Retrieval

【CVPR 2022】跨模态检索的协同双流视觉-语言前训练模型，COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal Retrieval

专知会员服务

13+阅读 · 2022年3月12日

【EMNLP 2019 最佳论文】信息瓶颈专门化单词嵌入（用于解析）（Specializing Word Embeddings（for Parsing）by Information Bottleneck）

【EMNLP 2019 最佳论文】信息瓶颈专门化单词嵌入（用于解析）（Specializing Word Embeddings（for Parsing）by Information Bottleneck）

专知会员服务

24+阅读 · 2019年11月20日

【MLA 2019】自然语言处理中的表示学习进展：从Transfomer到BERT，复旦大学邱锡鹏

【MLA 2019】自然语言处理中的表示学习进展：从Transfomer到BERT，复旦大学邱锡鹏

专知会员服务

100+阅读 · 2019年11月15日

【AAAI2020接受论文】隐式关系语言模型，CMU&微软，Latent Relation Language Models

【AAAI2020接受论文】隐式关系语言模型，CMU&微软，Latent Relation Language Models

专知会员服务

54+阅读 · 2019年11月12日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

赛尔原创 | EMNLP 2019 基于上下文感知的变分自编码器建模事件背景知识进行If-Then类型常识推理

赛尔原创 | EMNLP 2019 基于上下文感知的变分自编码器建模事件背景知识进行If-Then类型常识推理

哈工大SCIR

17+阅读 · 2019年9月23日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

44+阅读 · 2019年1月3日

meta learning 17年：MAML SNAIL

meta learning 17年：MAML SNAIL

CreateAMind

11+阅读 · 2019年1月2日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

李宏毅-201806-中文-Deep Reinforcement Learning精品课程分享

李宏毅-201806-中文-Deep Reinforcement Learning精品课程分享

深度学习与NLP

15+阅读 · 2018年6月20日

论文浅尝 | 嵌入常识知识的注意力 LSTM 模型用于特定目标的基于侧面的情感分析

论文浅尝 | 嵌入常识知识的注意力 LSTM 模型用于特定目标的基于侧面的情感分析

开放知识图谱

28+阅读 · 2018年6月11日

论文浅尝 | Improved Neural Relation Detection for KBQA

论文浅尝 | Improved Neural Relation Detection for KBQA

开放知识图谱

13+阅读 · 2018年1月21日

From Softmax to Sparsemax-ICML16（1）

From Softmax to Sparsemax-ICML16（1）

KingsGarden

74+阅读 · 2016年11月26日

不同发育起源MSCs在骨缺损修复重建过程中的作用及机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

纳米复合结构的可控构筑及其在LSPR免疫在线检测中的应用研究

国家自然科学基金

0+阅读 · 2015年12月31日

废液中铀酰类化合物超灵敏检测用SERS基底纳米结构的设计与构建

国家自然科学基金

0+阅读 · 2015年12月31日

人脑MRI数据特征提取方法的研究与应用

国家自然科学基金

0+阅读 · 2015年12月31日

力感应细胞源性外泌体及其microRNA货物介导的细胞间通讯在骨重建稳态中调控机制及干预策略的研究

国家自然科学基金

0+阅读 · 2015年12月31日

基于miRNA转染的血管化BMSC膜片的构建、血管化与骨再生能力及其机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

低氧预处理自体骨髓间充质干细胞修复关节软骨损伤的MRI活体示踪研究

国家自然科学基金

0+阅读 · 2014年12月31日

iPSC-MSCs复合活性大孔CPC用于牙槽骨缺损修复及转归机制的研究

国家自然科学基金

0+阅读 · 2014年12月31日

低氧下骨髓间充质干细胞复合细胞外基质膜片构建3D组织工程皮肤的研究

国家自然科学基金

0+阅读 · 2014年12月31日

MSM人群中HIV感染者生命质量评价及预警模型研究

国家自然科学基金

0+阅读 · 2014年12月31日

LLM Compression by Block Removal with Constrained Binary Optimization

Arxiv

0+阅读 · 6月17日

Latency Prediction for LLM Inference on NPU Systems

Arxiv

0+阅读 · 6月17日

Factorized Latent Reasoning for LLM-based Recommendation

Arxiv

0+阅读 · 6月12日

Revisiting Neural Processes via Fourier Transform and Volterra Series

Arxiv

0+阅读 · 6月11日

LLM-Guided Evolution for Medical Decision Pipelines

Arxiv

0+阅读 · 6月5日

From Layers to Submodules: Rethinking Granularity in Replacement-Based LLM Compression

Arxiv

0+阅读 · 6月1日

Neural Router: Semantic Content Matching for Agentic AI

Arxiv

0+阅读 · 5月25日

Neural-Actuarial Longevity Forecasting: Anchoring LSTMs for Explainable Risk Management

Arxiv

0+阅读 · 5月7日

A Survey on Neural Speech Synthesis

Arxiv

14+阅读 · 2021年6月30日

Latent Relation Language Models

Arxiv

21+阅读 · 2019年8月21日

VIP会员

文章信息

相关主题

最新内容

ICML 2026 Spotlight | SmoothSMoE：解析稀疏 MoE 路由不连续

ICML 2026 Spotlight | SmoothSMoE：解析稀疏 MoE 路由不连续

专知会员服务

0+阅读 · 今天14:40

综述 | 周期表视角下的大模型推理：范式、方法与失败模式

综述 | 周期表视角下的大模型推理：范式、方法与失败模式

专知会员服务

0+阅读 · 今天14:36

《廉价自杀式无人机战争的军事战略影响：乌克兰和伊朗案例研究》

《廉价自杀式无人机战争的军事战略影响：乌克兰和伊朗案例研究》

专知会员服务

7+阅读 · 今天2:06

《面向反无人机作战的联邦式可解释射频–光电/红外情报融合：边缘人工智能优化、电子战韧性及分布式监视验证》

《面向反无人机作战的联邦式可解释射频–光电/红外情报融合：边缘人工智能优化、电子战韧性及分布式监视验证》

专知会员服务

5+阅读 · 今天1:37

ICML 2026 | FR3D：解耦自车运动的未来动态三维重建世界模型

ICML 2026 | FR3D：解耦自车运动的未来动态三维重建世界模型

专知会员服务

3+阅读 · 6月17日

【伯克利博士论文】迈向可扩展与自我演进的大语言模型智能体

【伯克利博士论文】迈向可扩展与自我演进的大语言模型智能体

专知会员服务

5+阅读 · 6月17日

学习数据的几何：形状空间分析数学综述

学习数据的几何：形状空间分析数学综述

专知会员服务

4+阅读 · 6月17日

《现代防空系统综述：架构、传感器、拦截器及新兴威胁环境对基础设施受限防御环境的影响》2026最新长综述

《现代防空系统综述：架构、传感器、拦截器及新兴威胁环境对基础设施受限防御环境的影响》2026最新长综述

专知会员服务

7+阅读 · 6月17日

定向能反无人机系统最新发展动态

定向能反无人机系统最新发展动态

专知会员服务

7+阅读 · 6月17日

从燃煤战舰到算法战争：水面指挥的永恒要求

从燃煤战舰到算法战争：水面指挥的永恒要求

专知会员服务

4+阅读 · 6月17日

《短程弹道再入飞行器拦截时间中的一项异常现象》

《短程弹道再入飞行器拦截时间中的一项异常现象》

专知会员服务

6+阅读 · 6月17日

《基于回归方法与任务上下文的对抗环境动态战术网络报文优先级排序》

《基于回归方法与任务上下文的对抗环境动态战术网络报文优先级排序》

专知会员服务

6+阅读 · 6月17日

美智库《战术级指挥控制的迫切要求：构建弹性机动式指挥控制网络》报告

美智库《战术级指挥控制的迫切要求：构建弹性机动式指挥控制网络》报告

专知会员服务

5+阅读 · 6月17日

《韩国国防政策与军备出口：韩国安全与国防政策如何塑造其国防工业与军备出口格局》最新100页报告

《韩国国防政策与军备出口：韩国安全与国防政策如何塑造其国防工业与军备出口格局》最新100页报告

专知会员服务

4+阅读 · 6月17日

ICML 2026 | VOTP：用视频基础模型与最优传输，让离线偏好强化学习只需少量反馈

ICML 2026 | VOTP：用视频基础模型与最优传输，让离线偏好强化学习只需少量反馈

专知会员服务

6+阅读 · 6月16日

相关VIP内容

EMNLP 2024 | 基于知识编辑的大模型敏感知识擦除

EMNLP 2024 | 基于知识编辑的大模型敏感知识擦除

专知会员服务

22+阅读 · 2024年11月19日

ICLR2024｜Mol-Instructions: 面向大模型的大规模生物分子指令数据集

ICLR2024｜Mol-Instructions: 面向大模型的大规模生物分子指令数据集

专知会员服务

12+阅读 · 2024年2月10日

EMNLP2023：预训练模型的知识反刍

EMNLP2023：预训练模型的知识反刍

专知会员服务

32+阅读 · 2023年11月20日

LLM in Medical Domain: 大语言模型在医学领域的应用

LLM in Medical Domain: 大语言模型在医学领域的应用

专知会员服务

103+阅读 · 2023年6月17日

ICLR 2022 | BEIT论文解读：将MLM无监督预训练应用到CV领域

ICLR 2022 | BEIT论文解读：将MLM无监督预训练应用到CV领域

专知会员服务

33+阅读 · 2022年3月24日

【CVPR 2022】跨模态检索的协同双流视觉-语言前训练模型，COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal Retrieval

【CVPR 2022】跨模态检索的协同双流视觉-语言前训练模型，COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal Retrieval

专知会员服务

13+阅读 · 2022年3月12日

【EMNLP 2019 最佳论文】信息瓶颈专门化单词嵌入（用于解析）（Specializing Word Embeddings（for Parsing）by Information Bottleneck）

【EMNLP 2019 最佳论文】信息瓶颈专门化单词嵌入（用于解析）（Specializing Word Embeddings（for Parsing）by Information Bottleneck）

专知会员服务

24+阅读 · 2019年11月20日

【MLA 2019】自然语言处理中的表示学习进展：从Transfomer到BERT，复旦大学邱锡鹏

【MLA 2019】自然语言处理中的表示学习进展：从Transfomer到BERT，复旦大学邱锡鹏

专知会员服务

100+阅读 · 2019年11月15日

【AAAI2020接受论文】隐式关系语言模型，CMU&微软，Latent Relation Language Models

【AAAI2020接受论文】隐式关系语言模型，CMU&微软，Latent Relation Language Models

专知会员服务

54+阅读 · 2019年11月12日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

热门VIP内容

开通专知VIP会员享更多权益服务

综述 | 周期表视角下的大模型推理：范式、方法与失败模式

《面向反无人机作战的联邦式可解释射频–光电/红外情报融合：边缘人工智能优化、电子战韧性及分布式监视验证》

ICML 2026 Spotlight | SmoothSMoE：解析稀疏 MoE 路由不连续

《廉价自杀式无人机战争的军事战略影响：乌克兰和伊朗案例研究》

相关资讯

赛尔原创 | EMNLP 2019 基于上下文感知的变分自编码器建模事件背景知识进行If-Then类型常识推理

赛尔原创 | EMNLP 2019 基于上下文感知的变分自编码器建模事件背景知识进行If-Then类型常识推理

哈工大SCIR

17+阅读 · 2019年9月23日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

44+阅读 · 2019年1月3日

meta learning 17年：MAML SNAIL

meta learning 17年：MAML SNAIL

CreateAMind

11+阅读 · 2019年1月2日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

李宏毅-201806-中文-Deep Reinforcement Learning精品课程分享

李宏毅-201806-中文-Deep Reinforcement Learning精品课程分享

深度学习与NLP

15+阅读 · 2018年6月20日

论文浅尝 | 嵌入常识知识的注意力 LSTM 模型用于特定目标的基于侧面的情感分析

论文浅尝 | 嵌入常识知识的注意力 LSTM 模型用于特定目标的基于侧面的情感分析

开放知识图谱

28+阅读 · 2018年6月11日

论文浅尝 | Improved Neural Relation Detection for KBQA

论文浅尝 | Improved Neural Relation Detection for KBQA

开放知识图谱

13+阅读 · 2018年1月21日

From Softmax to Sparsemax-ICML16（1）

From Softmax to Sparsemax-ICML16（1）

KingsGarden

74+阅读 · 2016年11月26日

相关论文

LLM Compression by Block Removal with Constrained Binary Optimization

Arxiv

0+阅读 · 6月17日

Latency Prediction for LLM Inference on NPU Systems

Arxiv

0+阅读 · 6月17日

Factorized Latent Reasoning for LLM-based Recommendation

Arxiv

0+阅读 · 6月12日

Revisiting Neural Processes via Fourier Transform and Volterra Series

Arxiv

0+阅读 · 6月11日

LLM-Guided Evolution for Medical Decision Pipelines

Arxiv

0+阅读 · 6月5日

From Layers to Submodules: Rethinking Granularity in Replacement-Based LLM Compression

Arxiv

0+阅读 · 6月1日

Neural Router: Semantic Content Matching for Agentic AI

Arxiv

0+阅读 · 5月25日

Neural-Actuarial Longevity Forecasting: Anchoring LSTMs for Explainable Risk Management

Arxiv

0+阅读 · 5月7日

A Survey on Neural Speech Synthesis

Arxiv

14+阅读 · 2021年6月30日

Latent Relation Language Models

Arxiv

21+阅读 · 2019年8月21日

相关基金

不同发育起源MSCs在骨缺损修复重建过程中的作用及机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

纳米复合结构的可控构筑及其在LSPR免疫在线检测中的应用研究

国家自然科学基金

0+阅读 · 2015年12月31日

废液中铀酰类化合物超灵敏检测用SERS基底纳米结构的设计与构建

国家自然科学基金

0+阅读 · 2015年12月31日

人脑MRI数据特征提取方法的研究与应用

国家自然科学基金

0+阅读 · 2015年12月31日

力感应细胞源性外泌体及其microRNA货物介导的细胞间通讯在骨重建稳态中调控机制及干预策略的研究

国家自然科学基金

0+阅读 · 2015年12月31日

基于miRNA转染的血管化BMSC膜片的构建、血管化与骨再生能力及其机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

低氧预处理自体骨髓间充质干细胞修复关节软骨损伤的MRI活体示踪研究

国家自然科学基金

0+阅读 · 2014年12月31日

iPSC-MSCs复合活性大孔CPC用于牙槽骨缺损修复及转归机制的研究

国家自然科学基金

0+阅读 · 2014年12月31日

低氧下骨髓间充质干细胞复合细胞外基质膜片构建3D组织工程皮肤的研究

国家自然科学基金

0+阅读 · 2014年12月31日

MSM人群中HIV感染者生命质量评价及预警模型研究

国家自然科学基金

0+阅读 · 2014年12月31日

微信扫码咨询专知VIP会员