Retrieval-Augmented Generation (RAG) systems depend critically on document chunking quality for retrieving relevant context. Fixed chunking segments documents into uniform units irrespective of semantics or user intent, producing a precision-recall trade-off unresolvable by tuning chunk size alone. Semantic and agentic methods partially address these limitations but do not integrate user queries at the chunking stage. We present Query-Adaptive Semantic Chunking (QASC), which dynamically constructs chunks by integrating queries into segmentation through three mechanisms: cosine similarity scoring between sentence and query embeddings to identify seed sentences, contextual window expansion around seeds to preserve coherence, and chunk-level score aggregation to ensure holistic relevance. We evaluate QASC on 100 technical documents across 200 queries spanning four types, comparing against fixed chunking at five granularities, recursive splitting, semantic chunking, and agentic chunking. QASC achieves an F1-score of 0.85, a relative improvement of 18-27% over fixed chunking and 8-12% over semantic and agentic alternatives. Ablation studies confirm each component contributes meaningfully. Human evaluation by three annotators (Cohen kappa = 0.82) corroborates that QASC produces more relevant and coherent chunks than existing methods.


翻译:检索增强生成系统高度依赖文档分块质量以获取相关上下文。固定分块将文档切分为统一单元,忽略语义与用户意图,导致精确率与召回率之间难以通过调整分块粒度单独解决的权衡。语义分块与代理式方法虽部分缓解了上述局限,但未在分块阶段整合用户查询。我们提出查询自适应语义分块,通过三种机制将查询融入分割过程以动态构建分块:基于句子与查询嵌入的余弦相似度评分识别种子句,围绕种子句扩展上下文窗口以保持连贯性,以及通过分块级评分聚合确保整体相关性。我们在涵盖四类查询的100篇技术文档、200个查询上评估QASC,并与五种粒度下的固定分块、递归分割、语义分块及代理式分块进行对比。QASC的F1分数达0.85,较固定分块相对提升18-27%,较语义分块与代理式方法相对提升8-12%。消融实验证实各组件均有显著贡献。三名标注者的人工评估(Cohen kappa=0.82)进一步印证QASC较现有方法生成更相关且连贯的分块。

0
下载
关闭预览

相关内容

多模态检索增强生成综述
专知会员服务
40+阅读 · 2025年4月15日
定制化大型语言模型的图检索增强生成综述
专知会员服务
39+阅读 · 2025年1月28日
《大型语言模型中基于检索的文本生成》综述
专知会员服务
60+阅读 · 2024年4月18日
【WWW2024】元认知检索-增强大型语言模型
专知会员服务
50+阅读 · 2024年2月26日
【CVPR2023】基础模型驱动弱增量学习的语义分割
专知会员服务
18+阅读 · 2023年3月2日
搜索query意图识别的演进
DataFunTalk
13+阅读 · 2020年11月15日
【资源】领域自适应相关论文、代码分享
专知
32+阅读 · 2019年10月12日
用Attention玩转CV,一文总览自注意力语义分割进展
专栏 | NLP概述和文本自动分类算法详解
机器之心
12+阅读 · 2018年7月24日
干货|当深度学习遇见自动文本摘要,seq2seq+attention
机器学习算法与Python学习
10+阅读 · 2018年5月28日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员