In enterprise settings, efficiently retrieving relevant information from large and complex knowledge bases is essential for operational productivity and informed decision-making. This research presents a systematic empirical framework for metadata enrichment using large language models (LLMs) to enhance document retrieval in Retrieval-Augmented Generation (RAG) systems. Our approach employs a structured pipeline that dynamically generates meaningful metadata for document segments, substantially improving their semantic representations and retrieval accuracy. Through a controlled 3 X 3 experimental matrix, we compare three chunking strategies -- semantic, recursive, and naive -- and evaluate their interactions with three embedding techniques -- content-only, TF-IDF weighted, and prefix-fusion -- isolating the contribution of each component through ablation analysis. The results demonstrate that metadata-enriched approaches consistently outperform content-only baselines, with recursive chunking paired with TF-IDF weighted embeddings yielding 82.5% precision and naive chunking with prefix-fusion achieving the strongest ranking quality (NDCG 0.813). Our evaluation employs cross-encoder reranking for silver-standard ground truth generation, with statistical significance confirmed via Bonferroni-corrected paired t-tests. These findings confirm that metadata enrichment improves vector space organization and retrieval effectiveness while maintaining sub-30 ms P95 latency, providing a quantitative decision framework for deploying high-performance, scalable RAG systems in enterprise settings.


翻译:在企业场景中,从庞大复杂的知识库中高效检索相关信息,对于提升运营生产力和支持精准决策至关重要。本研究提出了一种系统化的实证框架,通过利用大语言模型(LLM)进行元数据增强来提升检索增强生成(RAG)系统中的文档检索能力。该方法采用结构化流水线,动态生成文档片段的语义化元数据,显著改善了其语义表示与检索准确率。通过受控的3X3实验矩阵,我们比较了三种分块策略——语义分块、递归分块与朴素分块,并评估了它们与三种嵌入技术的交互效应:纯内容嵌入、TF-IDF加权嵌入及前缀融合嵌入,同时借助消融分析分离了各组件的贡献。实验结果表明,元数据增强方法始终优于纯内容基线;其中递归分块结合TF-IDF加权嵌入的方案实现了82.5%的精确率,而朴素分块搭配前缀融合嵌入的方案则取得了最优排序质量(NDCG 0.813)。本研究采用交叉编码器重排序生成银标准真实标签,并通过经Bonferroni校正的配对t检验确认统计显著性。这些发现证实了元数据增强在优化向量空间组织与检索效能的同时,能将P95延迟控制在30毫秒以内,从而为企业部署高性能、可扩展的RAG系统提供了量化决策框架。

0
下载
关闭预览

相关内容

检索增强生成(RAG)技术,261页slides
专知会员服务
42+阅读 · 2025年10月16日
检索增强生成(RAG)与推理的协同作用:一项系统综述
专知会员服务
16+阅读 · 2025年4月27日
定制化大型语言模型的图检索增强生成综述
专知会员服务
39+阅读 · 2025年1月28日
智能体检索增强生成:关于智能体RAG的综述
专知会员服务
95+阅读 · 2025年1月21日
RAG 与 LLMs 的结合 - 迈向检索增强的大型语言模型综述
专知会员服务
101+阅读 · 2024年5月13日
【MIT博士论文】数据高效强化学习,176页pdf
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
18+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
25+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
8+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
VIP会员
最新内容
反制无人机:乌克兰提供的五点启示
专知会员服务
1+阅读 · 今天14:52
《各指挥层级均亟需红队能力》报告
专知会员服务
2+阅读 · 今天14:46
《远程传感对战机架次生成影响的仿真研究》80页
《航电任务系统框架(FAMOS)》50页报告
专知会员服务
4+阅读 · 9月22日
《对抗行动中的人工智能与自主性》智库报告
专知会员服务
7+阅读 · 9月22日
《从数据到胜利:战争中的分析优势之争》
专知会员服务
9+阅读 · 9月22日
战争不仅需要机器人:人类仍不可或缺
专知会员服务
5+阅读 · 9月21日
《描绘美国防部创新基础设施的未来蓝图》100页
专知会员服务
10+阅读 · 9月21日
相关基金
国家自然科学基金
18+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
25+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
8+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员