Query auto-completion (QAC) has been widely studied in the context of web search, yet remains underexplored for in-document search, which we term DocQAC. DocQAC aims to enhance search productivity within long documents by helping users craft faster, more precise queries, even for complex or hard-to-spell terms. While global historical queries are available to both WebQAC and DocQAC, DocQAC uniquely accesses document-specific context, including the current document's content and its specific history of user query interactions. To address this setting, we propose a novel adaptive trie-guided decoding framework that uses user query prefixes to softly steer language models toward high-quality completions. Our approach introduces an adaptive penalty mechanism with tunable hyperparameters, enabling a principled trade-off between model confidence and trie-based guidance. To efficiently incorporate document context, we explore retrieval-augmented generation (RAG) and lightweight contextual document signals such as titles, keyphrases, and summaries. When applied to encoder-decoder models like T5 and BART, our trie-guided framework outperforms strong baselines and even surpasses much larger instruction-tuned models such as LLaMA-3 and Phi-3 on seen queries across both seen and unseen documents. This demonstrates its practicality for real-world DocQAC deployments, where efficiency and scalability are critical. We evaluate our method on a newly introduced DocQAC benchmark derived from ORCAS, enriched with query-document pairs. We make both the DocQAC dataset (https://bit.ly/3IGEkbH) and code (https://github.com/rahcode7/DocQAC) publicly available.


翻译:查询自动补全(QAC)已在网络搜索领域得到广泛研究,但在文档内搜索场景中仍鲜有探索,我们将此问题称为DocQAC。DocQAC旨在通过帮助用户构建更快速、更精确的查询(即使针对复杂或难拼写术语),提升长文档中的搜索效率。尽管网络QAC和DocQAC均可利用全局历史查询,但DocQAC独有地访问文档特定上下文,包括当前文档内容及其用户查询交互的特定历史记录。针对这一场景,我们提出了一种新颖的自适应字典树引导解码框架,该框架利用用户查询前缀软性引导语言模型生成高质量补全。我们的方法引入了一种可调超参数的自适应惩罚机制,实现了模型置信度与字典树引导之间的原则性权衡。为高效融入文档上下文,我们探索了检索增强生成(RAG)及轻量级上下文文档信号(如标题、关键词和摘要)。当应用于T5和BART等编码器-解码器模型时,我们的字典树引导框架在已见查询上显著超越强基线模型,甚至优于LLaMA-3和Phi-3等更大规模的指令调优模型,且对已见和未见文档均有效。这证明了该方法在真实DocQAC部署中的实用性,其中效率与可扩展性至关重要。我们在基于ORCAS构建并增强查询-文档对的DocQAC基准上评估了该方法。我们已公开发布DocQAC数据集(https://bit.ly/3IGEkbH)及代码(https://github.com/rahcode7/DocQAC)。

0
下载
关闭预览

相关内容

【CMU博士论文】混合知识架构问答系统,150页pdf
专知会员服务
41+阅读 · 2023年12月14日
【2022新书】文本与知识库问答系统,208页pdf
专知会员服务
81+阅读 · 2022年11月14日
【知乎】超越Lexical:用于文本搜索引擎的语义检索框架
专知会员服务
22+阅读 · 2020年8月28日
Python推荐系统框架:RecQ
专知
12+阅读 · 2019年1月21日
干货|当深度学习遇见自动文本摘要,seq2seq+attention
机器学习算法与Python学习
10+阅读 · 2018年5月28日
TextInfoExp:自然语言处理相关实验(基于sougou数据集)
全球人工智能
12+阅读 · 2017年11月12日
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
18+阅读 · 2017年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
6+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
8+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
13+阅读 · 8月7日
相关VIP内容
【CMU博士论文】混合知识架构问答系统,150页pdf
专知会员服务
41+阅读 · 2023年12月14日
【2022新书】文本与知识库问答系统,208页pdf
专知会员服务
81+阅读 · 2022年11月14日
【知乎】超越Lexical:用于文本搜索引擎的语义检索框架
专知会员服务
22+阅读 · 2020年8月28日
相关资讯
Python推荐系统框架:RecQ
专知
12+阅读 · 2019年1月21日
干货|当深度学习遇见自动文本摘要,seq2seq+attention
机器学习算法与Python学习
10+阅读 · 2018年5月28日
TextInfoExp:自然语言处理相关实验(基于sougou数据集)
全球人工智能
12+阅读 · 2017年11月12日
相关基金
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
18+阅读 · 2017年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员