Structured prediction with large language models requires outputs that are label-accurate, ontology-constrained, structurally valid, and evidence-grounded under label imbalance and heterogeneous group difficulty. We present a unified framework for ontology-constrained generation. First, we introduce a modular prompt-engineering architecture combining XML-style structure, expert disambiguation rules, chain-of-thought reasoning, metadata-aware decision logic, schema contracts, and a self-validation gate. It targets recurrent in-context failures, including format drift, label ambiguity, evidence hallucination, and metadata-conditioned confusion. Second, we propose STaR-DRO, combining Tsallis mirror ascent, sparse entmax-style primal mapback, EMA-smoothed group-loss tracking, rescaled ascent signals, and bounded excess-only multipliers. Unlike conventional DRO, which relies on dense Shannon-entropy exponentiated-gradient updates, can introduce high-variance stochastic reweighting, assigns positive adversarial mass to groups that are not persistently hard, and incurs costs through simplex competition, STaR-DRO upweights only persistently hard groups without suppressing easier ones. We evaluate the framework on EPPC Miner, a clinically grounded high-stakes structured-prediction task requiring hierarchical label prediction and evidence-span extraction from patient-provider secure messages. Across 1B-70B Llama models, prompt engineering improves zero-shot extraction, yielding an average label F1 gain of +14.46 and a Span F1 gain of +17.40. Building on supervised fine-tuning, STaR-DRO further improves accuracy and robustness, increasing average label F1 by +1.08 and +2.20 while reducing mean groupwise validation cross-entropy by 21.3% and 14.8% relative to SFT and standard DRO, respectively. These results advance reliable automated communication mining for patient-centered clinical care analysis.


翻译:针对大语言模型的结构化预测,要求在标签不平衡及异质群体难度下,输出需满足标签准确、本体约束、结构有效且基于证据可验证。我们提出一个统一的框架用于本体约束生成。首先,引入模块化提示工程架构,结合XML风格结构、专家消歧规则、思维链推理、元数据感知决策逻辑、模式契约及自验证门控机制,旨在应对上下文中的重复性错误,包括格式漂移、标签歧义、证据幻觉及元数据条件混淆。其次,提出STaR-DRO,融合Tsallis镜像上升法、稀疏entmax风格原始映射回退、EMA平滑群体损失追踪、缩放上升信号及有界超额乘子。与依赖密集香农熵指数梯度更新的传统DRO不同——后者可能引入高方差随机重加权、对非持续困难群体赋予正对抗质量、并通过单纯形竞争产生成本——STaR-DRO仅对持续困难群体进行上加权,而不抑制较易群体。我们以EPPC Miner作为评估框架,这是一项临床高风险结构化预测任务,需从医患安全消息中完成层级标签预测及证据片段抽取。在1B-70B Llama模型上,提示工程提升零样本抽取性能,平均标签F1提升+14.46,片段F1提升+17.40。基于监督微调,STaR-DRO进一步优化准确性与鲁棒性,相较SFT与标准DRO,平均标签F1分别提升+1.08和+2.20,同时将群体验证交叉熵均值分别降低21.3%和14.8%。这些成果推进了以患者为中心的临床护理分析中可靠自动化通信挖掘的发展。

0
下载
关闭预览

相关内容

ICML 2026 | 面向视觉语言模型的语义鲁棒性认证
专知会员服务
10+阅读 · 6月21日
决策智能中的时间序列预测大模型
专知会员服务
34+阅读 · 1月7日
论文浅尝 | 基于事理图谱的脚本事件预测
开放知识图谱
10+阅读 · 2019年12月10日
GitHub超9千星:一个API调用27个NLP预训练模型
新智元
17+阅读 · 2019年7月22日
一文读懂目标检测:R-CNN、Fast R-CNN、Faster R-CNN、YOLO、SSD
七月在线实验室
11+阅读 · 2018年7月18日
推荐|上交大推出Texygen:文本生成模型的基准测试平台
回归预测&时间序列预测
GBASE数据工程部数据团队
44+阅读 · 2017年5月17日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
印度精确打击与指挥架构的断层
专知会员服务
4+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
6+阅读 · 7月20日
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
7+阅读 · 7月19日
《无人机蜂群通信技术研究》50页
专知会员服务
10+阅读 · 7月19日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员