Structured prediction with large language models requires outputs that are label-accurate, ontology-constrained, structurally valid, and evidence-grounded under label imbalance and heterogeneous group difficulty. We present a unified framework for ontology-constrained generation. First, we introduce a modular prompt-engineering architecture combining XML-style structure, expert disambiguation rules, chain-of-thought reasoning, metadata-aware decision logic, schema contracts, and a self-validation gate. It targets recurrent in-context failures, including format drift, label ambiguity, evidence hallucination, and metadata-conditioned confusion. Second, we propose STaR-DRO, combining Tsallis mirror ascent, sparse entmax-style primal mapback, EMA-smoothed group-loss tracking, rescaled ascent signals, and bounded excess-only multipliers. Unlike conventional DRO, which relies on dense Shannon-entropy exponentiated-gradient updates, can introduce high-variance stochastic reweighting, assigns positive adversarial mass to groups that are not persistently hard, and incurs costs through simplex competition, STaR-DRO upweights only persistently hard groups without suppressing easier ones. We evaluate the framework on EPPC Miner, a clinically grounded high-stakes structured-prediction task requiring hierarchical label prediction and evidence-span extraction from patient-provider secure messages. Across 1B-70B Llama models, prompt engineering improves zero-shot extraction, yielding an average label F1 gain of +14.46 and a Span F1 gain of +17.40. Building on supervised fine-tuning, STaR-DRO further improves accuracy and robustness, increasing average label F1 by +1.08 and +2.20 while reducing mean groupwise validation cross-entropy by 21.3% and 14.8% relative to SFT and standard DRO, respectively. These results advance reliable automated communication mining for patient-centered clinical care analysis.
翻译:针对大语言模型的结构化预测,要求在标签不平衡及异质群体难度下,输出需满足标签准确、本体约束、结构有效且基于证据可验证。我们提出一个统一的框架用于本体约束生成。首先,引入模块化提示工程架构,结合XML风格结构、专家消歧规则、思维链推理、元数据感知决策逻辑、模式契约及自验证门控机制,旨在应对上下文中的重复性错误,包括格式漂移、标签歧义、证据幻觉及元数据条件混淆。其次,提出STaR-DRO,融合Tsallis镜像上升法、稀疏entmax风格原始映射回退、EMA平滑群体损失追踪、缩放上升信号及有界超额乘子。与依赖密集香农熵指数梯度更新的传统DRO不同——后者可能引入高方差随机重加权、对非持续困难群体赋予正对抗质量、并通过单纯形竞争产生成本——STaR-DRO仅对持续困难群体进行上加权,而不抑制较易群体。我们以EPPC Miner作为评估框架,这是一项临床高风险结构化预测任务,需从医患安全消息中完成层级标签预测及证据片段抽取。在1B-70B Llama模型上,提示工程提升零样本抽取性能,平均标签F1提升+14.46,片段F1提升+17.40。基于监督微调,STaR-DRO进一步优化准确性与鲁棒性,相较SFT与标准DRO,平均标签F1分别提升+1.08和+2.20,同时将群体验证交叉熵均值分别降低21.3%和14.8%。这些成果推进了以患者为中心的临床护理分析中可靠自动化通信挖掘的发展。