Detecting Schwartz values in political text is difficult because implicit cues often depend on surrounding arguments and fine-grained distinctions between neighboring values. We study when context and explicit moral knowledge help sentence-level value detection. Using the ValuesML/Touché ValueEval format, we compare sentence, window, and full-document inputs; no-RAG and retrieval-augmented settings with a curated moral knowledge base; supervised DeBERTa-v3-base/large encoders; and zero-shot LLMs from 12B to 123B parameters. The results show that more context is not uniformly better: full-document context improves supervised DeBERTa encoders by 3.8-4.8 macro-F1 points over sentence-only input, but does not consistently help zero-shot LLMs. Retrieved moral knowledge is more consistently useful in matched comparisons, improving each tested model family and context condition under early fusion. However, scaling from DeBERTa-v3-base to large and from 12B to larger LLMs does not guarantee gains, and simple early fusion outperforms the tested late-fusion and cross-attention RAG variants for encoders. Per-value analyses show that context and retrieval help most for socially situated or conceptually confusable values. These findings suggest that value-sensitive NLP should evaluate context, knowledge, and model family jointly rather than treating longer inputs or larger models as universal improvements.
翻译:检测政治文本中的施瓦茨价值观具有挑战性,因为隐性线索往往依赖于周围论证及相邻价值观之间的细微区分。本研究系统探究了上下文与显性道德知识对句子级价值观检测的影响。基于ValuesML/Touché ValueEval格式,我们对比了句子、窗口与全文输入;无检索增强与结合 curated 道德知识库的检索增强设置;有监督的DeBERTa-v3-base/large编码器;以及从12B到123B参数量的零样本大型语言模型。结果表明,更多上下文并非总是更好:全文上下文能使有监督的DeBERTa编码器宏F1值比仅用句子输入提升3.8-4.8个百分点,但对零样本语言模型的提升并不稳定。在匹配比较中,检索到的道德知识更具一致性优势:采用早期融合方式时,每个测试模型系列与上下文条件均获提升。然而,从DeBERTa-v3-base扩展至large、从12B扩展至更大语言模型并不能保证增益,且对于编码器而言,简单早期融合优于所测试的晚期融合与交叉注意力检索增强变体。按价值观分类分析表明,上下文与检索对社交定位或概念易混淆的价值观帮助最大。这些发现提示,价值观敏感的NLP研究应联合评估上下文、知识与模型系列,而非将更长输入或更大模型视为普适改进方案。