Markedness in natural language is often associated with non-literal meanings in discourse. Differential Object Marking (DOM) in Korean is one instance of this phenomenon, where post-positional markers are selected based on both the semantic features of the noun phrases and the discourse features that are orthogonal to the semantic features. Previous work has shown that distributional models of language recover certain semantic features of words -- do these models capture implied discourse-level meanings as well? We evaluate whether a set of large language models are capable of associating discourse meanings with different object markings in Korean. Results suggest that discourse meanings of a grammatical marker can be more challenging to encode than that of a discourse marker.
翻译:自然语言中的标记性往往与话语中的非字面意义相关。韩语的区别性宾语标记(DOM)即为这一现象的实例,其后置标记的选择既取决于名词短语的语义特征,也取决于与语义特征正交的话语特征。已有研究表明,语言的分布模型能够复原词汇的特定语义特征——那么这些模型是否也能捕捉隐含的话语层面意义?本研究评估了一系列大型语言模型是否能够将韩语中不同宾语标记与话语意义相关联。结果表明,相较于话语标记,语法标记的话语意义可能更难被模型编码。