Coded language is an important part of human communication. It refers to cases where users intentionally encode meaning so that the surface text differs from the intended meaning and must be decoded to be understood. Current language models handle coded language poorly. Progress has been limited by the lack of real-world datasets and clear taxonomies. This paper introduces CodedLang, a dataset of 7,744 Chinese Google Maps reviews, including 900 reviews with span-level annotations of coded language. We developed a seven-class taxonomy that captures common encoding strategies, including phonetic, orthographic, and cross-lingual substitutions. We benchmarked language models on coded language detection, classification, and review rating prediction. Results show that even strong models can fail to identify or understand coded language. Because many coded expressions rely on pronunciation-based strategies, we further conducted a phonetic analysis of coded and decoded forms. Our code and dataset are publicly available. Together, our results highlight coded language as an important and underexplored challenge for real-world NLP systems.
翻译:隐语是人类交流的重要组成部分,指用户故意编码信息,使表面文本与真实含义不同,需解码才能理解。当前语言模型对隐语的处理效果不佳,而真实世界数据集和清晰分类体系的缺乏限制了相关研究进展。本文提出CodedLang数据集,包含7,744条中文谷歌地图评论,其中900条带有隐语的跨度标注。我们构建了一个七类别分类体系,涵盖语音替换、字形替换、跨语言替换等常见编码策略。我们在隐语检测、分类及评论评分预测任务上对语言模型进行了基准测试。结果表明,即使是强大的模型也可能无法识别或理解隐语。鉴于许多隐语表达依赖于发音策略,我们进一步对编码与解码形式进行了语音分析。我们的代码和数据集已公开。综合而言,本研究成果凸显了隐语作为现实世界NLP系统中重要且尚未充分探索的挑战。