We introduce CAPC-CG, the Chinese Adaptive Policy Communication (Central Government) Corpus, the first open dataset of Chinese policy directives annotated with a five-color taxonomy of clear and ambiguous language categories, building on Ang's theory of adaptive policy communication. Spanning 1949-2023, this corpus includes national laws, administrative regulations, and ministerial rules issued by China's top authorities. Each document is segmented into paragraphs, producing a total of 3.3 million units. Alongside the corpus, we release comprehensive metadata, a two-round labeling framework, and a gold-standard annotation set developed by expert and trained coders. Inter-annotator agreement achieves a Fleiss's kappa of K = 0.86 on directive labels, indicating high reliability for supervised modeling. We provide baseline classification results with several large language models (LLMs), together with our annotation codebook, and describe patterns from the dataset. This release aims to support downstream tasks and multilingual NLP research in policy communication.
翻译:我们提出CAPC-CG——中国适应性政策沟通(中央政府)语料库,这是首个基于Ang的适应性政策沟通理论、采用清晰与模糊语言类别的五色分类法标注的中国政策指令开放数据集。该语料库涵盖1949年至2023年期间中国最高权力机关颁布的国家法律、行政法规和部门规章。每条文档被切分为段落,共计生成330万个标注单元。除语料库外,我们同时发布完整的元数据、两轮标注框架以及由专家和培训编码员开发的金标准标注集。标注者间在指令标签上达到Fleiss's kappa系数K=0.86的一致性,表明其适用于监督建模的高可靠性。我们提供了基于多个大语言模型的基线分类结果,并附上标注编码手册,描述数据集中的统计模式。该资源旨在支持政策沟通领域的下游任务及多语言自然语言处理研究。