Text readability assessment has gained significant attention from researchers in various domains. However, the lack of exploration into corpus compatibility poses a challenge as different research groups utilize different corpora. In this study, we propose a novel evaluation framework, Cross-corpus text Readability Compatibility Assessment (CRCA), to address this issue. The framework encompasses three key components: (1) Corpus: CEFR, CLEC, CLOTH, NES, OSP, and RACE. Linguistic features, GloVe word vector representations, and their fusion features were extracted. (2) Classification models: Machine learning methods (XGBoost, SVM) and deep learning methods (BiLSTM, Attention-BiLSTM) were employed. (3) Compatibility metrics: RJSD, RRNSS, and NDCG metrics. Our findings revealed: (1) Validated corpus compatibility, with OSP standing out as significantly different from other datasets. (2) An adaptation effect among corpora, feature representations, and classification methods. (3) Consistent outcomes across the three metrics, validating the robustness of the compatibility assessment framework. The outcomes of this study offer valuable insights into corpus selection, feature representation, and classification methods, and it can also serve as a beginning effort for cross-corpus transfer learning.
翻译:文本可读性评估已引起各领域研究者的广泛关注。然而,不同研究团队采用不同语料库的现状导致语料库兼容性问题尚未得到充分探索。本研究提出一种新的评估框架——跨语料库文本可读性兼容性评估(CRCA)以解决该问题。该框架包含三个核心组件:(1)语料库:CEFR、CLEC、CLOTH、NES、OSP和RACE,提取了语言特征、GloVe词向量表示及其融合特征;(2)分类模型:采用机器学习方法(XGBoost、SVM)与深度学习方法(BiLSTM、Attention-BiLSTM);(3)兼容性指标:RJSD、RRNSS和NDCG指标。研究发现:(1)验证了语料库兼容性,其中OSP与其他数据集存在显著差异;(2)语料库、特征表示与分类方法之间存在适应效应;(3)三项指标结果一致,验证了兼容性评估框架的鲁棒性。本研究结果为语料库选择、特征表示和分类方法提供了重要见解,同时可作为跨语料库迁移学习的初步探索。