The rapid advancement of large language models has increasingly blurred the boundary between human-written and AI-generated text, raising societal risks such as misinformation dissemination, authorship ambiguity, and threats to intellectual property rights. These concerns highlight the urgent need for effective and reliable detection methods. While existing training-free approaches often achieve strong performance by aggregating token-level signals into a global score, they typically assume uniform token contributions, making them less robust under short sequences or localized token modifications. To address these limitations, we propose Exons-Detect, a training-free method for AI-generated text detection based on an exon-aware token reweighting perspective. Exons-Detect identifies and amplifies informative exonic tokens by measuring hidden-state discrepancy under a dual-model setting, and computes an interpretable translation score from the resulting importance-weighted token sequence. Empirical evaluations demonstrate that Exons-Detect achieves state-of-the-art detection performance and exhibits strong robustness to adversarial attacks and varying input lengths. In particular, it attains a 2.2\% relative improvement in average AUROC over the strongest prior baseline on DetectRL.
翻译:摘要:大语言模型的快速发展日益模糊了人写文本与AI生成文本的界限,引发信息传播错误、作者身份模糊及知识产权威胁等社会风险。这些担忧凸显了对有效且可靠检测方法的迫切需求。现有免训练方法虽常通过将令牌级信号聚合为全局评分实现优异性能,但通常假设令牌贡献均匀,导致其在短序列或局部令牌修改场景下鲁棒性不足。为解决上述局限,我们提出Exons-Detect——一种基于外显子感知令牌重加权视角的免训练AI文本检测方法。Exons-Detect通过双模型设置下的隐藏状态差异识别并放大信息性外显子令牌,继而从生成的加权令牌序列中计算可解释的翻译评分。实证评估表明,Exons-Detect实现了检测性能的最优水平,并对对抗攻击与可变输入长度展现出强鲁棒性。特别地,该方法在DetectRL基准测试上相较最强现有基线取得了平均AUROC 2.2%的相对提升。