As language models increasingly consume one another's outputs, covert influence -- a phenomenon where a sender's payload (the behavioral disposition it is conditioned to propagate) transfers to a receiver through carriers undetectable by humans -- becomes a growing risk. We characterize this risk across three interfaces: supervised fine-tuning, on-policy distillation, and in-context learning, and find that they vary in the scale of influence achievable without leaving behind human-visible traces. Using inference-time per-sample attribution scores, we study covert influence across all three interfaces with the ability to select carriers that amplify training-time influence, unlocking payload transfers that prior work could not achieve. We further provide evidence that covert influence with natural-language carriers is a distinct phenomenon from prior studies using number carriers, as the latter is more resistant to human detection and less portable across model families. Together, these results suggest that the risk surface for covert influence is broader than previously recognized, and we study pointwise attribution scoring methods as a tool to investigate and mitigate it.
翻译:随着语言模型日益频繁地消费彼此的生成输出,隐性影响——即发送者通过人类无法察觉的载体,将其条件化传播的行为倾向(payload)传递给接收者的现象——正成为一种日益增长的风险。我们通过三种接口刻画了这一风险:监督微调、在线策略蒸馏和上下文学习,并发现这些接口在实现不留下人类可见痕迹的影响力规模上存在差异。利用推理时逐样本归因分数,我们研究了所有三种接口下的隐性影响,并具备筛选出能放大训练时影响的载体的能力,从而解锁了先前工作中无法实现的负载传递。我们进一步提供了证据,表明使用自然语言载体的隐性影响与先前使用数字载体的研究是截然不同的现象:前者更不易被人类检测,且在跨模型家族间可迁移性较低。综合来看,这些结果表明隐性影响的风险面比先前认知更为广泛,我们研究了逐点归因评分方法作为调查和缓解该风险的工具。