Recommendation systems have become central gatekeepers of online information, shaping user behaviour across a wide range of activities. In response, users increasingly organize and coordinate to steer algorithmic outcomes toward diverse goals, such as promoting relevant content or limiting harmful material, relying on platform affordances -- such as likes, reviews, or ratings. While these mechanisms can serve beneficial purposes, they can also be leveraged for adversarial manipulation, particularly in systems where such feedback directly informs safety guarantees. In this paper, we study this vulnerability in recently proposed risk-controlling recommender systems, which use binary user feedback (e.g., "Not Interested") to provably limit exposure to unwanted content via conformal risk control. We empirically demonstrate that their reliance on aggregate feedback signals makes them inherently susceptible to coordinated adversarial user behaviour. Using data from a large-scale online video-sharing platform, we show that a small coordinated group (comprising only 1% of the user population) can induce up to a 20% degradation in nDCG for non-adversarial users by exploiting the affordances provided by risk-controlling recommender systems. We evaluate simple, realistic attack strategies that require little to no knowledge of the underlying recommendation algorithm and find that, while coordinated users can significantly harm overall recommendation quality, they cannot selectively suppress specific content groups through reporting alone. Finally, we propose a mitigation strategy that shifts guarantees from the group level to the user level, showing empirically how it can reduce the impact of adversarial coordinated behaviour while ensuring personalized safety for individuals.
翻译:推荐系统已成为在线信息的核心守门人,深刻影响着用户在各类活动中的行为。作为回应,用户正日益组织并协作,利用平台提供的功能(如点赞、评论或评分)引导算法结果趋向多元目标——例如推广相关内容或限制有害信息。尽管这些机制可服务于正当目的,但在反馈信号直接影响安全保证的系统中,它们也可能被用于对抗性操纵。本文针对近期提出的风险控制推荐系统研究这一漏洞:此类系统基于二元用户反馈(如"不感兴趣"),通过共形风险控制方法在概率意义上限制用户接触不想要的内容。我们通过实验证明,其对聚合反馈信号的依赖使其天然易受协同对抗性用户行为的攻击。基于某大型在线视频分享平台的数据,我们显示仅占用户总数1%的小规模协同群体可通过利用风险控制推荐系统的交互设计,导致非对抗性用户的nDCG指标下降高达20%。我们评估了多种简单、现实的攻击策略——这些策略几乎不需要对底层推荐算法有深入了解,并发现:虽然协同用户能显著损害整体推荐质量,但仅通过举报机制无法选择性压制特定内容类别。最后,我们提出一种将保障机制从群体层级迁移至用户层级的缓解策略,实证表明该方法能在确保个体化安全的同时降低对抗性协同行为的影响。