User retention is a key metric to measure long-term engagement in modern platforms. In real-time bidding (RTB) advertising system for user re-engagement, the retention model is required to predict future revisit probability at bidding time, before the user converts and consumes any content. Although post-conversion content, termed Onboarding Content, provides highly informative signals for retention prediction, directly using it in training causes severe feature leakage and creates a gap between training and serving. To address this issue, we propose OCARM, a two-stage distillation-aligned framework for Onboarding Content Augmented Retention Modeling, enabling the model to implicitly capture future content using only observable features during inference. In the first stage, we deliberately expose onboarding content to train a hierarchical encoder that produces teacher representations. In the second stage, a user encoder is aligned with the frozen teacher through distillation, allowing the model to approximate the inaccessible onboarding signals without leakage. Extensive offline experiments and online A/B tests demonstrate that our framework achieves consistent improvements in a real-world growth scenario.
翻译:用户留存是衡量现代平台长期参与度的关键指标。在面向用户重新激活的实时竞价(RTB)广告系统中,留存模型需在用户转化并消费内容前的竞价时刻预测其未来回访概率。尽管转化后内容(称为“新手引导内容”)为留存预测提供了高度信息化的信号,但直接将其用于训练会导致严重的特征泄露,并造成训练与线上服务之间的鸿沟。为解决此问题,我们提出OCARM——一种面向新手引导内容增强型留存建模的两阶段蒸馏对齐框架,使模型在推理时仅利用可观测特征便能隐式捕捉未来内容。在第一阶段,我们刻意引入新手引导内容以训练生成教师表示的分层编码器;在第二阶段,通过蒸馏将用户编码器与冻结的教师模型对齐,从而使模型能在无泄露情况下逼近不可访问的新手引导信号。大量离线实验与在线A/B测试表明,该框架在真实业务增长场景中均实现了一致性提升。