User modeling, which aims to capture users' characteristics or interests, heavily relies on task-specific labeled data and suffers from the data sparsity issue. Several recent studies tackled this problem by pre-training the user model on massive user behavior sequences with a contrastive learning task. Generally, these methods assume different views of the same behavior sequence constructed via data augmentation are semantically consistent, i.e., reflecting similar characteristics or interests of the user, and thus maximizing their agreement in the feature space. However, due to the diverse interests and heavy noise in user behaviors, existing augmentation methods tend to lose certain characteristics of the user or introduce noisy interests. Thus, forcing the user model to directly maximize the similarity between the augmented views may result in a negative transfer. To this end, we propose to replace the contrastive learning task with a new pretext task: Augmentation-Adaptive Self-Supervised Ranking (AdaptSSR), which alleviates the requirement of semantic consistency between the augmented views while pre-training a discriminative user model. Specifically, we adopt a multiple pairwise ranking loss which trains the user model to capture the similarity orders between the implicitly augmented view, the explicitly augmented view, and views from other users. We further employ an in-batch hard negative sampling strategy to facilitate model training. Moreover, considering the distinct impacts of data augmentation on different behavior sequences, we design an augmentation-adaptive fusion mechanism to automatically adjust the similarity order constraint applied to each sample based on the estimated similarity between the augmented views. Extensive experiments on both public and industrial datasets with six downstream tasks verify the effectiveness of AdaptSSR.
翻译:摘要:用户建模旨在捕捉用户的特征或兴趣,其高度依赖于特定任务的标注数据,并面临数据稀疏性问题。近期研究通过在大规模用户行为序列上采用对比学习任务预训练用户模型来应对这一挑战。通常,这些方法假设通过数据增强构建的同一行为序列的不同视角在语义上保持一致(即反映用户相似的特征或兴趣),从而在特征空间中最大化其一致性。然而,由于用户行为的多样性和大量噪声,现有增强方法可能丢失用户的某些特征或引入噪声兴趣。因此,强迫用户模型直接最大化增强视角之间的相似性可能导致负迁移。为此,我们提出用新的前置任务——增强自适应自监督排序(AdaptSSR)替代对比学习任务,该方法在预训练判别性用户模型时缓解了对增强视角间语义一致性的要求。具体而言,我们采用多对排序损失函数,训练用户模型捕捉隐式增强视角、显式增强视角及其他用户视角之间的相似性顺序。此外,我们引入批内硬负采样策略以促进模型训练。针对数据增强对不同行为序列的差异化影响,我们设计了增强自适应融合机制,基于增强视角间的估计相似性自动调整每个样本的相似性顺序约束。在包含六个下游任务的公共数据集与工业数据集上的大量实验验证了AdaptSSR的有效性。