The scaling laws for recommender systems have been increasingly validated, where MetaFormer-based architectures consistently benefit from increased model depth, hidden dimensionality, and user behavior sequence length. However, whether representation capacity scales proportionally with parameter growth remains largely unexplored. Prior studies on RankMixer reveal that the effective rank of token representations exhibits a damped oscillatory trajectory across layers, failing to increase consistently with depth and even degrading in deeper layers. Motivated by this observation, we propose \textbf{RankUp}, an architecture designed to mitigate representation collapse and enhance expressive capacity through randomized permutation splitting over sparse features, a multi-embedding paradigm, global token integration, crossed pretrained embedding tokens and task-specific token decoupling. RankUp has been fully deployed in large-scale production across Weixin Video Accounts, Official Accounts and Moments, yielding GMV improvements of 3.41\%, 4.81\% and 2.21\%, respectively.
翻译:推荐系统的规模定律已得到日益验证,基于MetaFormer的架构持续受益于模型深度、隐藏维度和用户行为序列长度的增长。然而,表示能力是否随参数增长成比例扩展仍未得到充分探索。先前关于RankMixer的研究揭示,token表示的秩在跨层时呈现阻尼振荡轨迹,未能随深度增加而持续提升,甚至在深层出现退化。受此发现启发,我们提出\textbf{RankUp}架构,通过稀疏特征的随机置换分割、多嵌入范式、全局token整合、跨预训练嵌入token以及任务特定token解耦,旨在缓解表示坍塌并增强表达能力。RankUp已在微信视频号、公众号和朋友圈的大规模生产环境中全面部署,分别实现了3.41%、4.81%和2.21%的GMV提升。