Large Language Models (LLMs) offer unprecedented potential for enhancing recommendation systems through their world knowledge and reasoning capabilities. However, existing approaches often rely on structured IDs or offline processing, limiting semantic richness, real-time adaptability, and user-facing interpretability. In this paper, we introduce a novel framework that enables real-time generation of LLM-based user interest personas for a large-scale commercial video recommendation platform. Our method generates natural-language user interest personas that address the exploitation-exploration trade-off by combining the summarization of existing interests with novel topics, directly during serving. To overcome the computational challenges of online LLM inference at a billion-user scale, we design a cost-efficient architecture leveraging knowledge distillation, asynchronous inference, and input optimization via semantically clustered video representations. Extensive offline evaluations, user studies, and live A/B tests demonstrate significant improvements in viewer value. This work bridges the gap between high-level semantic understanding and industrial-scale recommendation, paving the way for more dynamic, explainable, and satisfying personalized experiences.
翻译:大语言模型凭借其世界知识与推理能力,为增强推荐系统提供了前所未有的潜力。然而现有方法多依赖结构化标识符或离线处理,限制了语义丰富性、实时适应能力及面向用户的解释性。本文提出一种新型框架,可在大规模商业视频推荐平台中实时生成基于大语言模型的用户兴趣画像。该方法通过融合既有兴趣总结与新兴话题,直接在线推理阶段应对探索-利用权衡,生成自然语言形式的用户兴趣画像。为克服十亿级用户规模下大语言模型在线推理的计算挑战,我们设计了一种成本高效架构,通过知识蒸馏、异步推理及基于语义聚类视频表示的输入优化实现。广泛的离线评估、用户研究与在线A/B测试表明,该方法在观众价值方面实现了显著提升。本研究弥合了高层语义理解与工业级推荐之间的鸿沟,为构建更动态、可解释且令人满意的个性化体验铺平了道路。