VK-LSVD: A Large-Scale Industrial Dataset for Short-Video Recommendation

from arxiv, Accepted to The ACM Web Conference 2026 (WWW '26). Preprint of conference paper. 7 pages, 2 (7) figures, 4 tables. Dataset available at: https://huggingface.co/datasets/deepvk/VK-LSVD

Short-video recommendation presents unique challenges, such as modeling rapid user interest shifts from implicit feedback, but progress is constrained by a lack of large-scale open datasets that reflect real-world platform dynamics. To bridge this gap, we introduce the VK Large Short-Video Dataset (VK-LSVD), the largest publicly available industrial dataset of its kind. VK-LSVD offers an unprecedented scale of over 40 billion interactions from 10 million users and almost 20 million videos over six months, alongside rich features including content embeddings, diverse feedback signals, and contextual metadata. Our analysis supports the dataset's quality and diversity. The dataset's immediate impact is confirmed by its central role in the live VK RecSys Challenge 2025. VK-LSVD provides a vital, open dataset to use in building realistic benchmarks to accelerate research in sequential recommendation, cold-start scenarios, and next-generation recommender systems.

翻译：短视频推荐面临独特的挑战，例如从隐式反馈中建模用户兴趣的快速转移，但由于缺乏反映真实平台动态的大规模开放数据集，其进展受到限制。为弥补这一空白，我们推出了VK大规模短视频数据集（VK-LSVD），这是目前同类中最大的公开工业数据集。VK-LSVD提供了前所未有的规模，包含六个月内来自1000万用户与近2000万视频的超过400亿次交互，同时附带丰富的特征，包括内容嵌入、多样化的反馈信号以及上下文元数据。我们的分析验证了该数据集的质量与多样性。该数据集在正在进行的VK RecSys Challenge 2025中的核心作用，证实了其直接影响力。VK-LSVD为构建真实基准测试提供了一个至关重要的开放数据集，将加速序列推荐、冷启动场景以及下一代推荐系统的研究。

相关内容

数据集

关注 88

数据集，又称为资料集、数据集合或资料集合，是一种由数据所组成的集合。
Data set（或dataset）是一个数据的集合，通常以表格形式出现。每一列代表一个特定变量。每一行都对应于某一成员的数据集的问题。它列出的价值观为每一个变量，如身高和体重的一个物体或价值的随机数。每个数值被称为数据资料。对应于行数，该数据集的数据可能包括一个或多个成员。

探索长视频生成的最新趋势

专知会员服务

23+阅读 · 2024年12月30日

【CVPR2024】Koala: 关键帧条件化长视频语言模型

专知会员服务

13+阅读 · 2024年4月21日

【CVPR2024】MA-LMM: 内存增强的大型多模态模型，用于长期视频理解

专知会员服务

21+阅读 · 2024年4月9日

【CVPR2024】使用大型语言模型扩展视频摘要预训练

专知会员服务

22+阅读 · 2024年4月6日