Social media platforms are daily exhibiting millions of events. To preliminarily predict the mainstream public reaction to these events, we study trendy response prediction to automatically generate top-liked user replies to social media events. While previous works focus on generating responses without factoring in popularity, we propose Popularity-Aligned Language Models (PopALM) to distinguish responses liked by a larger audience through reinforcement learning. Recognizing the noisy labels from user "likes", we tailor-make curriculum learning in proximal policy optimization (PPO) to help models capture the essential samples for easy-to-hard training. In experiments, we build a large-scale Weibo dataset for trendy response prediction, and its results show that PopALM can help boost the performance of advanced language models.
翻译:社交媒体平台每天展示数百万计的事件。为初步预测公众对这些事件的主流反应,我们研究流行趋势回复预测,以自动生成社交媒体事件中获得最多点赞的用户回复。现有工作侧重于不考虑流行度的回复生成,我们提出流行度对齐语言模型(PopALM),通过强化学习区分受更广泛受众喜爱的回复。针对用户"点赞"中的噪声标签,我们定制了近端策略优化(PPO)中的课程学习,帮助模型从简单到困难的训练过程中捕捉关键样本。实验中,我们构建了用于流行趋势回复预测的大规模微博数据集,结果表明PopALM能够提升先进语言模型的性能。