成为VIP会员查看完整内容
VIP会员码认证
首页
主题
会员
服务
注册
·
登录
奖励模型
关注
0
综合
百科
VIP
热门
动态
论文
精华
TuneJury: An Open Metric for Improving Music Generation Preference Alignment
Arxiv
0+阅读 · 6月15日
The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Mathematical Reasoning
Arxiv
0+阅读 · 6月15日
Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions
Arxiv
0+阅读 · 6月14日
DriveReward: A Comprehensive Dataset and Generative Vision-Language Reward Model for Autonomous Driving
Arxiv
0+阅读 · 6月14日
AudioProcessBench: Benchmark for Identifying Process Errors in Audio-Grounded Reasoning
Arxiv
0+阅读 · 6月7日
StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents
Arxiv
0+阅读 · 6月12日
Understanding helpfulness and harmless tension in reward models
Arxiv
0+阅读 · 6月11日
Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding
Arxiv
0+阅读 · 6月10日
SARM2: Multi-Task Stage Aware Reward Modeling for Self Improving Robotic Manipulation
Arxiv
0+阅读 · 6月9日
Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output
Arxiv
0+阅读 · 6月9日
UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment
Arxiv
0+阅读 · 5月22日
One Token to Fool LLM-as-a-Judge
Arxiv
0+阅读 · 6月11日
Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill
Arxiv
0+阅读 · 6月2日
SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation
Arxiv
0+阅读 · 3月23日
Reward Models Can Improve Themselves: Reward-Guided Adversarial Failure Mode Discovery for Robust Reward Modeling
Arxiv
0+阅读 · 4月8日
参考链接
提示
微信扫码
咨询专知VIP会员与技术项目合作
(加微信请备注: "专知")
微信扫码咨询专知VIP会员
Top