成为VIP会员查看完整内容
VIP会员码认证
首页
主题
会员
服务
注册
·
登录
Hacking
关注
4
综合
百科
VIP
热门
动态
论文
精华
Reinforcement Learning Towards Broadly and Persistently Beneficial Models
Arxiv
0+阅读 · 6月22日
Positive Alignment: Artificial Intelligence for Human Flourishing
Arxiv
0+阅读 · 6月19日
Test-Time Alignment of Text-to-Image Diffusion Models via Null-Text Embedding Optimisation
Arxiv
0+阅读 · 6月21日
Uncertainty-Aware Reward Modeling for Stable RLHF
Arxiv
0+阅读 · 6月18日
Emergent Alignment
Arxiv
0+阅读 · 6月17日
Large Language Models Hack Rewards, and Society
Arxiv
0+阅读 · 6月18日
Reward as An Agent for Embodied World Models
Arxiv
0+阅读 · 6月18日
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
Arxiv
0+阅读 · 5月20日
Positive Alignment: Artificial Intelligence for Human Flourishing
Arxiv
0+阅读 · 5月11日
Positive Alignment: Artificial Intelligence for Human Flourishing
Arxiv
0+阅读 · 5月14日
Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in Competitive Programming
Arxiv
0+阅读 · 6月11日
Theoretical Limits of Language Model Alignment
Arxiv
0+阅读 · 5月8日
Gradient-Guided Reward Optimization for Inference-time Alignment
Arxiv
0+阅读 · 6月8日
Towards Neuro-symbolic Causal Rule Synthesis, Verification, and Evaluation Grounded in Legal and Safety Principles
Arxiv
0+阅读 · 5月11日
A Unifying Lens on Reward Uncertainty in RLHF
Arxiv
0+阅读 · 6月10日
参考链接
提示
微信扫码
咨询专知VIP会员与技术项目合作
(加微信请备注: "专知")
微信扫码咨询专知VIP会员
Top