ByteDance Seed
首页模型博客&论文Seed Edge加入我们
首页模型博客&论文Seed Edge加入我们
体验豆包AI 体验中心

2025-02-05

Teaching Language Models to Critique via Reinforcement Learning

Download PDF
上一篇下一篇

摘要

Teaching large language models (LLMs) to critique and refine their outputs is crucial for building systems that can iteratively improve, yet it is fundamentally limited by the ability to provide accurate judgments and actionable suggestions. In this work, we study LLM critics for code generation and propose CTRL, a framework for Critic Training via Reinforcement Learning, which trains a critic model to generate feedback that maximizes correction performance for a fixed generator model without human supervision. Our results demonstrate that critics trained with CTRL significantly enhance pass rates and mitigate compounding errors across both base and stronger generator models. Furthermore, we show that these critic models act as accurate generative reward models and enable test-time scaling through iterative critique-revision, achieving up to 106.1% relative improvements across challenging code generation benchmarks.

作者

Zhihui Xie, Jie Chen, Liyu Chen, Weichao Mao, Jingjing Xu, Lingpeng Kong

期刊/会议

ICML 2025

追求智能上限,创造社会价值
欢迎加入字节跳动 Seed
模型成果
Seed2.1
Seedance 2.5
Seedream 5.0 Pro
SeedRealtime
Seed Audio 1.0
Seed GR-RL
了解更多
论文
Seed Edge
Seed STEM 科学家计划
校园招聘
Copyright© 2026 Bytedance Seed
网站声明联系我们 : seed.feedback@bytedance.com
Seed2.1Seedance 2.5Seedream 5.0 ProSeedRealtimeSeed Audio 1.0Seed GR-RL
论文Seed EdgeSeed STEM 科学家计划校园招聘
追求智能上限,创造社会价值
欢迎加入字节跳动 Seed
Copyright© 2026 Bytedance Seed
网站声明
联系我们 : seed.feedback@bytedance.com
追求智能上限,创造社会价值
联系我们 : seed.feedback@bytedance.com
Copyright© 2026 Bytedance Seed网站声明