提出 AC2:用 critic 给 token 块打分,只需部分 rollout,训练快于 GRPO
RT @wen_kaiyue:你真的需要在 LLM 的 RL 中把每一次 rollout 都跑完吗?如果你把 critic 做得更好——并且信任它——就不必!
我们提出 Actor-Critic with Action Chunking(AC2):用学习到的 critic 给 token 块打分 => 只需要部分 rollout ⇒ 训练比 GRPO 更快!https://t.co/u9G9CUVx6L
我们提出 Actor-Critic with Action Chunking(AC2):用学习到的 critic 给 token 块打分 => 只需要部分 rollout ⇒ 训练比 GRPO 更快!https://t.co/u9G9CUVx6L