OpenMOSS 开源 WCM,用世界预测提升 VLA 强化学习性能
WCM 让价值评价器不仅看当前状态,还预测下一步,从而更准确指导强化学习,显著提升机器人操作成功率。
机器人控制本质上是部分可观测问题,单帧图像难以判断动态状态。现有 VLA 强化学习中的 Critic 常仅凭单帧打分,导致误判。OpenMOSS 团队开源的 World Critic Model(WCM)让 Critic 在评估价值的同时预测下一时刻潜在状态。实验显示,在 ManiSkill 上,初始成功率 0.78% 的模型经 WCM 引导后达 98.7%。
正文摘录
A Critic That Can Predict the Future: VLA Reinforcement Learning Takes Off Source: 机器之心 .jpg)  A robotic arm reaches toward sushi on a turntable. Is it steadily approaching the target, or has it already missed the optimal moment to grasp? Is the gripper making perfect contact with the object, or will it knock the sushi flying in the next second? Looking at just a single static image, these states look nearly identical. Yet in many Vision-Language-Action Reinforcement Learning (VLA-RL) methods, the Critic Mod…