动态

使用RL在形式化世界模型中生成证明携带策略

使用RL在形式化世界模型中生成证明携带策略
Andrew Curran
Very proud to have funded this work in my previous role @ARIA_research. I claimed that in environments with formal world-models, RL can be used to generate proof-carrying policies by just designing the right reward function, and this is a big theoretical and empirical validation. https://t.co/eJTW4PBQc0
动态Andrew Curran2026-06-10原文

相关内容