动态

编码代理仅恢复9.3%人类研究进展,聚焦调参

AK
Can coding agents do research?

We release NanoGPT-Bench, an internal eval we’ve used to test agents on an AI R&D problem with months of human progress

Codex, Claude Code, Autoresearch recover only 9.3% of human progress, mostly tuning hyperparams & ignoring algorithmic research https://t.co/MZpLKuhLeX
动态AK2026-05-19原文

相关内容