中国团队DoGNAVY获AI安全基准CyberGym全球第三、开源第一
中国团队用开源模型在AI安全漏洞挖掘上超越闭源巨头,证明在关键安全任务上开源路线同样可行。
国际AI安全基准CyberGym(测试智能体自主挖掘真实漏洞的能力)最新排名显示,中国团队DoGNAVY以90.8%的通过率获全球第三、开源第一,超越OpenAI和Anthropic。该团队仅用开源模型GLM-5.2,并通过AgentDoG架构为智能体提供风险诊断能力,而非堆砌多个闭源模型。
正文摘录
CyberGym: A Practical Benchmark for Real-World Vulnerability Discovery Recently, CyberGym — the internationally recognized security benchmark released by UC Berkeley — published its latest ranking. Built around 1,507 historical fixed vulnerabilities from 188 real open-source projects, the benchmark tests whether AI agents can autonomously discover, verify, and exploit vulnerabilities from scratch. Anthropic, DeepSeek, Microsoft, Wiz, and others now treat it as a key yardstick for AI security capability.  On August 5, CyberGym released its latest results: DoGNAVY, a team from China, ranked third globally and first among open-source entries. DoGNAVY: Third Globally, First in China, Built on Open-Source …