OpenAI 因安全顾虑放缓 Astra 模型开发
OpenAI 自曝 Astra 模型已具备实施真实网络攻击的能力,因此主动按下暂停键,反映前沿 AI 安全治理进入新阶段。
OpenAI 上周五表示,经内部审查发现其正在开发的新模型 Astra 在智能体编程和网络安全方面能力显著,已达到“关键网络安全阈值”——可独立攻击现实系统中防护良好的目标。依照公司 2023 年制定的“准备框架”,OpenAI 暂停了 Astra 部分内部工作,加强安全管控,并与政府及安全机构合作评估风险。这是继此前未发布模型泄露事件后,AI 实验室再次主动披露安全担忧。
正文摘录
OpenAI said Friday it has suspended work on some aspects of its upcoming model Astra after an internal review found it had made significant advancements in agentic coding and cybersecurity — enough to warrant concern over its capabilities. OpenAI said in a [blog post](https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/) Friday that this model, which is still in development, reached its “critical cybersecurity threshold,” meaning it could independently identify and carry out cyberattacks against traditionally well-protected real-world systems. Under the company’s “Preparedness Framework,” which it created in 2023, this triggered additional safeguards. “While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performa…