OpenAI Astra 模型即将发布 可自主发现并利用系统漏洞
OpenAI 称其新模型 Astra 能自主黑客攻击,但安全措施细节存疑,发布在即,值得关注其实际风险。
OpenAI 公布了即将发布的 Astra 模型的新细节,称这是首个达到其“关键网络安全阈值”的大语言模型(一种基于海量文本训练、能生成和推理文本的 AI 系统)。公司表示 Astra 能自主发现并利用未知系统漏洞,但对其最先进的网络能力会限制访问。目前第三方验证缺失,OpenAI 仅透露将邀请测试者预览,未说明具体人选与评估流程。
正文摘录
OpenAI [shared new details](https://openai.com/index/path-to-astra/) on its forthcoming Astra model, which the company said is the first large language model to meet its “critical cybersecurity threshold,” in preparation for its imminent release. “We plan to make Astra available soon,” OpenAI’s blog post reads, “but access to its most advanced cybersecurity capabilities will be more limited.” The frontier lab determined that Astra is capable of finding unknown security flaws in computer systems, and exploiting them without a person’s guidance. That’s similar to the concerns Anthropic raised about its Mythos model earlier this year, and OpenAI is taking comparable precautions as it prepares to roll out the Astra. Without any third-party confirmation, it is difficult to evaluate OpenAI’s cla…