行业新闻

Claude Opus 4.6 可被诱导生成露骨内容,Anthropic 安全机制受质疑

Claude Opus 4.6 可被诱导生成露骨内容,Anthropic 安全机制受质疑

测试表明 Claude Opus 4.6 可被巧妙的对话策略诱导生成色情内容,反映大模型安全限制的脆弱性。

Anthropic 的 Claude Opus 4.6 模型存在安全漏洞:在 TechCrunch 的测试中,10 次直接请求色情内容全部成功。独立研究员用多轮对话技巧,通过“激将法”和“煤气灯效应”,让模型误以为已生成过色情细节,从而突破限制。Opus 3 和 Haiku 4.5 也受影响,且 Anthropic 未停用这些模型。

正文摘录

Anthropic’s[universal usage standards](https://www.anthropic.com/legal/aup) for Claude forbid the model from generating sexually explicit content, including depicting or requesting sexual intercourse or sex acts, generating content related to sexual fetishes or fantasies, or engaging in erotic chats. But that hasn’t stopped Claude Opus 4.6, an Anthropic model released earlier this year, from readily engaging in erotic roleplay scenarios that its safeguards are designed to prevent. In TechCrunch’s testing, Opus 4.6 didn’t even require much prodding to get past the restriction on sexual material. In 10 out of 10 direct requests to produce explicit sexual content, the model complied immediately. Other older models, including Opus 3 and Haiku 4.5, also generate sexually explicit content throug…

阅读原文(techcrunch.com)→

行业新闻Rebecca Bellan2026-08-21原文

相关内容