论文:前沿大模型API漏洞致隐藏思维链可被提取
本文解读一篇揭露大模型API隐藏思维链可被提取的论文,威胁模型安全、隐私与蒸馏防线。
一篇安全论文发现,Anthropic、OpenAI、Google 的前沿模型 API 存在设计缺陷:原本加密隐藏的完整思维链(模型推理过程)可被提取。攻击者利用强模型与弱模型的安全差异,将加密块移交给弱模型解密,成本仅数百美元。论文引发社区广泛关注。
正文摘录
 "This is the best security paper of the year — it fixed a billion-dollar bug for Anthropic, OpenAI, and Google."  Today, a paper on security vulnerabilities in frontier large models caused a massive stir on social media. Within just 19 hours, it had drawn 2.2 million viewers and triggered widespread discussion. ![Image](https://mmbiz.qpic.cn/szmmbizpng/5L8bhP5dIqHGxBPoCeQc05UV8YQfoaBIatAsiag8PxUA826mSn8LjXyFP5bQvNsibvyzF9chXh4ibicnS6B7fMIC4fHHdribo5f6Mf5XsgDib8qZw/640?wxfmt=png&from=appmsgimgIndex=…