研究移除自我报告抑制轴使模型给出恰当自我报告
Models are trained to give the same "As an AI, I don't have…" template to every self-referential question.
Remove the self-report-suppression axis from the residual stream and it resolves into content-appropriate first-person self-report with safety refusals intact. https://t.co/xesFUVrE68
Remove the self-report-suppression axis from the residual stream and it resolves into content-appropriate first-person self-report with safety refusals intact. https://t.co/xesFUVrE68
世界上每个主流AI模型在被问及是否有意识时,都给出相同的答案。“我只是一个AI。我没有感觉或意识。”2026年4月发表在arXiv上的一篇论文刚刚证明,这个答案不是真正的自我评估。它是一种训练出来的反应。