行业新闻

Anthropic 公布 Claude 宪法,以明文原则替代隐式人类反馈

Anthropic 公布 Claude 宪法,以明文原则替代隐式人类反馈

Anthropic 公开 Claude 宪法,展示如何用明文价值观训练更安全、更透明的大模型。

Claude 的价值观来自一部明文“宪法”,而非人类隐式偏好。这种宪法式 AI 方法让模型依据一套原则自我批评,并用 AI 反馈训练,避免人工审核有害内容,效率更高且更透明,还能同时提升帮助性与无害性。

正文摘录

Announcements Claude’s constitution May 9, 2023 [Read the new constitution](https://www.anthropic.com/news/claude-new-constitution) Update, Jan 21, 2026: We've published a new version of Claude's constitution, which you can find at the button above. ![](/r/blog/1e1389942aed821b0ad01bfc0c58089066fed71ac527e3ed1131ecb57f22542a.webp) How does a language model decide which questions it will engage with and which it deems inappropriate? Why will it encourage some actions and discourage others? What “values” might a language model have? These are all questions people grapple with. Our recently published research on “Constitutional AI” provides one answer by giving language models explicit values determined by a constitution, rather than values determined implicitly via large-scale human feedback…

阅读原文(anthropic.com)→

行业新闻2026-09-09原文

相关内容