谷歌研究:GPT-5 和 Gemini 3 也会话到嘴边想不起来
谷歌用 450 万次测试证明:大模型答错常因“提取失败”而非“没学会”,并给出区分方法与基准。
谷歌一项被 ICML 2026 收录的研究发现,前沿大模型(如 GPT-5、Gemini 3)虽已把 95%–98% 的事实存入参数中,但直接提问时会有 26%–34% 答不上来;即使多思考一会儿,仍有 11%–12% 无法回忆。研究者用“知识画像”框架区分“没学会”和“想不起”,并发布包含 2150 个事实的 WikiProfile 基准。换句话说,模型答错很多时候不是没记住,而是“钥匙丢了”。
正文摘录
 Reported by Xin Zhiyuan  That name is right on the tip of your tongue. You see a familiar face but simply cannot recall the name. Psychology has a name for this phenomenon: the "tip-of-the-tongue" state. Now, in a study, Google has found that GPT-5 and Gemini 3 share the same affliction. Both models have stored 95%–98% of the test facts in their parameters. But when asked directly, they fail to produce 26%–34% of those facts. With thinking enabled for a while longer, 11%–12% still remain unrecallable. When an LLM fails to answer, in many cases it's not that the knowledge is absent from its "brain"—it has learned it, but cannot retrieve …