动态

(无标题)

(无标题)
Artificial Analysis
Claude Opus 4.6 takes the lead in GDPval-AA, surpassing GPT-5.2 in our benchmark of agentic real-world knowledge work tasks

We worked with @AnthropicAI to benchmark Claude Opus 4.6 ahead of launch - it reached an Elo of 1606 with adaptive thinking, nearly 150 points ahead of https://t.co/7HxffYYKSY
动态Andrew Curran2026-02-05原文

相关内容