Deepswe基准测试与真实感受匹配良好,新版本令人兴奋
deepswe bench continues to impress in terms of how well it matches real world qualitative feel
also with 5.6 being the first incremental version after a new post train, i think some brains are going to explode
very exciting times
also with 5.6 being the first incremental version after a new post train, i think some brains are going to explode
very exciting times
Claude Opus 4.8 has landed on DeepSWE Bench, posting a 58% Pass@1 and taking #2 overall behind GPT-5.5.
It continues a broader trend: slightly behind on raw score, but among the most reliable and efficient coding models across recent benchmarks. https://t.co/T9IWHGNo8d