动态

推出ViBench,开源AI编码基准,评估agent

推出ViBench,开源AI编码基准,评估agent
Amjad Masad
Most AI coding benchmarks miss what actually matters: how models perform at the application layer.

Introducing ViBench, an open-source benchmark for evaluating agents on end-to-end web application development. https://t.co/peTUqMIIgR
动态Amjad Masad2026-06-02原文

相关内容