推出ViBench,开源AI编码基准,评估agent
Most AI coding benchmarks miss what actually matters: how models perform at the application layer.
Introducing ViBench, an open-source benchmark for evaluating agents on end-to-end web application development. https://t.co/peTUqMIIgR
Introducing ViBench, an open-source benchmark for evaluating agents on end-to-end web application development. https://t.co/peTUqMIIgR