行业新闻

浙大提出 ProVisE 框架,让图像生成模型直接画出空间认知答案

用视觉协议让生成模型以画图方式回答空间问题,替代传统的坐标输出,更贴近人类直觉。

现有空间认知评测强迫 AI 用坐标或文字作答,这对图像生成模型并不自然。浙大团队提出 ProVisE,用视觉协议引导模型直接在图上标记、画轨迹或生成深度图,再自动解析为可评分的结果,并构建了覆盖 14 项子任务的 SpatialGen-Bench。

正文摘录

![](https://image.jiqizhixin.com/uploads/article/coverimage/9dbed617-be31-480d-9fc3-1f02edbe819d/03(2).jpg) ![图片](https://mmbiz.qpic.cn/szmmbizpng/5L8bhP5dIqElqZ0GcgINP4V78xIEARSSDI7pRpUQpZAhGK6rr8zoMJM3icYLohdHNhq93gll1YkHwsYaOO3tolCuL5QN5gAtJc46PnNgECQI/640?wxfmt=png&from=appmsgimgIndex=0) If you ask someone "Where is the cup?", they'd normally just point at it, rather than replying robotically: "The target is located within a bounding box spanning x-coordinates 345–412 and y-coordinates 512–600." Yet in existing spatial cognition benchmarks, we've been forcing AI to prove its spatial abilities through exactly this kind of counterintuitive coordinate-based output. ![图片](https://mmbiz.qpic.cn/szmmbizpng/5L8bhP5dIqHub3FiaNlBQu0oe6AyC3JwMYXAUYlvJCkzSCGOgqnKTUeYpFEGs8kqSo2pGchf1FDtXaJfj9ub4s…

阅读原文(jiqizhixin.com)→

行业新闻机器之心2026-08-08原文

相关内容