行业新闻

英伟达发布投机解码推理指南,核心方案多人主导

英伟达发布投机解码推理指南,核心方案多人主导

英伟达官方首次以选型表规范投机解码,指出最优解多源自华人团队,且硬件矩阵常熟正反向锁死模型架构。

英伟达发布大模型推理指南,对比6种投机解码方案:用轻量模型预测候选词,再由大模型验证,实现无损加速。方案中 EAGLE-3、MTP、DFlash、DSpark 多由华人团队提出,且提速上限受 GPU 物理矩阵的 Tile Size 约束,硬件常数正在反向决定模型架构设计。

正文摘录

NVIDIA's Official Inference Guide Is Here! 6x Lossless Speedup, and the Optimal Solutions Were All Written by Chinese Researchers Aiera Report ![](https://aiera.com.cn/wp-content/uploads/2026/09/aieraimg3b253ef414-25.png) Recently, NVIDIA published an official large model inference guide. But as you flip through it, it increasingly reads like a collection of papers by Chinese research teams. Here's how it happened. On September 2nd, NVIDIA's technical blog published a long article titled "Co-Designing AI Models Using Speculative Decoding for Faster LLM Inference." The core of the article is a single selection table that lays out the six most mainstream inference acceleration schemes side by side, down to training costs and applicable scenarios. ![](https://aiera.com.cn/wp-content/uploads/2…

阅读原文(aiera.com.cn)→

行业新闻新智元2026-09-04原文

相关内容