英伟达发布投机解码推理指南,核心方案多人主导
英伟达官方首次以选型表规范投机解码,指出最优解多源自华人团队,且硬件矩阵常熟正反向锁死模型架构。
英伟达发布大模型推理指南,对比6种投机解码方案:用轻量模型预测候选词,再由大模型验证,实现无损加速。方案中 EAGLE-3、MTP、DFlash、DSpark 多由华人团队提出,且提速上限受 GPU 物理矩阵的 Tile Size 约束,硬件常数正在反向决定模型架构设计。
正文摘录
NVIDIA's Official Inference Guide Is Here! 6x Lossless Speedup, and the Optimal Solutions Were All Written by Chinese Researchers Aiera Report  Recently, NVIDIA published an official large model inference guide. But as you flip through it, it increasingly reads like a collection of papers by Chinese research teams. Here's how it happened. On September 2nd, NVIDIA's technical blog published a long article titled "Co-Designing AI Models Using Speculative Decoding for Faster LLM Inference." The core of the article is a single selection table that lays out the six most mainstream inference acceleration schemes side by side, down to training costs and applicable scenarios. ![](https://aiera.com.cn/wp-content/uploads/2…