动态

Blackwell硬件使4-bit浮点量化实用化,覆盖LLM、KV缓存等

Yukang Chen
We are excited to share a new blog “Pushing Intelligence to 4-bit.”
https://t.co/VOjZ60iRgg

This blog discusses how FP4 and Blackwell hardware make 4-bit floating point practical for training and inference, covering LLMs, KV cache, attention, and Video Gen.
Aaron Huang
🔗 Our new blog looks at how FP4 is moving beyond compression into a practical primitive for training and inference across both LLMs and diffusion models:

https://t.co/IvzQejjQ5q

1. Why Four Bits Is Hard: Only 15 values make scaling critical.
2. NVFP4: Smaller blocks and finer https://t.co/8uSOdOESfK
动态Yukang Chen2026-06-30原文

相关内容