Blackwell硬件使4-bit浮点量化实用化,覆盖LLM、KV缓存等
We are excited to share a new blog “Pushing Intelligence to 4-bit.”
https://t.co/VOjZ60iRgg
This blog discusses how FP4 and Blackwell hardware make 4-bit floating point practical for training and inference, covering LLMs, KV cache, attention, and Video Gen.
https://t.co/VOjZ60iRgg
This blog discusses how FP4 and Blackwell hardware make 4-bit floating point practical for training and inference, covering LLMs, KV cache, attention, and Video Gen.
🔗 Our new blog looks at how FP4 is moving beyond compression into a practical primitive for training and inference across both LLMs and diffusion models:
https://t.co/IvzQejjQ5q
1. Why Four Bits Is Hard: Only 15 values make scaling critical.
2. NVFP4: Smaller blocks and finer https://t.co/8uSOdOESfK