NVIDIA发布NVFP4量化MiniMax-M3模型
NVIDIA just released an NVFP4-quantized MiniMax-M3 on Hugging Face
A 428B parameter multimodal MoE model
with a 1M-token context window,
now compressed to 4-bit precision
for 2x memory savings on Blackwell GPUs.
https://t.co/v3hlvqdfhe
A 428B parameter multimodal MoE model
with a 1M-token context window,
now compressed to 4-bit precision
for 2x memory savings on Blackwell GPUs.
https://t.co/v3hlvqdfhe