Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Hugging Face

Via r/LocalLlama
Thursday, Jun 4, 2026 · 11:48AM
Summary

Model Summary Total Parameters 550B (55B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 8x GB200/B200/GB300/B300, 16x H100, 8x H200 Supported Languages English, French, Spanish, Italian, Germa

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories