NVIDIA
Nemotron 3 Ultra 550B A55B
Nemotron 3 Ultra 550B A55B is an open-weights 550B-parameter model by NVIDIA. It is listed as in-progress on the local.ai leaderboard.
MoEopen weights
Parameters
550B
MoE · 55B active parameters
Benchmarks (local.ai snapshot)
HuggingFace metadata · updated 2026-06-10
Metadata refreshed from the HuggingFace API (snapshot 2026-08-15T15:53:42.405Z). Context length and long-form descriptions are the next pipeline stage.
Does it fit your hardware?
Detect or select your GPU first. Green = weights + context headroom, yellow = tight, grey = does not fit.
Q4_K_M 319GQ6_K 495GQ8_0 605GFP16 1100G| Quantization | ≈ File size | Note |
|---|---|---|
| IQ2_XXS | 148.5 GB | Extreme compression, largest quality loss |
| IQ3_XXS | 198 GB | Very small |
| Q3_K_M | 253 GB | Small |
| NVFP4 | 275 GB | NVIDIA 4-bit |
| Q4_K_S | 297 GB | Common |
| Q4_K_M | 319 GB | Most common sweet spot |
| Q5_K_M | 418 GB | Balanced |
| Q6_K | 495 GB | Near-lossless |
| Q8_0 | 605 GB | Near-lossless |
| FP16 | 1100 GB | Original weights |
Actual files on HuggingFace
| File | Size |
|---|---|
| model-00166-of-00225.safetensors | 4.66 GB |
| model-00198-of-00225.safetensors | 4.66 GB |
| model-00221-of-00225.safetensors | 4.66 GB |
| model-00038-of-00225.safetensors | 4.65 GB |
| model-00070-of-00225.safetensors | 4.65 GB |
| model-00102-of-00225.safetensors | 4.65 GB |
GGUF: unsloth/NVIDIA-Nemotron-3-Ultra-550B-A55B-GGUF · 7,145 downloads
| File | Size | |
|---|---|---|
| BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00001-of-00024.gguf | 45.15 GB | download |
| BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00002-of-00024.gguf | 42.61 GB | download |
| BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00003-of-00024.gguf | 42.61 GB | download |
| BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00004-of-00024.gguf | 42.35 GB | download |
| BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00005-of-00024.gguf | 42.61 GB | download |
| BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00006-of-00024.gguf | 42.35 GB | download |
| BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00007-of-00024.gguf | 42.61 GB | download |
| BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00008-of-00024.gguf | 42.35 GB | download |
| BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00009-of-00024.gguf | 42.61 GB | download |
| BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00010-of-00024.gguf | 42.35 GB | download |
| BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00011-of-00024.gguf | 42.61 GB | download |
| BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00012-of-00024.gguf | 42.35 GB | download |
GGUF: avar6/NVIDIA-Nemotron-3-Ultra-550B-A55B-Base-gguf · 448 downloads
| File | Size | |
|---|---|---|
| IQ4_XS/NVIDIA-Nemotron-3-Ultra-550B-A55B-Base-BF16-IQ4_XS.gguf-00001-of-00006.gguf | 44.83 GB | download |
| IQ4_XS/NVIDIA-Nemotron-3-Ultra-550B-A55B-Base-BF16-IQ4_XS.gguf-00002-of-00006.gguf | 45.59 GB | download |
| IQ4_XS/NVIDIA-Nemotron-3-Ultra-550B-A55B-Base-BF16-IQ4_XS.gguf-00003-of-00006.gguf | 45.50 GB | download |
| IQ4_XS/NVIDIA-Nemotron-3-Ultra-550B-A55B-Base-BF16-IQ4_XS.gguf-00004-of-00006.gguf | 45.59 GB | download |
| IQ4_XS/NVIDIA-Nemotron-3-Ultra-550B-A55B-Base-BF16-IQ4_XS.gguf-00005-of-00006.gguf | 45.39 GB | download |
| IQ4_XS/NVIDIA-Nemotron-3-Ultra-550B-A55B-Base-BF16-IQ4_XS.gguf-00006-of-00006.gguf | 31.71 GB | download |
What you save self-hosting this model
Uses the model's own API price when it exists; otherwise the closest commercial equivalent.
Waiting for hardware detection…
RunLocal presents third-party leaderboard data (local.ai / Exo Labs, 2026-08) with its own fit and cost estimates. Verify pricing, licenses and file sizes at the source before deploying a model.