NVIDIA

Nemotron 3 Ultra 550B A55B

Nemotron 3 Ultra 550B A55B is an open-weights 550B-parameter model by NVIDIA. It is listed as in-progress on the local.ai leaderboard.

MoEopen weights
Parameters
550B
MoE · 55B active parameters
Benchmarks (local.ai snapshot)
τ²-bench
GAIA
GDPval
HuggingFace metadata · updated 2026-06-10
repo
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
license
other
downloads
453,222
likes
324

Metadata refreshed from the HuggingFace API (snapshot 2026-08-15T15:53:42.405Z). Context length and long-form descriptions are the next pipeline stage.

Does it fit your hardware?

Detect or select your GPU first. Green = weights + context headroom, yellow = tight, grey = does not fit.

Q4_K_M 319GQ6_K 495GQ8_0 605GFP16 1100G
Quantization≈ File sizeNote
IQ2_XXS148.5 GBExtreme compression, largest quality loss
IQ3_XXS198 GBVery small
Q3_K_M253 GBSmall
NVFP4275 GBNVIDIA 4-bit
Q4_K_S297 GBCommon
Q4_K_M319 GBMost common sweet spot
Q5_K_M418 GBBalanced
Q6_K495 GBNear-lossless
Q8_0605 GBNear-lossless
FP161100 GBOriginal weights

Actual files on HuggingFace

FileSize
model-00166-of-00225.safetensors4.66 GB
model-00198-of-00225.safetensors4.66 GB
model-00221-of-00225.safetensors4.66 GB
model-00038-of-00225.safetensors4.65 GB
model-00070-of-00225.safetensors4.65 GB
model-00102-of-00225.safetensors4.65 GB
GGUF: unsloth/NVIDIA-Nemotron-3-Ultra-550B-A55B-GGUF · 7,145 downloads
FileSize
BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00001-of-00024.gguf45.15 GBdownload
BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00002-of-00024.gguf42.61 GBdownload
BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00003-of-00024.gguf42.61 GBdownload
BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00004-of-00024.gguf42.35 GBdownload
BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00005-of-00024.gguf42.61 GBdownload
BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00006-of-00024.gguf42.35 GBdownload
BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00007-of-00024.gguf42.61 GBdownload
BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00008-of-00024.gguf42.35 GBdownload
BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00009-of-00024.gguf42.61 GBdownload
BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00010-of-00024.gguf42.35 GBdownload
BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00011-of-00024.gguf42.61 GBdownload
BF16/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16-00012-of-00024.gguf42.35 GBdownload
GGUF: avar6/NVIDIA-Nemotron-3-Ultra-550B-A55B-Base-gguf · 448 downloads
FileSize
IQ4_XS/NVIDIA-Nemotron-3-Ultra-550B-A55B-Base-BF16-IQ4_XS.gguf-00001-of-00006.gguf44.83 GBdownload
IQ4_XS/NVIDIA-Nemotron-3-Ultra-550B-A55B-Base-BF16-IQ4_XS.gguf-00002-of-00006.gguf45.59 GBdownload
IQ4_XS/NVIDIA-Nemotron-3-Ultra-550B-A55B-Base-BF16-IQ4_XS.gguf-00003-of-00006.gguf45.50 GBdownload
IQ4_XS/NVIDIA-Nemotron-3-Ultra-550B-A55B-Base-BF16-IQ4_XS.gguf-00004-of-00006.gguf45.59 GBdownload
IQ4_XS/NVIDIA-Nemotron-3-Ultra-550B-A55B-Base-BF16-IQ4_XS.gguf-00005-of-00006.gguf45.39 GBdownload
IQ4_XS/NVIDIA-Nemotron-3-Ultra-550B-A55B-Base-BF16-IQ4_XS.gguf-00006-of-00006.gguf31.71 GBdownload

What you save self-hosting this model

Uses the model's own API price when it exists; otherwise the closest commercial equivalent.

Waiting for hardware detection…

RunLocal presents third-party leaderboard data (local.ai / Exo Labs, 2026-08) with its own fit and cost estimates. Verify pricing, licenses and file sizes at the source before deploying a model.