NVIDIA
Nemotron 3 Super 120B A12B
Nemotron 3 Super 120B A12B is an open-weights 120B-parameter model by NVIDIA. It holds an Intelligence score of 64.6 on the local.ai leaderboard (2026-08 snapshot).
MoEopen weights
64.6
local.ai Intelligence
Parameters
120B
MoE · 12B active parameters
Benchmarks (local.ai snapshot)
HuggingFace metadata · updated 2026-04-29
Metadata refreshed from the HuggingFace API (snapshot 2026-08-15T15:53:42.405Z). Context length and long-form descriptions are the next pipeline stage.
Does it fit your hardware?
Detect or select your GPU first. Green = weights + context headroom, yellow = tight, grey = does not fit.
Q4_K_M 69.6GQ6_K 108GQ8_0 132GFP16 240G| Quantization | ≈ File size | Note |
|---|---|---|
| IQ2_XXS | 32.4 GB | Extreme compression, largest quality loss |
| IQ3_XXS | 43.2 GB | Very small |
| Q3_K_M | 55.2 GB | Small |
| NVFP4 | 60 GB | NVIDIA 4-bit |
| Q4_K_S | 64.8 GB | Common |
| Q4_K_M | 69.6 GB | Most common sweet spot |
| Q5_K_M | 91.2 GB | Balanced |
| Q6_K | 108 GB | Near-lossless |
| Q8_0 | 132 GB | Near-lossless |
| FP16 | 240 GB | Original weights |
Actual files on HuggingFace
| File | Size |
|---|---|
| model-00049-of-00050.safetensors | 4.66 GB |
| model-00006-of-00050.safetensors | 4.65 GB |
| model-00001-of-00050.safetensors | 4.65 GB |
| model-00012-of-00050.safetensors | 4.65 GB |
| model-00018-of-00050.safetensors | 4.65 GB |
| model-00024-of-00050.safetensors | 4.65 GB |
GGUF: unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF · 22,619 downloads
| File | Size | |
|---|---|---|
| BF16/NVIDIA-Nemotron-3-Super-120B-A12B-BF16-00001-of-00005.gguf | 45.82 GB | download |
| BF16/NVIDIA-Nemotron-3-Super-120B-A12B-BF16-00002-of-00005.gguf | 44.61 GB | download |
| BF16/NVIDIA-Nemotron-3-Super-120B-A12B-BF16-00003-of-00005.gguf | 44.55 GB | download |
| BF16/NVIDIA-Nemotron-3-Super-120B-A12B-BF16-00004-of-00005.gguf | 44.61 GB | download |
| BF16/NVIDIA-Nemotron-3-Super-120B-A12B-BF16-00005-of-00005.gguf | 45.34 GB | download |
| MXFP4_MOE/NVIDIA-Nemotron-3-Super-120B-A12B-MXFP4_MOE-00001-of-00003.gguf | 0.01 GB | download |
| MXFP4_MOE/NVIDIA-Nemotron-3-Super-120B-A12B-MXFP4_MOE-00002-of-00003.gguf | 46.41 GB | download |
| MXFP4_MOE/NVIDIA-Nemotron-3-Super-120B-A12B-MXFP4_MOE-00003-of-00003.gguf | 30.01 GB | download |
| Q8_0/NVIDIA-Nemotron-3-Super-120B-A12B-Q8_0-00001-of-00004.gguf | 0.01 GB | download |
| Q8_0/NVIDIA-Nemotron-3-Super-120B-A12B-Q8_0-00002-of-00004.gguf | 45.64 GB | download |
| Q8_0/NVIDIA-Nemotron-3-Super-120B-A12B-Q8_0-00003-of-00004.gguf | 45.90 GB | download |
| Q8_0/NVIDIA-Nemotron-3-Super-120B-A12B-Q8_0-00004-of-00004.gguf | 28.10 GB | download |
GGUF: lmstudio-community/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF · 8,943 downloads
| File | Size | |
|---|---|---|
| NVIDIA-Nemotron-3-Super-120B-A12B-Q4_K_M-00001-of-00003.gguf | 36.51 GB | download |
| NVIDIA-Nemotron-3-Super-120B-A12B-Q4_K_M-00002-of-00003.gguf | 36.99 GB | download |
| NVIDIA-Nemotron-3-Super-120B-A12B-Q4_K_M-00003-of-00003.gguf | 6.64 GB | download |
| NVIDIA-Nemotron-3-Super-120B-A12B-Q6_K-00001-of-00003.gguf | 36.26 GB | download |
| NVIDIA-Nemotron-3-Super-120B-A12B-Q6_K-00002-of-00003.gguf | 36.52 GB | download |
| NVIDIA-Nemotron-3-Super-120B-A12B-Q6_K-00003-of-00003.gguf | 32.38 GB | download |
| NVIDIA-Nemotron-3-Super-120B-A12B-Q8_0-00001-of-00004.gguf | 36.77 GB | download |
| NVIDIA-Nemotron-3-Super-120B-A12B-Q8_0-00002-of-00004.gguf | 36.99 GB | download |
| NVIDIA-Nemotron-3-Super-120B-A12B-Q8_0-00003-of-00004.gguf | 37.12 GB | download |
| NVIDIA-Nemotron-3-Super-120B-A12B-Q8_0-00004-of-00004.gguf | 8.76 GB | download |
What you save self-hosting this model
Uses the model's own API price when it exists; otherwise the closest commercial equivalent.
Waiting for hardware detection…
RunLocal presents third-party leaderboard data (local.ai / Exo Labs, 2026-08) with its own fit and cost estimates. Verify pricing, licenses and file sizes at the source before deploying a model.