NVIDIA

Llama 3.1 Nemotron 70B

Llama 3.1 Nemotron 70B is an open-weights 70B-parameter model by NVIDIA. It is listed as in-progress on the local.ai leaderboard.

open weights
Parameters
70B
Benchmarks (local.ai snapshot)
τ²-bench
GAIA
GDPval
HuggingFace metadata · updated 2025-04-13
repo
nvidia/Llama-3.1-Nemotron-70B-Instruct-HF
license
llama3.1
downloads
10,551
likes
2,070

Metadata refreshed from the HuggingFace API (snapshot 2026-08-15T15:53:42.405Z). Context length and long-form descriptions are the next pipeline stage.

Does it fit your hardware?

Detect or select your GPU first. Green = weights + context headroom, yellow = tight, grey = does not fit.

Q4_K_M 40.6GQ6_K 63GQ8_0 77GFP16 140G
Quantization≈ File sizeNote
IQ2_XXS18.9 GBExtreme compression, largest quality loss
IQ3_XXS25.2 GBVery small
Q3_K_M32.2 GBSmall
NVFP435 GBNVIDIA 4-bit
Q4_K_S37.8 GBCommon
Q4_K_M40.6 GBMost common sweet spot
Q5_K_M53.2 GBBalanced
Q6_K63 GBNear-lossless
Q8_077 GBNear-lossless
FP16140 GBOriginal weights

Actual files on HuggingFace

FileSize
model-00008-of-00030.safetensors4.66 GB
model-00013-of-00030.safetensors4.66 GB
model-00018-of-00030.safetensors4.66 GB
model-00023-of-00030.safetensors4.66 GB
model-00028-of-00030.safetensors4.66 GB
model-00003-of-00030.safetensors4.66 GB
GGUF: bartowski/Llama-3.1-Nemotron-70B-Instruct-HF-GGUF · 2,030 downloads
FileSize
Llama-3.1-Nemotron-70B-Instruct-HF-IQ1_M.gguf15.60 GBdownload
Llama-3.1-Nemotron-70B-Instruct-HF-IQ2_M.gguf22.46 GBdownload
Llama-3.1-Nemotron-70B-Instruct-HF-IQ2_XS.gguf19.69 GBdownload
Llama-3.1-Nemotron-70B-Instruct-HF-IQ2_XXS.gguf17.79 GBdownload
Llama-3.1-Nemotron-70B-Instruct-HF-IQ3_M.gguf29.74 GBdownload
Llama-3.1-Nemotron-70B-Instruct-HF-IQ3_XXS.gguf25.58 GBdownload
Llama-3.1-Nemotron-70B-Instruct-HF-IQ4_XS.gguf35.30 GBdownload
Llama-3.1-Nemotron-70B-Instruct-HF-Q2_K.gguf24.56 GBdownload
Llama-3.1-Nemotron-70B-Instruct-HF-Q2_K_L.gguf25.52 GBdownload
Llama-3.1-Nemotron-70B-Instruct-HF-Q3_K_L.gguf34.59 GBdownload
Llama-3.1-Nemotron-70B-Instruct-HF-Q3_K_M.gguf31.91 GBdownload
Llama-3.1-Nemotron-70B-Instruct-HF-Q3_K_S.gguf28.79 GBdownload
GGUF: lmstudio-community/Llama-3.1-Nemotron-70B-Instruct-HF-GGUF · 1,785 downloads
FileSize
Llama-3.1-Nemotron-70B-Instruct-HF-Q3_K_L.gguf34.59 GBdownload
Llama-3.1-Nemotron-70B-Instruct-HF-Q4_K_M.gguf39.60 GBdownload
Llama-3.1-Nemotron-70B-Instruct-HF-Q6_K-00001-of-00002.gguf37.13 GBdownload
Llama-3.1-Nemotron-70B-Instruct-HF-Q6_K-00002-of-00002.gguf16.79 GBdownload
Llama-3.1-Nemotron-70B-Instruct-HF-Q8_0-00001-of-00002.gguf37.07 GBdownload
Llama-3.1-Nemotron-70B-Instruct-HF-Q8_0-00002-of-00002.gguf32.75 GBdownload

What you save self-hosting this model

Uses the model's own API price when it exists; otherwise the closest commercial equivalent.

Waiting for hardware detection…

RunLocal presents third-party leaderboard data (local.ai / Exo Labs, 2026-08) with its own fit and cost estimates. Verify pricing, licenses and file sizes at the source before deploying a model.