Meta

Llama 3.3 70B Instruct

Llama 3.3 70B Instruct is an open-weights 70B-parameter model by Meta. It is listed as in-progress on the local.ai leaderboard.

open weights
Parameters
70B
Benchmarks (local.ai snapshot)
τ²-bench
GAIA
GDPval
HuggingFace metadata · updated 2024-12-21
repo
meta-llama/Llama-3.3-70B-Instruct
license
llama3.3
downloads
320,776
likes
2,952

Metadata refreshed from the HuggingFace API (snapshot 2026-08-15T15:53:42.405Z). Context length and long-form descriptions are the next pipeline stage.

Does it fit your hardware?

Detect or select your GPU first. Green = weights + context headroom, yellow = tight, grey = does not fit.

Q4_K_M 40.6GQ6_K 63GQ8_0 77GFP16 140G
Quantization≈ File sizeNote
IQ2_XXS18.9 GBExtreme compression, largest quality loss
IQ3_XXS25.2 GBVery small
Q3_K_M32.2 GBSmall
NVFP435 GBNVIDIA 4-bit
Q4_K_S37.8 GBCommon
Q4_K_M40.6 GBMost common sweet spot
Q5_K_M53.2 GBBalanced
Q6_K63 GBNear-lossless
Q8_077 GBNear-lossless
FP16140 GBOriginal weights

Actual files on HuggingFace

FileSize
model-00008-of-00030.safetensors4.66 GB
model-00013-of-00030.safetensors4.66 GB
model-00018-of-00030.safetensors4.66 GB
model-00023-of-00030.safetensors4.66 GB
model-00028-of-00030.safetensors4.66 GB
model-00003-of-00030.safetensors4.66 GB
GGUF: MaziyarPanahi/Llama-3.3-70B-Instruct-GGUF · 169,345 downloads
FileSize
Llama-3.3-70B-Instruct.Q2_K.gguf24.56 GBdownload
Llama-3.3-70B-Instruct.Q3_K_L.gguf34.59 GBdownload
Llama-3.3-70B-Instruct.Q3_K_M.gguf31.91 GBdownload
Llama-3.3-70B-Instruct.Q3_K_S.gguf28.79 GBdownload
Llama-3.3-70B-Instruct.Q4_K_M.gguf39.60 GBdownload
Llama-3.3-70B-Instruct.Q4_K_S.gguf37.58 GBdownload
Llama-3.3-70B-Instruct.Q5_K_M.gguf46.52 GBdownload
Llama-3.3-70B-Instruct.Q5_K_S.gguf45.32 GBdownload
Llama-3.3-70B-Instruct.Q6_K.gguf-00001-of-00006.gguf10.59 GBdownload
Llama-3.3-70B-Instruct.Q6_K.gguf-00002-of-00006.gguf9.33 GBdownload
Llama-3.3-70B-Instruct.Q6_K.gguf-00003-of-00006.gguf9.16 GBdownload
Llama-3.3-70B-Instruct.Q6_K.gguf-00004-of-00006.gguf9.26 GBdownload
GGUF: bartowski/Llama-3.3-70B-Instruct-abliterated-GGUF · 61,272 downloads
FileSize
Llama-3.3-70B-Instruct-abliterated-IQ1_M.gguf15.60 GBdownload
Llama-3.3-70B-Instruct-abliterated-IQ2_M.gguf22.46 GBdownload
Llama-3.3-70B-Instruct-abliterated-IQ2_S.gguf20.71 GBdownload
Llama-3.3-70B-Instruct-abliterated-IQ2_XS.gguf19.69 GBdownload
Llama-3.3-70B-Instruct-abliterated-IQ2_XXS.gguf17.79 GBdownload
Llama-3.3-70B-Instruct-abliterated-IQ3_M.gguf29.74 GBdownload
Llama-3.3-70B-Instruct-abliterated-IQ3_XXS.gguf25.58 GBdownload
Llama-3.3-70B-Instruct-abliterated-IQ4_NL.gguf37.30 GBdownload
Llama-3.3-70B-Instruct-abliterated-IQ4_XS.gguf35.30 GBdownload
Llama-3.3-70B-Instruct-abliterated-Q2_K.gguf24.56 GBdownload
Llama-3.3-70B-Instruct-abliterated-Q2_K_L.gguf25.52 GBdownload
Llama-3.3-70B-Instruct-abliterated-Q3_K_L.gguf34.59 GBdownload

What you save self-hosting this model

Uses the model's own API price when it exists; otherwise the closest commercial equivalent.

Waiting for hardware detection…

RunLocal presents third-party leaderboard data (local.ai / Exo Labs, 2026-08) with its own fit and cost estimates. Verify pricing, licenses and file sizes at the source before deploying a model.