Local AI Local AI Intermediate

Pick the Right Quantization

Balance speed, quality, and VRAM for your hardware

53 of 66

What is quantization?

Quantization reduces the precision of model weights. A 16-bit weight becomes 4-bit or 5-bit. The model shrinks dramatically and runs faster, but loses a small amount of quality.

Common formats

FormatSize vs FP16QualitySpeed
Q4_K_M~25%GoodFastest
Q5_K_M~31%Very goodFast
Q8_0~50%ExcellentModerate
FP16100%BestSlowest

Picking a format

  • Q4_K_M for quick drafts and machines with 8 GB VRAM.
  • Q5_K_M for daily coding help on 12 GB VRAM.
  • Q8_0 when quality matters more than speed.
  • FP16 only if you have plenty of VRAM and want full fidelity.

Estimate VRAM

Rough formula for a Q4 model:

VRAM ≈ (params in billions × 1.2) + 1 GB overhead

So a 7B Q4 model needs about 9–10 GB total during generation.

Find quantized models

Search Hugging Face for

GGUF Q4_K_M

Download the ".gguf" file and load it directly in LM Studio.

Working out which model to run this on? See The Codex. Packaging it as a reusable skill? See The Armory.