Local AI
Local AI
Intermediate
Pick the Right Quantization
Balance speed, quality, and VRAM for your hardware
53 of 66
What is quantization?
Quantization reduces the precision of model weights. A 16-bit weight becomes 4-bit or 5-bit. The model shrinks dramatically and runs faster, but loses a small amount of quality.
Common formats
| Format | Size vs FP16 | Quality | Speed |
|---|---|---|---|
| Q4_K_M | ~25% | Good | Fastest |
| Q5_K_M | ~31% | Very good | Fast |
| Q8_0 | ~50% | Excellent | Moderate |
| FP16 | 100% | Best | Slowest |
Picking a format
- Q4_K_M for quick drafts and machines with 8 GB VRAM.
- Q5_K_M for daily coding help on 12 GB VRAM.
- Q8_0 when quality matters more than speed.
- FP16 only if you have plenty of VRAM and want full fidelity.
Estimate VRAM
Rough formula for a Q4 model:
VRAM ≈ (params in billions × 1.2) + 1 GB overhead
So a 7B Q4 model needs about 9–10 GB total during generation.
Find quantized models
Search Hugging Face for
GGUF Q4_K_M
Download the ".gguf" file and load it directly in LM Studio.