Quantization
The Core Idea
Quantization maps high-precision floating-point weights (FP16/BF16) to lower-precision integers (INT8, INT4).
- Storage: 16-bit → 4-bit = 4x smaller model file.
- Compute: Integer matrix multiplication (GEMM) is 2-4x faster and energy efficient on modern hardware (Tensor Cores).
- Bandwidth: Moving 4x less data from VRAM → Compute units is often the real speedup bottleneck.
Precision Levels Compared
| Format | Bits | Size (vs FP16) | Quality | Hardware Support | Use Case |
|---|---|---|---|---|---|
| FP16 / BF16 | 16 | 1.0x (Baseline) | Reference | All GPUs (Tensor Cores) | Training, Quality-critical |
| FP8 (E4M3/E5M2) | 8 | 0.5x | Near-lossless | H100, H200, Blackwell (Native) | Training, High-perf Inference |
| INT8 | 8 | 0.5x | Negligible Loss | All Modern GPUs (TensorRT, vLLM) | Production Standard |
| INT4 | 4 | 0.25x | Low Loss (1-2% PPL) | All GPUs (via GPTQ/AWQ kernels), Apple Silicon, Mobile | Consumer / Edge / Cost Savings |
| INT3 / INT2 | 3-2 | ~0.15x | Noticeable Degradation | Specialized Kernels | Extreme Compression |
Rule of Thumb: INT4 is the “Sweet Spot” for open-weight LLMs. Quality drop is barely perceptible for chat/coding; VRAM drops 4x.
How It Works: The Math
We map a continuous float range $[min, max]$ to discrete integer bins $[0, 2^N - 1]$.
1. Symmetric (Per-Tensor / Per-Channel)
$$ q = \text{round}(w / s) \quad \text{where } s = \frac{\max(|w|)}{2^{N-1} - 1} $$* Zero-point is always 0. Simple, fast. Used for Activations (INT8).
2. Asymmetric (Zero-Point)$$ q = \text{round}(w / s) + z \quad \text{where } z \text{ centers the range} $$
- Required for Weights (INT4) because weight distributions are rarely centered at 0.
3. Group Quantization (The Secret to INT4 Quality)
Instead of one scale for the whole layer (Per-Tensor) or row (Per-Channel), we group weights (e.g., Group Size 128).
- Row:
[w1, w2, ..., w4096]→ 1 scale (Bad for outliers). - Group:
[w1..w128], [w129..w256]...→ 32 scales per row. - Captures local outliers. Essential for INT4.
Post-Training Quantization (PTQ) Methods
You take a trained FP16 model and compress it without retraining.
1. RTN (Round-To-Nearest) - Naive
- Just clamp and round. Fails catastrophically at INT4.
2. GPTQ (Generative Pre-trained Transformer Quantization) - Classic Standard
- Paper: “GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers” (Frantar et al., 2022).
- Method: Layer-by-layer, uses Hessian (2nd order gradient) information to compensate rounding errors.
- Pros: High quality, widely supported (AutoGPTQ, ExLlamaV2).
- Cons: Slow calibration (hours for 70B). Requires calibration dataset.
3. AWQ (Activation-aware Weight Quantization) - Current Favorite
- Paper: “AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration” (Lin et al., 2023).
- Insight: Only 1% of weights (those multiplied by large activations) matter for accuracy. Protect those weights.
- Method: Scale weights before rounding based on activation magnitudes. No Hessian needed.
- Pros: Extremely fast (minutes), SOTA quality at INT4/INT3, no calibration data needed (or tiny sample).
- Support:
autoawq,llama.cpp(viallama-awq), vLLM, TGI.
4. GGUF / llama.cpp (K-Quants) - Apple Silicon / CPU / Hybrid
- Format: Single-file container (replaces GGML).
- Schemes:
Q4_K_M,Q5_K_M,Q8_0(Mixed precision: Important layers INT8/INT4, outliers FP16). - Pros: Runs on CPU/Metal (Mac) efficiently. Hybrid GPU/CPU offloading.
- Standard:
Q4_K_M(4-bit, medium) is the default recommendation.
Quantization Workflow (AWQ Example)
# 1. Install
pip install autoawq
# 2. Quantize (Python)
from awq import AutoAWQForCausalLM
from transformers import AutoTokenizer
model_path = "mistralai/Mistral-7B-v0.1"
quant_path = "Mistral-7B-v0.1-AWQ-INT4"
model = AutoAWQForCausalLM.from_pretrained(model_path, **{"low_cpu_mem_usage": True})
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
# Quantize: 4-bit, Group Size 128, Zero-Point True
model.quantize(tokenizer, quant_config={"zero_point": True, "q_group_size": 128, "w_bit": 4, "version": "GEMM"})
model.save_quantized(quant_path)
tokenizer.save_pretrained(quant_path)
Quantization Aware Training (QAT) & Fine-Tuning
If PTQ quality isn’t enough (e.g., INT3, INT2, or recovering lost benchmark points):
- QAT: Insert Fake Quantization nodes during training/finetuning. Model learns to be robust to quantization noise.
- QLoRA (Quantized LoRA): Fine-tune a frozen INT4 model using LoRA adapters in FP16/BF16.
- Paper: “QLoRA: Efficient Finetuning of Quantized LLMs” (Dettmers et al., 2023).
- Magic: Backprop through the frozen INT4 weights using a Double Quantization trick for optimizer states.
- Result: Finetune 65B model on 48 GB VRAM (vs 780 GB for FP16).
KV Cache Quantization
Quantizing weights saves static VRAM. Quantizing the KV Cache saves dynamic VRAM (scales with context).
- FP8 KV Cache: Native on H100 (TensorRT-LLM, vLLM). ~2x context length for same VRAM.
- INT4 KV Cache: Research stage (KIVI, GEAR). Significant quality risk (accumulation error).
- Head-wise / Layer-wise: Quantize less important heads/layers more aggressively.
Quality Evaluation Checklist
When you quantize a new model, verify:
- Perplexity (PPL): WikiText2, C4. Target: < 1-2% increase vs FP16.
- Downstream Benchmarks: MMLU, GSM8K, HumanEval, BBH.
- “Vibe Check”: Long conversation, coding, reasoning. Look for:
- Repetition loops.
- Loss of formatting (JSON, Markdown).
- Foreign language degradation.
- Outlier Sensitivity: Models with massive activation outliers (e.g., some MoE, OPT) quantize worse.
Summary: What Should You Use?
| Scenario | Recommendation |
|---|---|
| NVIDIA GPU (RTX 30/40, A10, A100) | AWQ INT4 (Group 128) via vLLM / TGI / ExLlamaV2 |
| Apple Silicon (Mac M1/M2/M3) | GGUF Q4_K_M via llama.cpp / Ollama / LM Studio |
| CPU Only / Hybrid Offload | GGUF Q4_K_M or Q5_K_M |
| H100 / Hopper Cluster | FP8 (Native) or AWQ INT4 for max throughput |
| Fine-tuning Custom Model | QLoRA (INT4 Base + LoRA Adapters) |
| Maximum Quality Required | FP16 / BF16 (No Quantization) |
Key Papers
- GPTQ (2022)
- AWQ (2023)
- QLoRA (2023)
- LLM.int8() (2022) – Emergent Outliers
- FP8 Formats (2022)
Related Pages
- Model Sizing - VRAM calculations for FP16/INT8/INT4.
- Model Architecture - MoE Active vs Total params sizing.