← Back to Glossary

Quantization

The Core Idea

Quantization maps high-precision floating-point weights (FP16/BF16) to lower-precision integers (INT8, INT4).


Precision Levels Compared

Format Bits Size (vs FP16) Quality Hardware Support Use Case
FP16 / BF16 16 1.0x (Baseline) Reference All GPUs (Tensor Cores) Training, Quality-critical
FP8 (E4M3/E5M2) 8 0.5x Near-lossless H100, H200, Blackwell (Native) Training, High-perf Inference
INT8 8 0.5x Negligible Loss All Modern GPUs (TensorRT, vLLM) Production Standard
INT4 4 0.25x Low Loss (1-2% PPL) All GPUs (via GPTQ/AWQ kernels), Apple Silicon, Mobile Consumer / Edge / Cost Savings
INT3 / INT2 3-2 ~0.15x Noticeable Degradation Specialized Kernels Extreme Compression

Rule of Thumb: INT4 is the “Sweet Spot” for open-weight LLMs. Quality drop is barely perceptible for chat/coding; VRAM drops 4x.


How It Works: The Math

We map a continuous float range $[min, max]$ to discrete integer bins $[0, 2^N - 1]$.

1. Symmetric (Per-Tensor / Per-Channel)

$$ q = \text{round}(w / s) \quad \text{where } s = \frac{\max(|w|)}{2^{N-1} - 1} $$* Zero-point is always 0. Simple, fast. Used for Activations (INT8).

2. Asymmetric (Zero-Point)$$ q = \text{round}(w / s) + z \quad \text{where } z \text{ centers the range} $$

3. Group Quantization (The Secret to INT4 Quality)

Instead of one scale for the whole layer (Per-Tensor) or row (Per-Channel), we group weights (e.g., Group Size 128).


Post-Training Quantization (PTQ) Methods

You take a trained FP16 model and compress it without retraining.

1. RTN (Round-To-Nearest) - Naive

2. GPTQ (Generative Pre-trained Transformer Quantization) - Classic Standard

3. AWQ (Activation-aware Weight Quantization) - Current Favorite

4. GGUF / llama.cpp (K-Quants) - Apple Silicon / CPU / Hybrid


Quantization Workflow (AWQ Example)

# 1. Install
pip install autoawq

# 2. Quantize (Python)
from awq import AutoAWQForCausalLM
from transformers import AutoTokenizer

model_path = "mistralai/Mistral-7B-v0.1"
quant_path = "Mistral-7B-v0.1-AWQ-INT4"

model = AutoAWQForCausalLM.from_pretrained(model_path, **{"low_cpu_mem_usage": True})
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)

# Quantize: 4-bit, Group Size 128, Zero-Point True
model.quantize(tokenizer, quant_config={"zero_point": True, "q_group_size": 128, "w_bit": 4, "version": "GEMM"})

model.save_quantized(quant_path)
tokenizer.save_pretrained(quant_path)

Quantization Aware Training (QAT) & Fine-Tuning

If PTQ quality isn’t enough (e.g., INT3, INT2, or recovering lost benchmark points):

  1. QAT: Insert Fake Quantization nodes during training/finetuning. Model learns to be robust to quantization noise.
  2. QLoRA (Quantized LoRA): Fine-tune a frozen INT4 model using LoRA adapters in FP16/BF16.
    • Paper: “QLoRA: Efficient Finetuning of Quantized LLMs” (Dettmers et al., 2023).
    • Magic: Backprop through the frozen INT4 weights using a Double Quantization trick for optimizer states.
    • Result: Finetune 65B model on 48 GB VRAM (vs 780 GB for FP16).

KV Cache Quantization

Quantizing weights saves static VRAM. Quantizing the KV Cache saves dynamic VRAM (scales with context).


Quality Evaluation Checklist

When you quantize a new model, verify:

  1. Perplexity (PPL): WikiText2, C4. Target: < 1-2% increase vs FP16.
  2. Downstream Benchmarks: MMLU, GSM8K, HumanEval, BBH.
  3. “Vibe Check”: Long conversation, coding, reasoning. Look for:
    • Repetition loops.
    • Loss of formatting (JSON, Markdown).
    • Foreign language degradation.
  4. Outlier Sensitivity: Models with massive activation outliers (e.g., some MoE, OPT) quantize worse.

Summary: What Should You Use?

Scenario Recommendation
NVIDIA GPU (RTX 30/40, A10, A100) AWQ INT4 (Group 128) via vLLM / TGI / ExLlamaV2
Apple Silicon (Mac M1/M2/M3) GGUF Q4_K_M via llama.cpp / Ollama / LM Studio
CPU Only / Hybrid Offload GGUF Q4_K_M or Q5_K_M
H100 / Hopper Cluster FP8 (Native) or AWQ INT4 for max throughput
Fine-tuning Custom Model QLoRA (INT4 Base + LoRA Adapters)
Maximum Quality Required FP16 / BF16 (No Quantization)

Key Papers