Model Sizing & VRAM Requirements
The Golden Rule: 2 Bytes per Parameter (FP16/BF16)
When running a Dense Large Language Model at standard 16-bit precision (FP16 or BF16), each parameter requires approximately 2 bytes of VRAM.
Formula:
Model Weights (GB) = Parameter Count (Billions) * 2
Worked Example: Llama-3 8B
Let’s calculate the footprint for Llama-3 8B running at FP16.
1. Model Weights
8 Billion Parameters × 2 bytes = ~16 GB VRAM
2. KV Cache (Context Memory)
Generative AI requires active memory to store the Key-Value (KV) Cache for the conversational context during inference. This grows linearly with context length and batch size.
- Rule of thumb: Allocate an additional 2–4 GB VRAM for the KV cache for typical serving scenarios (batch size 1-4, context 4k-8k).
- Long context (32k-128k) or large batch sizes can push this to 8-16 GB+.
3. Total Required VRAM
| Component | VRAM Estimate |
|---|---|
| Model Weights (FP16) | ~16 GB |
| KV Cache + Overhead | ~2–4 GB |
| Total Required | ~18–20 GB |
Hardware Fit: Llama-3 8B
- NVIDIA A10 (24 GB): Fits comfortably. Headroom for KV cache and OS/driver overhead.
- NVIDIA A100 (40/80 GB): Fits easily; allows larger batch sizes or longer context.
- RTX 3090 / 4090 (24 GB): Fits comfortably (Consumer favorite).
- RTX 3080 / 4080 (16 GB): Too tight / OOM. Weights alone take 16 GB, leaving 0 bytes for KV cache.
Critical Note: This assumes standard context lengths (4k-8k). Running at maximum context length (e.g., 128k) on a 24 GB card will likely OOM (Out of Memory) due to KV cache explosion.
The Quantization Multiplier
You can drastically reduce the Model Weights footprint by quantizing (compressing) the model. The KV Cache size is usually unaffected (stays FP16/BF16) unless specifically quantized (KV Cache Quantization).
| Precision | Bits/Param | Multiplier | Llama-3 8B Weights | Fits on? |
|---|---|---|---|---|
| FP16 / BF16 | 16 bits | 1.0x (Baseline) | 16 GB | A10, 24GB Consumer |
| INT8 | 8 bits | 0.5x | ~8 GB | RTX 3080 (10GB), T4 (16GB) |
| INT4 (GPTQ/AWQ/GGUF) | 4 bits | 0.25x | ~4 GB | RTX 3060 (12GB), M1/M2/M3 Mac (8GB+), Mobile |
| INT3 / INT2 | 3-2 bits | ~0.15x | ~2.5 GB | Extreme compression, quality loss |
Updated Totals with Quantization (Llama-3 8B)
| Precision | Weights | + KV Cache (4GB) | Total | Viable Hardware |
|---|---|---|---|---|
| FP16 | 16 GB | + 4 GB | 20 GB | 24 GB GPU |
| INT8 | 8 GB | + 4 GB | 12 GB | 16 GB GPU (RTX 3080, A10G) |
| INT4 | 4 GB | + 4 GB | 8 GB | 12 GB GPU (RTX 3060, Laptop 4070), MacBook |
Quick Reference Cheatsheet (Model Weights Only)
| Model Size | FP16 (16-bit) | INT8 (8-bit) | INT4 (4-bit) |
|---|---|---|---|
| 1B | 2 GB | 1 GB | 0.5 GB |
| 3B | 6 GB | 3 GB | 1.5 GB |
| 7B / 8B | 16 GB | 8 GB | 4 GB |
| 13B / 14B | 26 GB | 13 GB | 6.5 GB |
| 32B / 34B | 64 GB | 32 GB | 16 GB |
| 70B / 72B | 140 GB | 70 GB | 35 GB |
| 120B (MoE Active) | 240 GB | 120 GB | 60 GB |
MoE Note: For MoE models (Mixtral, DeepSeek), use Active Params for the weight calculation, but you must load Total Params into VRAM (unless using offloading).
Beyond Weights: The Full Deployment Budget
When planning a production deployment, add these overheads:
- Framework Overhead: PyTorch/vLLM/TGI runtime (~1-2 GB).
- Activation Memory: Intermediate activations during forward pass (scales with batch size × sequence length).
- CUDA Context: ~0.5–1 GB per process.
- OS / Display: 1-2 GB if running a desktop environment.
Safe Formula for Production:
Total VRAM = (Model_Weights_Quantized) + (KV_Cache_Estimate) + 4 GB (Overhead Buffer)
Next Steps
- See Quantization for deep dive on INT4, GPTQ, AWQ, GGUF, and calibration.
- See Model Architecture for MoE sizing (Active vs Total params).