← Back to Glossary

Model Sizing & VRAM Requirements

The Golden Rule: 2 Bytes per Parameter (FP16/BF16)

When running a Dense Large Language Model at standard 16-bit precision (FP16 or BF16), each parameter requires approximately 2 bytes of VRAM.

Formula: Model Weights (GB) = Parameter Count (Billions) * 2


Worked Example: Llama-3 8B

Let’s calculate the footprint for Llama-3 8B running at FP16.

1. Model Weights

8 Billion Parameters × 2 bytes = ~16 GB VRAM

2. KV Cache (Context Memory)

Generative AI requires active memory to store the Key-Value (KV) Cache for the conversational context during inference. This grows linearly with context length and batch size.

3. Total Required VRAM

Component VRAM Estimate
Model Weights (FP16) ~16 GB
KV Cache + Overhead ~2–4 GB
Total Required ~18–20 GB

Hardware Fit: Llama-3 8B

Critical Note: This assumes standard context lengths (4k-8k). Running at maximum context length (e.g., 128k) on a 24 GB card will likely OOM (Out of Memory) due to KV cache explosion.


The Quantization Multiplier

You can drastically reduce the Model Weights footprint by quantizing (compressing) the model. The KV Cache size is usually unaffected (stays FP16/BF16) unless specifically quantized (KV Cache Quantization).

Precision Bits/Param Multiplier Llama-3 8B Weights Fits on?
FP16 / BF16 16 bits 1.0x (Baseline) 16 GB A10, 24GB Consumer
INT8 8 bits 0.5x ~8 GB RTX 3080 (10GB), T4 (16GB)
INT4 (GPTQ/AWQ/GGUF) 4 bits 0.25x ~4 GB RTX 3060 (12GB), M1/M2/M3 Mac (8GB+), Mobile
INT3 / INT2 3-2 bits ~0.15x ~2.5 GB Extreme compression, quality loss

Updated Totals with Quantization (Llama-3 8B)

Precision Weights + KV Cache (4GB) Total Viable Hardware
FP16 16 GB + 4 GB 20 GB 24 GB GPU
INT8 8 GB + 4 GB 12 GB 16 GB GPU (RTX 3080, A10G)
INT4 4 GB + 4 GB 8 GB 12 GB GPU (RTX 3060, Laptop 4070), MacBook

Quick Reference Cheatsheet (Model Weights Only)

Model Size FP16 (16-bit) INT8 (8-bit) INT4 (4-bit)
1B 2 GB 1 GB 0.5 GB
3B 6 GB 3 GB 1.5 GB
7B / 8B 16 GB 8 GB 4 GB
13B / 14B 26 GB 13 GB 6.5 GB
32B / 34B 64 GB 32 GB 16 GB
70B / 72B 140 GB 70 GB 35 GB
120B (MoE Active) 240 GB 120 GB 60 GB

MoE Note: For MoE models (Mixtral, DeepSeek), use Active Params for the weight calculation, but you must load Total Params into VRAM (unless using offloading).


Beyond Weights: The Full Deployment Budget

When planning a production deployment, add these overheads:

  1. Framework Overhead: PyTorch/vLLM/TGI runtime (~1-2 GB).
  2. Activation Memory: Intermediate activations during forward pass (scales with batch size × sequence length).
  3. CUDA Context: ~0.5–1 GB per process.
  4. OS / Display: 1-2 GB if running a desktop environment.

Safe Formula for Production:

Total VRAM = (Model_Weights_Quantized) + (KV_Cache_Estimate) + 4 GB (Overhead Buffer)

Next Steps