Model Architecture
The Fundamental Split: Dense vs. Sparse
Modern Large Language Models (LLMs) generally fall into two architectural paradigms regarding how they process parameters during inference: Dense and Sparse.
1. Dense Models (The Standard)
In a Dense Model, every single parameter is activated for every single input token.
- Examples: GPT-3, LLaMA 1/2/3, BERT, Mistral 7B.
- Mechanism: Input $\rightarrow$ Embedding $\rightarrow$ All Layers (Full Matrix Multiplication) $\rightarrow$ Output.
- Compute Cost: Fixed and predictable. FLOPs $\approx 2 \times$ Params $\times$ Tokens.
Pros
- Simplicity: Easy to train, distribute (Data Parallelism), and optimize.
- Quality: Generally highest quality per parameter count because the full model capacity is always applied.
- Hardware Friendly: Matrix multiplications (GEMMs) are highly optimized on GPUs/TPUs.
Cons
- Scaling Wall: To get smarter, you must activate more parameters. Training/inference cost grows linearly with model size.
- Memory Bound: You must fit the entire model in VRAM to run inference.
2. Sparse Models (Conditional Computation)
In a Sparse Model, only a subset of parameters is activated for any given input token. The total parameter count can be massive (Trillions), but the active parameter count stays manageable.
- Key Concept: Conditional Computation — “Use only the parts of the brain relevant to the current thought.”
- Active Params $\ll$ Total Params.
Pros
- Decoupled Scaling: You can increase capacity (Total Params / Knowledge Storage) without increasing compute (Active Params / FLOPs) proportionally.
- Specialization: Different parts of the network can specialize in different domains (languages, coding, reasoning).
Cons
- Training Instability: Routing collapse (all tokens go to one expert), load balancing issues.
- Communication Overhead: Experts often live on different GPUs (Expert Parallelism), requiring high-bandwidth interconnects (NVLink/InfiniBand).
- Inference Complexity: Requires dynamic scheduling; harder to batch efficiently.
3. Mixture of Experts (MoE) — The Dominant Sparse Architecture
Mixture of Experts (MoE) is the specific implementation of sparsity used in almost all major sparse LLMs today (GPT-4, Mixtral, DeepSeek-V2/V3, Nemotron 3 Ultra, DBRX).
Core Components
- Experts: Independent Feed-Forward Networks (FFNs). Usually 8, 16, 32, or 64 per layer.
- Think of them as specialized sub-networks (e.g., “Python Expert”, “French Expert”, “Math Expert”).
- Router (Gate): A small neural network (usually a linear layer + Softmax) that takes the token embedding and outputs probabilities over experts.
- Top-k Routing: Select the $k$ highest-probability experts (usually $k=1$ or $k=2$).
The Forward Pass (Per Token, Per MoE Layer)
# Pseudocode
router_logits = router(token_embedding) # Shape: [num_experts]
router_probs = softmax(router_logits, dim=-1)
top_k_weights, top_k_indices = topk(router_probs, k=2)
output = 0
for weight, expert_idx in zip(top_k_weights, top_k_indices):
expert_output = experts[expert_idx](token_embedding)
output += weight * expert_output
Key Configuration: Granularity
- Token-level: Every token chooses its own experts (Standard in Mixtral, DeepSeek). High flexibility.
- Sequence-level: All tokens in a sequence share experts (Rare now).
Load Balancing: The Critical Trick
If the router always picks Expert #3, other experts never learn. We add an Auxiliary Loss to the training objective:
$$ L_{total} = L_{LM} + \lambda \cdot L_{balance} $$
Where $L_{balance}$ encourages uniform expert utilization (e.g., minimizing coefficient of variation of expert loads).
Popular MoE Architectures Comparison
| Model | Total Params | Active Params | Experts/Layer | Top-k | Notable Feature |
|---|---|---|---|---|---|
| Mixtral 8x7B | 46.7B | 12.9B | 8 | 2 | Shared Attention, No Shared Experts |
| Mixtral 8x22B | 141B | 39B | 8 | 2 | Larger experts |
| DeepSeek-V2 | 236B | 21B | 160 (Shared+Routed) | 6 | MHA $\rightarrow$ MLA, Shared Experts |
| DeepSeek-V3 | 671B | 37B | 256 (Shared+Routed) | 8 | FP8 Training, MTP |
| DBRX | 132B | 36B | 16 | 4 | Fine-grained experts (many small) |
| GPT-4 (Rumored) | ~1.8T | ~220B | 16 | 2 | Massive scale |
Note on “8x7B” Naming: Mixtral 8x7B has 8 experts of ~7B params each. It is not an ensemble of 8 models. It is one model with sparse layers.
MoE Innovations: Shared Experts & MLA
Shared Experts (DeepSeek, OLMoE)
Always-active experts (usually 1 or 2 per layer) that process every token.
- Why? Guarantees a baseline of “general knowledge” processing for every token, preventing the router from ignoring fundamental features.
- Reduces routing variance.
Multi-Head Latent Attention (MLA) — DeepSeek-V2/V3
Not strictly MoE, but critical for MoE scaling.
- Problem: KV Cache grows with context length. In MoE, you need many GPUs (Expert Parallelism), replicating KV Cache kills memory.
- Solution: Compress Keys/Values into a Latent Vector before RoPE. Drastically reduces KV Cache size (93% reduction vs standard MHA).
- Enables massive context windows (128k+) on MoE models.
Summary: When to Use Which?
| Scenario | Recommendation |
|---|---|
| Training from scratch, limited GPU comms | Dense (LLaMA, Mistral) |
| Max Quality per Active FLOP | Dense |
| Serving massive models on limited VRAM | MoE (Quantized Mixtral/DeepSeek) |
| Training Frontier Model ($>$100B Active) | MoE (Only viable path) |
| Multilingual / Multi-domain | MoE (Natural specialization) |
Further Reading
- “Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer” (Shazeer et al., 2017) — The seminal paper.
- “Mixtral of Experts” (Mistral AI, 2024) — Open weights MoE reference.
- “DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model” (2024) — Modern MoE + MLA deep dive.