← Back to Glossary

Model Architecture

The Fundamental Split: Dense vs. Sparse

Modern Large Language Models (LLMs) generally fall into two architectural paradigms regarding how they process parameters during inference: Dense and Sparse.


1. Dense Models (The Standard)

In a Dense Model, every single parameter is activated for every single input token.

Pros

Cons


2. Sparse Models (Conditional Computation)

In a Sparse Model, only a subset of parameters is activated for any given input token. The total parameter count can be massive (Trillions), but the active parameter count stays manageable.

Pros

Cons


3. Mixture of Experts (MoE) — The Dominant Sparse Architecture

Mixture of Experts (MoE) is the specific implementation of sparsity used in almost all major sparse LLMs today (GPT-4, Mixtral, DeepSeek-V2/V3, Nemotron 3 Ultra, DBRX).

Core Components

  1. Experts: Independent Feed-Forward Networks (FFNs). Usually 8, 16, 32, or 64 per layer.
    • Think of them as specialized sub-networks (e.g., “Python Expert”, “French Expert”, “Math Expert”).
  2. Router (Gate): A small neural network (usually a linear layer + Softmax) that takes the token embedding and outputs probabilities over experts.
  3. Top-k Routing: Select the $k$ highest-probability experts (usually $k=1$ or $k=2$).

The Forward Pass (Per Token, Per MoE Layer)

# Pseudocode
router_logits = router(token_embedding)           # Shape: [num_experts]
router_probs = softmax(router_logits, dim=-1)
top_k_weights, top_k_indices = topk(router_probs, k=2)

output = 0
for weight, expert_idx in zip(top_k_weights, top_k_indices):
    expert_output = experts[expert_idx](token_embedding)
    output += weight * expert_output

Key Configuration: Granularity

Load Balancing: The Critical Trick

If the router always picks Expert #3, other experts never learn. We add an Auxiliary Loss to the training objective:

$$ L_{total} = L_{LM} + \lambda \cdot L_{balance} $$

Where $L_{balance}$ encourages uniform expert utilization (e.g., minimizing coefficient of variation of expert loads).


Model Total Params Active Params Experts/Layer Top-k Notable Feature
Mixtral 8x7B 46.7B 12.9B 8 2 Shared Attention, No Shared Experts
Mixtral 8x22B 141B 39B 8 2 Larger experts
DeepSeek-V2 236B 21B 160 (Shared+Routed) 6 MHA $\rightarrow$ MLA, Shared Experts
DeepSeek-V3 671B 37B 256 (Shared+Routed) 8 FP8 Training, MTP
DBRX 132B 36B 16 4 Fine-grained experts (many small)
GPT-4 (Rumored) ~1.8T ~220B 16 2 Massive scale

Note on “8x7B” Naming: Mixtral 8x7B has 8 experts of ~7B params each. It is not an ensemble of 8 models. It is one model with sparse layers.


MoE Innovations: Shared Experts & MLA

Shared Experts (DeepSeek, OLMoE)

Always-active experts (usually 1 or 2 per layer) that process every token.

Multi-Head Latent Attention (MLA) — DeepSeek-V2/V3

Not strictly MoE, but critical for MoE scaling.


Summary: When to Use Which?

Scenario Recommendation
Training from scratch, limited GPU comms Dense (LLaMA, Mistral)
Max Quality per Active FLOP Dense
Serving massive models on limited VRAM MoE (Quantized Mixtral/DeepSeek)
Training Frontier Model ($>$100B Active) MoE (Only viable path)
Multilingual / Multi-domain MoE (Natural specialization)

Further Reading