LLM-D (Large Language Model Distributed)
What is LLM-D?
LLM-D (Large Language Model Distributed) is an open architecture and community initiative focused on disaggregated, distributed LLM inference. It moves beyond single-node serving (like standard vLLM or TGI) to enable running massive models—or high-throughput workloads—across multiple GPUs and multiple nodes using standard cloud-native tooling (Kubernetes, gRPC, etcd).
Core Philosophy: Treat LLM inference as a distributed systems problem, not just a model serving problem.
The Problem: Monolithic Serving Limits
Standard serving engines (vLLM, TGI, TensorRT-LLM) typically run on a single node (1-8 GPUs). This creates hard ceilings:
- Model Size Ceiling: Model must fit in VRAM of one node (NVLink domain).
- Throughput Ceiling: Limited by single-node memory bandwidth & compute.
- Scaling Granularity: Scale up = buy bigger node (expensive). Scale out = replicate whole model (inefficient for large models).
- Resource Fragmentation: Hard to pack mixed workloads (prefill vs decode) efficiently.
The LLM-D Solution: Disaggregated Inference
LLM-D splits the inference pipeline into independent, scalable services connected via high-speed networking (NVLink, NVSwitch, InfiniBand, RoCE).
1. Prefill-Decode Disaggregation (The Big Win)
- Prefill Phase: Compute-heavy, memory-bandwidth bound, highly parallelizable. Processes prompt tokens.
- Decode Phase: Memory-bound, sequential, latency-sensitive. Generates output tokens one by one.
LLM-D Architecture:
- Prefill Pool: Stateless workers optimized for throughput (large batches, KV cache offload to CPU/Disk).
- Decode Pool: Stateful workers optimized for latency (KV cache resident in GPU VRAM).
- Router / Scheduler: Directs incoming requests to Prefill, transfers KV cache to Decode pool.
Result: 3-10x better GPU utilization. Prefill GPUs stay 100% busy; Decode GPUs serve many concurrent users with low latency.
2. KV Cache Transfer Protocol
The critical technical challenge: Moving the KV Cache (Key-Value tensors) from Prefill workers to Decode workers instantly.
- Protocol: Custom high-performance gRPC / RDMA / Mooncake (ByteDance) / LMCache backends.
- Zero-Copy: GPU $ ightarrow$ NIC $ ightarrow$ GPU (GPUDirect RDMA) avoiding CPU/system RAM bounce.
- Chunking/Compression: Optional FP8/INT8 quantization of KV cache for transfer.
3. Distributed KV Cache (LMCache / Mooncake Integration)
- External KV Store: Offload cold/overflow KV cache to CPU RAM, NVMe, or distributed object store (S3/MinIO).
- Cache Reuse: Prefix Caching across requests (shared system prompts, few-shot examples) becomes trivial when KV cache is externalized.
- Persistence: Survives worker restarts / rolling updates.
Core Components (The Stack)
| Layer | Technology | Role |
|---|---|---|
| Orchestration | Kubernetes (K8s) | Scheduling, scaling, self-healing, networking |
| Service Mesh | gRPC / Envoy / Cilium | Low-latency service-to-service comms |
| State Coordination | etcd / Redis | Worker discovery, load metrics, KV cache metadata |
| Inference Engine | vLLM (Modified) | Core execution engine (P/D workers) |
| KV Transport | Mooncake / LMCache / XConnector | High-speed KV cache transfer |
| Autoscaling | KEDA / Custom Metrics | Scale Prefill/Decode pools independently based on queue depth |
Key Concepts
PD Disaggregation (Prefill-Decode Separation)
The architectural pattern where the Prefill (prompt processing) and Decode (token generation) stages run on separate GPU pools.
Chunked Prefill / Paged Attention
vLLM’s Paged Attention is a prerequisite. It allows KV cache to be non-contiguous (pages), making it transferable and swappable.
KV Cache as a First-Class Citizen
In LLM-D, the KV Cache is not an implementation detail of the engine—it is a distributed data object with metadata (model version, sequence ID, token positions, topology).
Heterogeneous Clusters
LLM-D enables running Prefill on H100s (high FLOPs) and Decode on A100s/L4s (high VRAM/$), or mixing GPU generations.
Comparison: LLM-D vs Standard Serving
| Feature | Standard vLLM / TGI (Single Node) | LLM-D (Distributed) |
|---|---|---|
| Max Model Size | Limited by Node VRAM (e.g., 8x H100 = 640GB) | Cluster Scale (Petabytes aggregate VRAM) |
| Scaling | Replicate whole model (Data Parallel) | Disaggregated Scale (Independent P/D scaling) |
| GPU Utilization | ~40-60% (Decode bound) | 80-95% (Workload isolation) |
| Infrastructure | VMs / Bare Metal | Kubernetes Native |
| Ops Complexity | Low | High (Distributed systems expertise needed) |
| Latency (TTFT) | Low (Local) | Slightly Higher (Network hop) |
| Throughput (TPS) | Limited by single node | Linear Scale-out |
| Prefix Caching | Local only | Global / Distributed |
When to Use LLM-D
✅ Use LLM-D if:
- Serving MoE models (DeepSeek-V3, Mixtral) where Expert Parallelism spans nodes.
- High Throughput requirements (>10k concurrent users).
- Long Context (128k-1M+) requiring KV cache offloading.
- Running on Kubernetes with existing MLOps/GitOps pipelines.
- Need Cost Optimization via heterogeneous GPU pools / spot instances.
❌ Stick to vLLM/TGI Single Node if:
- Model fits on 1-8 GPUs (Llama-3 70B, Mistral Large).
- Low/Medium throughput.
- Team lacks Kubernetes / Distributed Systems expertise.
- Ultra-low Latency (TTFT) is the only metric that matters.
Ecosystem & Projects
- LLM-D Project (CNCF Sandbox / Incubation track): The governance body defining APIs.
- vLLM: Primary reference engine implementation (P/D disaggregation support added v0.6+).
- KubeAI / KServe / KubeRay: Higher-level operators managing LLM-D deployments.
- Mooncake (ByteDance): High-perf KV cache transfer library.
- LMCache (MIT/HuggingFace): Distributed KV cache layer with prefix caching.
- XConnector (Alibaba): Unified communication layer for disaggregated inference.
- Dynamo (NVIDIA): Next-gen distributed inference runtime (successor to Triton/TensorRT-LLM concepts).
Getting Started (Conceptual)
# Hypothetical LLM-D Deployment Spec (KubeAI style)
apiVersion: kubeai.org/v1
kind: ModelDeployment
metadata:
name: deepseek-v3
spec:
model: deepseek-ai/DeepSeek-V3
disaggregation:
prefill:
replicas: 4
resources:
nvidia.com/gpu: 8 # H100 nodes
autoscaling:
minReplicas: 2
maxReplicas: 20
decode:
replicas: 16
resources:
nvidia.com/gpu: 4 # A100/H100 nodes
kvCache:
offload: true
storage: nvme
networking:
kvTransfer: mooncake # or lmcache