← Back to Glossary

LLM-D (Large Language Model Distributed)

What is LLM-D?

LLM-D (Large Language Model Distributed) is an open architecture and community initiative focused on disaggregated, distributed LLM inference. It moves beyond single-node serving (like standard vLLM or TGI) to enable running massive models—or high-throughput workloads—across multiple GPUs and multiple nodes using standard cloud-native tooling (Kubernetes, gRPC, etcd).

Core Philosophy: Treat LLM inference as a distributed systems problem, not just a model serving problem.


The Problem: Monolithic Serving Limits

Standard serving engines (vLLM, TGI, TensorRT-LLM) typically run on a single node (1-8 GPUs). This creates hard ceilings:

  1. Model Size Ceiling: Model must fit in VRAM of one node (NVLink domain).
  2. Throughput Ceiling: Limited by single-node memory bandwidth & compute.
  3. Scaling Granularity: Scale up = buy bigger node (expensive). Scale out = replicate whole model (inefficient for large models).
  4. Resource Fragmentation: Hard to pack mixed workloads (prefill vs decode) efficiently.

The LLM-D Solution: Disaggregated Inference

LLM-D splits the inference pipeline into independent, scalable services connected via high-speed networking (NVLink, NVSwitch, InfiniBand, RoCE).

1. Prefill-Decode Disaggregation (The Big Win)

LLM-D Architecture:

Result: 3-10x better GPU utilization. Prefill GPUs stay 100% busy; Decode GPUs serve many concurrent users with low latency.

2. KV Cache Transfer Protocol

The critical technical challenge: Moving the KV Cache (Key-Value tensors) from Prefill workers to Decode workers instantly.

3. Distributed KV Cache (LMCache / Mooncake Integration)


Core Components (The Stack)

Layer Technology Role
Orchestration Kubernetes (K8s) Scheduling, scaling, self-healing, networking
Service Mesh gRPC / Envoy / Cilium Low-latency service-to-service comms
State Coordination etcd / Redis Worker discovery, load metrics, KV cache metadata
Inference Engine vLLM (Modified) Core execution engine (P/D workers)
KV Transport Mooncake / LMCache / XConnector High-speed KV cache transfer
Autoscaling KEDA / Custom Metrics Scale Prefill/Decode pools independently based on queue depth

Key Concepts

PD Disaggregation (Prefill-Decode Separation)

The architectural pattern where the Prefill (prompt processing) and Decode (token generation) stages run on separate GPU pools.

Chunked Prefill / Paged Attention

vLLM’s Paged Attention is a prerequisite. It allows KV cache to be non-contiguous (pages), making it transferable and swappable.

KV Cache as a First-Class Citizen

In LLM-D, the KV Cache is not an implementation detail of the engine—it is a distributed data object with metadata (model version, sequence ID, token positions, topology).

Heterogeneous Clusters

LLM-D enables running Prefill on H100s (high FLOPs) and Decode on A100s/L4s (high VRAM/$), or mixing GPU generations.


Comparison: LLM-D vs Standard Serving

Feature Standard vLLM / TGI (Single Node) LLM-D (Distributed)
Max Model Size Limited by Node VRAM (e.g., 8x H100 = 640GB) Cluster Scale (Petabytes aggregate VRAM)
Scaling Replicate whole model (Data Parallel) Disaggregated Scale (Independent P/D scaling)
GPU Utilization ~40-60% (Decode bound) 80-95% (Workload isolation)
Infrastructure VMs / Bare Metal Kubernetes Native
Ops Complexity Low High (Distributed systems expertise needed)
Latency (TTFT) Low (Local) Slightly Higher (Network hop)
Throughput (TPS) Limited by single node Linear Scale-out
Prefix Caching Local only Global / Distributed

When to Use LLM-D

✅ Use LLM-D if:

❌ Stick to vLLM/TGI Single Node if:


Ecosystem & Projects


Getting Started (Conceptual)

# Hypothetical LLM-D Deployment Spec (KubeAI style)
apiVersion: kubeai.org/v1
kind: ModelDeployment
metadata:
  name: deepseek-v3
spec:
  model: deepseek-ai/DeepSeek-V3
  disaggregation:
    prefill:
      replicas: 4
      resources:
        nvidia.com/gpu: 8  # H100 nodes
      autoscaling:
        minReplicas: 2
        maxReplicas: 20
    decode:
      replicas: 16
      resources:
        nvidia.com/gpu: 4  # A100/H100 nodes
      kvCache:
        offload: true
        storage: nvme
  networking:
    kvTransfer: mooncake  # or lmcache

Key Resources