LWS PD Disaggregation for vLLM and SGLang
Deploy SGLang and vLLM Prefill/Decode roles with LWS DisaggregatedSet, NIXL startup commands, validation steps, and production safeguards.
Prefill-decode disaggregation is not just the same inference container copied into two Pods. A working system has at least three parts: Prefill instances process the input and create KV Cache, Decode instances receive that cache and generate tokens, and a router sends each request through both stages in order. LWS manages grouped Pods, role lifecycle, and topology. vLLM, SGLang, and their routers still own KV transfer and request orchestration.
This guide starts with one Prefill instance, one Decode instance, and one GPU per role. Both frameworks use NIXL so that the lifecycle and validation flow can be compared directly. The result is a verifiable lab deployment, not a production manifest to copy unchanged.
Decide whether PD disaggregation is justified
Good candidates include:
- Long prompts dominate and Prefill queueing visibly increases TTFT.
- New long prompts interrupt active Decode work and make ITL/TPOT unstable.
- Prefill and Decode need different GPU types, parallelism, or replica counts.
- The platform already records TTFT, ITL, queue depth, KV transfer latency, and failures.
Do not start here when:
- A small model with short prompts already meets latency and throughput targets on one GPU.
- Node-to-node connectivity, RDMA/UCX, or NIXL has not been verified.
- The only observable signal is HTTP 200, with no proof of which role handled the request or whether KV transfer succeeded.
The vLLM documentation makes the same boundary clear: disaggregated prefilling primarily lets operators control TTFT and ITL independently. It does not guarantee higher throughput and currently has feature-combination limits. Measure real traffic first; a more complicated topology is not automatically a faster one.
What LWS owns in this architecture
LWS v0.9.0 exposes two relevant APIs:
LeaderWorkerSetmanages one distributed role as a leader/worker Pod group.DisaggregatedSetcomposes Prefill, Decode, and other roles into one inference topology; each role is still backed by an independentLeaderWorkerSet.
That boundary matters. LWS does not transfer KV Cache or interpret OpenAI requests. It keeps grouped instances together during creation, rollout, and recovery, and exposes group information such as LWS_LEADER_ADDRESS, LWS_GROUP_SIZE, and LWS_WORKER_INDEX to the Pods.
Install and run preflight checks
As of September 3, 2026, the latest stable LWS release is v0.9.0. The commands below target that release. Before changing versions, inspect the installed CRD with kubectl explain instead of copying manifests from the main branch documentation.
The current SGLang LWS guide still deploys Prefill and Decode as two independent LeaderWorkerSet resources. This article uses the DisaggregatedSet included in LWS v0.9.0 to compose those roles into one topology. Run kubectl explain disaggregatedset.spec before deployment to confirm that the installed CRD and controller versions match.
helm install lws \
oci://registry.k8s.io/lws/charts/lws \
--version 0.9.0 \
--namespace lws-system \
--create-namespace \
--wait
kubectl api-resources | grep -E 'LeaderWorkerSet|DisaggregatedSet'
kubectl get pods -n lws-system
kubectl get nodes -o custom-columns='NODE:.metadata.name,GPU:.status.allocatable.nvidia\.com/gpu'
The LWS 0.9 chart installs and enables DisaggregatedSet by default; no extra feature flag is required.
For an existing LWS release, do not run helm upgrade alone: Helm does not update CRDs stored in a chart’s crds/ directory. Apply the matching LeaderWorkerSet and DisaggregatedSet CRDs explicitly as documented in the LWS chart README before upgrading the controller.
The example assumes:
- A
pd-demonamespace exists. - A PVC named
model-cachecontains the model at/models/Qwen2.5-7B-Instruct; it supports concurrent read-only mounts across nodes, such asReadOnlyManyorReadWriteMany. WithReadWriteOnce, use node-local model caches or separate PVCs for the two roles. - Prefill and Decode nodes carry
inference-role=prefillandinference-role=decodelabels. - The image contains compatible CUDA, UCX, NIXL, and inference-framework versions.
kubectl create namespace pd-demo
kubectl label node <prefill-node> inference-role=prefill
kubectl label node <decode-node> inference-role=decode
Do not use a floating latest tag in production. Replace <your-registry>/sglang-nixl:<pinned-version> below with a tested image pinned by digest.
Start SGLang Prefill and Decode with DisaggregatedSet
Save this as sglang-pd.yaml:
apiVersion: disaggregatedset.x-k8s.io/v1
kind: DisaggregatedSet
metadata:
name: sglang-pd
namespace: pd-demo
spec:
roles:
- name: prefill
spec:
replicas: 1
leaderWorkerTemplate:
size: 1
restartPolicy: RecreateGroupOnPodRestart
workerTemplate:
metadata:
labels:
app: sglang-pd
spec:
nodeSelector:
inference-role: prefill
hostNetwork: true
dnsPolicy: ClusterFirstWithHostNet
containers:
- name: server
image: <your-registry>/sglang-nixl:<pinned-version>
command: ["bash", "-lc"]
args:
- |
exec python3 -m sglang.launch_server \
--model-path /models/Qwen2.5-7B-Instruct \
--host 0.0.0.0 \
--port 30000 \
--disaggregation-mode prefill \
--disaggregation-transfer-backend nixl \
--disaggregation-bootstrap-port 8998 \
--mem-fraction-static 0.80
env:
- name: UCX_NET_DEVICES
value: all
ports:
- name: http
containerPort: 30000
- name: bootstrap
containerPort: 8998
startupProbe:
tcpSocket:
port: http
failureThreshold: 120
periodSeconds: 10
readinessProbe:
tcpSocket:
port: http
periodSeconds: 10
resources:
limits:
nvidia.com/gpu: "1"
volumeMounts:
- name: model
mountPath: /models
readOnly: true
- name: dshm
mountPath: /dev/shm
volumes:
- name: model
persistentVolumeClaim:
claimName: model-cache
- name: dshm
emptyDir:
medium: Memory
sizeLimit: 8Gi
- name: decode
spec:
replicas: 1
leaderWorkerTemplate:
size: 1
restartPolicy: RecreateGroupOnPodRestart
workerTemplate:
metadata:
labels:
app: sglang-pd
spec:
nodeSelector:
inference-role: decode
hostNetwork: true
dnsPolicy: ClusterFirstWithHostNet
containers:
- name: server
image: <your-registry>/sglang-nixl:<pinned-version>
command: ["bash", "-lc"]
args:
- |
exec python3 -m sglang.launch_server \
--model-path /models/Qwen2.5-7B-Instruct \
--host 0.0.0.0 \
--port 30000 \
--disaggregation-mode decode \
--disaggregation-transfer-backend nixl \
--mem-fraction-static 0.85
env:
- name: UCX_NET_DEVICES
value: all
ports:
- name: http
containerPort: 30000
startupProbe:
tcpSocket:
port: http
failureThreshold: 120
periodSeconds: 10
readinessProbe:
tcpSocket:
port: http
periodSeconds: 10
resources:
limits:
nvidia.com/gpu: "1"
volumeMounts:
- name: model
mountPath: /models
readOnly: true
- name: dshm
mountPath: /dev/shm
volumes:
- name: model
persistentVolumeClaim:
claimName: model-cache
- name: dshm
emptyDir:
medium: Memory
sizeLimit: 8Gi
---
apiVersion: v1
kind: Service
metadata:
name: sglang-prefill
namespace: pd-demo
spec:
selector:
disaggregatedset.x-k8s.io/name: sglang-pd
disaggregatedset.x-k8s.io/role: prefill
leaderworkerset.sigs.k8s.io/worker-index: "0"
ports:
- name: http
port: 30000
targetPort: http
- name: bootstrap
port: 8998
targetPort: bootstrap
---
apiVersion: v1
kind: Service
metadata:
name: sglang-decode
namespace: pd-demo
spec:
selector:
disaggregatedset.x-k8s.io/name: sglang-pd
disaggregatedset.x-k8s.io/role: decode
leaderworkerset.sigs.k8s.io/worker-index: "0"
ports:
- name: http
port: 30000
targetPort: http
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: sglang-pd-router
namespace: pd-demo
spec:
replicas: 1
selector:
matchLabels:
app: sglang-pd-router
template:
metadata:
labels:
app: sglang-pd-router
spec:
containers:
- name: router
image: <your-registry>/sglang-nixl:<pinned-version>
command: ["python3", "-m", "sglang_router.launch_router"]
args:
- --pd-disaggregation
- --prefill
- http://sglang-prefill:30000
- --decode
- http://sglang-decode:30000
- --host
- 0.0.0.0
- --port
- "8000"
ports:
- name: http
containerPort: 8000
readinessProbe:
tcpSocket:
port: http
periodSeconds: 5
---
apiVersion: v1
kind: Service
metadata:
name: sglang-gateway
namespace: pd-demo
spec:
selector:
app: sglang-pd-router
ports:
- name: http
port: 8000
targetPort: http
The static Services are intentional for the first functional test. They select Pods across revisions, so a rolling update can mix incompatible Prefill and Decode versions. Production routing must consume the revision-aware Services generated by DisaggregatedSet, or discover backend pairs using disaggregatedset.x-k8s.io/revision instead of copying these selectors unchanged.
The example uses hostNetwork to simplify NIXL/UCX validation and node labels to keep Prefill and Decode on different nodes. Before adding replicas, enforce Pod anti-affinity or move to an RDMA CNI/dedicated network; otherwise instances placed on the same node will compete for the same port.
Validate resources before load testing
kubectl apply -f sglang-pd.yaml
kubectl get disaggregatedset -n pd-demo
kubectl get leaderworkerset -n pd-demo \
-l disaggregatedset.x-k8s.io/name=sglang-pd
kubectl get service -n pd-demo \
-l disaggregatedset.x-k8s.io/name=sglang-pd \
-L disaggregatedset.x-k8s.io/role,disaggregatedset.x-k8s.io/revision
kubectl get pod -n pd-demo -o wide
kubectl get endpointslice -n pd-demo \
-l kubernetes.io/service-name=sglang-prefill
kubectl get endpointslice -n pd-demo \
-l kubernetes.io/service-name=sglang-decode
kubectl wait --for=condition=Ready pod \
-n pd-demo -l app=sglang-pd --timeout=30m
kubectl rollout status deployment/sglang-pd-router \
-n pd-demo --timeout=5m
A Running Pod does not prove that the chain works. Confirm that both Services have Ready endpoints, Prefill and Decode use the same model and tokenizer, and NIXL/UCX logs contain no connection or memory-registration errors.
Send a request and prove that both roles handled it
kubectl port-forward -n pd-demo svc/sglang-gateway 8000:8000
Send a request with a reasonably long prompt from another terminal:
curl -sS http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "/models/Qwen2.5-7B-Instruct",
"messages": [{
"role": "user",
"content": "Explain in five steps how a Kubernetes controller drives actual state toward desired state."
}],
"max_tokens": 128,
"temperature": 0
}'
Inspect all three log planes:
kubectl logs -n pd-demo -f deployment/sglang-pd-router
kubectl logs -n pd-demo \
-l 'disaggregatedset.x-k8s.io/name=sglang-pd,disaggregatedset.x-k8s.io/role=prefill' \
--prefix --since=5m
kubectl logs -n pd-demo \
-l 'disaggregatedset.x-k8s.io/name=sglang-pd,disaggregatedset.x-k8s.io/role=decode' \
--prefix --since=5m
The acceptance criterion is not merely receiving text. Build the complete evidence chain:
- The router accepts the request and selects a Prefill backend first.
- Prefill processes the prompt and initializes KV transfer.
- Decode receives the request-specific KV information and generates tokens.
- No bootstrap timeout, KV transfer timeout, UCX endpoint error, or memory-registration error appears.
vLLM: reuse the same LWS topology
The Kubernetes topology does not need to change for vLLM. Replace the Prefill and Decode container commands and Service ports. Both roles must pin compatible model, tokenizer, --block-size, vLLM, and NIXL versions.
For cross-node transfer, do not leave the NIXL side channel on its default localhost address. Inject a routable Pod IP into both containers with the Downward API. Because this example uses hostNetwork, the Pod IP is the node IP:
env:
- name: POD_IP
valueFrom:
fieldRef:
fieldPath: status.podIP
- name: UCX_NET_DEVICES
value: all
Prefill:
export VLLM_NIXL_SIDE_CHANNEL_HOST="${POD_IP}"
export VLLM_NIXL_SIDE_CHANNEL_PORT=5600
exec vllm serve /models/Qwen2.5-7B-Instruct \
--host 0.0.0.0 \
--port 8100 \
--block-size 128 \
--max-model-len 8192 \
--gpu-memory-utilization 0.85 \
--enforce-eager \
--kv-transfer-config \
'{"kv_connector":"NixlConnector","kv_role":"kv_producer","kv_load_failure_policy":"fail"}'
Decode:
export VLLM_NIXL_SIDE_CHANNEL_HOST="${POD_IP}"
export VLLM_NIXL_SIDE_CHANNEL_PORT=5600
exec vllm serve /models/Qwen2.5-7B-Instruct \
--host 0.0.0.0 \
--port 8200 \
--block-size 128 \
--max-model-len 8192 \
--gpu-memory-utilization 0.85 \
--enforce-eager \
--kv-transfer-config \
'{"kv_connector":"NixlConnector","kv_role":"kv_consumer","kv_load_failure_policy":"fail"}'
The validation setup sets kv_load_failure_policy=fail explicitly. Decode then fails when remote KV loading fails instead of letting a successful HTTP response hide a transfer problem. --enforce-eager also reduces variables during first validation; reassess CUDA Graph and throughput settings for the target model after correctness is established.
The vLLM repository includes toy_proxy_server.py for verifying the ordered P/D flow. Copy that script into a lab image built from the same vLLM revision and start it with:
python3 /workspace/vllm/tests/v1/kv_connector/nixl_integration/toy_proxy_server.py \
--host 0.0.0.0 \
--port 8000 \
--prefiller-hosts vllm-prefill \
--prefiller-ports 8100 \
--decoder-hosts vllm-decode \
--decoder-ports 8200
Check the number of discovered instances, then send an OpenAI-compatible request:
curl -sS http://127.0.0.1:8000/healthcheck
curl -sS http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "/models/Qwen2.5-7B-Instruct",
"messages": [{"role":"user","content":"Explain the performance difference between Prefill and Decode."}],
"max_tokens": 128,
"temperature": 0
}'
Forward the Decode metrics port separately. Confirm that transferred bytes increase without a matching increase in failed transfers:
kubectl port-forward -n pd-demo svc/vllm-decode 8200:8200
In another terminal:
curl -sS http://127.0.0.1:8200/metrics | \
grep -E 'vllm:nixl_(bytes_transferred|num_failed_transfers|num_kv_expired_reqs)'
Compare the counters before and after at least one request. vllm:nixl_bytes_transferred should increase; vllm:nixl_num_failed_transfers should not. If the image does not expose these metrics, verify the vLLM revision and that NIXL Connector is actually enabled instead of trusting response text alone.
toy_proxy_server.py is a functional test helper, not a production router. It first sends a short generation request to Prefill, reads the returned kv_transfer_params, and passes them to Decode. Production deployments should use the vLLM Production Stack PD routing mode or an equivalent router with health checks, balancing, retries, authentication, and revision-aware discovery.
Map multi-node roles to LWS
When one role needs tensor parallelism across nodes, increase that role’s leaderWorkerTemplate.size. SGLang can consume the group information injected by LWS directly:
python3 -m sglang.launch_server \
--model-path /models/DeepSeek-R1 \
--disaggregation-mode prefill \
--dist-init-addr "${LWS_LEADER_ADDRESS}:5000" \
--nnodes "${LWS_GROUP_SIZE}" \
--node-rank "${LWS_WORKER_INDEX}" \
--tp-size 16
size is the number of Pods in one role group, not the total GPU count. Total TP still depends on GPUs per Pod and framework arguments. For multi-node vLLM, generate an entrypoint for the selected distributed backend; SGLang’s --nnodes flags cannot be copied directly into vllm serve.
Diagnose the common failure modes
| Symptom | Check first |
|---|---|
| DisaggregatedSet exists but Pods are Pending | GPU requests, node selectors, taints/tolerations, PVC binding, and image pulls. |
| Prefill and Decode are Ready but the router returns 5xx | Service endpoints, model name, framework versions, and whether the router preserves KV parameters. |
| Prefill completes but transfer times out | NIXL/UCX device selection, RDMA CNI, ports, firewall rules, and Pod-IP reachability. |
| Decode runs out of memory | KV Cache headroom, maximum context, concurrency, and memory-utilization settings. |
| Failures appear only during rollout | Static Services are mixing Prefill and Decode revisions. |
| HTTP 200 succeeds but PD cannot be proven | Correlate router, Prefill, Decode logs, and KV transfer metrics. |
Minimum production gate
- Pin image digests and record Prefill, Decode, and router versions together.
- Keep model, tokenizer, KV layout, block size, and transfer backend compatible.
- Route Prefill and Decode by matching revision; never mix versions during rollout.
- Scale Prefill from input-token queue and TTFT; scale Decode from active sequences, KV use, and ITL.
- Record TTFT, ITL/TPOT, end-to-end latency, KV transfer latency, timeout rate, and OOMs.
- Drill Prefill Pod, Decode Pod, router Pod, and network failures; measure failure scope and recovery time.
- Add authentication, TLS, rate limits, and request-size limits before exposing the gateway.
The value of PD disaggregation is not the existence of two role names. It is the ability to configure and scale Prefill and Decode around different bottlenecks. LWS makes those distributed roles easier to manage on Kubernetes; stable production still depends on KV transfer, revision-aware routing, and observable evidence.
References
- LWS v0.9.0 release
- LWS v0.9.0 Helm chart and CRD upgrade notes
- LWS DisaggregatedSet concepts
- DisaggregatedSet roles
- LWS labels, annotations, and environment variables
- SGLang PD Disaggregation
- SGLang LWS PD Deployment
- vLLM Disaggregated Prefilling
- vLLM NIXL validation proxy
- vLLM Production Stack PD Disaggregation
- NIXL