PD-Disaggregated Inference Deployment with sgl-project/rbg
Use RoleBasedGroup to run SGLang Prefill/Decode disaggregation as one operable unit: routing, startup dependencies, multi-node tensor parallelism, KV transfer, and coordinated rollouts.
After splitting Prefill and Decode, the hard part is the coupling between them: the Router must wait until Prefill/Decode are ready, scaling must keep both sides balanced, and a rollout must not leave one role stranded. Traditional StatefulSets/Deployments push all of this into external scripts. sgl-project/rbg models “a group of roles = one inference service” as a Kubernetes CRD, giving PD disaggregated deployments a single operable unit.
Snapshot: 2026-09-21, based on RBG v0.7.0 (
v1alpha2API) and SGLang v0.5.9 examples. The project evolves fast; verify against the official docs before production.
What RBG solves
The core abstraction in RBG is the RoleBasedGroup: one CR that declares multiple Roles (router, prefill, decode), each with its own replicas, template, and lifecycle policy. It delivers the three things PD disaggregation needs most:
- Startup ordering: the
dependenciesfield declares role dependencies, so decode waits for prefill and the router waits for both, preventing the router from sending traffic to uninitialized engines. - Service discovery: the controller generates deterministic names like
<group>-<role>-<index>.s-<group>-<role>, so the router can point directly at Prefill/Decode instance addresses without maintaining headless Service assembly logic. - Unit-level operations: rolling updates, scaling, and failure recovery are coordinated at the RBG level, and
CoordinatedPolicyadditionally controlsmaxSkewand rollout progression across roles.
Minimal PD disaggregation on single nodes
The minimal topology has three roles: Router (SGLang Model Gateway), Prefill, and Decode. Using Qwen3-0.6B as an example, the essential configuration:
apiVersion: workloads.x-k8s.io/v1alpha2
kind: RoleBasedGroup
metadata:
name: sglang-pd-inference
spec:
roles:
- name: router
replicas: 1
standalonePattern:
template:
spec:
containers:
- name: router
image: lmsysorg/sglang-router:v0.2.4
command:
- python3
- -m
- sglang_router.launch_router
- --pd-disaggregation
- --prefill
- "http://sglang-pd-inference-prefill-0.s-sglang-pd-inference-prefill:8000"
- --decode
- "http://sglang-pd-inference-decode-0.s-sglang-pd-inference-decode:8000"
- --host
- "0.0.0.0"
- --port
- "8000"
ports:
- name: http
containerPort: 8000
- name: prefill
replicas: 1
standalonePattern:
template:
spec:
containers:
- name: prefill
image: lmsysorg/sglang:v0.5.9
command:
- python3
- -m
- sglang.launch_server
- --model-path
- "Qwen/Qwen3-0.6B"
- --host
- "0.0.0.0"
- --port
- "8000"
- --disaggregation-mode
- "prefill"
- --tp-size
- "1"
resources:
requests:
nvidia.com/gpu: "1"
- name: decode
replicas: 1
standalonePattern:
template:
spec:
containers:
- name: decode
image: lmsysorg/sglang:v0.5.9
command:
- python3
- -m
- sglang.launch_server
- --model-path
- "Qwen/Qwen3-0.6B"
- --host
- "0.0.0.0"
- --port
- "8000"
- --disaggregation-mode
- "decode"
- --tp-size
- "1"
resources:
requests:
nvidia.com/gpu: "1"
The full example is pd-disagg-standalone.yaml, which also includes a Memory emptyDir for /dev/shm (KV transfer needs shared memory for intermediate computation), readiness/liveness probes, and the InPlaceIfPossible rolling update strategy.
Common pitfalls:
- Shared memory is mandatory. PD disaggregation KV transfer depends on
/dev/shm; the example allocates a 30Gi Memory emptyDir. Without it, large-context requests fail on shared memory first, not on GPU memory. - Give probes enough initial delay. SGLang needs time to load model weights; the example uses 60s for readiness and 120s for liveness. Scale these to actual load times for larger models, or pods will restart repeatedly during startup.
- Router addresses are deterministic. The
<name>-<role>-<index>.s-<name>-<role>format comes from RBG service discovery; renaming the RBG means updating the router arguments too.
Multi-node tensor parallelism: the LeaderWorker pattern
When the decode side needs a larger TP degree, switch from standalonePattern to leaderWorkerPattern. Each instance is one Leader plus N Workers, and RBG injects environment variables for distributed initialization:
- name: decode
replicas: 4
leaderWorkerPattern:
size: 2 # 1 leader + 1 worker, tp-size=2
restartPolicyConfig:
type: None
template:
spec:
containers:
- name: sglang
command:
- python3
- -m
- sglang.launch_server
- --disaggregation-mode
- "decode"
- --tp-size
- "2"
- --dist-init-addr
- $(RBG_LWP_LEADER_ADDRESS):6379
- --nnodes
- $(RBG_LWP_GROUP_SIZE)
- --node-rank
- $(RBG_LWP_WORKER_INDEX)
resources:
requests:
nvidia.com/gpu: "1"
The pd-disagg-leader-worker.yaml example configures Prefill as 2 instances x TP4 and Decode as 4 instances x TP2, with equal GPU totals on both sides (8 GPUs each). This reflects a basic constraint of PD disaggregation: the compute ratio between the two sides must match the workload profile. Prefill is compute-bound, Decode is memory-bandwidth-bound, and long-output workloads usually need more Decode instances.
RBG_LWP_LEADER_ADDRESS, RBG_LWP_GROUP_SIZE, and RBG_LWP_WORKER_INDEX are topology variables injected by the controller, so there is no need to write initContainers that assemble IPs. This reliability is the core value RBG offers over raw primitives.
KV transfer backend: Mooncake
SGLang’s built-in KV transfer works for basic cases; for high-concurrency cross-node scenarios, use the Mooncake Transfer Engine. RBG ships a ready-made example, and the only difference is two startup arguments:
# Prefill side
- --disaggregation-mode
- prefill
- --disaggregation-transfer-backend
- mooncake
The Decode side only needs --disaggregation-mode=decode and no transfer backend argument. See sgl-pd-disagg-with-mooncake-te.yaml for the full example.
Note that this example depends on an externally deployed Mooncake service (mooncake-store/standalone-mooncake-store.yaml), not a sidecar. In production, Mooncake Master availability and network bandwidth become new failure domains for PD disaggregation and must be monitored.
Coordinated rollouts and scaling
RBG’s InPlaceIfPossible strategy updates Pods in place when only the image or resource limits change, without recreating the instance, which matters for PD disaggregation because recreating a Decode instance means losing the KV cache it holds.
rolloutStrategy:
type: RollingUpdate
rollingUpdate:
type: InPlaceIfPossible
maxUnavailable: 1
inPlaceUpdateStrategy:
gracePeriodSeconds: 30
Scaling uses scalingAdapter to automatically create a RoleBasedGroupScalingAdapter CR that plugs directly into HPA. CoordinatedPolicy handles cross-role coordination: it controls maxSkew between Prefill and Decode scaling so one side does not scale first and cause KV transfer backlog.
The community also maintains rbg-planner, an SLA-driven autoscaler with ARIMA-based load prediction that adjusts Prefill/Decode replicas against TTFT/ITL targets, a good fit for teams that do not want to hand-write HPA metric rules.
NVIDIA Dynamo runtime
If your team already orchestrates inference with NVIDIA Dynamo, RBG provides the dynamo/pd-disagg.yaml example: a Processor (Dynamo frontend) handles request routing while Prefill/Decode use the nvcr.io/nvidia/ai-dynamo/sglang-runtime image. Dynamo requires external etcd and NATS for service discovery, which adds two operational components compared with the plain SGLang Router, in exchange for Dynamo ecosystem scheduling and acceleration (such as Model Express P2P weight distribution).
Comparison with LWS DisaggregatedSet
If you have read this site’s LWS-based PD disaggregation guide, you probably want to know how the two compare. Both solve “grouped Pod management”, but at different abstraction levels:
| Dimension | RBG (RoleBasedGroup) | LWS (DisaggregatedSet) |
|---|---|---|
| API maturity | v1alpha2 (v0.7.0); API still alpha |
LWS v0.9.0; LeaderWorkerSet API established earlier in the ecosystem |
| Role modeling | Multiple roles declared in one CR, native multi-role topology | DisaggregatedSet composes multiple LWS resources; each role is still an independent LeaderWorkerSet underneath |
| Startup dependencies | Native dependencies field for startup order |
No built-in dependency declaration; relies on Pod readiness order and external orchestration |
| Service discovery | Deterministic names <group>-<role>-<index>.s-<group>-<role> |
Independent Service per role; revision-aware routing requires pair discovery by label |
| Coordinated rollout/scaling | CoordinatedPolicy controls maxSkew and progression |
No cross-role coordination primitive; mixed-version rollout risk must be handled manually |
| In-place updates | Native InPlaceIfPossible |
Primarily Pod recreation; KV cache lost with the Pod |
| Ecosystem integration | Official examples for Mooncake, Dynamo, rbg-planner | Mature vLLM/SGLang official deployment paths, complete NIXL transfer examples |
| Community size | Newer, contributors concentrated in the sgl-project org | kubernetes-sigs project with a broader community and audit surface |
Choose RBG when you need cross-role coordination (ratio-aware scaling, coordinated rollouts), want startup dependencies expressed declaratively, or plan to use rbg-planner for SLA-driven autoscaling. RBG builds these PD-specific operational problems into the API.
Choose LWS when you prioritize stability and community backing, are already on the vLLM/SGLang official LWS deployment path, or use NIXL for KV transfer without needing cross-role coordination primitives. The LeaderWorkerSet API is more mature, and there is more field experience to consult when things break.
A pragmatic decision rule: if your pain is “how do roles coordinate” (startup order, ratios, rollouts), RBG’s abstraction fits better; if the pain is only “how do I manage multi-node TP as a group”, LWS is safer. The two are not mutually exclusive, RBG itself borrows and reuses LWS code.
Pre-launch checklist
- GPU nodes have CUDA drivers installed and
nvidia.com/gpuresources are available -
/dev/shmMemory emptyDir is configured, sized by model context and concurrency - Router Prefill/Decode addresses match the RBG name
- Probe initial delays match the model load time
- For multi-node TP, verify inter-node network bandwidth for KV transfer (prefer RDMA)
- When using Mooncake, deploy the Mooncake service before the RBG and monitor it
- Confirm the rolling update strategy is
InPlaceIfPossibleto avoid unnecessary KV cache loss - Wire HPA metrics to
RoleBasedGroupScalingAdapter, or evaluate rbg-planner
PD disaggregation is not “split one Deployment into two”; it introduces a new problem domain of inter-role coordination. RBG’s value is turning that problem domain into a declarative API: startup ordering, service discovery, coordinated upgrades, and ratio-aware scaling all live in one CR instead of scattered scripts and operational conventions. If your PD disaggregated deployment is still glued together with bash scripts around StatefulSets, RBG is worth a trial.