After splitting Prefill and Decode, the hard part is the coupling between them: the Router must wait until Prefill/Decode are ready, scaling must keep both sides balanced, and a rollout must not leave one role stranded. Traditional StatefulSets/Deployments push all of this into external scripts. sgl-project/rbg models “a group of roles = one inference service” as a Kubernetes CRD, giving PD disaggregated deployments a single operable unit.

Snapshot: 2026-09-21, based on RBG v0.7.0 (v1alpha2 API) and SGLang v0.5.9 examples. The project evolves fast; verify against the official docs before production.

What RBG solves

The core abstraction in RBG is the RoleBasedGroup: one CR that declares multiple Roles (router, prefill, decode), each with its own replicas, template, and lifecycle policy. It delivers the three things PD disaggregation needs most:

  • Startup ordering: the dependencies field declares role dependencies, so decode waits for prefill and the router waits for both, preventing the router from sending traffic to uninitialized engines.
  • Service discovery: the controller generates deterministic names like <group>-<role>-<index>.s-<group>-<role>, so the router can point directly at Prefill/Decode instance addresses without maintaining headless Service assembly logic.
  • Unit-level operations: rolling updates, scaling, and failure recovery are coordinated at the RBG level, and CoordinatedPolicy additionally controls maxSkew and rollout progression across roles.

Minimal PD disaggregation on single nodes

The minimal topology has three roles: Router (SGLang Model Gateway), Prefill, and Decode. Using Qwen3-0.6B as an example, the essential configuration:

apiVersion: workloads.x-k8s.io/v1alpha2
kind: RoleBasedGroup
metadata:
  name: sglang-pd-inference
spec:
  roles:
    - name: router
      replicas: 1
      standalonePattern:
        template:
          spec:
            containers:
              - name: router
                image: lmsysorg/sglang-router:v0.2.4
                command:
                  - python3
                  - -m
                  - sglang_router.launch_router
                  - --pd-disaggregation
                  - --prefill
                  - "http://sglang-pd-inference-prefill-0.s-sglang-pd-inference-prefill:8000"
                  - --decode
                  - "http://sglang-pd-inference-decode-0.s-sglang-pd-inference-decode:8000"
                  - --host
                  - "0.0.0.0"
                  - --port
                  - "8000"
                ports:
                  - name: http
                    containerPort: 8000

    - name: prefill
      replicas: 1
      standalonePattern:
        template:
          spec:
            containers:
              - name: prefill
                image: lmsysorg/sglang:v0.5.9
                command:
                  - python3
                  - -m
                  - sglang.launch_server
                  - --model-path
                  - "Qwen/Qwen3-0.6B"
                  - --host
                  - "0.0.0.0"
                  - --port
                  - "8000"
                  - --disaggregation-mode
                  - "prefill"
                  - --tp-size
                  - "1"
                resources:
                  requests:
                    nvidia.com/gpu: "1"

    - name: decode
      replicas: 1
      standalonePattern:
        template:
          spec:
            containers:
              - name: decode
                image: lmsysorg/sglang:v0.5.9
                command:
                  - python3
                  - -m
                  - sglang.launch_server
                  - --model-path
                  - "Qwen/Qwen3-0.6B"
                  - --host
                  - "0.0.0.0"
                  - --port
                  - "8000"
                  - --disaggregation-mode
                  - "decode"
                  - --tp-size
                  - "1"
                resources:
                  requests:
                    nvidia.com/gpu: "1"

The full example is pd-disagg-standalone.yaml, which also includes a Memory emptyDir for /dev/shm (KV transfer needs shared memory for intermediate computation), readiness/liveness probes, and the InPlaceIfPossible rolling update strategy.

Common pitfalls:

  • Shared memory is mandatory. PD disaggregation KV transfer depends on /dev/shm; the example allocates a 30Gi Memory emptyDir. Without it, large-context requests fail on shared memory first, not on GPU memory.
  • Give probes enough initial delay. SGLang needs time to load model weights; the example uses 60s for readiness and 120s for liveness. Scale these to actual load times for larger models, or pods will restart repeatedly during startup.
  • Router addresses are deterministic. The <name>-<role>-<index>.s-<name>-<role> format comes from RBG service discovery; renaming the RBG means updating the router arguments too.

Multi-node tensor parallelism: the LeaderWorker pattern

When the decode side needs a larger TP degree, switch from standalonePattern to leaderWorkerPattern. Each instance is one Leader plus N Workers, and RBG injects environment variables for distributed initialization:

    - name: decode
      replicas: 4
      leaderWorkerPattern:
        size: 2   # 1 leader + 1 worker, tp-size=2
        restartPolicyConfig:
          type: None
        template:
          spec:
            containers:
              - name: sglang
                command:
                  - python3
                  - -m
                  - sglang.launch_server
                  - --disaggregation-mode
                  - "decode"
                  - --tp-size
                  - "2"
                  - --dist-init-addr
                  - $(RBG_LWP_LEADER_ADDRESS):6379
                  - --nnodes
                  - $(RBG_LWP_GROUP_SIZE)
                  - --node-rank
                  - $(RBG_LWP_WORKER_INDEX)
                resources:
                  requests:
                    nvidia.com/gpu: "1"

The pd-disagg-leader-worker.yaml example configures Prefill as 2 instances x TP4 and Decode as 4 instances x TP2, with equal GPU totals on both sides (8 GPUs each). This reflects a basic constraint of PD disaggregation: the compute ratio between the two sides must match the workload profile. Prefill is compute-bound, Decode is memory-bandwidth-bound, and long-output workloads usually need more Decode instances.

RBG_LWP_LEADER_ADDRESS, RBG_LWP_GROUP_SIZE, and RBG_LWP_WORKER_INDEX are topology variables injected by the controller, so there is no need to write initContainers that assemble IPs. This reliability is the core value RBG offers over raw primitives.

KV transfer backend: Mooncake

SGLang’s built-in KV transfer works for basic cases; for high-concurrency cross-node scenarios, use the Mooncake Transfer Engine. RBG ships a ready-made example, and the only difference is two startup arguments:

                  # Prefill side
                  - --disaggregation-mode
                  - prefill
                  - --disaggregation-transfer-backend
                  - mooncake

The Decode side only needs --disaggregation-mode=decode and no transfer backend argument. See sgl-pd-disagg-with-mooncake-te.yaml for the full example.

Note that this example depends on an externally deployed Mooncake service (mooncake-store/standalone-mooncake-store.yaml), not a sidecar. In production, Mooncake Master availability and network bandwidth become new failure domains for PD disaggregation and must be monitored.

Coordinated rollouts and scaling

RBG’s InPlaceIfPossible strategy updates Pods in place when only the image or resource limits change, without recreating the instance, which matters for PD disaggregation because recreating a Decode instance means losing the KV cache it holds.

      rolloutStrategy:
        type: RollingUpdate
        rollingUpdate:
          type: InPlaceIfPossible
          maxUnavailable: 1
          inPlaceUpdateStrategy:
            gracePeriodSeconds: 30

Scaling uses scalingAdapter to automatically create a RoleBasedGroupScalingAdapter CR that plugs directly into HPA. CoordinatedPolicy handles cross-role coordination: it controls maxSkew between Prefill and Decode scaling so one side does not scale first and cause KV transfer backlog.

The community also maintains rbg-planner, an SLA-driven autoscaler with ARIMA-based load prediction that adjusts Prefill/Decode replicas against TTFT/ITL targets, a good fit for teams that do not want to hand-write HPA metric rules.

NVIDIA Dynamo runtime

If your team already orchestrates inference with NVIDIA Dynamo, RBG provides the dynamo/pd-disagg.yaml example: a Processor (Dynamo frontend) handles request routing while Prefill/Decode use the nvcr.io/nvidia/ai-dynamo/sglang-runtime image. Dynamo requires external etcd and NATS for service discovery, which adds two operational components compared with the plain SGLang Router, in exchange for Dynamo ecosystem scheduling and acceleration (such as Model Express P2P weight distribution).

Comparison with LWS DisaggregatedSet

If you have read this site’s LWS-based PD disaggregation guide, you probably want to know how the two compare. Both solve “grouped Pod management”, but at different abstraction levels:

Dimension RBG (RoleBasedGroup) LWS (DisaggregatedSet)
API maturity v1alpha2 (v0.7.0); API still alpha LWS v0.9.0; LeaderWorkerSet API established earlier in the ecosystem
Role modeling Multiple roles declared in one CR, native multi-role topology DisaggregatedSet composes multiple LWS resources; each role is still an independent LeaderWorkerSet underneath
Startup dependencies Native dependencies field for startup order No built-in dependency declaration; relies on Pod readiness order and external orchestration
Service discovery Deterministic names <group>-<role>-<index>.s-<group>-<role> Independent Service per role; revision-aware routing requires pair discovery by label
Coordinated rollout/scaling CoordinatedPolicy controls maxSkew and progression No cross-role coordination primitive; mixed-version rollout risk must be handled manually
In-place updates Native InPlaceIfPossible Primarily Pod recreation; KV cache lost with the Pod
Ecosystem integration Official examples for Mooncake, Dynamo, rbg-planner Mature vLLM/SGLang official deployment paths, complete NIXL transfer examples
Community size Newer, contributors concentrated in the sgl-project org kubernetes-sigs project with a broader community and audit surface

Choose RBG when you need cross-role coordination (ratio-aware scaling, coordinated rollouts), want startup dependencies expressed declaratively, or plan to use rbg-planner for SLA-driven autoscaling. RBG builds these PD-specific operational problems into the API.

Choose LWS when you prioritize stability and community backing, are already on the vLLM/SGLang official LWS deployment path, or use NIXL for KV transfer without needing cross-role coordination primitives. The LeaderWorkerSet API is more mature, and there is more field experience to consult when things break.

A pragmatic decision rule: if your pain is “how do roles coordinate” (startup order, ratios, rollouts), RBG’s abstraction fits better; if the pain is only “how do I manage multi-node TP as a group”, LWS is safer. The two are not mutually exclusive, RBG itself borrows and reuses LWS code.

Pre-launch checklist

  • GPU nodes have CUDA drivers installed and nvidia.com/gpu resources are available
  • /dev/shm Memory emptyDir is configured, sized by model context and concurrency
  • Router Prefill/Decode addresses match the RBG name
  • Probe initial delays match the model load time
  • For multi-node TP, verify inter-node network bandwidth for KV transfer (prefer RDMA)
  • When using Mooncake, deploy the Mooncake service before the RBG and monitor it
  • Confirm the rolling update strategy is InPlaceIfPossible to avoid unnecessary KV cache loss
  • Wire HPA metrics to RoleBasedGroupScalingAdapter, or evaluate rbg-planner

PD disaggregation is not “split one Deployment into two”; it introduces a new problem domain of inter-role coordination. RBG’s value is turning that problem domain into a declarative API: startup ordering, service discovery, coordinated upgrades, and ratio-aware scaling all live in one CR instead of scattered scripts and operational conventions. If your PD disaggregated deployment is still glued together with bash scripts around StatefulSets, RBG is worth a trial.