PD-Disaggregated Inference Deployment with sgl-project/rbg
Use RoleBasedGroup to run SGLang Prefill/Decode disaggregation as one operable unit: routing, startup dependencies, multi-node tensor parallelism, KV transfer, and coordinated rollouts.
6 practical articles on GPU, covering setup, troubleshooting, and production decisions.
Use RoleBasedGroup to run SGLang Prefill/Decode disaggregation as one operable unit: routing, startup dependencies, multi-node tensor parallelism, KV transfer, and coordinated rollouts.
Calculate Serverless versus dedicated GPU break-even using effective GPU hours, cold starts, idle capacity, operations, and SLOs, then plan hybrid capacity.
Compare queue backfill, NVIDIA Time-Slicing, MIG, and runtime sharing with Kubernetes configuration, validation, isolation, and rollback steps.
Compare KAI-Scheduler Reservation Pod accounting with HAMi CUDA API-level memory and compute enforcement, including their boundary versus hardware isolation such as MIG.
A review of the hetGPU research prototype that separates paper design, conceptual pseudocode, and current cross-vendor GPU engineering capability.
Review tkestack/gpu-manager device registration, vCUDA resources, library injection, topology allocation, CUDA limits, and migration choices.
All posts in reverse chronological order.
Use RoleBasedGroup to run SGLang Prefill/Decode disaggregation as one operable unit: routing, startup dependencies, multi-node tensor parallelism, KV transfer, and coordinated rollouts.
Calculate Serverless versus dedicated GPU break-even using effective GPU hours, cold starts, idle capacity, operations, and SLOs, then plan hybrid capacity.
Compare queue backfill, NVIDIA Time-Slicing, MIG, and runtime sharing with Kubernetes configuration, validation, isolation, and rollback steps.
Compare KAI-Scheduler Reservation Pod accounting with HAMi CUDA API-level memory and compute enforcement, including their boundary versus hardware isolation such as MIG.
A review of the hetGPU research prototype that separates paper design, conceptual pseudocode, and current cross-vendor GPU engineering capability.
Review tkestack/gpu-manager device registration, vCUDA resources, library injection, topology allocation, CUDA limits, and migration choices.