PD-Disaggregated Inference Deployment with sgl-project/rbg
Use RoleBasedGroup to run SGLang Prefill/Decode disaggregation as one operable unit: routing, startup dependencies, multi-node tensor parallelism, KV transfer, and coordinated rollouts.
Use RoleBasedGroup to run SGLang Prefill/Decode disaggregation as one operable unit: routing, startup dependencies, multi-node tensor parallelism, KV transfer, and coordinated rollouts.
Use better Kubernetes probes by choosing the right signal, tuning thresholds, and avoiding false restarts, traffic drops, and noisy rollouts.
Use a reproducible program, /proc, and strace to inspect how glibc malloc uses brk, mmap, and madvise, and why RSS may stay high after free.
Safely inspect a live Pod without baking debugging tools into production images.
A concrete rollout path for Kubernetes NetworkPolicy: start with default deny, whitelist DNS and key dependencies, and avoid breaking production traffic.
Calculate Serverless versus dedicated GPU break-even using effective GPU hours, cold starts, idle capacity, operations, and SLOs, then plan hybrid capacity.
A curated reading track for Kubernetes.
A curated reading track for Systems.
A curated reading track for GPU.
Understand how StorageClass enables dynamic provisioning in Kubernetes, how default classes work, and how to choose the right storage policy.
Learn how PersistentVolumes and PersistentVolumeClaims work in Kubernetes, how binding happens, and how to troubleshoot storage lifecycle issues.
Use emptyDir, memory-backed volumes, and generic ephemeral volumes for caches, build spaces, and sidecar data while managing scheduling, eviction, and cleanup.
Learn the core Kubernetes volume types, what data survives Pod restarts, and how to choose between temporary and persistent storage.
Understand when to use ConfigMap or Secret in Kubernetes, how they reach Pods, and which practices reduce config drift and secret exposure.
Learn the essentials of running MySQL on Kubernetes, including StatefulSets, persistent storage, Services, and operational tradeoffs.
A concrete guide to stateful applications on Kubernetes, covering storage choices, stable identities, rollout concerns, and failure handling.
Learn practical canary release patterns in Kubernetes, how to reduce rollout risk, and which signals to watch before promoting traffic.
Understand declarative configuration in Kubernetes, why desired state matters, and how apply, diff, and reconciliation shape safe operations.
Learn how Kubernetes Namespaces organize resources, scope policies and quotas, and support safer multi-team or multi-environment clusters.
Learn how Kubernetes Services provide stable networking for Pods, how service types differ, and how to troubleshoot selectors, endpoints, and traffic flow.
Learn how Deployments and ReplicaSets work together in Kubernetes, how rolling updates happen, and how to debug rollout and selector problems.
Understand what a Pod really is in Kubernetes, how Pods are scheduled and restarted, and which commands help you debug them.
Learn when to use K3s, how to install it quickly, and how to run core Kubernetes workloads on a lightweight cluster for labs and edge environments.
Build a repeatable local Kubernetes cluster with Minikube, then verify networking, image loading, and reset workflows.
A concrete introduction to Kubernetes: what problems it solves, what it does not solve, and how desired state and controllers work in real clusters.
Understand Kubernetes architecture from the control plane to worker nodes, including the API server, scheduler, controllers, kubelet, and reconciliation loops.
A step-by-step Kubernetes learning path covering core concepts, workloads, networking, storage, and troubleshooting so you can study in the right order.
Understand Pods, Deployments, Services, and Namespaces through the way Kubernetes actually reconciles and fails, then verify changes with real commands.