PD-Disaggregated Inference Deployment with sgl-project/rbg
Use RoleBasedGroup to run SGLang Prefill/Decode disaggregation as one operable unit: routing, startup dependencies, multi-node tensor parallelism, KV transfer, and coordinated rollouts.
Use RoleBasedGroup to run SGLang Prefill/Decode disaggregation as one operable unit: routing, startup dependencies, multi-node tensor parallelism, KV transfer, and coordinated rollouts.
Use better Kubernetes probes by choosing the right signal, tuning thresholds, and avoiding false restarts, traffic drops, and noisy rollouts.
Use a reproducible program, /proc, and strace to inspect how glibc malloc uses brk, mmap, and madvise, and why RSS may stay high after free.
Safely inspect a live Pod without baking debugging tools into production images.
A concrete rollout path for Kubernetes NetworkPolicy: start with default deny, whitelist DNS and key dependencies, and avoid breaking production traffic.
Calculate Serverless versus dedicated GPU break-even using effective GPU hours, cold starts, idle capacity, operations, and SLOs, then plan hybrid capacity.
A curated reading track for Kubernetes.
A curated reading track for Systems.
A curated reading track for GPU.
Use RoleBasedGroup to run SGLang Prefill/Decode disaggregation as one operable unit: routing, startup dependencies, multi-node tensor parallelism, KV transfer, and coordinated rollouts.
Use better Kubernetes probes by choosing the right signal, tuning thresholds, and avoiding false restarts, traffic drops, and noisy rollouts.
Use a reproducible program, /proc, and strace to inspect how glibc malloc uses brk, mmap, and madvise, and why RSS may stay high after free.
Safely inspect a live Pod without baking debugging tools into production images.
A concrete rollout path for Kubernetes NetworkPolicy: start with default deny, whitelist DNS and key dependencies, and avoid breaking production traffic.
Calculate Serverless versus dedicated GPU break-even using effective GPU hours, cold starts, idle capacity, operations, and SLOs, then plan hybrid capacity.
Combine Deployment rollingUpdate settings with PodDisruptionBudgets to keep availability during upgrades and node maintenance.
Deploy SGLang and vLLM Prefill/Decode roles with LWS DisaggregatedSet, NIXL startup commands, validation steps, and production safeguards.
Compare queue backfill, NVIDIA Time-Slicing, MIG, and runtime sharing with Kubernetes configuration, validation, isolation, and rollback steps.
How CPU/memory requests and limits actually affect scheduling, throttling, OOMKills, and autoscaling.
Use a reproducible C program and GDB session to inspect SysV AMD64 arguments, call/ret, stack frames, unwinding, PIE addresses, and optimization.
Learn concrete Kubernetes RBAC least-privilege patterns, how to reduce overbroad permissions, and which checks catch risky role bindings before incidents.
A walkthrough into how Linux dynamically links shared libraries at runtime — PIC, GOT, PLT, lazy binding, gdb tracing, and the real cost of -fPIC.
Build and inspect a reproducible Linux binary to understand Stack Canaries, FORTIFY, NX, Stack Clash protection, ASLR, PIE, RELRO, and CET.
Use a reproducible glibc malloc experiment to separate real leaks, allocator caching, fragmentation, and RSS retention across Arenas, Chunks, Bins, and Tcache.
A September 3, 2026 snapshot of five self-hosted agent frameworks, comparing deployment, extensions, security boundaries, and operational cost.
Compare KAI-Scheduler Reservation Pod accounting with HAMi CUDA API-level memory and compute enforcement, including their boundary versus hardware isolation such as MIG.
Separate container packaging, Kubernetes orchestration, and OpenStack IaaS so teams choose the layer that matches their actual operating problem.
A review of the hetGPU research prototype that separates paper design, conceptual pseudocode, and current cross-vendor GPU engineering capability.
Debug Linux cgroup V1 and V2 with cpu.stat, memory.events, PSI, and Kubernetes resource settings to explain CPU throttling, reclaim, and OOM kills.
Review tkestack/gpu-manager device registration, vCUDA resources, library injection, topology allocation, CUDA limits, and migration choices.
Understand ELF files from sections and segments to relocations and dynamic linking, with concrete examples for debugging Linux binaries and loader issues.
A concrete Kubernetes troubleshooting playbook for Pending Pods, CrashLoopBackOff, readiness failures, networking issues, and node-level problems.
How to make autoscaling predictable: right requests, sane HPA behavior, VPA recommendations, and capacity-aware cluster scaling.
Learn how liveness, readiness, and startup probes work in Kubernetes, what each one should check, and how to avoid restart loops and false failures.
Use Helm to deploy a MySQL cluster on Kubernetes while understanding chart defaults, persistence, networking, and production tradeoffs.
Learn how kubectl port-forward works, when to use it for debugging, and how it differs from Services, Ingress, and production traffic paths.
Understand MySQL source-replica topology on Kubernetes and use MySQL 8.4 commands to verify replication status, routing, and failover.
Understand when to use a headless Service in Kubernetes, how DNS works without a virtual IP, and why it matters for StatefulSets and peer discovery.
Learn when to use a StatefulSet in Kubernetes, how stable Pod identity works, and why ordering and persistent storage matter.