Kubernetes Probes: Liveness, Readiness, and Startup
Use better Kubernetes probes by choosing the right signal, tuning thresholds, and avoiding false restarts, traffic drops, and noisy rollouts.
10 practical articles on Kubernetes / Operations Tips, covering setup, troubleshooting, and production decisions.
Use better Kubernetes probes by choosing the right signal, tuning thresholds, and avoiding false restarts, traffic drops, and noisy rollouts.
Safely inspect a live Pod without baking debugging tools into production images.
A concrete rollout path for Kubernetes NetworkPolicy: start with default deny, whitelist DNS and key dependencies, and avoid breaking production traffic.
Combine Deployment rollingUpdate settings with PodDisruptionBudgets to keep availability during upgrades and node maintenance.
Deploy SGLang and vLLM Prefill/Decode roles with LWS DisaggregatedSet, NIXL startup commands, validation steps, and production safeguards.
How CPU/memory requests and limits actually affect scheduling, throttling, OOMKills, and autoscaling.
All posts in reverse chronological order.
Use better Kubernetes probes by choosing the right signal, tuning thresholds, and avoiding false restarts, traffic drops, and noisy rollouts.
Safely inspect a live Pod without baking debugging tools into production images.
A concrete rollout path for Kubernetes NetworkPolicy: start with default deny, whitelist DNS and key dependencies, and avoid breaking production traffic.
Combine Deployment rollingUpdate settings with PodDisruptionBudgets to keep availability during upgrades and node maintenance.
Deploy SGLang and vLLM Prefill/Decode roles with LWS DisaggregatedSet, NIXL startup commands, validation steps, and production safeguards.
How CPU/memory requests and limits actually affect scheduling, throttling, OOMKills, and autoscaling.
Learn concrete Kubernetes RBAC least-privilege patterns, how to reduce overbroad permissions, and which checks catch risky role bindings before incidents.
Separate container packaging, Kubernetes orchestration, and OpenStack IaaS so teams choose the layer that matches their actual operating problem.
A concrete Kubernetes troubleshooting playbook for Pending Pods, CrashLoopBackOff, readiness failures, networking issues, and node-level problems.
How to make autoscaling predictable: right requests, sane HPA behavior, VPA recommendations, and capacity-aware cluster scaling.