Most outages during “routine deploys” come from a mismatch between:
- how many Pods you allow to go down during rollout
- how many can be evicted during disruptions
- whether you actually have enough replicas and capacity
Start with the failure shape
Rollout safety failures rarely announce themselves as “PDB is wrong.” They usually look like this:
- A rollout gets stuck: new Pods cannot become ready, old Pods cannot be removed, and the Deployment sits in a half-updated state.
- Node maintenance removes too many replicas: the PDB does not match the Pods you meant to protect, or the workload has too few replicas.
- Scaling appears successful but traffic still drops: readiness goes green before the application is actually ready.
- Everything is conservative but deploys are painfully slow:
maxUnavailable: 0is safe, but there is not enough surge capacity.
I separate the controls before changing anything. Deployment settings control rollout replacement. PDB controls voluntary disruption. Readiness controls whether traffic should reach a Pod. Mixing those together usually leads to changing the wrong setting.
Common misreads
- PDB does not protect against involuntary failures such as node crashes or process exits.
- PDB also does not constrain Pod deletions performed by Deployment or StatefulSet controllers; their rollout strategies govern those updates.
maxUnavailable: 0needs spare capacity; otherwise it can turn safety into a stuck rollout.- A PDB with the wrong selector can exist while protecting no real workload Pods.
- If readiness is wrong, rollout settings and PDB cannot prevent early traffic exposure.
Start with explicit availability numbers
For a three-replica API, a conservative starting point is:
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1
maxUnavailable: 0
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: api-pdb
namespace: app
spec:
minAvailable: 2
unhealthyPodEvictionPolicy: AlwaysAllow
selector:
matchLabels:
app: api
The Deployment may create one extra Pod during rollout, while the PDB independently allows one voluntary disruption. The Deployment controller does not consult that PDB when replacing Pods. This setup works only when the cluster has room for the surge Pod and the selector matches the workload you meant to protect.
unhealthyPodEvictionPolicy: AlwaysAllow is Stable starting with Kubernetes 1.31. It lets a drain evict a persistently unready Pod instead of letting that Pod block node maintenance forever; healthy Pods remain budget-protected. Stateful systems that retain the default IfHealthyBudget must accept the corresponding maintenance-blocking risk.
In policy/v1, selector: {} selects every Pod in the namespace, while selector: null selects none. Use explicit labels in production so an empty selector cannot unexpectedly cover unrelated workloads.
Rollout math: make the numbers explicit
When you set maxSurge and maxUnavailable, you’re defining how many Pods can exist and how many can be down during an update.
Example: 8 replicas, maxSurge: 25%, maxUnavailable: 0
maxSurgeallows up to 2 extra Pods (25% of 8 = 2)maxUnavailable: 0means Kubernetes tries not to reduce available Pods below 8
Deployment percentages are rounded differently: maxSurge rounds up, while maxUnavailable rounds down. Convert percentages into actual Pod counts before approving a small-replica rollout.
This only works if:
- the cluster has capacity to schedule the surge Pods
- readiness gates actually represent “safe to receive traffic”
If the cluster can’t schedule the surge (common when requests are high), the rollout stalls.
PDB protects against voluntary disruptions only
This is one of the most misunderstood parts of PDB.
PDB helps with:
kubectl drain- platform node upgrades that cordon+drain nodes
- some automated maintenance workflows
Only workflows that use the Eviction API honor the budget. Node-pressure eviction, direct Pod deletion, and controller-driven rollout deletion do not.
PDB does not protect against:
- node crashes
- kernel OOMs
- container crashes due to bugs
- network partitions or zone outages
So you still need:
- enough replicas
- spreading across nodes/zones
- good health checks and graceful shutdown
PDB percentages round up. With 7 replicas and maxUnavailable: "30%", the budget can allow 3 unavailable Pods, which is more than 30% in practice. Prefer integers for small workloads and validate the result with a drain exercise.
Spread your replicas (or PDB won’t save you)
If all replicas land on the same node, a single node drain breaks availability even with a PDB.
Prefer topologySpreadConstraints (modern approach) or pod anti-affinity.
Example: spread across zones:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: api
This improves both:
- rollout stability (new Pods spread correctly)
- disruption resilience (maintenance doesn’t remove all replicas at once)
Graceful shutdown is part of “availability”
Even with good rollout settings, you can see errors if old Pods terminate abruptly.
Make sure:
- readiness flips to “not ready” quickly on shutdown
terminationGracePeriodSecondsis long enough- optional
preStophook helps drain
This matters for:
- long-lived HTTP keep-alives
- gRPC streams
- background workers processing jobs
Deployment knobs that reduce rollout risk
minReadySeconds
Ensures a Pod stays ready for a minimum time before it’s considered “available”. This reduces flip-flopping readiness during warmup.
progressDeadlineSeconds
Controls when Kubernetes marks a rollout as failed. Helpful for alerting and automation.
revisionHistoryLimit
Keeps old ReplicaSets for rollback. Keep enough history to safely undo.
Interaction with HPA and Cluster Autoscaler
Rollouts often create temporary extra Pods (surge). If the cluster lacks headroom, Cluster Autoscaler may add nodes—but:
- provisioning time can slow rollouts
- quotas and scale-up limits can block surge Pods
If HPA is active and traffic is high, you can also get competing behaviors:
- HPA scales up for load
- rollout creates surge Pods
Practical suggestions:
- roll out during lower-traffic windows when possible
- ensure node pools have buffer or autoscaler is configured well
- ensure requests are realistic (HPA uses requests in utilization calculations)
StatefulSets: similar goals, different mechanics
StatefulSets roll out in order (pod-0, pod-1, …). For stateful systems:
- PDB still helps for voluntary disruptions
- but you must understand whether the app can tolerate sequential restarts
For databases, follow operator guidance and validate replication/leader behavior.
StatefulSet controller updates do not use the Eviction API, so a PDB does not block the StatefulSet’s own rolling deletion either. Use the StatefulSet update strategy, partitioning, and application quorum rules for rollout safety; use PDB for external voluntary disruptions such as drains.
Runbook: when a rollout stalls
Check status, events, then individual Pods:
kubectl rollout status deploy/<name> -n <ns>
kubectl describe deploy/<name> -n <ns>
kubectl get pdb -n <ns>
kubectl describe pdb -n <ns> <pdb>
kubectl get events -n <ns> --sort-by=.lastTimestamp | rg -n "FailedScheduling|Insufficient"
kubectl describe pod -n <ns> <pod>
kubectl logs -n <ns> <pod> -c <container> --previous
If the release cannot continue, rollback is usually more valuable than waiting:
kubectl rollout undo deploy/<name> -n <ns>
Watch the right signals
Do not stop at “rollout status = success”. Watch error rate, p95/p99 latency, CPU throttling, memory pressure, readiness failures, and restart loops. If these regress, pause or roll back early.
How kubectl drain interacts with PDB (what you’ll see)
When you drain a node, Kubernetes will try to evict Pods. For Pods covered by a PDB:
- if evicting would violate the budget, the eviction is blocked
- the drain command may “hang” (it’s waiting for enough Pods to be available elsewhere)
This is not a bug—it’s the budget doing its job. But it means you should:
- test node drains in staging
- ensure your workloads can reschedule quickly (requests, node selectors, tolerations)
- avoid over-constraining placement (too strict affinity rules)
Do not use kubectl drain --disable-eviction to bypass a blocked PDB unless the availability impact is explicitly accepted. That flag switches to direct deletion, so the PDB no longer applies.
Canary and blue/green: when rollingUpdate isn’t enough
Deployments and PDBs are foundational, but some changes are high risk:
- schema migrations
- dependency upgrades
- major config changes
For these, consider progressive delivery:
- canary (shift 1%, then 10%, then 50%)
- blue/green (swap traffic between two stable environments)
Even if you don’t have a full progressive delivery platform, you can approximate canary by:
- creating a second Deployment with a small replica count
- routing a small portion of traffic (ingress rules, header-based routing)
The key idea is to limit blast radius while you validate the new version.
Practical baselines
Three-replica stateless API
- Deployment:
maxUnavailable: 0,maxSurge: 1 - PDB:
minAvailable: 2 - spread replicas across nodes or zones
Service with ten or more replicas
- Deployment: start with
maxUnavailable: 10%,maxSurge: 10% - PDB: use a percentage budget aligned with maintenance capacity
- stop or roll back automatically when error or latency budgets regress
Final acceptance criteria
Rollout and PDB settings work only when replica count, failure-domain spreading, readiness, graceful shutdown, and surge capacity all agree. Validate those conditions with a node-drain drill before treating the policy as production-ready.
Three rollout decisions
Q: What does a PDB protect against? A: Voluntary disruptions (drains, upgrades), not node crashes.
Q: Why is my rollout stuck?
A: A PDB does not block Deployment controller rollouts. Check Pending surge Pods, readiness, image pulls, and progressDeadlineSeconds; adjust PDB only when an eviction or drain is blocked.
Q: How should I set surge/unavailable?
A: Set maxSurge and maxUnavailable from the minimum Ready Pod count required during rollout. Calculate PDB separately from the voluntary disruptions allowed during node maintenance.