Kubernetes Requests and Limits: Scheduling, Throttling, OOM
How CPU/memory requests and limits actually affect scheduling, throttling, OOMKills, and autoscaling.
Resource settings are one of the fastest ways to make a cluster feel “stable” (or mysteriously broken).
Start with three symptoms
When requests and limits are wrong, the incident usually does not announce itself as a resource-settings problem. It shows up as one of these patterns:
- Pods stay Pending: the scheduler cannot fit requested resources, even though node dashboards still show idle CPU.
- P99 jumps while CPU usage looks modest: CPU limits trigger throttling, so the app is not “underused”; it is being held down by CFS quota.
- Pods get OOMKilled or nodes hit memory pressure: memory request/limit values do not match real peaks, and eviction or restart behavior takes over.
I split the investigation into two phases. Scheduling problems start with requests. Runtime problems start with limits, QoS, throttling, and OOMKilled events. Mixing those together usually leads to blaming HPA, node pools, and the application in the wrong order.
Common misreads
kubectl topshows current usage, not what the scheduler reserved.- CPU limits are not free safety rails; they can turn load into latency.
- HPA CPU utilization is usually calculated against requests, so wrong requests produce wrong scaling behavior.
- Node
Capacityis not schedulable capacity.Allocatableis the number that matters.
A baseline that is easy to measure
resources:
requests:
cpu: "250m"
memory: "256Mi"
limits:
memory: "512Mi"
The scheduler places Pods from their final requests. If a resource has a limit but no request, and no admission default has already supplied one, Kubernetes copies that limit into the request. Inspect the Pod stored by the API server instead of assuming an omitted request is always zero.
This baseline gives the scheduler a usable CPU and memory estimate, adds a memory boundary, and avoids introducing CPU throttling before it has been measured.
It is a starting point, not a universal recommendation. Replace the numbers after observing the workload under representative traffic.
Requests, limits, and QoS classes (why it matters)
Kubernetes assigns each Pod a QoS class based on the resources you set. QoS affects eviction priority when a node is under pressure (especially memory pressure).
QoS classes in practice
- Guaranteed: Every container sets CPU+memory requests and limits, and
requests == limitsfor both. - Burstable: You set at least some requests/limits, but the Pod is not Guaranteed.
- BestEffort: No requests/limits set at all.
QoS affects OOM scores and node-pressure decisions, but eviction is not a fixed BestEffort-then-Burstable-then-Guaranteed queue. Kubelet also considers whether usage exceeds requests, Pod priority, and relative usage above requests. Guaranteed Pods are not immune; interpret QoS together with measured usage and PriorityClass.
Quick check:
kubectl get pod -n <ns> <pod> -o jsonpath='{.status.qosClass}{"\n"}'
CPU: scheduling vs throttling (and why limits can hurt latency)
CPU is “compressible”: if there isn’t enough CPU, processes usually slow down rather than crash.
- Requests.cpu affects placement (bin packing): the scheduler uses requests.
- Limits.cpu affects runtime enforcement: CPU quota can throttle the container.
Without a CPU limit, a container can use spare node CPU. During contention, its CPU request also influences relative CPU share. A request is not an exclusive-core guarantee unless the platform uses additional CPU Manager isolation.
This is why a common production baseline is:
- set CPU requests for every container
- avoid CPU limits unless you have a clear reason (hard fairness, strict multi-tenancy policies, or capacity control)
A real-world symptom of too-low CPU limits
Your application is “healthy” but p95 latency spikes during traffic bursts. Metrics show CPU usage “below the limit”, yet response time is bad. Often the missing piece is CPU throttling (CFS quota), which may not be obvious unless you monitor throttling counters.
If you must set CPU limits, consider:
limits.cpu>=requests.cpu(always)- enough headroom for bursty code paths, GC, TLS handshakes, and cold caches
HPA uses requests as the denominator
For CPU utilization targets, HPA usually compares current CPU usage with CPU requests:
- requests too high make utilization look low, so scaling starts late
- requests too low make ordinary traffic look saturated, so scaling becomes noisy
If a relevant container has no CPU request, HPA cannot calculate its utilization and may omit the metric or dampen scaling. A missing request is not an “infinitely sensitive” autoscaling configuration.
Fix the request baseline before tuning HPA stabilization windows or scaling policies.
Memory: protect the node, but design for OOMKills
Memory is not compressible, and Linux enforces memory limits reactively. A container may temporarily exceed its limit before the kernel detects pressure and invokes OOM handling. If the container’s main process is killed, Kubernetes commonly records OOMKilled; whether it restarts still depends on restart policy.
Treat memory limits as:
- a safety boundary for the node and other workloads
- a failure scenario your app must tolerate (restarts, warmups, cache rebuilds)
Confirm the termination reason instead of guessing:
kubectl describe pod -n <ns> <pod> | rg -n "OOMKilled|Killed|Exit Code"
On cgroup v2 nodes, inspect throttling and memory events directly from the container:
kubectl exec -n <ns> <pod> -c <container> -- cat /sys/fs/cgroup/cpu.stat
kubectl exec -n <ns> <pod> -c <container> -- cat /sys/fs/cgroup/memory.events
Growth in nr_throttled or throttled_usec is direct evidence that the CPU limit affects execution. oom and oom_kill in memory.events identify cgroup memory pressure. cgroup v1 uses different files and fields.
Language runtime alignment
If your runtime can “think” it has more memory than the container allows, you’ll get surprises:
- JVM: align
-Xmxwithlimits.memory(leave headroom for native memory) - Node: align
--max-old-space-size - Go: use peak measurements and consider
GOMEMLIMIT, while leaving room for non-heap memory
Node allocatable is not capacity
Capacity is physical capacity. Allocatable is what Kubernetes can offer to Pods after system reservations. Nodes can still hit memory pressure when requests are far below real usage, workloads spike together, or DaemonSet requests understate their actual consumption.
kubectl describe node <node> | rg -n "Capacity|Allocatable|Allocated resources"
kubectl get events -A --sort-by=.lastTimestamp | rg -n "Evicted|MemoryPressure|OOM"
Sidecars and initContainers: the hidden tax
Many production Pods include sidecars:
- service mesh proxies (Envoy)
- log forwarders
- security agents
These containers need their own requests/limits. If you size only the “main” container but ignore sidecars, the Pod’s total resource footprint can be significantly larger than expected.
Traditional init containers run sequentially. For scheduling, the effective Pod request is generally the larger of the sum of regular-container requests and the largest individual init-container request, plus Pod overhead. Native sidecars declared as init containers with restartPolicy: Always follow different accounting rules. Understated migration, warmup, or download peaks can make Pods unschedulable or overload nodes during a large rollout.
Memory-backed emptyDir volumes use tmpfs, and written data counts toward the memory usage of the writing container. Caches, model files, or logs placed in emptyDir.medium: Memory can therefore trigger the memory limit.
Guardrails: LimitRange and ResourceQuota
To prevent “someone forgot requests” from entering production, many teams enforce guardrails:
LimitRange (defaults + min/max)
LimitRange can provide defaults and also enforce min/max per container. This is useful when teams forget to set requests (which leads to BestEffort/Burstable chaos).
apiVersion: v1
kind: LimitRange
metadata:
name: defaults
namespace: app
spec:
limits:
- type: Container
defaultRequest:
cpu: 100m
memory: 128Mi
default:
memory: 512Mi
min:
cpu: 25m
memory: 64Mi
max:
cpu: "2"
memory: 4Gi
ResourceQuota (namespace budget)
ResourceQuota prevents a namespace from consuming unlimited cluster capacity and can enforce request/limit usage.
Newer capabilities: Pod budgets and in-place resize
Starting with Kubernetes 1.34, PodLevelResources is Beta and enabled by default. It allows CPU, memory, and hugepage budgets at Pod scope so containers in the same Pod can share unused capacity. Before adopting it, verify the cluster version, feature gate, monitoring, and chargeback systems all understand the Pod-level fields.
Starting with Kubernetes 1.35, In-place Pod Resize is Stable. Some container CPU and memory requests or limits can change without recreating the Pod. That does not make every resize immediate or unconditional; check .status.resize, .status.containerStatuses[].allocatedResources, and available node capacity during rollout.
How to pick numbers (a concrete sizing loop)
Think in iterations:
- Start simple: set CPU+memory requests; set memory limits; skip CPU limits initially.
- Measure in real traffic conditions:
- peak memory (p95/p99)
- CPU usage and throttling (if limits exist)
- restart/OOMKilled rate
- HPA behavior (scale timing and stability)
- Adjust:
- requests too high => slow scaling + wasted bin packing
- requests too low => noisy neighbors + unstable autoscaling
- memory limit too low => OOMKills; too high => reduced density
- Automate recommendations:
- VPA is excellent for recommendations even if you don’t auto-apply
A quick decision table
| Goal | CPU request | CPU limit | Memory request | Memory limit |
|---|---|---|---|---|
| General web service | required | optional | required | required |
| Strict multi-tenant | required | required | required | required |
| Batch job | required | optional | required | required |
Final rule of thumb
If you do only one thing:
Set requests for CPU+memory, set memory limits, and only set CPU limits when you can measure and accept the throttling trade-offs.
Pre-rollout checklist
- Every application, sidecar, and long-running init container has measured requests.
- Init-container peaks, Pod overhead, and memory-backed
emptyDirusage are included. - Memory limits leave room for runtime overhead and real P95/P99 peaks.
- CPU limits exist only where throttling is understood and acceptable.
- HPA targets use realistic requests rather than placeholder values.
- Node Allocatable and DaemonSet overhead are included in capacity checks.
Three resource-sizing decisions
Q: Should requests equal limits? A: Only if you need strict guarantees. For most services, Burstable (requests < limits) is fine.
Q: Why do Pods get OOMKilled? A: The memory limit is too low for peak usage. Raise limits or reduce memory spikes.
Q: Do limits affect scheduling? A: Scheduling uses final requests. If only a limit is specified, Kubernetes may copy it into the request, so the limit can indirectly change placement.