Kubernetes Requests and Limits, Explained by What Breaks
Missing requests get a pod evicted under pressure. Missing limits get it throttled at the worst moment. Each one fails differently, on a different timeline.
Most explanations of Kubernetes resources describe the mechanism accurately and still leave you unable to set a number. Requests and limits are not really about CPU and memory arithmetic — they are about which of two bad outcomes you would rather have when the node runs out.
Once you can name the outcome, the value mostly writes itself.
The one paragraph version
A request is what the scheduler reserves. It is not what the container gets; it is what the node promises it can always have. A limit is the ceiling the kernel enforces. Requests shape placement, limits shape runtime behaviour, and they fail on completely different timelines.
What happens when requests are missing
Set no CPU request and the pod is scheduled as if it needs nothing. On an idle node that is correct and free. Under pressure it becomes a fairness contest, and your pod can lose — throttled to a fraction of what it needs while the node stays nominally "ready."
Set no memory request and it is worse, because memory cannot be borrowed back. The kubelet sees the node as having no room reserved for you, so when memory tightens, your pod is the eviction candidate. It gets killed, and kubectl describe pod reports Evicted with a reason that reads like a capacity problem rather than a configuration one.
Neither shows up in a happy-path load test. Both appear the day the node is busy.
What happens when limits are missing
No CPU limit means the container can burst as high as the node allows — which is often what you want, until one noisy pod makes everybody else's tail latency worse and nobody can point at the culprit.
No memory limit means the cgroup does not stop it. The pod keeps allocating until the kernel's OOM killer picks a victim, and the victim is chosen by memory consumption rather than importance. From the outside this looks like a crash, not a quota breach:
$ kubectl describe pod api-7f9c # last state
State: Terminated
Reason: OOMKilled
Exit Code: 137
An OOMKill is a limit event you never set. Exit code 137 with no application-level error almost always means the kernel made the decision for you.
The combination that actually bites
Requests without limits gives you guaranteed placement and unbounded bursting — fine for batch work, dangerous for anything latency-sensitive. Limits without requests gives you a ceiling nobody reserved, so you can be throttled and evicted.
The pair that holds up under pressure is a request you can justify from measurement and a limit with headroom above it. The usual starting point is to request what you observe at the 95th percentile under normal load and limit at roughly twice that, then adjust from actual metrics rather than from a guess made during onboarding.
Why this surfaces during a rollout, not a deploy
Resource settings almost never fail at the moment you change them. They fail when something else makes the node busy — a batch job in the same namespace, a traffic spike, a second replica landing during a rolling update. That is why a deployment that looked healthy in staging dies in production: staging ran one replica on a quiet node.
The symptom is nearly always one of the same pair. Memory pressure produces an eviction or an OOMKill partway through the rollout, and CPU contention produces a pod that never passes its readiness probe, so the update stalls a fraction of the way across. Both look like application bugs on a dashboard and like capacity problems in kubectl describe pod.
Rollout strategy decides how many of these land at once; requests and limits decide whether the node can absorb the ones that do. And when a pod does die this way, the exit code and last state in pod triage are what separate "the container crashed" from "the kernel removed it."
Summary
Requests decide where you live and whether you get evicted; limits decide whether you burst or get throttled or get OOMKilled. Set both, from measurement rather than instinct, and give the limit real headroom above the request. The failure modes are silent under light load and loud under heavy load, which is exactly backwards from where most teams go looking for them.
SDP Clouds Team
DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.
More about us →