SDP Clouds
Kubernetes·4 min read

CrashLoopBackOff and Friends: Triage a Failing Pod in Five Minutes

Pending, ImagePullBackOff, OOMKilled, CrashLoopBackOff — four states, four meanings, and the kubectl commands that tell you which one you actually have.


A pod that isn't running gives you a status word and very little else. CrashLoopBackOff is the famous one, but it is only one of four states, and picking the wrong one sends you reading application logs when the real problem is a subnet.

Here's the triage order I use, in the order the cluster reports it.

First: what does the scheduler say?

bash
kubectl describe pod api-7f9c | sed -n '/Events/,$p'

Read the events before anything else. They are the cluster telling you what it decided, in plain language.

Pending means the scheduler hasn't placed it. Almost always one of three things: not enough CPU or memory available against the requested values, no node matches the node selector or taint, or the PVC never attached. The events will say which — Insufficient cpu, didn't match Pod's node affinity/selector, or a volume that's stuck in ContainerCreating.

A Pending pod is not an application problem. Nothing has run yet.

Second: did it pull the image?

ImagePullBackOff and ErrImagePull are the same failure at different back-off intervals. The cluster couldn't fetch the image: wrong tag, private registry with no imagePullSecret, or a mirror that doesn't exist in this region.

bash
kubectl describe pod api-7f9c | grep -A2 Failed

Failed to pull image "...": not found versus unauthorized tells you whether to fix the tag or the credentials. Retrying will not help either — the back-off will just get longer.

Third: is it dying, or is it being killed?

This is the fork that wastes the most time. Two very different failures produce the same CrashLoopBackOff status.

The container exits on its own — a missing env var, a config file it won't parse, a migration that fails on boot, a port already bound inside the container.

The container is killed — OOMKilled means the process exceeded its memory limit and the kernel ended it. The exit code tells you: 137 is SIGKILL (usually the OOM killer), 1 is the process deciding to fail.

bash
kubectl get pod api-7f9c -o jsonpath='{.status.containerStatuses[*].lastState}'
kubectl logs api-7f9c --previous

--previous is the one people forget. Without it you get the logs of a container that already restarted and therefore explains nothing.

The order of operations

  1. Describe → events. Establishes Pending vs. everything else.
  2. Get lastState. Distinguishes application exit from OOM kill.
  3. Logs --previous. What did it say before it died?
  4. If both are empty, check the readiness/liveness probes — a probe with the wrong path or an initial delay shorter than your app's boot time will kill a perfectly healthy process forever.
  5. If it started fine and later broke, look at limits. A Java app with a 512Mi limit and a default heap will OOMKill roughly on schedule.

Where the limits bite

Requests and limits are the usual culprit for the slow-burn version: the pod runs, gets busy, gets killed, restarts, repeats. A memory limit below what the process needs in production is a restart timer, not a safeguard.

The other common one is a liveness probe pointing at a downstream dependency. If your health endpoint checks the database, a database blip becomes a full pod restart — which is exactly the cascade the security baseline warns about when probes and policies are set without thought.

Requests cause a different, quieter failure. Ask for more CPU than any node can give and the pod simply never schedules — no error in the application, no logs, just a Pending that has been sitting there since the deploy. kubectl describe will name the resource it ran out of, and the fix is either lowering the request or adding a node. Guessing here is how teams end up with four-node clusters where each pod has reserved enough CPU for eight.

The five-minute checklist

  • Pending → scheduler: resources, selectors, volumes.
  • ImagePullBackOff → tag or credentials; back-off means retrying won't fix it.
  • CrashLoopBackOff → lastState exit code, then logs --previous.
  • OOMKilled → raise the limit or fix the heap; it's config, not code.
  • All quiet, restarts anyway → probes, almost always the probe.

Summary

Status, then events, then exit code, then previous logs. Four states, four different owners — the scheduler, the registry, your process, or your limits. Diagnose in that order and most "Kubernetes is broken" tickets resolve in under five minutes without a single YAML edit.

#kubernetes#debugging#troubleshooting#devops

SDP Clouds Team

DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.

More about us →

Related articles