Learn / Kubernetes survival kit / Debugging CrashLoopBackOff, OOMKilled, and Pending pods
Debugging CrashLoopBackOff, OOMKilled, and Pending pods
The three failure states you'll actually hit, what each one means, and the exact commands to diagnose them.
The one command that answers all three
Before anything else, for any failing Pod:
kubectl describe pod <pod-name>
Read the Events section at the bottom first. Kubernetes records the actual reason there in plain language - “Insufficient cpu,” “OOMKilled,” “Back-off restarting failed container” - far more directly than guessing from the Pod’s status alone. This single command is the starting point for every failure below; the rest is about interpreting what it tells you.
CrashLoopBackOff: the container keeps exiting
The container starts, then exits (crashes, or exits non-zero on purpose), and Kubernetes retries with increasing backoff delay. The backoff itself is not the problem - it’s Kubernetes correctly detecting a container that can’t stay running. The actual cause is almost always inside the application:
kubectl logs <pod-name> # current container's logs, if any were written before it died
kubectl logs <pod-name> --previous # the PREVIOUS crashed container's logs - usually the one you actually need
kubectl describe pod <pod-name> # confirm the exit code and reason in Events / Last State
Common real causes: a missing required environment variable or config file, a database the app
can’t reach on startup (and doesn’t retry), a bad command/entrypoint in the image, or an
unhandled exception at startup. --previous is the flag people forget - once a container has
already restarted, plain kubectl logs shows the new (possibly still-starting) container,
not the one that actually crashed.
OOMKilled: hit the memory limit
If a container exceeds its configured memory limit, the kernel’s OOM killer terminates it -
this shows up as OOMKilled in kubectl describe pod’s last-state reason. Two genuinely
different root causes produce the identical symptom:
- The limit is legitimately too low for what the application actually needs under real load (not the light load it was tested with).
- A real memory leak - usage climbs over time and would eventually exceed any reasonable limit.
You can’t tell which from the OOMKilled event alone - you have to look at actual memory usage over time:
kubectl top pod <pod-name> # current snapshot (needs metrics-server installed)
# or a Grafana/CloudWatch/etc. dashboard tracking the container's memory over its recent lifetime
A flat usage line that simply exceeds the limit points at raising resources.limits.memory. A
steadily climbing line, especially one that resets right after each restart and climbs the
same way again, points at a leak in the application - raising the limit there only delays the
next OOMKill, it doesn’t fix it.
Pending: never got scheduled
A Pod stuck in Pending hasn’t been assigned to a node at all - so kubectl logs has nothing
to show (there’s no running container yet). Go straight to kubectl describe pod’s Events,
which names the scheduler’s actual blocker:
- “Insufficient cpu” / “Insufficient memory” - no node currently has enough free capacity to fit the Pod’s requested resources. Fix: scale the cluster, lower the request, or free up capacity elsewhere.
- An unsatisfied
nodeSelectoror affinity rule - the Pod requires a node with a label (e.g. a specific instance type or zone) that no current node has. - An unbound
PersistentVolumeClaim- the Pod needs a volume that hasn’t been provisioned yet; check the PVC’s own status separately (kubectl describe pvc <name>).
See the kubectl cheatsheet for the full set of get/describe/logs
commands referenced across this lesson.
Key takeaways
- CrashLoopBackOff means the container keeps starting and exiting - the fix is in the application/config, not the cluster; kubectl logs (with --previous) and kubectl describe pod are the first two commands, always.
- OOMKilled means the container hit its memory limit and was killed by the kernel - the fix is either raising the memory limit or fixing an actual leak, and you can only tell which from watching real usage over time.
- A Pod stuck Pending usually means the scheduler can't place it - check kubectl describe pod's Events section first, which names the exact reason (insufficient resources, an unsatisfied node selector/affinity, an unbound PersistentVolumeClaim).
- kubectl describe pod's Events section is the single most useful piece of information for all three failure types - read it before anything else.
Quick check
3 questions - see how much stuck.