Synthetic Industry

Troubleshooting guide · updated 2026-10-11

Kubernetes pod will not start: CrashLoopBackOff and ImagePullBackOff are different problems

How to read container state, previous logs and events to tell a crashing container from an image that cannot be pulled, and which manifest or values change fits each.

Read the state before the name

The words CrashLoopBackOff and ImagePullBackOff are summaries that kubectl shows. The container state underneath tells you more. Use the pod description to see the state, the reason, the restart count and the events at the bottom. A container that started, ran and exited is Terminated, with an exit code and a reason; one that never started is Waiting. Kubernetes documents CrashLoopBackOff as a display field showing that the back-off delay is in effect, not as a pod phase, so do not confuse the two. That one distinction already sends you to the right half of this guide.

  • Waiting with an image error: the image never started.
  • Terminated with an exit code: the container ran and exited.
  • Restart count rising: the back-off is in effect.

A container that runs and exits

When a container crashes, read the log of the previous instance, because the current one has not run long enough to say anything. The kubectl logs command has a previous flag for that, and a container flag when the pod has several. Kubernetes lists the usual causes: an error in the application, a configuration mistake such as a wrong environment variable or a missing file, too little memory or CPU, and probes that fail. The back-off starts at ten seconds, doubles with each restart and is capped at five minutes, resetting after a container has run for ten minutes without trouble, which is why a pod can look stuck for minutes between attempts.

  • Run the logs command with the previous flag.
  • Check the exit reason for a memory kill.
  • Compare the environment and mounted files with what the program expects.

An image that cannot be pulled

ImagePullBackOff means the container could not start because Kubernetes could not pull the image. The documented reasons include an invalid image name and a private registry without a pull secret. Work through them in order: is the name and tag spelled correctly, was that tag actually pushed, can the node reach the registry, and does the pod reference a pull secret. A pull secret must be in the same namespace as the pod and must be of the docker config type; one created in another namespace is invisible to the pod. Tags can be changed in the registry, while a digest is fixed, and using the latest tag in production makes it hard to know what is running.

  • Name and tag exactly as pushed.
  • Pull secret in the pod's own namespace.
  • Prefer a version tag or digest to latest.

Spec mistakes that are silently ignored

A manifest with a misspelled field can be accepted, with the field ignored, and the pod runs without the setting you thought you gave it. Validate the manifest on apply, and compare what you wrote with what the API server stored. With Helm, render the templates and run the chart linter with its strict option before installing, and remember that rendering locally cannot perform server-side checks such as whether an API version exists. Many CrashLoopBackOff cases are a value that never reached the container because of a typo.

  • Validate on apply and compare with the stored object.
  • Render the chart and lint with strict.
  • Look for a key that is missing from the stored object.

What this fix does not reach

A fix in a manifest or values file cannot cure a cluster that has no capacity, a node that is unhealthy, or a bug in your code. If the log shows an application error, the next step is a bug fix with a regression test. If the pod is Pending and was never scheduled, the cluster administrator needs to read the scheduling events. The fixed Kubernetes job covers one Deployment from plain manifests or one Helm release, works in a test namespace, shows the cause with the event and log evidence, and ends with a completed rollout and a rollback rehearsal.

Sources and limits

  • Kubernetes: pod lifecycle Checked 2026-10-11.
    • Container states are Waiting, Running and Terminated; restartPolicy defaults to Always.
    • Restarts use an exponential back-off of 10s, 20s, 40s, capped at 300 seconds, reset after 10 minutes of running without problems.
    • Listed causes of CrashLoopBackOff include application errors, configuration errors, resource constraints and failing probes; investigation starts with kubectl logs and kubectl describe pod.
  • Kubernetes: images Checked 2026-10-11.
    • ImagePullBackOff means the image could not be pulled, for example for an invalid name or a private registry without an imagePullSecret; the delay is capped at 300 seconds.
    • imagePullSecrets must exist in the same namespace as the pod and be of type dockercfg or dockerconfigjson.
    • Tags are mutable; digests are immutable, and avoiding :latest in production is advised.
  • Kubernetes: debug Pods Checked 2026-10-11.
    • A Waiting pod is most commonly an image pull failure: check the image name, that it was pushed, and that it can be pulled manually.
    • A misspelled field in a spec can be silently ignored; validate the manifest.
  • kubectl logs reference Checked 2026-10-11.
    • --previous prints the log of the previous instance of a container, useful after a crash.