Synthetic Industry

Troubleshooting guide · updated 2026-10-11

Pods Running but not Ready, and rollouts that never finish: probes and the progress deadline

What liveness, readiness and startup probes each do, why a liveness probe can cause an outage, how a Deployment reports a stalled rollout, and how to roll back safely.

Three probes with three different jobs

A liveness probe decides whether a container should be restarted. A readiness probe decides whether it should receive traffic, and when it fails the pod's address is removed from the endpoints of every matching Service while the container keeps running. A startup probe covers slow starts: until it succeeds, Kubernetes does not run the liveness or readiness probes. Because the three do different things, giving them the same check without thought produces bad results. A readiness failure is a pause; a liveness failure is a restart; a startup failure is a restart after a long wait.

  • Readiness: stop sending traffic, do not restart.
  • Liveness: restart a container that cannot recover.
  • Startup: protect slow starts from liveness restarts.

When a liveness probe causes the outage

Kubernetes warns that liveness probes must indicate an unrecoverable failure of the application itself, and that careless ones cause cascading failures: containers restart under load, requests fail and the remaining pods take more load. The documentation's pattern puts checks of required back-end services in the readiness probe, so the pod leaves rotation instead of being killed. If a pod restarts just as the app is busy or while a database is unreachable, look at the liveness probe first. The fix is usually to make it cheaper, give it a longer timeout and threshold, or add a startup probe for a slow start.

  • Do not check external services in the liveness probe.
  • Use a startup probe for a slow start rather than a long initial delay.
  • A startup probe's budget is its failure threshold times its period.

Why a rollout stalls

A Deployment replaces pods gradually, and by default may take down 25% of the desired pods and add 25% extra during the update. If the new pods never become Ready, the old ones keep serving up to that limit and the rollout waits. The Deployment controller reports a stall only after its progress deadline, 600 seconds by default, by adding a Progressing condition with the reason ProgressDeadlineExceeded. Kubernetes takes no other action; the rollout status command exits with an error and nothing rolls back unless a higher-level tool does it. The documented causes include quota, readiness probe failures, image pull errors, permissions, limit ranges and application misconfiguration.

  • Check the new ReplicaSet's pods, not the old ones.
  • Read the Deployment's conditions with the describe command.
  • A paused Deployment does not hit the deadline.

Rolling back safely

The rollout undo command returns a Deployment to its previous revision, or to a named one, and only changes to the pod template create revisions, so a scale change is not a revision. Rollback depends on retained history, which defaults to ten old ReplicaSets; setting it to zero removes the ability. With Helm, the release has its own history and rollback command, and the Helm 4 documentation for the upgrade command lists a rollback-on-failure option that reverts to the last successful release if an upgrade fails. Know which of the two owns your Deployment and use only that, because mixing a manual change with a Helm release makes the next upgrade surprising.

  • Check revision history before you need it.
  • Use the tool that owns the object: kubectl for plain manifests, Helm for a release.
  • Rehearse the rollback in a test namespace.

Acceptance for the fix

A completed fix shows a finished rollout in a test namespace with all desired replicas Ready, a Service with endpoints that answers a real request, no restarts during a ten-minute observation, a diff limited to the cause found and a recorded rollback. Never remove a probe just to make pods green, and never raise a limit to hide a leak. The fixed job states the cause with the event and log evidence, and says plainly when the cause is in the code or the cluster.

Sources and limits

  • Kubernetes: liveness, readiness and startup probes Checked 2026-10-11.
    • Readiness failure removes the Pod from Service endpoints; liveness failure restarts the container; no liveness or readiness probe runs until a startup probe succeeds.
    • Liveness probes must be configured carefully; incorrect ones can cause cascading failures, and back-end checks suit the readiness probe.
  • Kubernetes: configure probes Checked 2026-10-11.
    • A startup probe budget is failureThreshold times periodSeconds; readiness probes run for the container's whole lifecycle.
  • Kubernetes: Deployments Checked 2026-10-11.
    • progressDeadlineSeconds defaults to 600 and a missed deadline adds a Progressing condition with reason ProgressDeadlineExceeded; Kubernetes only reports it.
    • rollout undo returns to the previous revision; revisionHistoryLimit defaults to 10.
    • maxUnavailable and maxSurge default to 25%.
  • Helm: upgrade Checked 2026-10-11.
    • --rollback-on-failure reverts to the last successful release; --reuse-values and --reset-values differ; rightmost -f file wins.
  • Helm: rollback Checked 2026-10-11.
    • helm rollback returns a release to an earlier revision; revision numbers come from helm history.