Three probes with three different jobs
A liveness probe decides whether a container should be restarted. A readiness probe decides whether it should receive traffic, and when it fails the pod's address is removed from the endpoints of every matching Service while the container keeps running. A startup probe covers slow starts: until it succeeds, Kubernetes does not run the liveness or readiness probes. Because the three do different things, giving them the same check without thought produces bad results. A readiness failure is a pause; a liveness failure is a restart; a startup failure is a restart after a long wait.
- Readiness: stop sending traffic, do not restart.
- Liveness: restart a container that cannot recover.
- Startup: protect slow starts from liveness restarts.
When a liveness probe causes the outage
Kubernetes warns that liveness probes must indicate an unrecoverable failure of the application itself, and that careless ones cause cascading failures: containers restart under load, requests fail and the remaining pods take more load. The documentation's pattern puts checks of required back-end services in the readiness probe, so the pod leaves rotation instead of being killed. If a pod restarts just as the app is busy or while a database is unreachable, look at the liveness probe first. The fix is usually to make it cheaper, give it a longer timeout and threshold, or add a startup probe for a slow start.
- Do not check external services in the liveness probe.
- Use a startup probe for a slow start rather than a long initial delay.
- A startup probe's budget is its failure threshold times its period.
Why a rollout stalls
A Deployment replaces pods gradually, and by default may take down 25% of the desired pods and add 25% extra during the update. If the new pods never become Ready, the old ones keep serving up to that limit and the rollout waits. The Deployment controller reports a stall only after its progress deadline, 600 seconds by default, by adding a Progressing condition with the reason ProgressDeadlineExceeded. Kubernetes takes no other action; the rollout status command exits with an error and nothing rolls back unless a higher-level tool does it. The documented causes include quota, readiness probe failures, image pull errors, permissions, limit ranges and application misconfiguration.
- Check the new ReplicaSet's pods, not the old ones.
- Read the Deployment's conditions with the describe command.
- A paused Deployment does not hit the deadline.
Rolling back safely
The rollout undo command returns a Deployment to its previous revision, or to a named one, and only changes to the pod template create revisions, so a scale change is not a revision. Rollback depends on retained history, which defaults to ten old ReplicaSets; setting it to zero removes the ability. With Helm, the release has its own history and rollback command, and the Helm 4 documentation for the upgrade command lists a rollback-on-failure option that reverts to the last successful release if an upgrade fails. Know which of the two owns your Deployment and use only that, because mixing a manual change with a Helm release makes the next upgrade surprising.
- Check revision history before you need it.
- Use the tool that owns the object: kubectl for plain manifests, Helm for a release.
- Rehearse the rollback in a test namespace.
Acceptance for the fix
A completed fix shows a finished rollout in a test namespace with all desired replicas Ready, a Service with endpoints that answers a real request, no restarts during a ten-minute observation, a diff limited to the cause found and a recorded rollback. Never remove a probe just to make pods green, and never raise a limit to hide a leak. The fixed job states the cause with the event and log evidence, and says plainly when the cause is in the code or the cluster.
Sources and limits
- Kubernetes: liveness, readiness and startup probes Checked 2026-10-11.
- Readiness failure removes the Pod from Service endpoints; liveness failure restarts the container; no liveness or readiness probe runs until a startup probe succeeds.
- Liveness probes must be configured carefully; incorrect ones can cause cascading failures, and back-end checks suit the readiness probe.
- Kubernetes: configure probes Checked 2026-10-11.
- A startup probe budget is failureThreshold times periodSeconds; readiness probes run for the container's whole lifecycle.
- Kubernetes: Deployments Checked 2026-10-11.
- progressDeadlineSeconds defaults to 600 and a missed deadline adds a Progressing condition with reason ProgressDeadlineExceeded; Kubernetes only reports it.
- rollout undo returns to the previous revision; revisionHistoryLimit defaults to 10.
- maxUnavailable and maxSurge default to 25%.
- Helm: upgrade Checked 2026-10-11.
- --rollback-on-failure reverts to the last successful release; --reuse-values and --reset-values differ; rightmost -f file wins.
- Helm: rollback Checked 2026-10-11.
- helm rollback returns a release to an earlier revision; revision numbers come from helm history.