Synthetic Industry

Platform · updated 2026-10-11

Kubernetes: find out whether the manifest, the image or the cluster is at fault

How to separate a Deployment problem you can fix in a manifest or Helm values from application bugs and cluster faults, and how a one-Deployment repair is accepted.

Three sources of a failing Deployment

A Deployment that will not roll out has one of three broad causes. The manifest or Helm values may be wrong, with a bad command, a missing environment variable, a wrong image tag, a missing pull secret or a probe that is too strict. The application may be wrong, crashing for a reason in its own code. Or the cluster may be wrong, with no capacity, an unhealthy node or a policy that blocks the pod. Only the first is fixed by editing the manifest. The events on the pod and the previous container log usually tell you which.

  • Manifest or values: the first thing to check.
  • Application: the log shows an error from your code.
  • Cluster: the pod is Pending or events name a node or policy.

Where a one-Deployment repair fits

The fixed repair is scoped to one Deployment created from plain manifests or one Helm release, worked on in a test namespace or a copy of the configuration. It reads the events, container states and previous logs, makes the smallest change in the manifest or values, renders it with the tool's own checks and runs the rollout in the test namespace. It never applies to your production cluster, and it does not hold your credentials. A person on your side applies the change.

  • One Deployment, not a fleet.
  • A test namespace, a least-privilege account and a copy of the files.
  • You apply and verify in production.

Evidence of a fix

Acceptance is observable: a completed rollout, all desired replicas Ready, endpoints on the Service, an unchanged restart count over ten minutes, a diff that changes only the cause and a recorded rollback. A green pod list alone is not enough; a loosened probe can make pods look ready while the app is broken, so the test includes a real request through the Service.

  • Rollout status succeeds.
  • The Service answers a real request.
  • Rollback rehearsed.

What it does not promise

The job does not promise capacity, availability or that the application behaves under load. It does not administer nodes, networking, ingress controllers or admission policies. It does not fix application code. If the evidence shows the cause is outside the manifest, the job stops and says so, and you do not pay for the fixed scope.

Sources and limits

  • Kubernetes: pod lifecycle Checked 2026-10-11.
    • Container states, restartPolicy and the CrashLoopBackOff back-off are documented.
  • Kubernetes: Deployments Checked 2026-10-11.
    • The progress deadline, rollout status and rollout undo are documented.
  • Helm: lint Checked 2026-10-11.
    • helm lint reports errors and warnings and --strict turns warnings into failures.
  • Helm: template Checked 2026-10-11.
    • helm template renders locally and does not perform server-side checks.