Container Healthchecks, Restart Policies, and the PID 1 Problem
restart=always is not a health check. HEALTHCHECK, backoff, and why your container ignores SIGTERM — the three things that decide whether self-healing actually works.
The container was running. The process inside it had been wedged for four days.
docker ps showed it green, the orchestrator showed it healthy, and the application answered nothing. The restart policy never fired, because restarting requires a failure and a process that is alive but stuck isn't one.
Three mechanisms decide whether a container heals itself: a health check, a restart policy, and correct signal handling. Most setups have one of the three and assume it's enough.
The health check: what "healthy" should mean
A HEALTHCHECK asks the container a question on an interval. If the answers stop making sense, it's marked unhealthy.
HEALTHCHECK --interval=30s --timeout=5s --start-period=10s --retries=3 \
CMD curl -fsS http://localhost:8080/healthz || exit 1
interval/timeout— too aggressive and you load a struggling service; too lax and you detect trouble late. 30s/5s is sane.start-period— grace for boot. Without it, a service taking 20 seconds to listen gets marked unhealthy before it ever becomes healthy.retries— consecutive failures before declaring unhealthy. Three is enough; one turns a blip into an outage.
The check should be cheap and local. Hitting a downstream dependency means a database hiccup marks every app instance unhealthy, and you've built a cascade.
curl -fsS http://localhost:8080/healthz # good: is this process answering?
curl -fsS https://payments.example.com/status # bad: is the whole world fine?
The restart policy: what happens after a failure
This is where most people stop, because it's one word in a Compose file.
| Policy | Behaviour |
|---|---|
no | Never restart. Default. |
always | Always, including when the Docker daemon restarts. |
unless-stopped | Always, but not if you manually stopped it. |
on-failure | Only non-zero exit, and optionally capped. |
Here's the catch: a restart policy is not a health check, and a health check does not trigger a restart. Docker marks a container unhealthy; Compose won't restart it for that. The orchestrator will, if you're on Swarm or Kubernetes with the probe wired up. On plain Compose, an unhealthy container just sits there — which is the four-day-stuck container from the top of this post. No non-zero exit, so on-failure never fired; no orchestrator watching, so nobody restarted it.
The reliable pattern is simpler than a supervisor sidecar: make your process exit when it can't do its job. A liveness endpoint that returns failure, and a server that shuts down on repeated failure, self-reports.
The PID 1 problem: why your container ignores SIGTERM
docker stop my-container # ...waits the full 10 seconds, then SIGKILL
Ten seconds of grace, then a hard kill, and any in-flight work is gone. The reason is almost always one of these:
1. Shell form runs a shell as PID 1.
CMD node server.js # shell form — /bin/sh -c is PID 1
CMD ["node", "server.js"] # exec form — node is PID 1
In the shell form the shell receives SIGTERM, and shells don't forward signals they haven't been told about — so the child never learns it's being stopped. Use the exec form. This is the single most common cause of slow, dirty shutdowns.
2. The process doesn't handle SIGTERM at all. Node, Python, Go — none exit gracefully by default. Stop accepting connections, drain, close, exit 0.
3. Zombie processes accumulate. PID 1 reaps orphans, and most applications aren't written to do it, so defunct processes pile up until you hit a PID limit. Docker's --init flag, or tini as the entrypoint, fixes it:
ENTRYPOINT ["tini", "--", "node", "server.js"]
Verify with docker top that PID 1 is your app or tini, not /bin/sh.
The three together
None substitutes for another: health check says is it working, restart policy says what do we do when it exits, signal handling says does it exit cleanly when asked.
Get all three and self-healing is real. Miss one and you get the failure mode at the top — green in the dashboard, dead to users.
On Kubernetes the same three map to livenessProbe/readinessProbe, restart policies are cluster-managed, and the signal requirement is identical. That's covered in pod triage from the other side.
Summary
A restart policy restarts crashed processes; only a health check notices a process that's alive but useless. Add one, make it local and cheap, give boot room with start-period, and — most importantly — use the exec form and handle SIGTERM so shutdown is a drain rather than a SIGKILL. The container that heals itself is the one that can accurately report when it isn't working.
SDP Clouds Team
DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.
More about us →