Intro
When a container crashes in production at 3 a.m., the difference between a minor blip and a full outage often comes down to one setting: the restart policy. Docker restart policies determine whether a container automatically starts again after it exits, and under what conditions. Get them wrong, and you either have services that stay down silently or containers that restart endlessly, masking deeper problems.
This article is a practical operations checklist for using Docker restart policies in production. It is written for developers, DevOps consultants, and startup teams who need to move from an observed problem to a verified result. You will find concrete commands, expected outputs, configuration snippets, and recovery guidance. The goal is operational safety: observe before changing, limit the blast radius, use placeholders instead of secrets, verify the result, and document how to recover if the expected state is not reached.
We will cover the four restart policies (no, on-failure, unless-stopped, always), show how to set them via Docker CLI and Docker Compose, explain how to inspect and test them, highlight common mistakes, and provide a checklist you can use in production.
Understanding Docker Restart Policies
Docker restart policies control whether a container is automatically restarted after it stops or exits. The policy is attached to the container at creation time and applies to both normal exits and crashes, depending on the policy.
The four built-in policies are:
| Policy | Behavior | Typical Use Case |
|---|---|---|
no | Never restart automatically. | One-off tasks, debugging containers. |
on-failure[:max-retries] | Restart only if the container exits with a non-zero code. Optionally limit the number of retries. | Services that may crash but should not retry forever. |
always | Always restart, even if the container stops cleanly. Also restarts on daemon startup. | Long-running services that must always be up. |
unless-stopped | Always restart, unless the container was manually stopped by an operator. Restarts on daemon startup if it was running before. | Services that should stay up unless deliberately stopped. |
Understanding these policies is the foundation. The next sections give you a systematic way to inspect, set, and verify them in your environment.
Version and Environment Inventory
Before touching any restart policy, you need to know what you are working with. A version and environment inventory is the first step in any safe operation. It names the relevant component, the supported version range, prerequisites, and a read-only observation.
Check Docker version. Restart policies have existed for years, but syntax and behavior can vary slightly. Run:
docker version --format '{{.Server.Version}}'
Expected output example:
24.0.5
Also check the API version if you use a client library:
docker version --format '{{.Server.APIVersion}}'
List running containers and their status. Use a formatted output to see names, status, and ports:
docker ps --format "table {{.Names}}\t{{.Status}}\t{{.Ports}}"
Example output:
NAMES STATUS PORTS
web Up 3 hours 0.0.0.0:80->80/tcp
api Restarting (1) 2 seconds ago 0.0.0.0:8080->8080/tcp
Notice the api container shows Restarting (1) 2 seconds ago. That indicates it has crashed once and is being restarted by a policy. This is a key observation.
Inspect a container's restart policy. Use docker inspect to see the current policy and other details:
docker inspect <container-name> --format '{{.HostConfig.RestartPolicy.Name}} {{.HostConfig.RestartPolicy.MaximumRetryCount}}'
For a container with on-failure:5, the output would be:
on-failure 5
Check logs for crash patterns. Read recent logs to understand why a container is restarting:
docker logs <container-name> --tail 100
For Docker Compose projects, use docker compose ps and docker compose logs similarly. Compose services may have restart policies defined in the compose file, which override defaults.
Data persistence check. When changing restart policies or recreating containers, ensure data persists. Confirm where files are stored:
docker inspect <container-name> --format '{{json .Mounts}}'
Look for Type: volume (named volume) or Type: bind (host path). A named volume like app_data:/var/lib/app is managed by Docker and survives container recreation. A bind mount like ./data:/var/lib/app maps a host directory directly, which is useful for development but can cause permission or portability issues if the path differs across machines.
Practical test: Stop the container, recreate it without changing the image, and verify the application still sees expected files. If data disappears, the service was writing to the container's writable layer instead of a volume. That is a separate issue to fix before relying on restart policies.
With this inventory, you know your versions, current policies, and data layout. Now you can safely configure policies.
Safe Configuration Path
Changing a restart policy is usually low-risk, but doing it carelessly can cause unexpected behavior. The safe configuration path involves choosing the right policy, applying it idempotently, and verifying the change.
Choosing the right policy
Ask two questions:
- Should the container ever restart on its own?
- Should it restart even after a clean stop?
Use this decision guide:
| Situation | Recommended Policy |
|---|---|
| One-off batch job or interactive shell | no |
| Cron-like container with internal retry | no or on-failure |
| Web server or API that must always run | always or unless-stopped |
| Service that should not restart after a deliberate stop (e.g., maintenance) | unless-stopped |
| Service that should retry only on crash, with a limit to avoid crash loops | on-failure:5 |
Applying a policy via Docker CLI
For an existing container, you cannot change the restart policy directly. You must recreate the container with the new policy. Use docker update for some resource limits, but restart policy is not one of them.
Example: recreate a container named web with unless-stopped:
# Stop and remove the existing container (careful: ensure data is on a volume)
docker rm -f web
# Run a new container with the policy
docker run -d --name web --restart unless-stopped -p 80:80 nginx:alpine
Verify:
docker inspect web --format '{{.HostConfig.RestartPolicy.Name}}'
Output:
unless-stopped
Applying a policy via Docker Compose
In a docker-compose.yml file, set the restart key under a service:
services:
web:
image: nginx:alpine
restart: unless-stopped
ports:
- "80:80"
Apply with:
docker compose up -d
Compose will recreate containers that need the new policy. To force recreation:
docker compose up -d --force-recreate
Testing the policy locally
Do not deploy an untested policy. Simulate a crash and observe behavior.
For on-failure, create a container that exits with error after a few seconds:
docker run -d --name test-fail --restart on-failure:2 alpine sh -c "sleep 5; exit 1"
Watch the status with:
docker ps -a --filter name=test-fail
After about 15 seconds, you should see:
CONTAINER ID STATUS NAMES
abc123 Exited (1) 2 minutes ago test-fail
The container attempted restart twice (because of :2) and then stopped permanently. Check logs to confirm:
docker logs test-fail
For always, create a container that exits immediately:
docker run -d --name test-always --restart always alpine sh -c "exit 0"
Within a few seconds, docker ps -a will show it in a restart loop:
CONTAINER ID STATUS NAMES
abc456 Restarting (0) 1 second ago test-always
Stop it manually:
docker stop test-always
Then check: it should not be restarted because the daemon sees it as manually stopped. But if you restart the Docker daemon, always will restart it, while unless-stopped will not (because it was stopped manually before daemon restart).
Test unless-stopped by stopping a container, restarting Docker daemon, and confirming the container remains stopped.
These tests give you confidence in the policy behavior before production.
Verification and Diagnostics
After setting a policy, you need to verify it works as expected. This section gives a systematic way to observe, diagnose, and confirm restart behavior.
Observing current restart state
Use docker ps to see if any container is currently restarting:
docker ps --filter "status=restarting"
Also check how many restarts a container has had:
docker inspect <container-name> --format '{{.RestartCount}}'
For example, if a container has restarted 12 times in an hour, something is wrong. Compare with logs:
docker logs --tail 50 <container-name>
Diagnosing crash loops
A container in a restart loop (e.g., Restarting (1) 5 seconds ago) indicates the policy is doing its job but the app is failing repeatedly. To break the loop temporarily and investigate:
- Stop the container (this overrides
alwaysandunless-stoppeduntil daemon restart):
docker stop <container-name>
- Inspect the container's filesystem or run a one-off command to debug. For example, run an interactive shell in the same image:
docker run -it --rm --entrypoint sh <image>
- Check logs from the stopped container without restarting:
docker logs <container-name>
- Once fixed, start the container again:
docker start <container-name>
Verifying policy persistence across daemon restarts
Restart policies interact with the Docker daemon. After a daemon restart, containers with always are started regardless of their state before daemon stop. Containers with unless-stopped are started only if they were running before daemon stop.
To test this in a controlled way:
- Create two containers: one with
always, one withunless-stopped. - Manually stop both.
- Restart Docker daemon (
sudo systemctl restart dockeron Linux). - Check
docker ps -a:
alwayscontainer will be running.unless-stoppedcontainer will remain exited.
Expected output after daemon restart:
CONTAINER ID STATUS NAMES
def789 Up 10 seconds test-always
ghi012 Exited (0) test-unless-stopped
This is an important operational distinction.
Monitoring restart events
For production, you should monitor restart counts and events. Docker emits events that you can watch:
docker events --filter container=<container-name> --filter event=die
This streams events like:
2023-09-01T10:00:00.123456789Z container die abc123 (exitCode=1, image=nginx:alpine, name=web)
You can also use docker inspect --format in a cron or monitoring script to alert when restart count increases rapidly.
Failure Modes and Recovery
Even with correct policies, failures happen. This section covers common failure modes with restart policies and how to recover.
Crash loop due to application bug. The policy restarts the container, but it keeps crashing. This is not solved by the policy alone; you must fix the app. Recovery: stop the container to break the loop, debug with logs and a shell, fix the code or configuration, and redeploy. Use on-failure with a retry limit to avoid endless loops that consume CPU and hide issues.
Data loss caused by recreating container. If you recreate a container to change its restart policy without proper volumes, you may lose data written inside the container. Recovery: implement a volume or bind mount before recreating, and test data persistence as described earlier. If data is already lost, restore from backup if available.
Restart policy too aggressive. A always policy on a container that exits immediately can cause a tight restart loop, filling logs and consuming resources. Recovery: stop the container, change the command to keep it running or fix the exit reason, then start again. Consider on-failure with a retry limit.
Restart policy not applied as expected. You think you changed the policy, but the container still shows the old one. This often happens because you used docker update (which does not support restart policy) or edited a compose file but did not recreate the container. Recovery: check the actual policy with docker inspect, then recreate the container with the correct policy.
Daemon restart causing unexpected container starts. If you have containers with always, a planned daemon restart will start them even if they were intentionally stopped. This can surprise operators. Recovery: use unless-stopped for containers that should respect manual stops.
Restart policy interferes with orchestration. In Docker Swarm or Kubernetes, the container restart policy is managed by the orchestrator; setting restart in the container spec may conflict. Recovery: use the orchestrator's restart mechanism (e.g., Swarm's restart_policy under service spec, Kubernetes pod restartPolicy). Do not set container-level policies for orchestrated workloads.
Each failure mode should be documented in your runbook with specific symptoms and recovery steps.
Operations Checklist
Here is a production-focused checklist for managing Docker restart policies. Each item includes the responsible role and review frequency.
1. Inventory and Audit (Owner: Platform Engineer, Reviewed Quarterly)
- [ ] List all running containers and their restart policies:
docker ps --format "table {{.Names}}\t{{.Status}}\t{{.HostConfig.RestartPolicy.Name}}"
Note: {{.HostConfig.RestartPolicy.Name}} may not be directly available in docker ps; use docker inspect in a loop:
for c in $(docker ps -q); do echo -n "$c "; docker inspect $c --format '{{.Name}} {{.HostConfig.RestartPolicy.Name}}'; done
- [ ] Verify that each critical service has an appropriate policy (e.g.,
unless-stoppedfor long-running,on-failure:5for crash-prone but recoverable). - [ ] Check restart counts for abnormal values:
docker inspect <container> --format '{{.RestartCount}}'
If count > 10 in 24 hours, investigate.
2. Policy Selection (Owner: Application Owner, Reviewed with Every Deployment)
- [ ] For each service, decide between
no,on-failure,always,unless-stoppedbased on expected lifecycle. - [ ] Document the rationale in the service runbook.
- [ ] For
on-failure, set a reasonable retry limit to prevent endless loops.
3. Implementation (Owner: DevOps Engineer, Executed per Change)
- [ ] For Docker CLI, recreate container with
--restart <policy>. - [ ] For Compose, add
restart:to service definition and rundocker compose up -d. - [ ] Use
docker inspectto confirm the policy is applied.
4. Testing (Owner: QA/Release Manager, Executed per Release)
- [ ] Perform a local restart test: create a container with the policy, simulate failure or manual stop, observe behavior.
- [ ] Test daemon restart interaction for
alwaysandunless-stopped. - [ ] Verify data persistence across container recreation.
5. Monitoring (Owner: SRE/On-call, Continuous)
- [ ] Set up alerts for rapid restart count increase (e.g., via
docker eventsstream or monitoring tool). - [ ] Log all restart events for post-incident analysis.
- [ ] Include restart policy checks in regular health checks.
6. Recovery Runbook (Owner: Incident Commander, Reviewed Monthly)
- [ ] Document steps to stop a crash loop.
- [ ] Document steps to recover from accidental data loss.
- [ ] Document how to change a policy on a running service without causing downtime.
This checklist should be integrated into your existing operational workflows.
Common Pitfalls and How to Avoid Them
Even experienced Docker users make mistakes with restart policies. Here are the most common ones and how to avoid them.
Pitfall 1: Using always for short-lived jobs. A batch job that exits successfully will be restarted indefinitely, consuming resources. Use no or on-failure for such containers. If you need a periodic job, use an external scheduler (e.g., cron on host) to run a new container each time.
Pitfall 2: Forgetting that always overrides manual stops on daemon restart. You stop a container for maintenance, then restart Docker, and it comes back. Use unless-stopped if you want manual stops to persist across daemon restarts.
Pitfall 3: Not setting a retry limit for on-failure. Without a limit, a container can restart forever, hiding the root cause. Set on-failure:5 or similar to give up after a few attempts, forcing investigation.
Pitfall 4: Trying to change restart policy with docker update. docker update can change resource limits but not restart policy. You must recreate the container. Always verify with docker inspect.
Pitfall 5: Ignoring restart count spikes. A high restart count is a symptom of an unstable service. Monitor it and alert on abnormal increases.
Pitfall 6: Applying container-level restart policies in orchestrated environments. In Docker Swarm or Kubernetes, the orchestrator manages restarts. Setting container-level policies can conflict. Use the orchestrator's restart specifications.
Pitfall 7: Not testing data persistence before changing restart policy. If you recreate a container to change policy and haven't set up volumes, you may lose data. Always test with volumes.
By being aware of these pitfalls, you can avoid most restart policy-related issues.
Conclusion
Docker restart policies are a small but critical part of production container operations. A well-configured policy ensures your services recover from crashes automatically without masking deeper problems. This article provided a practical checklist and examples for setting, verifying, and troubleshooting restart policies.
Remember the operational safety principles: observe before changing, limit the blast radius, use placeholders instead of secrets, verify the result, and document recovery steps. Apply the checklist to your environment, tailor it to your services, and review it regularly.
As a next step, pick one container in your production environment, verify its restart policy against the guidelines in this article, and run a local test to confirm expected behavior. Then extend the audit to all critical services.
With these practices, you will be better prepared for those 3 a.m. crashes.