Introduction
Containers promise repeatable environments, but production incidents still happen: a container exits unexpectedly, an upgrade corrupts data, or a host failure takes down critical services. Debugging and recovery depend on knowing exactly where data lives, which commands reveal the current state, and how to restore without making things worse.
This guide provides a practical workflow for Docker container debugging, backup, restore, and rollback. You will learn how to inventory your environment, inspect containers and volumes, capture configuration, back up state safely, and validate recovery. Every section includes concrete commands, expected output, and decision points.
We focus on operational safety: observe before changing, limit the blast radius, protect secrets, and verify every step. Whether you are a developer, DevOps engineer, or technical startup team, these practices apply to single hosts and multi-node deployments alike.
Version and Environment Inventory
Before touching a container, know what you are running. Start with a read-only inventory of Docker versions, running containers, storage drivers, and compose projects. This creates a baseline for troubleshooting and recovery.
Check Docker and Container Runtime
docker version
Expected output includes client and server versions, plus the OS and architecture. Verify the API version matches between client and server. If they mismatch, update the client or daemon.
docker system info
Look for the storage driver (e.g., overlay2), logging driver, and number of containers/images. The storage driver affects how filesystem snapshots and volume backups work.
List Running Containers
docker ps --format "table {{.Names}}\t{{.Status}}\t{{.Ports}}"
This shows container names, uptime, and port mappings. Note any container restarting repeatedly, which indicates a crash loop.
For Compose projects:
docker compose ps
This lists services defined in docker-compose.yml and their health. If a service is missing, confirm you are in the correct directory or set -f to the compose file path.
Inspect a Container
docker inspect <container_name_or_id>
Inspect reveals mounts, environment variables, network settings, and health checks. Use it to verify data paths before backup or restore. For example:
docker inspect my_app --format '{{json .Mounts}}'
This returns JSON showing volume names, bind mount paths, and read-write flags. If "RW": false, the mount is read-only and may not need backup.
Identify Data Locations
Containers can write to three places: container writable layer, named volumes, or bind mounts. The writable layer is ephemeral and lost when the container is removed. Named volumes are managed by Docker and persist independently. Bind mounts map host directories and are portable if relative paths are consistent.
For any stateful service (databases, queues, file uploads), confirm data is on a volume or bind mount. Example:
docker inspect my_postgres --format '{{range .Mounts}}{{println .Name " -> " .Destination}}{{end}}'
Expected output:
postgres_data -> /var/lib/postgresql/data
If the destination is inside the container but has no corresponding volume or bind mount, data is in the writable layer. Move it immediately.
Document Environment
Create a simple inventory file for each host:
docker version > inventory.txt
docker system info >> inventory.txt
docker ps -a >> inventory.txt
Store this outside the container. During an incident, this file helps you identify which version and configuration was in use.
Safe Configuration Path
Configuration drift causes many failures. Before modifying a container, capture its current configuration and understand the smallest change that fixes the issue. Always separate observation from intervention.
Capture Current Configuration
For a running container, export its configuration:
docker inspect <container> > container_config.json
For Compose, archive the entire project:
tar czf compose_project_backup.tar.gz docker-compose.yml .env secrets/
Store the archive safely. This lets you restore the exact configuration if a change breaks something.
Modify Configuration Safely
Avoid docker exec to edit files inside containers. Changes are lost on recreation. Instead, use environment variables, configuration files mounted from volumes or bind mounts, or rebuild the image.
If a container uses a bind-mounted config file, edit it on the host:
nano ./config/app.conf
Then restart the container:
docker restart <container>
If the service crashes due to a bad config, revert the file and restart again.
For environment variable changes, edit the Compose file or the docker run command:
docker run -d --name my_app -e DATABASE_URL=postgres://user:pass@db:5432/mydb my_app:latest
After change, verify with docker inspect that the variable is correct.
Manage Secrets
Never hardcode secrets in Compose files or command lines. Use Docker secrets (Swarm) or environment files excluded from version control:
# .env file (not committed)
DB_PASSWORD=secretpassword
In Compose:
services:
db:
image: postgres:16
environment:
POSTGRES_PASSWORD: ${DB_PASSWORD}
Run docker compose config to see the resolved configuration without exposing secrets in logs. Redact sensitive values when sharing diagnostics.
Test Configuration Changes
For any change, apply it to a staging container first. For example, if you need to change a bind mount path, spin up a test container:
docker run -d --name test_app -v ./new_data:/var/lib/app my_app:latest
Verify functionality, then apply the same change to production. Keep a rollback plan: the previous Compose file or command.
Verification and Diagnostics
Diagnosing a failing container requires systematic checks: logs, resource usage, health checks, and network connectivity. Move from symptoms to root cause without changing state prematurely.
Examine Logs
docker logs <container> --tail 100
Add -f to follow logs in real time. For specific timestamps:
docker logs --since 2025-01-01T00:00:00Z <container>
If the container crashed, logs may show an exception or exit code. Check exit code with:
docker inspect <container> --format '{{.State.ExitCode}}'
Exit code 137 often means OOM kill; code 1 may be application error.
Check Resource Usage
docker stats --no-stream
This shows CPU, memory, and network I/O per container. If a container is exceeding memory limit, increase limit if appropriate or fix the leak.
Health Checks
Many images define health checks. View health status:
docker inspect <container> --format '{{.State.Health.Status}}'
If unhealthy, inspect the last output:
docker inspect <container> --format '{{json .State.Health}}'
For services without health checks, test manually. For a web app:
docker exec <container> curl -f http://localhost:8080/ || echo "FAILED"
Network Diagnostics
Verify container can reach dependencies:
docker exec <container> ping -c 3 db
Check DNS resolution:
docker exec <container> nslookup db
If DNS fails, check the container is on the correct network:
docker network inspect <network_name>
Filesystem Checks
Inspect data directory size inside container:
docker exec <container> du -sh /var/lib/app
If the disk is full, logs or data may have filled the volume. Clean up or expand.
Backup and Restore of Container Data
Backups are only as good as the restore process. This section covers backup techniques for container volumes, bind mounts, and configuration. Always test restore on a separate environment before relying on it.
Backup a Named Volume
Docker does not have a built-in volume backup command, but you can use a temporary container to copy data from a volume to a tarball on the host.
Stop writes if possible, or ensure low activity. For a database, use its own backup tool instead (e.g., pg_dump), but for general volumes:
docker run --rm -v my_volume:/data -v $(pwd):/backup alpine tar czf /backup/my_volume_backup.tar.gz -C /data .
This runs an Alpine container with two mounts: the named volume at /data and the current host directory at /backup. The tar command archives the volume contents to the host.
Backup a Bind Mount
Simply archive the host directory:
tar czf bind_mount_backup.tar.gz ./data
For large directories, use rsync:
rsync -av --delete ./data /backup/location/
Backup Container Configuration
Always export container configurations and Compose files:
docker inspect <container> > container_config.json
For all containers:
docker ps -a --format '{{.Names}}' | while read name; do docker inspect "$name" > "${name}_config.json"; done
Database-Specific Backups
If running a database in a container, use its native backup tool for consistency:
PostgreSQL:
docker exec -t my_postgres pg_dump -U postgres mydb > postgres_backup.sql
MySQL:
docker exec my_mysql mysqldump -u root -psecret mydb > mysql_backup.sql
Schedule these with cron and store offsite.
Restore a Named Volume
Restore from a tar archive:
docker run --rm -v my_volume:/data -v $(pwd):/backup alpine sh -c "rm -rf /data/* && tar xzf /backup/my_volume_backup.tar.gz -C /data"
Restore a Database
PostgreSQL:
docker exec -i my_postgres psql -U postgres mydb < postgres_backup.sql
Validate the Restore
After restore, run application smoke tests. For a web app, check health endpoint:
curl -f http://localhost:8080/health
Compare file counts or checksums before and after backup if possible.
Failure Modes and Recovery
Understanding common failure modes helps you plan recovery. Table 1 lists typical container failures, their symptoms, and recovery actions.
| Failure mode | Symptom | Root cause | Recovery action |
|---|---|---|---|
| Crash loop | Container restarting repeatedly | Application error, bad config, missing dependency | Check logs, fix config or revert, restart |
| OOM kill | Exit code 137 | Memory limit exceeded | Increase memory limit, optimize app, restart |
| Disk full | Writes fail, container stops | Volume or host disk full | Clean up old data, expand disk, restart |
| Network unreachable | Can't connect to other services | Wrong network, DNS failure | Check network attach, DNS config |
| Data loss | Data missing after container recreation | Wrote to container layer, not volume | Restore from backup, move to volume |
| Broken config | Service crashes after config change | Typo or invalid setting | Revert config file, restart |
| Version incompatibility | API errors after Docker update | Client/server mismatch | Update client or daemon to match |
Rollback Strategies
For images, keep previous tags:
docker tag my_app:latest my_app:prev
docker build -t my_app:latest .
# If new image fails
docker tag my_app:prev my_app:latest
docker service update --force my_app (Swarm) or docker-compose up -d
For configuration, maintain versioned config files:
cp app.conf app.conf.bak.20250101
To revert, copy the backup and restart.
Disaster Recovery Planning
Define Recovery Time Objective (RTO) and Recovery Point Objective (RPO) for each service. For a critical database, RPO might be 5 minutes (frequent backups) and RTO 30 minutes. Test recovery quarterly.
Keep backups offsite and encrypt sensitive data. Use automation to schedule backups.
Operations Checklist
Use this checklist before, during, and after any container change or incident. Each item includes an owner and review frequency.
Pre-Change Checklist (Owner: DevOps Engineer, reviewed before every change)
- [ ] Record current state:
docker ps -a,docker system info, and relevant logs. - [ ] Confirm data location: volume or bind mount, not container layer.
- [ ] Capture configuration: export
docker inspector archive Compose files. - [ ] Identify rollback plan: previous image tag, config file backup, or volume snapshot.
- [ ] Notify stakeholders if change may cause downtime.
Post-Change Verification (Owner: Application Developer, reviewed after every change)
- [ ] Check container health:
docker inspect <container> --format '{{.State.Health.Status}}' - [ ] Verify service endpoint returns 200 or expected response.
- [ ] Confirm data integrity: compare row counts, file counts, or checksums.
- [ ] Watch logs for errors for at least 10 minutes:
docker logs -f --since 10m <container> - [ ] Document the change and its outcome in incident log.
Weekly Backup Verification (Owner: Operations Lead, reviewed weekly)
- [ ] Run backup for all critical volumes and databases.
- [ ] Check backup file sizes and timestamps; ensure they are recent.
- [ ] Test restore one backup to a staging environment.
- [ ] Verify restored data integrity with application queries.
- [ ] Update backup documentation if procedures changed.
Monthly Security Review (Owner: Security Engineer, reviewed monthly)
- [ ] Scan images for vulnerabilities:
docker scan <image> - [ ] Audit container capabilities: no unnecessary privileges.
- [ ] Review secrets management: ensure no plaintext secrets in configs.
- [ ] Rotate credentials used by containers.
- [ ] Verify network policies isolate services appropriately.
Quarterly Disaster Recovery Drill (Owner: Site Reliability Engineer, reviewed quarterly)
- [ ] Simulate host failure: restore volumes and databases from backup to new host.
- [ ] Measure time to restore against RTO.
- [ ] Check data freshness against RPO.
- [ ] Update runbooks based on drill findings.
- [ ] Train team members on recovery procedures.
Common Pitfalls and How to Avoid Them
Several mistakes recur in container operations. Each pitfall below includes why it happens and how to avoid or recover.
Pitfall 1: Storing Data in the Container Writable Layer
Why it happens: Developers use the container as a full VM and write to default paths without configuring volumes.
Avoidance: Always define volumes or bind mounts for stateful data in Dockerfile or Compose file.
Recovery: If data is still accessible, copy it out immediately:
docker cp my_container:/var/lib/app ./recovered_data
Then recreate container with proper volume.
Pitfall 2: Backing Up a Live Database Without Consistency
Why it happens: Using tar on volume data while database is writing can produce corrupt backup.
Avoidance: Use database-native backup tools or stop the database before tar.
Recovery: If backup is corrupt, restore from previous valid backup or use replication logs.
Pitfall 3: Not Testing Restores
Why it happens: Backups appear successful, so teams assume restore works.
Avoidance: Schedule regular restore drills. Automate a weekly restore to staging.
Recovery: If restore fails during incident, contact vendor support or use secondary backup.
Pitfall 4: Hardcoding Secrets in Compose Files
Why it happens: Convenience during development.
Avoidance: Use environment variables, Docker secrets, or a secret manager.
Recovery: Rotate exposed secrets immediately. Scan code repos for secrets.
Pitfall 5: Ignoring Resource Limits
Why it happens: Containers run fine in test but may consume unlimited resources in production.
Avoidance: Set CPU and memory limits in Compose:
services:
app:
image: my_app
deploy:
resources:
limits:
cpus: '0.5'
memory: 512M
Recovery: If container is OOM killed, increase limit appropriately and optimize application.
Pitfall 6: Skipping Version Compatibility Checks
Why it happens: Docker client and daemon updated independently.
Avoidance: Keep client and daemon on compatible versions; check docker version before major changes.
Recovery: Downgrade or upgrade as needed; use official compatibility matrix.
Conclusion
Debugging and recovering Docker containers requires a disciplined approach: know your environment, change carefully, verify outcomes, and plan for failure. By following the commands and checklists in this guide, you reduce downtime and data loss.
Start with the version and environment inventory to understand your current state. Capture configuration before modifications, use systematic diagnostics, and implement robust backup and restore procedures. Rehearse disaster recovery regularly.
Assign clear owners to each checklist item and review frequency. The operations checklist keeps your team accountable and prepared. Avoid common pitfalls by learning from them. Containerization offers agility, but operational rigor ensures reliability.
As a next step, run the environment inventory on one host, identify a stateful container without a volume, and fix it before you need a backup.