## Introduction

Containers promise repeatable environments, but production incidents still happen: a container exits unexpectedly, an upgrade corrupts data, or a host failure takes down critical services. Debugging and recovery depend on knowing exactly where data lives, which commands reveal the current state, and how to restore without making things worse.

This guide provides a practical workflow for Docker container debugging, backup, restore, and rollback. You will learn how to inventory your environment, inspect containers and volumes, capture configuration, back up state safely, and validate recovery. Every section includes concrete commands, expected output, and decision points.

We focus on operational safety: observe before changing, limit the blast radius, protect secrets, and verify every step. Whether you are a developer, DevOps engineer, or technical startup team, these practices apply to single hosts and multi-node deployments alike.

## Version and Environment Inventory

Before touching a container, know what you are running. Start with a read-only inventory of Docker versions, running containers, storage drivers, and compose projects. This creates a baseline for troubleshooting and recovery.

### Check Docker and Container Runtime

```bash
docker version
```

Expected output includes client and server versions, plus the OS and architecture. Verify the API version matches between client and server. If they mismatch, update the client or daemon.

```bash
docker system info
```

Look for the storage driver (e.g., overlay2), logging driver, and number of containers/images. The storage driver affects how filesystem snapshots and volume backups work.

### List Running Containers

```bash
docker ps --format "table {{.Names}}\t{{.Status}}\t{{.Ports}}"
```

This shows container names, uptime, and port mappings. Note any container restarting repeatedly, which indicates a crash loop.

For Compose projects:

```bash
docker compose ps
```

This lists services defined in `docker-compose.yml` and their health. If a service is missing, confirm you are in the correct directory or set `-f` to the compose file path.

### Inspect a Container

```bash
docker inspect <container_name_or_id>
```

Inspect reveals mounts, environment variables, network settings, and health checks. Use it to verify data paths before backup or restore. For example:

```bash
docker inspect my_app --format '{{json .Mounts}}'
```

This returns JSON showing volume names, bind mount paths, and read-write flags. If `"RW": false`, the mount is read-only and may not need backup.

### Identify Data Locations

Containers can write to three places: container writable layer, named volumes, or bind mounts. The writable layer is ephemeral and lost when the container is removed. Named volumes are managed by Docker and persist independently. Bind mounts map host directories and are portable if relative paths are consistent.

For any stateful service (databases, queues, file uploads), confirm data is on a volume or bind mount. Example:

```bash
docker inspect my_postgres --format '{{range .Mounts}}{{println .Name " -> " .Destination}}{{end}}'
```

Expected output:

```
postgres_data -> /var/lib/postgresql/data
```

If the destination is inside the container but has no corresponding volume or bind mount, data is in the writable layer. Move it immediately.

### Document Environment

Create a simple inventory file for each host:

```bash
docker version > inventory.txt
docker system info >> inventory.txt
docker ps -a >> inventory.txt
```

Store this outside the container. During an incident, this file helps you identify which version and configuration was in use.

## Safe Configuration Path

Configuration drift causes many failures. Before modifying a container, capture its current configuration and understand the smallest change that fixes the issue. Always separate observation from intervention.

### Capture Current Configuration

For a running container, export its configuration:

```bash
docker inspect <container> > container_config.json
```

For Compose, archive the entire project:

```bash
tar czf compose_project_backup.tar.gz docker-compose.yml .env secrets/
```

Store the archive safely. This lets you restore the exact configuration if a change breaks something.

### Modify Configuration Safely

Avoid `docker exec` to edit files inside containers. Changes are lost on recreation. Instead, use environment variables, configuration files mounted from volumes or bind mounts, or rebuild the image.

If a container uses a bind-mounted config file, edit it on the host:

```bash
nano ./config/app.conf
```

Then restart the container:

```bash
docker restart <container>
```

If the service crashes due to a bad config, revert the file and restart again.

For environment variable changes, edit the Compose file or the `docker run` command:

```bash
docker run -d --name my_app -e DATABASE_URL=postgres://user:pass@db:5432/mydb my_app:latest
```

After change, verify with `docker inspect` that the variable is correct.

### Manage Secrets

Never hardcode secrets in Compose files or command lines. Use Docker secrets (Swarm) or environment files excluded from version control:

```bash
# .env file (not committed)
DB_PASSWORD=secretpassword
```

In Compose:

```yaml
services:
  db:
    image: postgres:16
    environment:
      POSTGRES_PASSWORD: ${DB_PASSWORD}
```

Run `docker compose config` to see the resolved configuration without exposing secrets in logs. Redact sensitive values when sharing diagnostics.

### Test Configuration Changes

For any change, apply it to a staging container first. For example, if you need to change a bind mount path, spin up a test container:

```bash
docker run -d --name test_app -v ./new_data:/var/lib/app my_app:latest
```

Verify functionality, then apply the same change to production. Keep a rollback plan: the previous Compose file or command.

## Verification and Diagnostics

Diagnosing a failing container requires systematic checks: logs, resource usage, health checks, and network connectivity. Move from symptoms to root cause without changing state prematurely.

### Examine Logs

```bash
docker logs <container> --tail 100
```

Add `-f` to follow logs in real time. For specific timestamps:

```bash
docker logs --since 2025-01-01T00:00:00Z <container>
```

If the container crashed, logs may show an exception or exit code. Check exit code with:

```bash
docker inspect <container> --format '{{.State.ExitCode}}'
```

Exit code 137 often means OOM kill; code 1 may be application error.

### Check Resource Usage

```bash
docker stats --no-stream
```

This shows CPU, memory, and network I/O per container. If a container is exceeding memory limit, increase limit if appropriate or fix the leak.

### Health Checks

Many images define health checks. View health status:

```bash
docker inspect <container> --format '{{.State.Health.Status}}'
```

If unhealthy, inspect the last output:

```bash
docker inspect <container> --format '{{json .State.Health}}'
```

For services without health checks, test manually. For a web app:

```bash
docker exec <container> curl -f http://localhost:8080/ || echo "FAILED"
```

### Network Diagnostics

Verify container can reach dependencies:

```bash
docker exec <container> ping -c 3 db
```

Check DNS resolution:

```bash
docker exec <container> nslookup db
```

If DNS fails, check the container is on the correct network:

```bash
docker network inspect <network_name>
```

### Filesystem Checks

Inspect data directory size inside container:

```bash
docker exec <container> du -sh /var/lib/app
```

If the disk is full, logs or data may have filled the volume. Clean up or expand.

## Backup and Restore of Container Data

Backups are only as good as the restore process. This section covers backup techniques for container volumes, bind mounts, and configuration. Always test restore on a separate environment before relying on it.

### Backup a Named Volume

Docker does not have a built-in volume backup command, but you can use a temporary container to copy data from a volume to a tarball on the host.

Stop writes if possible, or ensure low activity. For a database, use its own backup tool instead (e.g., `pg_dump`), but for general volumes:

```bash
docker run --rm -v my_volume:/data -v $(pwd):/backup alpine tar czf /backup/my_volume_backup.tar.gz -C /data .
```

This runs an Alpine container with two mounts: the named volume at `/data` and the current host directory at `/backup`. The `tar` command archives the volume contents to the host.

### Backup a Bind Mount

Simply archive the host directory:

```bash
tar czf bind_mount_backup.tar.gz ./data
```

For large directories, use `rsync`:

```bash
rsync -av --delete ./data /backup/location/
```

### Backup Container Configuration

Always export container configurations and Compose files:

```bash
docker inspect <container> > container_config.json
```

For all containers:

```bash
docker ps -a --format '{{.Names}}' | while read name; do docker inspect "$name" > "${name}_config.json"; done
```

### Database-Specific Backups

If running a database in a container, use its native backup tool for consistency:

PostgreSQL:

```bash
docker exec -t my_postgres pg_dump -U postgres mydb > postgres_backup.sql
```

MySQL:

```bash
docker exec my_mysql mysqldump -u root -psecret mydb > mysql_backup.sql
```

Schedule these with cron and store offsite.

### Restore a Named Volume

Restore from a tar archive:

```bash
docker run --rm -v my_volume:/data -v $(pwd):/backup alpine sh -c "rm -rf /data/* && tar xzf /backup/my_volume_backup.tar.gz -C /data"
```

### Restore a Database

PostgreSQL:

```bash
docker exec -i my_postgres psql -U postgres mydb < postgres_backup.sql
```

### Validate the Restore

After restore, run application smoke tests. For a web app, check health endpoint:

```bash
curl -f http://localhost:8080/health
```

Compare file counts or checksums before and after backup if possible.

## Failure Modes and Recovery

Understanding common failure modes helps you plan recovery. Table 1 lists typical container failures, their symptoms, and recovery actions.

| Failure mode | Symptom | Root cause | Recovery action |
|---|---|---|---|
| Crash loop | Container restarting repeatedly | Application error, bad config, missing dependency | Check logs, fix config or revert, restart |
| OOM kill | Exit code 137 | Memory limit exceeded | Increase memory limit, optimize app, restart |
| Disk full | Writes fail, container stops | Volume or host disk full | Clean up old data, expand disk, restart |
| Network unreachable | Can't connect to other services | Wrong network, DNS failure | Check network attach, DNS config |
| Data loss | Data missing after container recreation | Wrote to container layer, not volume | Restore from backup, move to volume |
| Broken config | Service crashes after config change | Typo or invalid setting | Revert config file, restart |
| Version incompatibility | API errors after Docker update | Client/server mismatch | Update client or daemon to match |

### Rollback Strategies

For images, keep previous tags:

```bash
docker tag my_app:latest my_app:prev
docker build -t my_app:latest .
# If new image fails
docker tag my_app:prev my_app:latest
docker service update --force my_app (Swarm) or docker-compose up -d
```

For configuration, maintain versioned config files:

```bash
cp app.conf app.conf.bak.20250101
```

To revert, copy the backup and restart.

### Disaster Recovery Planning

Define Recovery Time Objective (RTO) and Recovery Point Objective (RPO) for each service. For a critical database, RPO might be 5 minutes (frequent backups) and RTO 30 minutes. Test recovery quarterly.

Keep backups offsite and encrypt sensitive data. Use automation to schedule backups.

## Operations Checklist

Use this checklist before, during, and after any container change or incident. Each item includes an owner and review frequency.

### Pre-Change Checklist (Owner: DevOps Engineer, reviewed before every change)

- [ ] Record current state: `docker ps -a`, `docker system info`, and relevant logs.
- [ ] Confirm data location: volume or bind mount, not container layer.
- [ ] Capture configuration: export `docker inspect` or archive Compose files.
- [ ] Identify rollback plan: previous image tag, config file backup, or volume snapshot.
- [ ] Notify stakeholders if change may cause downtime.

### Post-Change Verification (Owner: Application Developer, reviewed after every change)

- [ ] Check container health: `docker inspect <container> --format '{{.State.Health.Status}}'`
- [ ] Verify service endpoint returns 200 or expected response.
- [ ] Confirm data integrity: compare row counts, file counts, or checksums.
- [ ] Watch logs for errors for at least 10 minutes: `docker logs -f --since 10m <container>`
- [ ] Document the change and its outcome in incident log.

### Weekly Backup Verification (Owner: Operations Lead, reviewed weekly)

- [ ] Run backup for all critical volumes and databases.
- [ ] Check backup file sizes and timestamps; ensure they are recent.
- [ ] Test restore one backup to a staging environment.
- [ ] Verify restored data integrity with application queries.
- [ ] Update backup documentation if procedures changed.

### Monthly Security Review (Owner: Security Engineer, reviewed monthly)

- [ ] Scan images for vulnerabilities: `docker scan <image>`
- [ ] Audit container capabilities: no unnecessary privileges.
- [ ] Review secrets management: ensure no plaintext secrets in configs.
- [ ] Rotate credentials used by containers.
- [ ] Verify network policies isolate services appropriately.

### Quarterly Disaster Recovery Drill (Owner: Site Reliability Engineer, reviewed quarterly)

- [ ] Simulate host failure: restore volumes and databases from backup to new host.
- [ ] Measure time to restore against RTO.
- [ ] Check data freshness against RPO.
- [ ] Update runbooks based on drill findings.
- [ ] Train team members on recovery procedures.

## Common Pitfalls and How to Avoid Them

Several mistakes recur in container operations. Each pitfall below includes why it happens and how to avoid or recover.

### Pitfall 1: Storing Data in the Container Writable Layer

**Why it happens:** Developers use the container as a full VM and write to default paths without configuring volumes.

**Avoidance:** Always define volumes or bind mounts for stateful data in Dockerfile or Compose file.

**Recovery:** If data is still accessible, copy it out immediately:

```bash
docker cp my_container:/var/lib/app ./recovered_data
```

Then recreate container with proper volume.

### Pitfall 2: Backing Up a Live Database Without Consistency

**Why it happens:** Using `tar` on volume data while database is writing can produce corrupt backup.

**Avoidance:** Use database-native backup tools or stop the database before tar.

**Recovery:** If backup is corrupt, restore from previous valid backup or use replication logs.

### Pitfall 3: Not Testing Restores

**Why it happens:** Backups appear successful, so teams assume restore works.

**Avoidance:** Schedule regular restore drills. Automate a weekly restore to staging.

**Recovery:** If restore fails during incident, contact vendor support or use secondary backup.

### Pitfall 4: Hardcoding Secrets in Compose Files

**Why it happens:** Convenience during development.

**Avoidance:** Use environment variables, Docker secrets, or a secret manager.

**Recovery:** Rotate exposed secrets immediately. Scan code repos for secrets.

### Pitfall 5: Ignoring Resource Limits

**Why it happens:** Containers run fine in test but may consume unlimited resources in production.

**Avoidance:** Set CPU and memory limits in Compose:

```yaml
services:
  app:
    image: my_app
    deploy:
      resources:
        limits:
          cpus: '0.5'
          memory: 512M
```

**Recovery:** If container is OOM killed, increase limit appropriately and optimize application.

### Pitfall 6: Skipping Version Compatibility Checks

**Why it happens:** Docker client and daemon updated independently.

**Avoidance:** Keep client and daemon on compatible versions; check `docker version` before major changes.

**Recovery:** Downgrade or upgrade as needed; use official compatibility matrix.

## Conclusion

Debugging and recovering Docker containers requires a disciplined approach: know your environment, change carefully, verify outcomes, and plan for failure. By following the commands and checklists in this guide, you reduce downtime and data loss.

Start with the version and environment inventory to understand your current state. Capture configuration before modifications, use systematic diagnostics, and implement robust backup and restore procedures. Rehearse disaster recovery regularly.

Assign clear owners to each checklist item and review frequency. The operations checklist keeps your team accountable and prepared. Avoid common pitfalls by learning from them. Containerization offers agility, but operational rigor ensures reliability.

As a next step, run the environment inventory on one host, identify a stateful container without a volume, and fix it before you need a backup.