## Intro

Debugging a misconfigured Docker container is rarely about finding one magic command. It is about building a repeatable sequence: observe the current state, identify the smallest possible change, apply it in a controlled way, verify the result, and know how to roll back if things get worse. This guide walks through common configuration mistakes - networking, mounts, environment variables, health checks, and resource limits - using real commands and expected outputs so you can move from symptom to solution quickly.

We will focus on practical, production-oriented debugging. The goal is operational safety: never change a running container blindly, protect secrets, limit your blast radius, and always have a recovery path documented before you need it.

## Version and Environment Inventory

Before you touch anything, know exactly what you are working with. Run these read-only commands and capture the output in your incident notes with a timestamp.

**Check the Docker version and daemon info:**
```bash
docker version
```
Expected output includes the client and server versions. For example:
```
Client: Docker Engine - Community
 Version:           24.0.7
 API version:       1.43
 Go version:        go1.20.10
 Git commit:        311b9ff
 Built:             Thu Oct 26 09:08:15 2023
 OS/Arch:           linux/amd64
 Context:           default

Server: Docker Engine - Community
 Engine:
  Version:          24.0.7
```
If the client and server versions differ significantly, some features may behave unexpectedly. Upgrade or downgrade to match the team standard.

**List running containers with key metadata:**
```bash
docker ps --format "table {{.Names}}\t{{.Status}}\t{{.Ports}}\t{{.Image}}"
```
Example result:
```
NAMES                STATUS                    PORTS                     IMAGE
web-app              Up 2 hours (unhealthy)   0.0.0.0:8080->80/tcp      nginx:1.25
postgres-db          Up 5 hours                0.0.0.0:5432->5432/tcp    postgres:15
```
Note the health status. An "unhealthy" container is a clue, not a verdict - it may still be serving traffic while failing its health check.

**Inspect a specific container for mounts, networks, environment variables, and health configuration:**
```bash
docker inspect web-app
```
This produces a large JSON object. Filter the parts you need:
```bash
docker inspect web-app --format '{{json .Mounts}}' | jq
docker inspect web-app --format '{{json .NetworkSettings.Networks}}' | jq
docker inspect web-app --format '{{json .Config.Env}}' | jq
docker inspect web-app --format '{{json .Config.Healthcheck}}' | jq
```
Each command returns structured data. For example, the mounts output for a container with a named volume and a bind mount might look like:
```json
[
  {
    "Type": "volume",
    "Name": "web-app-data",
    "Source": "/var/lib/docker/volumes/web-app-data/_data",
    "Destination": "/var/www/html",
    "Mode": "",
    "RW": true,
    "Propagation": "rprivate"
  },
  {
    "Type": "bind",
    "Source": "/home/user/config",
    "Destination": "/etc/app/config",
    "Mode": "ro",
    "RW": false,
    "Propagation": "rprivate"
  }
]
```
Pay close attention to the `RW` field: a read-only bind mount will cause writes to fail with a permission error if the application expects to modify those files.

**For Compose projects**, get a project-wide view:
```bash
docker compose ps
docker compose config --services
docker compose config --volumes
```
These affirm the services defined in your compose file and whether Docker sees the volumes you expect.

### Data Persistence Before You Change Anything

One of the most common container misconfigurations is losing data on container recreation. Before debugging further, confirm where persistent data lives:

- Named volume: managed by Docker, easy to reuse. Example Compose snippet:
```yaml
services:
  app:
    image: myapp:1.2
    volumes:
      - app_data:/var/lib/app
volumes:
  app_data:
```
- Bind mount: maps a host path directly. Useful for development but can fail in production if the host path does not exist on all machines.

Run a restart test in a staging environment:
1. Stop the container: `docker stop web-app`
2. Remove it: `docker rm web-app`
3. Recreate it from the same compose file or run command.
4. Check if the data is still present: `docker exec web-app ls /var/lib/app`

If files are missing, the application was probably writing to the container's writable layer instead of a volume or mount. Fix the Dockerfile or compose file to mount a volume at that path.

## Safe Configuration Path

When you have identified a likely configuration error, follow a surgical change process. Never edit files inside a running container with `docker exec vi` - that change disappears on container recreation and is not tracked.

### Use Environment Variables Correctly

A frequent mistake is passing secrets via plain environment variables in a `docker run` command or compose file, where they are visible in `docker inspect` and process lists.

**Bad practice:**
```bash
docker run -e DATABASE_PASSWORD=supersecret myapp
```
`docker inspect` will show this password in clear text.

**Better: use Docker secrets (for Swarm) or a secrets file mounted as a read-only volume.**
Example compose file with a secret file mount:
```yaml
services:
  app:
    image: myapp:1.2
    secrets:
      - db_password
secrets:
  db_password:
    file: ./secrets/db_password.txt
```
Inside the container, the secret appears at `/run/secrets/db_password`. The application reads from that path.

### Change One Variable at a Time

If you suspect an environment variable is wrong, change only that one variable and restart the container. Do not bundle unrelated tweaks.

For a single container:
```bash
docker stop web-app
docker rm web-app
docker run -d --name web-app \
  -e APP_ENV=production \
  -e DATABASE_URL=postgres://db:5432/app \
  -p 8080:80 \
  myapp:1.2
```
Then check logs and health:
```bash
docker logs web-app --tail 50
docker inspect web-app --format '{{.State.Health.Status}}'
```
If the health status is not "healthy" within a minute, you know the new environment variable contributed to the problem.

For Compose, edit `docker-compose.yml`, then recreate only the affected service:
```bash
docker compose up -d --force-recreate app
```
`--force-recreate` ensures the container is rebuilt with the new configuration rather than reusing the old one.

### Rollback Plan

Before applying a change, write the exact rollback command in your notes. For a Compose service, the previous configuration is in your version control - simply check out the previous commit and run `docker compose up -d` again. For a manual `docker run`, you can inspect the container's current configuration before removal with:
```bash
docker inspect web-app --format '{{json .Config}}' > web-app-config-backup.json
```
Then, if needed, reconstruct the run command from that JSON. Better yet, keep all container definitions in a compose file or script under version control so rollback is reliable.

## Verification and Diagnostics

After any change, verify the container behaves as expected. Do not rely on "it started" - check actual functionality and resource usage.

### Check Logs for Error Patterns

```bash
docker logs web-app --tail 200
```
Look for recurring exceptions, connection refused messages, or invalid configuration errors. For a container that crashes immediately after start, use:
```bash
docker logs web-app
```
If the container has restarted multiple times, add timestamps:
```bash
docker logs --timestamps web-app
```

For a container that exits immediately, run it in the foreground with a shell override to debug startup issues:
```bash
docker run --rm -it myapp:1.2 /bin/sh
```
Then manually run the application's startup command to see the error.

### Test Network Connectivity

From inside a running container, test DNS resolution and TCP connectivity:
```bash
docker exec web-app ping -c 3 database
```
Expected output if DNS works:
```
PING database (172.18.0.2) 56(84) bytes of data.
64 bytes from database (172.18.0.2): icmp_seq=1 ttl=64 time=0.089 ms
```
If the name does not resolve, the container may not be on the same user-defined network. Check networks:
```bash
docker network ls
docker network inspect bridge
```
For a container attached to a specific network, ensure the Compose services share that network.

Check port binding conflicts:
```bash
docker ps --format "table {{.Names}}\t{{.Ports}}"
```
If another container already uses the host port you want, you will see a port binding error in the logs.

### Health Checks and Readiness

A misconfigured health check can mark a healthy container as unhealthy, triggering unnecessary restarts. Inspect the health check definition:
```bash
docker inspect web-app --format '{{json .Config.Healthcheck}}' | jq
```
Example output:
```json
{
  "Test": ["CMD-SHELL", "curl -f http://localhost/ || exit 1"],
  "Interval": 30000000000,
  "Timeout": 5000000000,
  "Retries": 3,
  "StartPeriod": 60000000000
}
```
The intervals are in nanoseconds. A start period of 60 seconds (60,000,000,000 ns) may be too short if your application takes longer to initialize. Adjust in your Dockerfile or compose file:
```yaml
healthcheck:
  test: ["CMD", "curl", "-f", "http://localhost/"]
  interval: 30s
  timeout: 5s
  retries: 3
  start_period: 120s
```
After updating, recreate the container and monitor the health status over time:
```bash
watch -n 5 'docker inspect web-app --format "{{.State.Health.Status}}"'
```

## Failure Modes and Recovery

Let's examine specific configuration mistakes and how to recover from each.

### Mistake 1: Missing Volume Mount Leads to Data Loss

**Symptom:** After recreating the container, all user-uploaded files are gone.
**Why it happens:** The application wrote to a directory inside the container filesystem, not to a mounted volume.
**Recovery:** Check the old container's filesystem (if still available) with `docker cp` to retrieve files before removing it. Then update your compose file to add a named volume at the application's data path. Recreate the container and verify persistence with a restart test.

### Mistake 2: Environment Variable Typos Cause Connection Failures

**Symptom:** The application cannot connect to the database, but the database container is running.
**Why it happens:** A typo in the environment variable name, e.g., `DATABSE_URL` instead of `DATABASE_URL`.
**Recovery:** Inspect the container's environment:
```bash
docker exec web-app env | sort
```
Compare with what the application expects. Fix the compose file or Dockerfile, then recreate the container. Use a configuration library that validates required environment variables at startup to fail fast.

### Mistake 3: Incorrect Port Mapping Prevents External Access

**Symptom:** Service is running but cannot be reached from the host.
**Why it happens:** Either the wrong container port is mapped, or the host port is already in use.
**Recovery:** Check actual port mappings:
```bash
docker port web-app
```
Compare with the application's listening port (check Dockerfile EXPOSE or startup logs). Update the port mapping and recreate. Verify with `curl http://localhost:8080`.

### Mistake 4: Resource Limits Not Set Causes Container to Kill Processes

**Symptom:** The container is killed with exit code 137 (SIGKILL) under load.
**Why it happens:** The container exceeded memory limit (if set) or the host ran out of memory.
**Recovery:** Inspect current resource usage:
```bash
docker stats --no-stream
```
Set appropriate memory and CPU limits in your compose file:
```yaml
services:
  app:
    image: myapp:1.2
    deploy:
      resources:
        limits:
          cpus: '0.50'
          memory: 512M
        reservations:
          cpus: '0.25'
          memory: 256M
```
Recreate and run load tests to confirm stability.

### Mistake 5: Using the Wrong Base Image or Tag

**Symptom:** Application behaves differently in production than in development, or dependencies are missing.
**Why it happens:** The Dockerfile uses a mutable tag like `latest`, or the base image differs between environments.
**Recovery:** Pin the base image to a specific digest or version tag in your Dockerfile. Rebuild the image, push to registry, and redeploy. Use `docker image inspect` to verify the image ID matches across environments.

## Common Pitfalls and How to Avoid Them

Beyond the specific mistakes above, here are general pitfalls that trip up teams.

### Editing Files Inside the Container

Running `docker exec -it web-app bash` and editing a config file with `vi` or `sed` is tempting for a quick fix. But the change is lost on container recreation and is not visible to other team members. Instead, mount a configuration file or use environment variables, then recreate.

### Ignoring Logs Until Something Breaks

Logs are the first place to look, yet many teams do not collect them centrally. Use Docker logging drivers to forward logs to an aggregation service. Example compose snippet:
```yaml
services:
  app:
    image: myapp:1.2
    logging:
      driver: json-file
      options:
        max-size: "10m"
        max-file: "3"
```
For a central syslog server, switch driver to `syslog` and specify address. This ensures logs are available even if the container is removed.

### Not Using a `.dockerignore` File

Without a `.dockerignore`, the build context may include the local `.git` directory, node_modules, or secrets, making the image unnecessarily large and potentially leaking sensitive data. Create a `.dockerignore` file at the root of your build context:
```
.git
node_modules
*.log
.env
secrets
```

### Overusing `latest` Tag

As mentioned, `latest` makes builds non-reproducible. Always use semantic version tags or digests in production deployment files. For local development, `latest` is acceptable but avoid it in CI/CD pipelines.

### Skipping Health Checks

Without a health check, Docker cannot tell if your application is truly ready. Define a health check that tests a critical endpoint, not just that the process is running. Use the `start_period` to allow slow-starting applications enough time.

## Operations Checklist

Use this checklist before and after any configuration change to a containerized service. Assign each item to a single accountable person (not a team) and revisit the list during every release or incident review.

| Step | Action | Owner | Frequency |
|------|--------|-------|-----------|
| 1 | Record current state: `docker ps`, `docker logs`, `docker inspect` output | On-call engineer | At incident start |
| 2 | Identify the change needed and its blast radius | App owner (e.g., Priya Shah, Engineering Lead) | During change planning |
| 3 | Back up critical data: `docker cp <container>:/path /backup/` if necessary | Database administrator | Before destructive change |
| 4 | Apply the smallest possible change (one env var, one mount, one port) | On-call engineer | At change time |
| 5 | Verify functionality: check logs, health status, and actual endpoint response | QA engineer | Immediately after change |
| 6 | Document the change and rollback procedure in the runbook | Service owner | Within 24 hours |
| 7 | Run a restart test in staging to confirm data persistence and config | App owner | Before next deploy |
| 8 | Review the incident timeline and update runbook if needed | Engineering manager | Weekly incident review |

This checklist ensures no step is skipped when under pressure. The owners are illustrative - replace with actual names and roles in your organization.

## Conclusion

Debugging Docker container configuration mistakes is a disciplined process, not a guessing game. Start with a complete inventory of the environment, understand the current state, make one minimal change at a time, and verify with concrete commands and expected outputs. Always have a rollback plan before you change anything.

The most reliable workflows make failure visible through logs and health checks, protect sensitive values with secrets management, limit changes to the intended resource, and define recovery verification before an incident forces a decision. Apply the checklist and pitfalls in this guide to build safer, more resilient container deployments. As a next step, pick one service you manage, run through the Operations Checklist, and document at least one recovery procedure in your team's runbook.

A reliable technical workflow makes failure visible, protects sensitive values, limits changes to the intended resource, and defines recovery verification before an incident forces the decision.