Introduction
Docker image monitoring is essential for maintaining reliable containerized applications. This guide provides a practical approach to monitoring Docker images, setting up alerts, and responding to incidents. We'll cover version and environment inventory, safe configuration paths, verification and diagnostics, failure modes, and an operations checklist, with concrete commands and examples.
This article is aimed at developers, DevOps consultants, and technical startup teams who need to move from observing a problem to verifying a resolution. The goal is operational safety: observe before changing, limit the blast radius, protect sensitive information, verify results, and document recovery procedures.
Version and Environment Inventory
Before making any changes, establish a clear picture of your Docker environment. Start by checking the Docker version and daemon status. Run docker version to see client and server versions, and docker info to get daemon details like storage driver, logging driver, and number of containers. This helps ensure compatibility with monitoring tools.
For observing running containers, use docker ps --format "table {{.Names}}\t{{.Status}}\t{{.Ports}}" to list containers with names, status, and ports. This read-only command is safe for production.
If a container is misbehaving, examine its logs: docker logs <container_name> --tail 100 shows the last 100 lines. For deeper inspection, use docker inspect <container_name> to view mounts, networks, environment variables, and health status.
For Compose projects, use docker compose ps to see service status, docker compose logs -f <service> to stream logs, and docker compose exec <service> sh to enter a shell without rebuilding the image.
When data persistence matters, identify where files are stored before changing anything. A named volume like app_data:/var/lib/app is managed by Docker and typically easier to reuse across rebuilds. A bind mount such as ./data:/var/lib/app maps a host directory into the container, useful for development but can cause permission and portability issues if the path differs across machines.
Perform a restart test in a staging environment: stop the container, recreate it, and verify the application still accesses expected data. If data disappears, the service likely writes to the container filesystem instead of a volume or mount.
Example verification:
docker run -d --name test-app -v app_data:/var/lib/app myapp:latest
docker stop test-app
docker rm test-app
docker run -d --name test-app -v app_data:/var/lib/app myapp:latest
docker exec test-app ls /var/lib/app
Expected output should include your application files, confirming persistence.
Safe Configuration Path
Configuring Docker monitoring involves setting up metrics collection, logging, and alerting without compromising security. Always separate observation from intervention.
Start by defining the monitoring scope. Decide which containers and images are critical. Use labels to organize: docker run -d --label env=prod --label app=web myapp:latest. This aids in filtering metrics.
For metrics, enable the Docker daemon metrics endpoint. Edit /etc/docker/daemon.json:
{
"metrics-addr": "127.0.0.1:9323",
"experimental": true
}
Restart Docker: sudo systemctl restart docker. Test with curl http://127.0.0.1:9323/metrics to see Prometheus-formatted metrics.
For logs, use Docker's logging drivers. Configure the JSON file driver with rotation to prevent disk exhaustion:
{
"log-driver": "json-file",
"log-opts": {
"max-size": "10m",
"max-file": "3"
}
}
Apply globally in daemon.json or per container with --log-opt.
Set up a basic alert using a tool like cAdvisor combined with Prometheus and Alertmanager. Run cAdvisor to collect container metrics:
docker run -d \
--name=cadvisor \
-p 8080:8080 \
-v /:/rootfs:ro \
-v /var/run:/var/run:ro \
-v /sys:/sys:ro \
-v /var/lib/docker/:/var/lib/docker:ro \
gcr.io/cadvisor/cadvisor:latest
Then configure Prometheus to scrape cAdvisor and define alert rules in alert.rules:
groups:
- name: container_alerts
rules:
- alert: ContainerHighCPU
expr: sum(rate(container_cpu_usage_seconds_total{name!=""}[5m])) by (name) > 0.8
for: 5m
labels:
severity: warning
annotations:
summary: "Container {{ $labels.name }} high CPU usage"
Route alerts via Alertmanager to email, Slack, or other channels.
Protect sensitive values: never hardcode passwords in configuration files. Use Docker secrets or environment variables. For example, use docker secret create db_password ./password.txt and reference in service as db_password.
Before applying changes, test in a non-production environment. Use docker compose config to validate Compose file syntax.
Verification and Diagnostics
After setting up monitoring, verify that metrics and alerts work as expected.
Check if metrics are being scraped by Prometheus: open Prometheus UI (default port 9090), go to Status > Targets, and ensure cAdvisor target is up. Query a metric like container_memory_usage_bytes to see data.
For alerts, simulate a condition. For example, run a container that generates CPU load:
docker run -d --name stress --rm alpine sh -c "while true; do :; done"
Wait for the alert to fire (check Alertmanager UI or configured channel). Then stop the container and confirm alert resolves.
Diagnose image-specific issues using docker image inspect <image> to view metadata, layers, and environment. Check for vulnerabilities with a scanner like Trivy:
trivy image --severity HIGH,CRITICAL myapp:latest
This scans the image and lists vulnerabilities, helping you decide whether to update or patch.
For container health, configure a healthcheck in Dockerfile or run command:
HEALTHCHECK CMD curl --fail http://localhost/health || exit 1
Then docker ps shows health status. Use docker inspect --format='{{.State.Health.Status}}' <container> to get current status.
If metrics don't appear, check connectivity: from Prometheus container, curl cAdvisor endpoint. Ensure network connectivity and correct ports. Check daemon metrics endpoint is accessible from Prometheus host.
Use docker events to stream real-time events: docker events --filter 'event=die' to monitor container exits. This can help correlate alerts with incidents.
Failure Modes and Recovery
Containers fail for various reasons: application crashes, resource exhaustion, misconfiguration, or host issues. Monitoring should alert on these failures.
Common failure modes:
- Container exits unexpectedly: alert on
container_last_seenmetric or usedocker events. - High CPU/memory: alert on usage thresholds.
- Image pull failures: monitor registry availability and authentication.
- Disk space exhaustion from logs or volumes: alert on host disk usage.
- Network issues: container cannot reach dependent services.
Recovery procedures should be documented. For a crashed container, verify restart policy: docker inspect --format='{{.HostConfig.RestartPolicy.Name}}' <container>. If not set to always or unless-stopped, update with docker update --restart unless-stopped <container>.
For resource exhaustion, identify culprit: docker stats --no-stream shows CPU, memory, and I/O per container. Increase limits or optimize application. For memory leaks, consider restarting periodically or fix code.
If disk is full, clean up unused images and volumes: docker system prune -a --volumes (with caution, this deletes unused data). Set up log rotation as described earlier.
For image pull failures, check registry credentials: docker login and ensure image exists and tag correct. docker image ls to verify local images.
Create a runbook with steps for each failure mode. Include commands to run, expected outputs, and rollback steps. For example, to rollback a bad image deployment, tag previous version and redeploy:
docker tag myapp:previous myapp:latest
docker stack deploy -c docker-compose.yml myapp
Regularly test recovery in staging. Simulate failures (kill container, fill disk) and practice runbook.
Operations Checklist
Use this checklist for ongoing Docker image monitoring operations. Assign an owner to each item (e.g., DevOps engineer) and review frequency (e.g., weekly).
- [ ] Verify all critical containers are running and healthy (
docker ps). Owner: DevOps engineer. Frequency: daily. - [ ] Check metrics dashboards for anomalies (CPU, memory, network). Owner: On-call engineer. Frequency: shift.
- [ ] Ensure alerting rules are up-to-date and routing correctly. Owner: Monitoring admin. Frequency: monthly.
- [ ] Scan images for vulnerabilities with Trivy or similar. Owner: Security team. Frequency: weekly.
- [ ] Review and test recovery runbooks. Owner: DevOps lead. Frequency: quarterly.
- [ ] Monitor host disk and memory usage. Owner: System admin. Frequency: daily.
- [ ] Check daemon logs for errors (
journalctl -u docker). Owner: DevOps engineer. Frequency: daily. - [ ] Validate backup of persistent volumes. Owner: Backup admin. Frequency: daily.
- [ ] Review monitoring tooling performance and scalability. Owner: Architect. Frequency: monthly.
Each item should have a defined resolution path. For example, if disk usage exceeds 80%, alert and clean up logs or increase disk.
Use automation where possible. Schedule scripts with cron or CI/CD pipelines to run checks and send reports.
Common Pitfalls and How to Avoid Them
- Ignoring image size bloat: Large images slow deployments and increase attack surface. Monitor image size with
docker imagesand use multi-stage builds. Avoid installing unnecessary packages.
- Not setting resource limits: Without limits, a runaway container can starve others. Set limits in Compose or run:
docker run --memory=512m --cpus=1 myapp. Alert on usage approaching limits.
- Overlooking log management: Unmanaged logs can fill disks. Always configure log rotation and ship logs to a central system.
- Hardcoding secrets in images: Secrets in images are exposed if image is shared. Use Docker secrets or environment variables, and scan images for secrets.
- Not monitoring image pull failures: Pull failures can cause deployment outages. Monitor registry and set up alerts on pull errors.
- Ignoring base image vulnerabilities: Outdated base images can have known exploits. Regularly rebuild images with updated base and scan.
- Lack of healthchecks: Without healthchecks, Docker cannot detect unhealthy containers. Define healthchecks in Dockerfile or Compose.
- Not testing recovery: Untested runbooks fail in real incidents. Simulate failures and practice.
Conclusion
Effective Docker image monitoring requires a systematic approach: inventory your environment, configure safe monitoring, verify functionality, prepare for failure, and follow an operations checklist. By implementing the practices and commands in this guide, you can ensure your containerized applications run reliably and securely.
Start with one low-risk step: set up basic metrics scraping with cAdvisor and Prometheus. Then add alerts for critical conditions. As you gain confidence, expand to comprehensive monitoring with dashboards and automated responses.
Remember: monitoring is not just about collecting data; it's about making informed decisions quickly. A well-monitored Docker environment allows you to detect issues early and recover with minimal downtime.