Intro
Kubernetes Deployment monitoring and alerts are critical for keeping applications healthy, but many teams stop at checking if Pods are running. A Deployment can appear healthy while rolling out broken code, exhausting CPU, or silently failing readiness probes. This guide gives you a practical, command-first approach to monitoring Deployments, setting meaningful alerts, and recovering quickly when things go wrong.
We will walk through a realistic example: the web Deployment running the nginx:1.25 image, scaled to three replicas. You will learn how to inspect its current state, configure safe changes, verify rollouts, diagnose failures, and build a monitoring stack using kubectl, metrics-server, and Prometheus. Every command includes expected output and what to do when the result differs.
By the end, you will have an operations checklist you can adapt to your own cluster, plus common pitfalls and recovery patterns that save hours during an incident.
1. Version and Environment Inventory
Before you can monitor a Deployment, you need to know what is running and whether your tools support it. This step is all about read-only observation. No changes yet.
Check your cluster version
kubectl version --short
Expected output:
Client Version: v1.28.2
Kustomize Version: v5.0.4-0.20230601165947-6ce0bf390ce3
Server Version: v1.28.3
Why it matters: Deployment features like maxSurge and maxUnavailable work differently across versions. If your client is older than the server, some flags may be unavailable. For this guide, we assume Kubernetes 1.25 or later.
Check the metrics API
Monitoring Deployments requires metrics. See if metrics-server is installed:
kubectl get apiservices | grep metrics
Expected output:
v1beta1.metrics.k8s.io kube-system/metrics-server True 24h
If the status is not True, install metrics-server before continuing. Without it, kubectl top will return error: Metrics API not available.
Inspect the Deployment
List your Deployments in the target namespace:
kubectl get deployments -n production
Expected output:
NAME READY UP-TO-DATE AVAILABLE AGE
web 3/3 3 3 5d
Now look at the full spec:
kubectl get deployment web -n production -o yaml
Key fields to record:
spec.replicas(current desired count)spec.strategy.type(RollingUpdate or Recreate)spec.template.spec.containers[].image(image and tag)spec.template.spec.containers[].resources(limits and requests)spec.template.spec.containers[].readinessProbeandlivenessProbe
Create a local inventory file with timestamps. Example:
# Deployment inventory 2025-04-12 10:00 UTC
Deployment: web
Namespace: production
Replicas: 3
Strategy: RollingUpdate (maxSurge 25%, maxUnavailable 25%)
Image: nginx:1.25
CPU request/limit: 100m / 500m
Memory request/limit: 128Mi / 256Mi
Readiness probe: HTTP GET / on port 80, initialDelay 5s, period 10s
Liveness probe: HTTP GET /healthz on port 80, initialDelay 15s, period 20s
Record the timestamp. This inventory becomes your baseline for comparison after any change.
2. Safe Configuration Path
Now that you know the current state, you can make a scoped change safely. The principle: change one thing at a time, with a rollback plan.
Before changing: capture rollback point
Save the current Deployment manifest:
kubectl get deployment web -n production -o yaml > web-deployment-backup.yaml
If something goes wrong, restore with:
kubectl apply -f web-deployment-backup.yaml
Example change: update the container image
Suppose you need to update to nginx:1.26. Use the imperative command for a single change:
kubectl set image deployment/web nginx=nginx:1.26 -n production
Expected output:
deployment.apps/web image updated
Alternatively, edit the manifest directly:
kubectl edit deployment web -n production
Change only the image tag, save, and exit. The Deployment controller will start a rolling update.
Verify the rollout before assuming success
kubectl rollout status deployment/web -n production
Expected output while rolling out:
Waiting for deployment "web" rollout to finish: 1 out of 3 new replicas have been updated...
On success:
deployment "web" successfully rolled out
If the rollout hangs, inspect the new Pods immediately. Do not apply another change until you understand why (see Failure Modes).
Keep local tests small
Before exposing the Deployment via a LoadBalancer or Ingress, verify it locally:
kubectl port-forward deployment/web 8080:80 -n production
Then open http://localhost:8080 in a browser or use curl.
Expected result: HTTP 200 and the nginx welcome page. If you get connection refused, the Pod may not be listening on port 80 or the Service selector is wrong.
3. Verification and Diagnostics
After a change, you need to confirm the Deployment is healthy from multiple angles: Pods, ReplicaSets, resource usage, and events.
Check Pod status and distribution
kubectl get pods -n production -l app=web -o wide
Expected output:
NAME READY STATUS RESTARTS AGE IP NODE
web-7d9f8c5b6-abcde 1/1 Running 0 2m 10.0.1.12 node-1
web-7d9f8c5b6-fghij 1/1 Running 0 2m 10.0.2.34 node-2
web-7d9f8c5b6-klmno 1/1 Running 0 2m 10.0.3.56 node-3
Check that Pods are spread across nodes if you have anti-affinity rules. If all Pods are on one node, a node failure will take down all replicas.
Describe a Pod for events and probes
Pick one Pod:
kubectl describe pod web-7d9f8c5b6-abcde -n production
Look for:
Events:section for image pull errors, probe failures, or scheduling issuesConditions:forReadystatusContainers:readiness and liveness probe results
Example event indicating a failed readiness probe:
Warning Unhealthy 45s (x3 over 75s) kubelet Readiness probe failed: HTTP probe failed with statuscode: 503
This means the application inside the container is not ready to serve traffic, even if the container is running.
Check logs
For a running Pod:
kubectl logs web-7d9f8c5b6-abcde -n production
For a crashed or restarted container, check the previous instance:
kubectl logs web-7d9f8c5b6-abcde -n production --previous
Check current resource usage
kubectl top pods -n production -l app=web
Expected output (example):
NAME CPU(cores) MEMORY(bytes)
web-7d9f8c5b6-abcde 45m 120Mi
web-7d9f8c5b6-fghij 52m 118Mi
web-7d9f8c5b6-klmno 48m 122Mi
Compare against requests/limits. If CPU usage is near the limit, the Pod may be throttled, causing latency. If memory usage is near the limit, the Pod may be OOM-killed soon.
Verify ReplicaSet history
kubectl get replicasets -n production -l app=web
Expected output:
NAME DESIRED CURRENT READY AGE
web-7d9f8c5b6 3 3 3 2m
web-6c8b7d9f4e 0 0 0 5d
The old ReplicaSet with 0 replicas is kept for rollback. You can roll back to it with kubectl rollout undo deployment/web.
4. Failure Modes and Recovery
Deployments fail in predictable ways. Here are the most common failure modes, how to diagnose them, and how to recover.
4.1 CrashLoopBackOff
Symptom: Pod status shows CrashLoopBackOff.
Diagnose:
kubectl get pods -n production -l app=web
kubectl logs <pod-name> -n production --previous
kubectl describe pod <pod-name> -n production
The logs usually show the application error (e.g., missing config, wrong database URL). The describe output will show the exit code and events.
Recovery: Fix the underlying cause (config, image, command). If it was caused by a recent change, roll back:
kubectl rollout undo deployment/web -n production
4.2 ImagePullBackOff
Symptom: Pod status shows ImagePullBackOff or ErrImagePull.
Diagnose:
kubectl describe pod <pod-name> -n production | grep -A5 Events
Look for messages like Failed to pull image "nginx:1.26": rpc error: code = NotFound or unauthorized: authentication required.
Recovery: Correct the image tag, or configure image pull secrets if using a private registry:
kubectl create secret docker-registry regcred \
--docker-server=myregistry.example.com \
--docker-username=myuser \
--docker-password=mypassword \
[email protected] \
-n production
Then add imagePullSecrets to the Deployment spec.
4.3 Rollout stuck
Symptom: kubectl rollout status hangs, and new Pods are not becoming ready.
Diagnose:
kubectl get pods -n production -l app=web
If new Pods are in Pending state, check scheduling events:
kubectl describe pod <new-pod-name> -n production
Common reasons: insufficient CPU/memory on nodes, node selectors that do not match any node, or persistent volume claims not bound.
Recovery: Adjust resource requests, or scale down the Deployment to fit available capacity, or remove/adjust node selectors. If it is a readiness probe failing, the Pod will run but not become Ready; check the probe and application logs.
4.4 Deployment not updating (no new ReplicaSet)
Symptom: After changing the spec, no new Pods are created.
Diagnose:
kubectl get deployment web -n production -o yaml
Check if the change was actually applied. Sometimes kubectl apply fails silently or you edited the wrong object.
Also check for a paused Deployment:
kubectl get deployment web -n production -o jsonpath='{.spec.paused}'`
If output is true, unpause:
kubectl rollout resume deployment/web -n production
4.5 Zero replicas after scale down
Symptom: You scaled to zero, but now the Deployment shows 0/0 and you cannot access the app.
Diagnose:
kubectl get deployment web -n production
Expected: READY 0/0.
Recovery: Scale back up:
kubectl scale deployment web --replicas=3 -n production
Then monitor rollout status.
5. Monitoring and Alerts
Manual checks are not enough. You need continuous monitoring and alerts that trigger before users notice. This section gives you a minimal Prometheus-based setup with example alerts for Deployments.
5.1 Metrics to collect
Prometheus with kube-state-metrics exposes Deployment-level metrics. Install kube-state-metrics and configure Prometheus to scrape it. Key metrics:
kube_deployment_spec_replicas: desired number of replicaskube_deployment_status_replicas_available: currently available replicaskube_deployment_status_replicas_unavailable: currently unavailable replicaskube_deployment_status_condition: Deployment conditions (Available, Progressing, ReplicaFailure)kube_deployment_metadata_generation: generation of the Deployment speckube_deployment_status_observed_generation: generation observed by the controller
5.2 Example alert rules
Create a Prometheus alert rule file deployment-alerts.yaml:
groups:
- name: deployment
rules:
- alert: DeploymentReplicasMismatch
expr: |
kube_deployment_spec_replicas{namespace="production"}
!=
kube_deployment_status_replicas_available{namespace="production"}
for: 5m
labels:
severity: warning
annotations:
summary: "Deployment {{ $labels.deployment }} has mismatched replicas"
description: "Desired {{ $labels.deployment }} replicas: {{ $value }} available, expected {{ $labels.kube_deployment_spec_replicas }}. Check rollout status."
- alert: DeploymentNotProgressing
expr: |
kube_deployment_status_condition{namespace="production", condition="Progressing", status="false"} == 1
for: 10m
labels:
severity: critical
annotations:
summary: "Deployment {{ $labels.deployment }} is not progressing"
description: "Deployment {{ $labels.deployment }} has been failing to progress for 10 minutes. Check events and pod logs."
- alert: DeploymentReplicaFailure
expr: |
kube_deployment_status_condition{namespace="production", condition="ReplicaFailure", status="true"} == 1
for: 5m
labels:
severity: critical
annotations:
summary: "Deployment {{ $labels.deployment }} has replica failure"
description: "Deployment {{ $labels.deployment }} cannot create or maintain replicas. Check resource quotas, node capacity, and image pull errors."
Apply to Prometheus and reload. Test by intentionally scaling beyond capacity or changing the image to an invalid tag. You should see the alert fire within the for duration.
5.3 Dashboard example (Grafana)
Create a simple Grafana dashboard with these panels:
- Deployment Replicas (Desired vs Available): Query
kube_deployment_spec_replicasandkube_deployment_status_replicas_availablefor the selected Deployment, visualize as two series. - Deployment Rollout Status: Use
kube_deployment_status_condition{condition="Progressing"}to show 1 or 0 as a stat panel with green/red thresholds. - Pod Restarts: Query
rate(kube_pod_container_status_restarts_total{namespace="production"}[5m])and alert if > 0.2. - CPU and Memory Usage: From metrics-server or node exporter, show per-Pod and sum per Deployment with limit lines.
5.4 Incident response runbook for Deployment alerts
When a Deployment alert fires, follow this order:
- Acknowledge and note the time, alert name, and affected Deployment.
- Check rollout status:
kubectl rollout status deployment/<name> -n <namespace>. - Check Pods:
kubectl get pods -n <namespace> -l app=<label>and look for CrashLoopBackOff, ImagePullBackOff, Pending. - Check events:
kubectl describe deployment <name> -n <namespace>and read the last 10 events. - Check metrics:
kubectl top podsand compare with limits. - If recent change caused it, roll back:
kubectl rollout undo deployment/<name> -n <namespace>. - If capacity issue, scale down or add nodes.
- After recovery, document the root cause and set a reminder to review in 24 hours.
Assign a single owner: the on-call engineer responding to the page owns the incident until resolved and is responsible for writing the post-mortem within 48 hours.
6. Common Pitfalls and How to Avoid Them
Here are mistakes teams frequently make when monitoring Deployments, with prevention and recovery tips.
Pitfall 1: Monitoring only Pods, not Deployments
Why it happens: Pods are the visible unit; teams assume if Pods are running, the Deployment is healthy. How to avoid: Monitor Deployment conditions (Available, Progressing, ReplicaFailure) and alert on those, not just Pod status. A Deployment can have running Pods but be failing to progress (e.g., stuck rollout, old ReplicaSet still serving). Recovery: Use the alert DeploymentNotProgressing above and check kubectl rollout status immediately.
Pitfall 2: Ignoring readiness probe failures during rollout
Why it happens: Teams often only set liveness probes, or configure readiness probes too loosely. The Deployment controller then considers new Pods ready even if they cannot serve traffic, leading to partial outages. How to avoid: Set appropriate readiness probes and use kubectl rollout status to block CI/CD until rollout succeeds. Use minReadySeconds to give probes time to stabilize. Recovery: If a rollout is in progress and Pods are not ready, investigate the readiness endpoint. Roll back if necessary.
Pitfall 3: Not having a rollback plan
Why it happens: Teams apply changes directly and hope for the best. How to avoid: Always save the previous manifest (kubectl get deployment <name> -o yaml > backup.yaml) before changing. Use version control for manifests. Recovery: If change fails, kubectl rollout undo deployment/<name> or kubectl apply -f backup.yaml.
Pitfall 4: Alert fatigue with too many noisy alerts
Why it happens: Teams create alerts for every minor metric fluctuation. How to avoid: Use for clauses (e.g., 5 minutes) to avoid transient alerts. Set meaningful thresholds based on historical baselines. Route alerts to appropriate channels (warning vs critical). Recovery: Tune alert rules regularly; each alert should have a clear action. If no action is needed, remove the alert.
Pitfall 5: Monitoring without metrics-server or Prometheus
Why it happens: Some clusters do not have metrics-server installed, or teams rely on manual kubectl top. How to avoid: Ensure metrics-server is installed and functioning. For production, set up Prometheus and kube-state-metrics for historical and condition-based alerting. Recovery: If metrics API is missing, install metrics-server; if alerts are missing, install kube-state-metrics and configure rules.
7. Operations Checklist
Use this checklist for every Deployment change or incident. Assign a single accountable owner for each major step (example: Priya Shah, Engineering Lead).
Pre-change checklist
- [ ] Record current Deployment spec and timestamp (Owner: Priya Shah)
- [ ] Save backup manifest:
kubectl get deployment <name> -n <ns> -o yaml > backup.yaml(Owner: on-call engineer) - [ ] Verify metrics-server is available:
kubectl get apiservices | grep metrics(Owner: Priya Shah) - [ ] Identify single change and its expected effect (Owner: change requester)
- [ ] Set rollback criteria (e.g., if rollout not complete in 10 minutes, roll back) (Owner: Priya Shah)
Post-change verification
- [ ] Run
kubectl rollout status deployment/<name> -n <ns>and confirm success (Owner: on-call engineer) - [ ] Check Pods:
kubectl get pods -n <ns> -l app=<label>all Running and Ready (Owner: on-call engineer) - [ ] Check logs for errors (Owner: developer)
- [ ] Check
kubectl top podsfor resource usage within limits (Owner: on-call engineer) - [ ] Verify traffic via local port-forward or Service endpoint (Owner: QA engineer)
- [ ] Confirm alerts not firing (Owner: monitoring admin)
Incident recovery checklist
- [ ] Acknowledge alert and note time (Owner: on-call engineer)
- [ ] Check rollout status and recent Events (Owner: on-call engineer)
- [ ] Determine if recent change caused incident (Owner: on-call engineer)
- [ ] Roll back if necessary:
kubectl rollout undo deployment/<name> -n <ns>(Owner: on-call engineer) - [ ] Scale down or add capacity if resource-related (Owner: on-call engineer)
- [ ] Document root cause and post-mortem within 48 hours (Owner: incident commander)
- [ ] Schedule review of incident and alert thresholds in next weekly ops meeting (Owner: Priya Shah)
Review frequency: The checklist and alert thresholds should be reviewed monthly, or after any major incident.
Conclusion
Kubernetes Deployment monitoring is more than checking kubectl get pods. It requires a systematic approach: know your version and environment, change configurations safely, verify rollouts with the right commands, anticipate common failures, and set up automated alerts that tell you when something is wrong.
The practical examples in this guide—backup manifests, rollout status checks, Prometheus alert rules, and a detailed checklist—give you a solid foundation. Start by implementing the three basic alerts (ReplicasMismatch, NotProgressing, ReplicaFailure) and the operations checklist. Test your alerts by intentionally breaking a Deployment in a staging environment. Then practice rollbacks. When an incident occurs, you will have the muscle memory to respond quickly and safely.
Remember: the goal is not just to detect problems, but to make recovery fast and predictable. Keep your monitoring focused on actionable signals, assign clear ownership, and review your approach regularly.