## Intro

Kubernetes Deployment monitoring and alerts are critical for keeping applications healthy, but many teams stop at checking if Pods are running. A Deployment can appear healthy while rolling out broken code, exhausting CPU, or silently failing readiness probes. This guide gives you a practical, command-first approach to monitoring Deployments, setting meaningful alerts, and recovering quickly when things go wrong.

We will walk through a realistic example: the `web` Deployment running the `nginx:1.25` image, scaled to three replicas. You will learn how to inspect its current state, configure safe changes, verify rollouts, diagnose failures, and build a monitoring stack using kubectl, metrics-server, and Prometheus. Every command includes expected output and what to do when the result differs.

By the end, you will have an operations checklist you can adapt to your own cluster, plus common pitfalls and recovery patterns that save hours during an incident.

## 1. Version and Environment Inventory

Before you can monitor a Deployment, you need to know what is running and whether your tools support it. This step is all about read-only observation. No changes yet.

### Check your cluster version

```bash
kubectl version --short
```

Expected output:
```text
Client Version: v1.28.2
Kustomize Version: v5.0.4-0.20230601165947-6ce0bf390ce3
Server Version: v1.28.3
```

Why it matters: Deployment features like `maxSurge` and `maxUnavailable` work differently across versions. If your client is older than the server, some flags may be unavailable. For this guide, we assume Kubernetes 1.25 or later.

### Check the metrics API

Monitoring Deployments requires metrics. See if metrics-server is installed:

```bash
kubectl get apiservices | grep metrics
```

Expected output:
```text
v1beta1.metrics.k8s.io                 kube-system/metrics-server   True        24h
```

If the status is not `True`, install metrics-server before continuing. Without it, `kubectl top` will return `error: Metrics API not available`.

### Inspect the Deployment

List your Deployments in the target namespace:

```bash
kubectl get deployments -n production
```

Expected output:
```text
NAME   READY   UP-TO-DATE   AVAILABLE   AGE
web    3/3     3            3           5d
```

Now look at the full spec:

```bash
kubectl get deployment web -n production -o yaml
```

Key fields to record:
- `spec.replicas` (current desired count)
- `spec.strategy.type` (RollingUpdate or Recreate)
- `spec.template.spec.containers[].image` (image and tag)
- `spec.template.spec.containers[].resources` (limits and requests)
- `spec.template.spec.containers[].readinessProbe` and `livenessProbe`

Create a local inventory file with timestamps. Example:

```text
# Deployment inventory 2025-04-12 10:00 UTC
Deployment: web
Namespace: production
Replicas: 3
Strategy: RollingUpdate (maxSurge 25%, maxUnavailable 25%)
Image: nginx:1.25
CPU request/limit: 100m / 500m
Memory request/limit: 128Mi / 256Mi
Readiness probe: HTTP GET / on port 80, initialDelay 5s, period 10s
Liveness probe: HTTP GET /healthz on port 80, initialDelay 15s, period 20s
```

Record the timestamp. This inventory becomes your baseline for comparison after any change.

## 2. Safe Configuration Path

Now that you know the current state, you can make a scoped change safely. The principle: change one thing at a time, with a rollback plan.

### Before changing: capture rollback point

Save the current Deployment manifest:

```bash
kubectl get deployment web -n production -o yaml > web-deployment-backup.yaml
```

If something goes wrong, restore with:

```bash
kubectl apply -f web-deployment-backup.yaml
```

### Example change: update the container image

Suppose you need to update to `nginx:1.26`. Use the imperative command for a single change:

```bash
kubectl set image deployment/web nginx=nginx:1.26 -n production
```

Expected output:
```text
deployment.apps/web image updated
```

Alternatively, edit the manifest directly:

```bash
kubectl edit deployment web -n production
```

Change only the image tag, save, and exit. The Deployment controller will start a rolling update.

### Verify the rollout before assuming success

```bash
kubectl rollout status deployment/web -n production
```

Expected output while rolling out:
```text
Waiting for deployment "web" rollout to finish: 1 out of 3 new replicas have been updated...
```

On success:
```text
deployment "web" successfully rolled out
```

If the rollout hangs, inspect the new Pods immediately. Do not apply another change until you understand why (see Failure Modes).

### Keep local tests small

Before exposing the Deployment via a LoadBalancer or Ingress, verify it locally:

```bash
kubectl port-forward deployment/web 8080:80 -n production
```

Then open `http://localhost:8080` in a browser or use curl.

Expected result: HTTP 200 and the nginx welcome page. If you get connection refused, the Pod may not be listening on port 80 or the Service selector is wrong.

## 3. Verification and Diagnostics

After a change, you need to confirm the Deployment is healthy from multiple angles: Pods, ReplicaSets, resource usage, and events.

### Check Pod status and distribution

```bash
kubectl get pods -n production -l app=web -o wide
```

Expected output:
```text
NAME                   READY   STATUS    RESTARTS   AGE   IP           NODE
web-7d9f8c5b6-abcde   1/1     Running   0          2m    10.0.1.12    node-1
web-7d9f8c5b6-fghij   1/1     Running   0          2m    10.0.2.34    node-2
web-7d9f8c5b6-klmno   1/1     Running   0          2m    10.0.3.56    node-3
```

Check that Pods are spread across nodes if you have anti-affinity rules. If all Pods are on one node, a node failure will take down all replicas.

### Describe a Pod for events and probes

Pick one Pod:

```bash
kubectl describe pod web-7d9f8c5b6-abcde -n production
```

Look for:
- `Events:` section for image pull errors, probe failures, or scheduling issues
- `Conditions:` for `Ready` status
- `Containers:` readiness and liveness probe results

Example event indicating a failed readiness probe:
```text
Warning  Unhealthy  45s (x3 over 75s)  kubelet  Readiness probe failed: HTTP probe failed with statuscode: 503
```

This means the application inside the container is not ready to serve traffic, even if the container is running.

### Check logs

For a running Pod:

```bash
kubectl logs web-7d9f8c5b6-abcde -n production
```

For a crashed or restarted container, check the previous instance:

```bash
kubectl logs web-7d9f8c5b6-abcde -n production --previous
```

### Check current resource usage

```bash
kubectl top pods -n production -l app=web
```

Expected output (example):
```text
NAME                   CPU(cores)   MEMORY(bytes)
web-7d9f8c5b6-abcde   45m          120Mi
web-7d9f8c5b6-fghij   52m          118Mi
web-7d9f8c5b6-klmno   48m          122Mi
```

Compare against requests/limits. If CPU usage is near the limit, the Pod may be throttled, causing latency. If memory usage is near the limit, the Pod may be OOM-killed soon.

### Verify ReplicaSet history

```bash
kubectl get replicasets -n production -l app=web
```

Expected output:
```text
NAME              DESIRED   CURRENT   READY   AGE
web-7d9f8c5b6     3         3         3       2m
web-6c8b7d9f4e     0         0         0       5d
```

The old ReplicaSet with 0 replicas is kept for rollback. You can roll back to it with `kubectl rollout undo deployment/web`.

## 4. Failure Modes and Recovery

Deployments fail in predictable ways. Here are the most common failure modes, how to diagnose them, and how to recover.

### 4.1 CrashLoopBackOff

Symptom: Pod status shows `CrashLoopBackOff`.

Diagnose:

```bash
kubectl get pods -n production -l app=web
kubectl logs <pod-name> -n production --previous
kubectl describe pod <pod-name> -n production
```

The logs usually show the application error (e.g., missing config, wrong database URL). The describe output will show the exit code and events.

Recovery: Fix the underlying cause (config, image, command). If it was caused by a recent change, roll back:

```bash
kubectl rollout undo deployment/web -n production
```

### 4.2 ImagePullBackOff

Symptom: Pod status shows `ImagePullBackOff` or `ErrImagePull`.

Diagnose:

```bash
kubectl describe pod <pod-name> -n production | grep -A5 Events
```

Look for messages like `Failed to pull image "nginx:1.26": rpc error: code = NotFound` or `unauthorized: authentication required`.

Recovery: Correct the image tag, or configure image pull secrets if using a private registry:

```bash
kubectl create secret docker-registry regcred \
  --docker-server=myregistry.example.com \
  --docker-username=myuser \
  --docker-password=mypassword \
  --docker-email=myemail@example.com \
  -n production
```

Then add `imagePullSecrets` to the Deployment spec.

### 4.3 Rollout stuck

Symptom: `kubectl rollout status` hangs, and new Pods are not becoming ready.

Diagnose:

```bash
kubectl get pods -n production -l app=web
```

If new Pods are in `Pending` state, check scheduling events:

```bash
kubectl describe pod <new-pod-name> -n production
```

Common reasons: insufficient CPU/memory on nodes, node selectors that do not match any node, or persistent volume claims not bound.

Recovery: Adjust resource requests, or scale down the Deployment to fit available capacity, or remove/adjust node selectors. If it is a readiness probe failing, the Pod will run but not become Ready; check the probe and application logs.

### 4.4 Deployment not updating (no new ReplicaSet)

Symptom: After changing the spec, no new Pods are created.

Diagnose:

```bash
kubectl get deployment web -n production -o yaml
```

Check if the change was actually applied. Sometimes `kubectl apply` fails silently or you edited the wrong object.

Also check for a paused Deployment:

```bash
kubectl get deployment web -n production -o jsonpath='{.spec.paused}'`
```

If output is `true`, unpause:

```bash
kubectl rollout resume deployment/web -n production
```

### 4.5 Zero replicas after scale down

Symptom: You scaled to zero, but now the Deployment shows 0/0 and you cannot access the app.

Diagnose:

```bash
kubectl get deployment web -n production
```

Expected: `READY 0/0`.

Recovery: Scale back up:

```bash
kubectl scale deployment web --replicas=3 -n production
```

Then monitor rollout status.

## 5. Monitoring and Alerts

Manual checks are not enough. You need continuous monitoring and alerts that trigger before users notice. This section gives you a minimal Prometheus-based setup with example alerts for Deployments.

### 5.1 Metrics to collect

Prometheus with kube-state-metrics exposes Deployment-level metrics. Install kube-state-metrics and configure Prometheus to scrape it. Key metrics:

- `kube_deployment_spec_replicas`: desired number of replicas
- `kube_deployment_status_replicas_available`: currently available replicas
- `kube_deployment_status_replicas_unavailable`: currently unavailable replicas
- `kube_deployment_status_condition`: Deployment conditions (Available, Progressing, ReplicaFailure)
- `kube_deployment_metadata_generation`: generation of the Deployment spec
- `kube_deployment_status_observed_generation`: generation observed by the controller

### 5.2 Example alert rules

Create a Prometheus alert rule file `deployment-alerts.yaml`:

```yaml
groups:
- name: deployment
  rules:
  - alert: DeploymentReplicasMismatch
    expr: |
      kube_deployment_spec_replicas{namespace="production"}
      !=
      kube_deployment_status_replicas_available{namespace="production"}
    for: 5m
    labels:
      severity: warning
    annotations:
      summary: "Deployment {{ $labels.deployment }} has mismatched replicas"
      description: "Desired {{ $labels.deployment }} replicas: {{ $value }} available, expected {{ $labels.kube_deployment_spec_replicas }}. Check rollout status."

  - alert: DeploymentNotProgressing
    expr: |
      kube_deployment_status_condition{namespace="production", condition="Progressing", status="false"} == 1
    for: 10m
    labels:
      severity: critical
    annotations:
      summary: "Deployment {{ $labels.deployment }} is not progressing"
      description: "Deployment {{ $labels.deployment }} has been failing to progress for 10 minutes. Check events and pod logs."

  - alert: DeploymentReplicaFailure
    expr: |
      kube_deployment_status_condition{namespace="production", condition="ReplicaFailure", status="true"} == 1
    for: 5m
    labels:
      severity: critical
    annotations:
      summary: "Deployment {{ $labels.deployment }} has replica failure"
      description: "Deployment {{ $labels.deployment }} cannot create or maintain replicas. Check resource quotas, node capacity, and image pull errors."
```

Apply to Prometheus and reload. Test by intentionally scaling beyond capacity or changing the image to an invalid tag. You should see the alert fire within the `for` duration.

### 5.3 Dashboard example (Grafana)

Create a simple Grafana dashboard with these panels:

1. **Deployment Replicas (Desired vs Available)**: Query `kube_deployment_spec_replicas` and `kube_deployment_status_replicas_available` for the selected Deployment, visualize as two series.
2. **Deployment Rollout Status**: Use `kube_deployment_status_condition{condition="Progressing"}` to show 1 or 0 as a stat panel with green/red thresholds.
3. **Pod Restarts**: Query `rate(kube_pod_container_status_restarts_total{namespace="production"}[5m])` and alert if > 0.2.
4. **CPU and Memory Usage**: From metrics-server or node exporter, show per-Pod and sum per Deployment with limit lines.

### 5.4 Incident response runbook for Deployment alerts

When a Deployment alert fires, follow this order:

1. **Acknowledge** and note the time, alert name, and affected Deployment.
2. **Check rollout status**: `kubectl rollout status deployment/<name> -n <namespace>`.
3. **Check Pods**: `kubectl get pods -n <namespace> -l app=<label>` and look for CrashLoopBackOff, ImagePullBackOff, Pending.
4. **Check events**: `kubectl describe deployment <name> -n <namespace>` and read the last 10 events.
5. **Check metrics**: `kubectl top pods` and compare with limits.
6. **If recent change caused it, roll back**: `kubectl rollout undo deployment/<name> -n <namespace>`.
7. **If capacity issue, scale down or add nodes**.
8. **After recovery, document the root cause** and set a reminder to review in 24 hours.

Assign a single owner: the on-call engineer responding to the page owns the incident until resolved and is responsible for writing the post-mortem within 48 hours.

## 6. Common Pitfalls and How to Avoid Them

Here are mistakes teams frequently make when monitoring Deployments, with prevention and recovery tips.

### Pitfall 1: Monitoring only Pods, not Deployments

Why it happens: Pods are the visible unit; teams assume if Pods are running, the Deployment is healthy.
How to avoid: Monitor Deployment conditions (`Available`, `Progressing`, `ReplicaFailure`) and alert on those, not just Pod status. A Deployment can have running Pods but be failing to progress (e.g., stuck rollout, old ReplicaSet still serving).
Recovery: Use the alert `DeploymentNotProgressing` above and check `kubectl rollout status` immediately.

### Pitfall 2: Ignoring readiness probe failures during rollout

Why it happens: Teams often only set liveness probes, or configure readiness probes too loosely. The Deployment controller then considers new Pods ready even if they cannot serve traffic, leading to partial outages.
How to avoid: Set appropriate readiness probes and use `kubectl rollout status` to block CI/CD until rollout succeeds. Use `minReadySeconds` to give probes time to stabilize.
Recovery: If a rollout is in progress and Pods are not ready, investigate the readiness endpoint. Roll back if necessary.

### Pitfall 3: Not having a rollback plan

Why it happens: Teams apply changes directly and hope for the best.
How to avoid: Always save the previous manifest (`kubectl get deployment <name> -o yaml > backup.yaml`) before changing. Use version control for manifests.
Recovery: If change fails, `kubectl rollout undo deployment/<name>` or `kubectl apply -f backup.yaml`.

### Pitfall 4: Alert fatigue with too many noisy alerts

Why it happens: Teams create alerts for every minor metric fluctuation.
How to avoid: Use `for` clauses (e.g., 5 minutes) to avoid transient alerts. Set meaningful thresholds based on historical baselines. Route alerts to appropriate channels (warning vs critical).
Recovery: Tune alert rules regularly; each alert should have a clear action. If no action is needed, remove the alert.

### Pitfall 5: Monitoring without metrics-server or Prometheus

Why it happens: Some clusters do not have metrics-server installed, or teams rely on manual `kubectl top`.
How to avoid: Ensure metrics-server is installed and functioning. For production, set up Prometheus and kube-state-metrics for historical and condition-based alerting.
Recovery: If metrics API is missing, install metrics-server; if alerts are missing, install kube-state-metrics and configure rules.

## 7. Operations Checklist

Use this checklist for every Deployment change or incident. Assign a single accountable owner for each major step (example: Priya Shah, Engineering Lead).

### Pre-change checklist
- [ ] Record current Deployment spec and timestamp (Owner: Priya Shah)
- [ ] Save backup manifest: `kubectl get deployment <name> -n <ns> -o yaml > backup.yaml` (Owner: on-call engineer)
- [ ] Verify metrics-server is available: `kubectl get apiservices | grep metrics` (Owner: Priya Shah)
- [ ] Identify single change and its expected effect (Owner: change requester)
- [ ] Set rollback criteria (e.g., if rollout not complete in 10 minutes, roll back) (Owner: Priya Shah)

### Post-change verification
- [ ] Run `kubectl rollout status deployment/<name> -n <ns>` and confirm success (Owner: on-call engineer)
- [ ] Check Pods: `kubectl get pods -n <ns> -l app=<label>` all Running and Ready (Owner: on-call engineer)
- [ ] Check logs for errors (Owner: developer)
- [ ] Check `kubectl top pods` for resource usage within limits (Owner: on-call engineer)
- [ ] Verify traffic via local port-forward or Service endpoint (Owner: QA engineer)
- [ ] Confirm alerts not firing (Owner: monitoring admin)

### Incident recovery checklist
- [ ] Acknowledge alert and note time (Owner: on-call engineer)
- [ ] Check rollout status and recent Events (Owner: on-call engineer)
- [ ] Determine if recent change caused incident (Owner: on-call engineer)
- [ ] Roll back if necessary: `kubectl rollout undo deployment/<name> -n <ns>` (Owner: on-call engineer)
- [ ] Scale down or add capacity if resource-related (Owner: on-call engineer)
- [ ] Document root cause and post-mortem within 48 hours (Owner: incident commander)
- [ ] Schedule review of incident and alert thresholds in next weekly ops meeting (Owner: Priya Shah)

Review frequency: The checklist and alert thresholds should be reviewed monthly, or after any major incident.

## Conclusion

Kubernetes Deployment monitoring is more than checking `kubectl get pods`. It requires a systematic approach: know your version and environment, change configurations safely, verify rollouts with the right commands, anticipate common failures, and set up automated alerts that tell you when something is wrong.

The practical examples in this guide—backup manifests, rollout status checks, Prometheus alert rules, and a detailed checklist—give you a solid foundation. Start by implementing the three basic alerts (ReplicasMismatch, NotProgressing, ReplicaFailure) and the operations checklist. Test your alerts by intentionally breaking a Deployment in a staging environment. Then practice rollbacks. When an incident occurs, you will have the muscle memory to respond quickly and safely.

Remember: the goal is not just to detect problems, but to make recovery fast and predictable. Keep your monitoring focused on actionable signals, assign clear ownership, and review your approach regularly.