## Intro

Etcd is the distributed key-value store at the heart of every Kubernetes cluster. It holds all cluster state, including Pods, Services, ConfigMaps, and Secrets. If etcd becomes unavailable or loses data, the entire control plane can fail. Monitoring and alerting for etcd are therefore critical for operational safety.

This guide focuses on practical configuration and upgrade of etcd monitoring and alerts in Kubernetes. It is written for developers, DevOps consultants, and technical startup teams who operate their own clusters and need to move from observing a problem to verifying a fix.

You will learn how to:

- Inventory your etcd version and deployment topology.
- Expose and collect etcd metrics.
- Build dashboards that show key health indicators.
- Configure alerts that catch failure before users notice.
- Upgrade etcd safely with monitoring in place.
- Respond to incidents with a clear diagnostic path.

Every section includes concrete commands, expected output, failure signals, and recovery decisions. The goal is operational safety: observe before changing, limit the blast radius, use placeholders instead of secrets, verify the result, and document recovery before an incident forces it.

## Version and Environment Inventory

Before touching monitoring or upgrading etcd, you must know exactly what you are running. Start with a read-only inventory of the cluster and the etcd component.

Identify the etcd version:

```bash
kubectl get pods -n kube-system -l component=etcd -o jsonpath='{.items[0].spec.containers[0].image}'
```

Expected output example:

```
registry.k8s.io/etcd:3.5.9-0
```

Determine the etcd deployment topology. Most kubeadm clusters run etcd as a static Pod on each control-plane node. List the etcd Pods:

```bash
kubectl get pods -n kube-system -l component=etcd -o wide
```

Expected output example:

```
NAME                READY   STATUS    RESTARTS   AGE   IP              NODE
etcd-control-plane-1   1/1     Running   0          10m   192.168.1.10   control-plane-1
etcd-control-plane-2   1/1     Running   0          10m   192.168.1.11   control-plane-2
etcd-control-plane-3   1/1     Running   0          10m   192.168.1.12   control-plane-3
```

Check prerequisites:

- `kubectl` access to the cluster with permissions to get Pods in `kube-system`.
- Ability to `exec` into the etcd container or access the node filesystem if needed.
- Network access from your monitoring stack to the etcd metrics endpoint.

Perform a read-only health check using `etcdctl` inside the Pod:

```bash
kubectl exec -n kube-system etcd-control-plane-1 -- etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/peer.crt --key=/etc/kubernetes/pki/etcd/peer.key endpoint health
```

Expected healthy output:

```
https://127.0.0.1:2379 is healthy: successfully committed proposal: took = 2.1ms
```

If the cluster uses an external etcd (not static Pods), adjust commands accordingly. Record the etcd version, number of members, and health status in your operations log with timestamps. This baseline is essential before any change.

## Safe Configuration Path

Configuring etcd monitoring requires enabling the metrics endpoint and ensuring your Prometheus or other scraper can reach it securely.

Etcd exposes metrics on port 2379 by default when started with the `--listen-metrics-urls` flag. Check the current etcd Pod spec for the metrics URL:

```bash
kubectl get pod -n kube-system etcd-control-plane-1 -o yaml | grep -A2 'listen-metrics-urls'
```

Expected output may include:

```yaml
    - --listen-metrics-urls=http://0.0.0.0:2381
```

If the flag is missing, etcd still exposes metrics on port 2379, but that port is also used for client requests. It is safer to separate metrics traffic on a dedicated port (e.g., 2381) to avoid interference.

To enable a dedicated metrics port on a static Pod managed by kubeadm, edit the etcd manifest on the control-plane node:

```bash
sudo vi /etc/kubernetes/manifests/etcd.yaml
```

Add or modify the container command arguments:

```yaml
spec:
  containers:
  - command:
    - etcd
    - --listen-metrics-urls=http://0.0.0.0:2381
    - --metrics=extensive
```

The `--metrics=extensive` option provides additional metrics like gRPC request histograms, which are useful for latency analysis.

After saving the manifest, kubelet will restart the etcd Pod automatically. Verify the new port is listening:

```bash
kubectl exec -n kube-system etcd-control-plane-1 -- sh -c 'curl -s http://localhost:2381/metrics | head -5'
```

Expected output:

```
# HELP etcd_server_has_leader Whether or not a leader exists. 1 is leader exists.
# TYPE etcd_server_has_leader gauge
etcd_server_has_leader 1
```

Now configure Prometheus to scrape this endpoint. If you use the Prometheus Operator, create a ServiceMonitor targeting the etcd Pods. If using a raw Prometheus config, add a scrape job:

```yaml
scrape_configs:
  - job_name: 'etcd'
    scheme: https
    tls_config:
      ca_file: /etc/prometheus/secrets/etcd-ca.crt
      cert_file: /etc/prometheus/secrets/etcd-client.crt
      key_file: /etc/prometheus/secrets/etcd-client.key
    static_configs:
      - targets: ['192.168.1.10:2379', '192.168.1.11:2379', '192.168.1.12:2379']
```

If you enabled the dedicated metrics port, use that port instead. Ensure the certificates have the correct Subject Alternative Names for the node IPs.

Test the scrape manually from the Prometheus Pod:

```bash
kubectl exec -n monitoring prometheus-0 -- wget -qO- https://192.168.1.10:2379/metrics --ca-certificate=/etc/prometheus/secrets/etcd-ca.crt | head
```

Expected output shows Prometheus metrics. If you get a certificate error, check the certificate validity and SANs.

## Verification and Diagnostics

After enabling metrics collection, verify that Prometheus is receiving etcd metrics and that they are not empty.

Query Prometheus for the `up` metric for etcd targets:

```
up{job="etcd"}
```

Expected result: all etcd targets should show `1`.

Check for any metric sample, e.g., `etcd_server_has_leader`:

```
etcd_server_has_leader
```

Expected result: the value is `1` on the leader node and `0` on followers. If the metric is absent, check the scrape configuration and network policies.

Diagnose common issues:

- **Target down in Prometheus**: Verify the endpoint is reachable from Prometheus Pod. Check firewall rules or Kubernetes NetworkPolicy.
- **Certificate errors**: Ensure the CA and client certs in Prometheus match the etcd server certs. `etcd` uses mutual TLS by default in kubeadm clusters.
- **No metrics on port 2381**: Confirm the `--listen-metrics-urls` flag was applied. Check etcd Pod logs for startup errors:

```bash
kubectl logs -n kube-system etcd-control-plane-1
```

You can also use `etcdctl` to check cluster health and member list to ensure no split brain:

```bash
kubectl exec -n kube-system etcd-control-plane-1 -- etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key member list
```

Expected output lists all members with their client URLs and peer URLs.

If metrics are flowing, you can now build dashboards. For Grafana, import a dashboard that uses etcd metrics. A popular choice is the etcd dashboard from the Kubernetes Mixins or the Grafana community. Key panels to include:

- **Leader changes**: `increase(etcd_server_leader_changes_seen_total[1h])` should be low. Frequent changes indicate network issues.
- **Raft proposal duration**: `histogram_quantile(0.99, sum(rate(etcd_disk_wal_fsync_duration_seconds_bucket[5m])) by (le))` should be less than 25ms.
- **Disk operation duration**: `histogram_quantile(0.99, sum(rate(etcd_disk_backend_commit_duration_seconds_bucket[5m])) by (le))` less than 50ms.
- **gRPC request rate**: `sum(rate(grpc_server_handled_total[5m])) by (grpc_method)` to see request load.

## Failure Modes and Recovery

Etcd failures can be subtle. Knowing the common failure modes helps you react quickly.

### Leader Election Storms

If the etcd leader changes frequently (more than once per hour), it indicates network latency or packet loss between members. This can cause write timeouts for the API server.

**Diagnosis**: Check `increase(etcd_server_leader_changes_seen_total[1h])`. Also check network latency between nodes with `ping` or `iperf`.

**Recovery**: Investigate network performance. If caused by high disk I/O on the leader, reduce load or move etcd to dedicated disks. If network issue persists, consider adding a faster network or adjusting etcd heartbeat interval (not recommended without deep knowledge).

### Disk Space Exhaustion

Etcd stores data in a directory. If the disk fills up, etcd may stop accepting writes or crash.

**Diagnosis**: Check disk usage on the node:

```bash
df -h /var/lib/etcd
```

Expected output example:

```
Filesystem      Size  Used Avail Use% Mounted on
/dev/sda1       100G   95G   5G  95% /var/lib/etcd
```

If usage is above 80%, plan for compaction and defragmentation.

**Recovery**: Compact the etcd history to remove old revisions:

```bash
kubectl exec -n kube-system etcd-control-plane-1 -- etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key compact $(etcdctl ... get / --prefix --keys-only | wc -l)
```

Better to use the current revision minus some safety margin. Then defragment:

```bash
kubectl exec -n kube-system etcd-control-plane-1 -- etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key defrag
```

Run defrag on each member sequentially to avoid quorum loss. Schedule regular compaction via a CronJob if your etcd version does not auto-compact.

### Quorum Loss

If more than (n/2) members fail, the cluster loses quorum and cannot process writes. Read-only requests may still work depending on configuration.

**Diagnosis**: Check `etcd_server_has_leader` metric; if it is 0 on all nodes, quorum is lost. Also run `etcdctl endpoint status` to see leader and raft term.

**Recovery**: Restore the failed members as soon as possible. If the majority is permanently lost, you must perform an etcd disaster recovery from a snapshot. Always take regular snapshots:

```bash
kubectl exec -n kube-system etcd-control-plane-1 -- etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key snapshot save /var/lib/etcd/snapshot.db
```

Store snapshots off-node. Test restoration in a non-production environment regularly.

## Upgrading Etcd with Monitoring in Place

Upgrading etcd requires careful planning. Monitor the cluster during the upgrade to catch issues early.

**Pre-upgrade checklist:**

- Verify current version and cluster health (see Version and Environment Inventory).
- Take a snapshot of the etcd data.
- Check compatibility: Kubernetes supports etcd 3.5.x for recent versions; do not jump major versions without testing.
- Notify stakeholders and schedule a maintenance window if necessary.

For kubeadm clusters, upgrading etcd is part of the control-plane upgrade. First, upgrade kubeadm:

```bash
sudo apt-get update && sudo apt-get install -y kubeadm=1.28.0-00
```

Then run the upgrade plan:

```bash
sudo kubeadm upgrade plan
```

This will show the recommended etcd version. Apply the upgrade on the first control-plane node:

```bash
sudo kubeadm upgrade apply v1.28.0
```

This upgrades etcd, kube-apiserver, kube-controller-manager, and kube-scheduler on that node. Monitor etcd metrics during the process. Watch for leader changes and quorum loss. The cluster should still function because of redundancy.

After the first node is upgraded, upgrade other control-plane nodes one by one:

```bash
sudo kubeadm upgrade node
```

Monitor etcd after each node. Check `etcd_server_has_leader` and `etcd_server_leader_changes_seen_total`. Ensure metrics are still scraped.

If you use an external etcd, follow the etcd documentation for in-place upgrades. Never upgrade more than one member at a time, and allow the cluster to stabilize between upgrades.

**Post-upgrade verification:**

- Check etcd version: `kubectl exec ... etcdctl version`
- Check cluster health: `etcdctl endpoint health`
- Check metrics in Prometheus: `up{job="etcd"}` is 1.
- Check dashboards for anomalies.

If an upgrade fails, you can roll back by restoring the snapshot if necessary. However, etcd upgrades are usually backward compatible within the same major version. Rolling back Kubernetes is more complex; always test in staging.

## Configuring Alerts

Set up Prometheus alerts for etcd. Use the following alert rules as a starting point. Adjust thresholds based on your environment.

Create an alert rule file (e.g., `etcd-alerts.yaml`) and load it into Prometheus.

```yaml
groups:
- name: etcd
  rules:
  - alert: EtcdNoLeader
    expr: etcd_server_has_leader == 0
    for: 5m
    labels:
      severity: critical
    annotations:
      summary: "etcd cluster has no leader"
      description: "etcd instance {{ $labels.instance }} has no leader for more than 5 minutes."

  - alert: EtcdHighFsyncDuration
    expr: histogram_quantile(0.99, sum(rate(etcd_disk_wal_fsync_duration_seconds_bucket[5m])) by (le)) > 0.025
    for: 10m
    labels:
      severity: warning
    annotations:
      summary: "etcd fsync latency high"
      description: "etcd instance {{ $labels.instance }} 99th percentile fsync duration is above 25ms for 10 minutes."

  - alert: EtcdHighCommitDuration
    expr: histogram_quantile(0.99, sum(rate(etcd_disk_backend_commit_duration_seconds_bucket[5m])) by (le)) > 0.05
    for: 10m
    labels:
      severity: warning
    annotations:
      summary: "etcd commit latency high"
      description: "etcd instance {{ $labels.instance }} 99th percentile commit duration is above 50ms for 10 minutes."

  - alert: EtcdLeaderChangesFrequent
    expr: increase(etcd_server_leader_changes_seen_total[1h]) > 3
    for: 5m
    labels:
      severity: warning
    annotations:
      summary: "etcd leader changes are frequent"
      description: "etcd cluster has had more than 3 leader changes in the last hour."

  - alert: EtcdMemberDown
    expr: up{job="etcd"} == 0
    for: 5m
    labels:
      severity: critical
    annotations:
      summary: "etcd member is down"
      description: "etcd instance {{ $labels.instance }} has been down for more than 5 minutes."
```

These alerts cover leadership, disk latency, and member availability. Integrate them with your notification system (e.g., PagerDuty, Slack). Define an escalation policy: critical alerts page on-call immediately, warnings can wait for business hours.

For each alert, assign an owner. For example:

- EtcdNoLeader: Priya Shah, Engineering Lead (primary), Ravi Kumar, SRE (secondary). Revisit alert thresholds monthly.
- EtcdHighFsyncDuration: Ravi Kumar, SRE. Revisit monthly.
- EtcdHighCommitDuration: Ravi Kumar, SRE. Revisit monthly.
- EtcdLeaderChangesFrequent: Priya Shah. Revisit weekly if triggered.
- EtcdMemberDown: On-call engineer. Immediate response.

Document these owners in your runbook.

## Common Pitfalls and How to Avoid Them

### 1. Monitoring the Wrong Port

Many users assume etcd metrics are on the client port 2379, but if you set a dedicated metrics port like 2381, your Prometheus scrape job must use that port. Pitfall: forget to update the scrape config after enabling `--listen-metrics-urls`. Always verify with a manual curl.

### 2. Certificate Mismatch

Etcd uses mutual TLS. If Prometheus scrapes with wrong or expired certs, the target shows as down. Pitfall: not renewing certificates before expiry. Use a certificate management tool (cert-manager) and set alerts on certificate expiry.

### 3. Ignoring Disk Latency

Etcd is sensitive to disk latency. Running etcd on network-attached storage or slow disks causes fsync issues. Pitfall: assuming any disk works. Use fast local SSD storage dedicated to etcd. Monitor `etcd_disk_wal_fsync_duration_seconds_bucket` and alert if p99 exceeds 25ms.

### 4. Overlooking Compaction and Defragmentation

Etcd keeps a history of all changes. Without compaction, the database grows and performance degrades. Pitfall: not scheduling compaction. Enable automatic compaction in etcd by setting `--auto-compaction-mode=periodic --auto-compaction-retention=1h` or create a CronJob to run `etcdctl compact` and `defrag` regularly.

### 5. Upgrading All Etcd Members Simultaneously

If you upgrade all etcd members at once, you risk quorum loss and data unavailability. Pitfall: not following rolling upgrade. Always upgrade one member at a time and wait for the cluster to become healthy before the next.

### 6. No Backup or Snapshot Strategy

If etcd data is lost, without snapshots recovery is impossible. Pitfall: not testing restores. Take snapshots regularly using a CronJob, store them off-cluster, and test the restore procedure in a staging environment quarterly.

## Operations Checklist

Use this checklist before and after any etcd configuration change or upgrade.

**Pre-change:**

- [ ] Record baseline metrics: leader, fsync duration, commit duration, member list.
- [ ] Take an etcd snapshot and verify its integrity.
- [ ] Notify team and schedule if needed.
- [ ] Ensure all members are healthy.
- [ ] Check available disk space.

**During change:**

- [ ] Apply change to one member only (if applicable).
- [ ] Monitor etcd metrics continuously.
- [ ] Watch for leader changes or quorum loss.
- [ ] Verify that Prometheus `up` remains 1.

**Post-change:**

- [ ] Verify etcd version and health.
- [ ] Check metrics and dashboards for anomalies.
- [ ] Confirm alerts are not firing unexpectedly.
- [ ] Document the change and outcome in the operations log.
- [ ] Revisit after 24 hours to check stability.

Assign an owner for each checklist execution: the on-call engineer is responsible for pre-change checks, the SRE lead approves the change, and the on-call engineer performs post-change verification. Review this checklist quarterly for necessary updates based on incidents.

## Conclusion

Etcd monitoring and alerting are not set-and-forget tasks. They require continuous attention, especially during upgrades and configuration changes. This guide provided a practical path from inventory to safe configuration, verification, alerting, and recovery.

By following the commands and examples, you can build a monitoring stack that catches etcd issues early and reduces downtime. Remember to:

- Version-scope every recommendation.
- Observe before changing.
- Limit blast radius.
- Use placeholders for secrets.
- Verify results.
- Document recovery procedures.

As a next step, choose one low-risk verification from this guide, such as checking metrics on port 2381 or creating a test alert. Record the current state, run the check, compare with expected output, and review dependencies like the Kube API Server and certificates.

A reliable etcd monitoring workflow makes failures visible, protects sensitive values, and defines recovery before an incident forces the decision.