Intro
Etcd is the distributed key-value store at the heart of every Kubernetes cluster. It holds all cluster state, including Pods, Services, ConfigMaps, and Secrets. If etcd becomes unavailable or loses data, the entire control plane can fail. Monitoring and alerting for etcd are therefore critical for operational safety.
This guide focuses on practical configuration and upgrade of etcd monitoring and alerts in Kubernetes. It is written for developers, DevOps consultants, and technical startup teams who operate their own clusters and need to move from observing a problem to verifying a fix.
You will learn how to:
- Inventory your etcd version and deployment topology.
- Expose and collect etcd metrics.
- Build dashboards that show key health indicators.
- Configure alerts that catch failure before users notice.
- Upgrade etcd safely with monitoring in place.
- Respond to incidents with a clear diagnostic path.
Every section includes concrete commands, expected output, failure signals, and recovery decisions. The goal is operational safety: observe before changing, limit the blast radius, use placeholders instead of secrets, verify the result, and document recovery before an incident forces it.
Version and Environment Inventory
Before touching monitoring or upgrading etcd, you must know exactly what you are running. Start with a read-only inventory of the cluster and the etcd component.
Identify the etcd version:
kubectl get pods -n kube-system -l component=etcd -o jsonpath='{.items[0].spec.containers[0].image}'
Expected output example:
registry.k8s.io/etcd:3.5.9-0
Determine the etcd deployment topology. Most kubeadm clusters run etcd as a static Pod on each control-plane node. List the etcd Pods:
kubectl get pods -n kube-system -l component=etcd -o wide
Expected output example:
NAME READY STATUS RESTARTS AGE IP NODE
etcd-control-plane-1 1/1 Running 0 10m 192.168.1.10 control-plane-1
etcd-control-plane-2 1/1 Running 0 10m 192.168.1.11 control-plane-2
etcd-control-plane-3 1/1 Running 0 10m 192.168.1.12 control-plane-3
Check prerequisites:
kubectlaccess to the cluster with permissions to get Pods inkube-system.- Ability to
execinto the etcd container or access the node filesystem if needed. - Network access from your monitoring stack to the etcd metrics endpoint.
Perform a read-only health check using etcdctl inside the Pod:
kubectl exec -n kube-system etcd-control-plane-1 -- etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/peer.crt --key=/etc/kubernetes/pki/etcd/peer.key endpoint health
Expected healthy output:
https://127.0.0.1:2379 is healthy: successfully committed proposal: took = 2.1ms
If the cluster uses an external etcd (not static Pods), adjust commands accordingly. Record the etcd version, number of members, and health status in your operations log with timestamps. This baseline is essential before any change.
Safe Configuration Path
Configuring etcd monitoring requires enabling the metrics endpoint and ensuring your Prometheus or other scraper can reach it securely.
Etcd exposes metrics on port 2379 by default when started with the --listen-metrics-urls flag. Check the current etcd Pod spec for the metrics URL:
kubectl get pod -n kube-system etcd-control-plane-1 -o yaml | grep -A2 'listen-metrics-urls'
Expected output may include:
- --listen-metrics-urls=http://0.0.0.0:2381
If the flag is missing, etcd still exposes metrics on port 2379, but that port is also used for client requests. It is safer to separate metrics traffic on a dedicated port (e.g., 2381) to avoid interference.
To enable a dedicated metrics port on a static Pod managed by kubeadm, edit the etcd manifest on the control-plane node:
sudo vi /etc/kubernetes/manifests/etcd.yaml
Add or modify the container command arguments:
spec:
containers:
- command:
- etcd
- --listen-metrics-urls=http://0.0.0.0:2381
- --metrics=extensive
The --metrics=extensive option provides additional metrics like gRPC request histograms, which are useful for latency analysis.
After saving the manifest, kubelet will restart the etcd Pod automatically. Verify the new port is listening:
kubectl exec -n kube-system etcd-control-plane-1 -- sh -c 'curl -s http://localhost:2381/metrics | head -5'
Expected output:
# HELP etcd_server_has_leader Whether or not a leader exists. 1 is leader exists.
# TYPE etcd_server_has_leader gauge
etcd_server_has_leader 1
Now configure Prometheus to scrape this endpoint. If you use the Prometheus Operator, create a ServiceMonitor targeting the etcd Pods. If using a raw Prometheus config, add a scrape job:
scrape_configs:
- job_name: 'etcd'
scheme: https
tls_config:
ca_file: /etc/prometheus/secrets/etcd-ca.crt
cert_file: /etc/prometheus/secrets/etcd-client.crt
key_file: /etc/prometheus/secrets/etcd-client.key
static_configs:
- targets: ['192.168.1.10:2379', '192.168.1.11:2379', '192.168.1.12:2379']
If you enabled the dedicated metrics port, use that port instead. Ensure the certificates have the correct Subject Alternative Names for the node IPs.
Test the scrape manually from the Prometheus Pod:
kubectl exec -n monitoring prometheus-0 -- wget -qO- https://192.168.1.10:2379/metrics --ca-certificate=/etc/prometheus/secrets/etcd-ca.crt | head
Expected output shows Prometheus metrics. If you get a certificate error, check the certificate validity and SANs.
Verification and Diagnostics
After enabling metrics collection, verify that Prometheus is receiving etcd metrics and that they are not empty.
Query Prometheus for the up metric for etcd targets:
up{job="etcd"}
Expected result: all etcd targets should show 1.
Check for any metric sample, e.g., etcd_server_has_leader:
etcd_server_has_leader
Expected result: the value is 1 on the leader node and 0 on followers. If the metric is absent, check the scrape configuration and network policies.
Diagnose common issues:
- Target down in Prometheus: Verify the endpoint is reachable from Prometheus Pod. Check firewall rules or Kubernetes NetworkPolicy.
- Certificate errors: Ensure the CA and client certs in Prometheus match the etcd server certs.
etcduses mutual TLS by default in kubeadm clusters. - No metrics on port 2381: Confirm the
--listen-metrics-urlsflag was applied. Check etcd Pod logs for startup errors:
kubectl logs -n kube-system etcd-control-plane-1
You can also use etcdctl to check cluster health and member list to ensure no split brain:
kubectl exec -n kube-system etcd-control-plane-1 -- etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key member list
Expected output lists all members with their client URLs and peer URLs.
If metrics are flowing, you can now build dashboards. For Grafana, import a dashboard that uses etcd metrics. A popular choice is the etcd dashboard from the Kubernetes Mixins or the Grafana community. Key panels to include:
- Leader changes:
increase(etcd_server_leader_changes_seen_total[1h])should be low. Frequent changes indicate network issues. - Raft proposal duration:
histogram_quantile(0.99, sum(rate(etcd_disk_wal_fsync_duration_seconds_bucket[5m])) by (le))should be less than 25ms. - Disk operation duration:
histogram_quantile(0.99, sum(rate(etcd_disk_backend_commit_duration_seconds_bucket[5m])) by (le))less than 50ms. - gRPC request rate:
sum(rate(grpc_server_handled_total[5m])) by (grpc_method)to see request load.
Failure Modes and Recovery
Etcd failures can be subtle. Knowing the common failure modes helps you react quickly.
Leader Election Storms
If the etcd leader changes frequently (more than once per hour), it indicates network latency or packet loss between members. This can cause write timeouts for the API server.
Diagnosis: Check increase(etcd_server_leader_changes_seen_total[1h]). Also check network latency between nodes with ping or iperf.
Recovery: Investigate network performance. If caused by high disk I/O on the leader, reduce load or move etcd to dedicated disks. If network issue persists, consider adding a faster network or adjusting etcd heartbeat interval (not recommended without deep knowledge).
Disk Space Exhaustion
Etcd stores data in a directory. If the disk fills up, etcd may stop accepting writes or crash.
Diagnosis: Check disk usage on the node:
df -h /var/lib/etcd
Expected output example:
Filesystem Size Used Avail Use% Mounted on
/dev/sda1 100G 95G 5G 95% /var/lib/etcd
If usage is above 80%, plan for compaction and defragmentation.
Recovery: Compact the etcd history to remove old revisions:
kubectl exec -n kube-system etcd-control-plane-1 -- etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key compact $(etcdctl ... get / --prefix --keys-only | wc -l)
Better to use the current revision minus some safety margin. Then defragment:
kubectl exec -n kube-system etcd-control-plane-1 -- etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key defrag
Run defrag on each member sequentially to avoid quorum loss. Schedule regular compaction via a CronJob if your etcd version does not auto-compact.
Quorum Loss
If more than (n/2) members fail, the cluster loses quorum and cannot process writes. Read-only requests may still work depending on configuration.
Diagnosis: Check etcd_server_has_leader metric; if it is 0 on all nodes, quorum is lost. Also run etcdctl endpoint status to see leader and raft term.
Recovery: Restore the failed members as soon as possible. If the majority is permanently lost, you must perform an etcd disaster recovery from a snapshot. Always take regular snapshots:
kubectl exec -n kube-system etcd-control-plane-1 -- etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key snapshot save /var/lib/etcd/snapshot.db
Store snapshots off-node. Test restoration in a non-production environment regularly.
Upgrading Etcd with Monitoring in Place
Upgrading etcd requires careful planning. Monitor the cluster during the upgrade to catch issues early.
Pre-upgrade checklist:
- Verify current version and cluster health (see Version and Environment Inventory).
- Take a snapshot of the etcd data.
- Check compatibility: Kubernetes supports etcd 3.5.x for recent versions; do not jump major versions without testing.
- Notify stakeholders and schedule a maintenance window if necessary.
For kubeadm clusters, upgrading etcd is part of the control-plane upgrade. First, upgrade kubeadm:
sudo apt-get update && sudo apt-get install -y kubeadm=1.28.0-00
Then run the upgrade plan:
sudo kubeadm upgrade plan
This will show the recommended etcd version. Apply the upgrade on the first control-plane node:
sudo kubeadm upgrade apply v1.28.0
This upgrades etcd, kube-apiserver, kube-controller-manager, and kube-scheduler on that node. Monitor etcd metrics during the process. Watch for leader changes and quorum loss. The cluster should still function because of redundancy.
After the first node is upgraded, upgrade other control-plane nodes one by one:
sudo kubeadm upgrade node
Monitor etcd after each node. Check etcd_server_has_leader and etcd_server_leader_changes_seen_total. Ensure metrics are still scraped.
If you use an external etcd, follow the etcd documentation for in-place upgrades. Never upgrade more than one member at a time, and allow the cluster to stabilize between upgrades.
Post-upgrade verification:
- Check etcd version:
kubectl exec ... etcdctl version - Check cluster health:
etcdctl endpoint health - Check metrics in Prometheus:
up{job="etcd"}is 1. - Check dashboards for anomalies.
If an upgrade fails, you can roll back by restoring the snapshot if necessary. However, etcd upgrades are usually backward compatible within the same major version. Rolling back Kubernetes is more complex; always test in staging.
Configuring Alerts
Set up Prometheus alerts for etcd. Use the following alert rules as a starting point. Adjust thresholds based on your environment.
Create an alert rule file (e.g., etcd-alerts.yaml) and load it into Prometheus.
groups:
- name: etcd
rules:
- alert: EtcdNoLeader
expr: etcd_server_has_leader == 0
for: 5m
labels:
severity: critical
annotations:
summary: "etcd cluster has no leader"
description: "etcd instance {{ $labels.instance }} has no leader for more than 5 minutes."
- alert: EtcdHighFsyncDuration
expr: histogram_quantile(0.99, sum(rate(etcd_disk_wal_fsync_duration_seconds_bucket[5m])) by (le)) > 0.025
for: 10m
labels:
severity: warning
annotations:
summary: "etcd fsync latency high"
description: "etcd instance {{ $labels.instance }} 99th percentile fsync duration is above 25ms for 10 minutes."
- alert: EtcdHighCommitDuration
expr: histogram_quantile(0.99, sum(rate(etcd_disk_backend_commit_duration_seconds_bucket[5m])) by (le)) > 0.05
for: 10m
labels:
severity: warning
annotations:
summary: "etcd commit latency high"
description: "etcd instance {{ $labels.instance }} 99th percentile commit duration is above 50ms for 10 minutes."
- alert: EtcdLeaderChangesFrequent
expr: increase(etcd_server_leader_changes_seen_total[1h]) > 3
for: 5m
labels:
severity: warning
annotations:
summary: "etcd leader changes are frequent"
description: "etcd cluster has had more than 3 leader changes in the last hour."
- alert: EtcdMemberDown
expr: up{job="etcd"} == 0
for: 5m
labels:
severity: critical
annotations:
summary: "etcd member is down"
description: "etcd instance {{ $labels.instance }} has been down for more than 5 minutes."
These alerts cover leadership, disk latency, and member availability. Integrate them with your notification system (e.g., PagerDuty, Slack). Define an escalation policy: critical alerts page on-call immediately, warnings can wait for business hours.
For each alert, assign an owner. For example:
- EtcdNoLeader: Priya Shah, Engineering Lead (primary), Ravi Kumar, SRE (secondary). Revisit alert thresholds monthly.
- EtcdHighFsyncDuration: Ravi Kumar, SRE. Revisit monthly.
- EtcdHighCommitDuration: Ravi Kumar, SRE. Revisit monthly.
- EtcdLeaderChangesFrequent: Priya Shah. Revisit weekly if triggered.
- EtcdMemberDown: On-call engineer. Immediate response.
Document these owners in your runbook.
Common Pitfalls and How to Avoid Them
1. Monitoring the Wrong Port
Many users assume etcd metrics are on the client port 2379, but if you set a dedicated metrics port like 2381, your Prometheus scrape job must use that port. Pitfall: forget to update the scrape config after enabling --listen-metrics-urls. Always verify with a manual curl.
2. Certificate Mismatch
Etcd uses mutual TLS. If Prometheus scrapes with wrong or expired certs, the target shows as down. Pitfall: not renewing certificates before expiry. Use a certificate management tool (cert-manager) and set alerts on certificate expiry.
3. Ignoring Disk Latency
Etcd is sensitive to disk latency. Running etcd on network-attached storage or slow disks causes fsync issues. Pitfall: assuming any disk works. Use fast local SSD storage dedicated to etcd. Monitor etcd_disk_wal_fsync_duration_seconds_bucket and alert if p99 exceeds 25ms.
4. Overlooking Compaction and Defragmentation
Etcd keeps a history of all changes. Without compaction, the database grows and performance degrades. Pitfall: not scheduling compaction. Enable automatic compaction in etcd by setting --auto-compaction-mode=periodic --auto-compaction-retention=1h or create a CronJob to run etcdctl compact and defrag regularly.
5. Upgrading All Etcd Members Simultaneously
If you upgrade all etcd members at once, you risk quorum loss and data unavailability. Pitfall: not following rolling upgrade. Always upgrade one member at a time and wait for the cluster to become healthy before the next.
6. No Backup or Snapshot Strategy
If etcd data is lost, without snapshots recovery is impossible. Pitfall: not testing restores. Take snapshots regularly using a CronJob, store them off-cluster, and test the restore procedure in a staging environment quarterly.
Operations Checklist
Use this checklist before and after any etcd configuration change or upgrade.
Pre-change:
- [ ] Record baseline metrics: leader, fsync duration, commit duration, member list.
- [ ] Take an etcd snapshot and verify its integrity.
- [ ] Notify team and schedule if needed.
- [ ] Ensure all members are healthy.
- [ ] Check available disk space.
During change:
- [ ] Apply change to one member only (if applicable).
- [ ] Monitor etcd metrics continuously.
- [ ] Watch for leader changes or quorum loss.
- [ ] Verify that Prometheus
upremains 1.
Post-change:
- [ ] Verify etcd version and health.
- [ ] Check metrics and dashboards for anomalies.
- [ ] Confirm alerts are not firing unexpectedly.
- [ ] Document the change and outcome in the operations log.
- [ ] Revisit after 24 hours to check stability.
Assign an owner for each checklist execution: the on-call engineer is responsible for pre-change checks, the SRE lead approves the change, and the on-call engineer performs post-change verification. Review this checklist quarterly for necessary updates based on incidents.
Conclusion
Etcd monitoring and alerting are not set-and-forget tasks. They require continuous attention, especially during upgrades and configuration changes. This guide provided a practical path from inventory to safe configuration, verification, alerting, and recovery.
By following the commands and examples, you can build a monitoring stack that catches etcd issues early and reduces downtime. Remember to:
- Version-scope every recommendation.
- Observe before changing.
- Limit blast radius.
- Use placeholders for secrets.
- Verify results.
- Document recovery procedures.
As a next step, choose one low-risk verification from this guide, such as checking metrics on port 2381 or creating a test alert. Record the current state, run the check, compare with expected output, and review dependencies like the Kube API Server and certificates.
A reliable etcd monitoring workflow makes failures visible, protects sensitive values, and defines recovery before an incident forces the decision.