E-NO
Kubernetes Configure and Upgrade Etcd monitoring 7 Min Read

Kubernetes Etcd Monitoring and Alerts: A Practical Configuration and Upgrade Guide

calendar_today Published: 2026-09-27
update Last Updated: 2026-09-27
analytics SEO Efficiency: 100%
Technical guide illustration for Kubernetes Etcd Monitoring and Alerts: A Practical Configuration and Upgrade Guide.

Intro

Etcd is the distributed key-value store at the heart of every Kubernetes cluster. It holds all cluster state, including Pods, Services, ConfigMaps, and Secrets. If etcd becomes unavailable or loses data, the entire control plane can fail. Monitoring and alerting for etcd are therefore critical for operational safety.

This guide focuses on practical configuration and upgrade of etcd monitoring and alerts in Kubernetes. It is written for developers, DevOps consultants, and technical startup teams who operate their own clusters and need to move from observing a problem to verifying a fix.

You will learn how to:

  • Inventory your etcd version and deployment topology.
  • Expose and collect etcd metrics.
  • Build dashboards that show key health indicators.
  • Configure alerts that catch failure before users notice.
  • Upgrade etcd safely with monitoring in place.
  • Respond to incidents with a clear diagnostic path.

Every section includes concrete commands, expected output, failure signals, and recovery decisions. The goal is operational safety: observe before changing, limit the blast radius, use placeholders instead of secrets, verify the result, and document recovery before an incident forces it.

Version and Environment Inventory

Before touching monitoring or upgrading etcd, you must know exactly what you are running. Start with a read-only inventory of the cluster and the etcd component.

Identify the etcd version:

kubectl get pods -n kube-system -l component=etcd -o jsonpath='{.items[0].spec.containers[0].image}'

Expected output example:

registry.k8s.io/etcd:3.5.9-0

Determine the etcd deployment topology. Most kubeadm clusters run etcd as a static Pod on each control-plane node. List the etcd Pods:

kubectl get pods -n kube-system -l component=etcd -o wide

Expected output example:

NAME                READY   STATUS    RESTARTS   AGE   IP              NODE
etcd-control-plane-1   1/1     Running   0          10m   192.168.1.10   control-plane-1
etcd-control-plane-2   1/1     Running   0          10m   192.168.1.11   control-plane-2
etcd-control-plane-3   1/1     Running   0          10m   192.168.1.12   control-plane-3

Check prerequisites:

  • kubectl access to the cluster with permissions to get Pods in kube-system.
  • Ability to exec into the etcd container or access the node filesystem if needed.
  • Network access from your monitoring stack to the etcd metrics endpoint.

Perform a read-only health check using etcdctl inside the Pod:

kubectl exec -n kube-system etcd-control-plane-1 -- etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/peer.crt --key=/etc/kubernetes/pki/etcd/peer.key endpoint health

Expected healthy output:

https://127.0.0.1:2379 is healthy: successfully committed proposal: took = 2.1ms

If the cluster uses an external etcd (not static Pods), adjust commands accordingly. Record the etcd version, number of members, and health status in your operations log with timestamps. This baseline is essential before any change.

Safe Configuration Path

Configuring etcd monitoring requires enabling the metrics endpoint and ensuring your Prometheus or other scraper can reach it securely.

Etcd exposes metrics on port 2379 by default when started with the --listen-metrics-urls flag. Check the current etcd Pod spec for the metrics URL:

kubectl get pod -n kube-system etcd-control-plane-1 -o yaml | grep -A2 'listen-metrics-urls'

Expected output may include:

    - --listen-metrics-urls=http://0.0.0.0:2381

If the flag is missing, etcd still exposes metrics on port 2379, but that port is also used for client requests. It is safer to separate metrics traffic on a dedicated port (e.g., 2381) to avoid interference.

To enable a dedicated metrics port on a static Pod managed by kubeadm, edit the etcd manifest on the control-plane node:

sudo vi /etc/kubernetes/manifests/etcd.yaml

Add or modify the container command arguments:

spec:
  containers:
  - command:
    - etcd
    - --listen-metrics-urls=http://0.0.0.0:2381
    - --metrics=extensive

The --metrics=extensive option provides additional metrics like gRPC request histograms, which are useful for latency analysis.

After saving the manifest, kubelet will restart the etcd Pod automatically. Verify the new port is listening:

kubectl exec -n kube-system etcd-control-plane-1 -- sh -c 'curl -s http://localhost:2381/metrics | head -5'

Expected output:

# HELP etcd_server_has_leader Whether or not a leader exists. 1 is leader exists.
# TYPE etcd_server_has_leader gauge
etcd_server_has_leader 1

Now configure Prometheus to scrape this endpoint. If you use the Prometheus Operator, create a ServiceMonitor targeting the etcd Pods. If using a raw Prometheus config, add a scrape job:

scrape_configs:
  - job_name: 'etcd'
    scheme: https
    tls_config:
      ca_file: /etc/prometheus/secrets/etcd-ca.crt
      cert_file: /etc/prometheus/secrets/etcd-client.crt
      key_file: /etc/prometheus/secrets/etcd-client.key
    static_configs:
      - targets: ['192.168.1.10:2379', '192.168.1.11:2379', '192.168.1.12:2379']

If you enabled the dedicated metrics port, use that port instead. Ensure the certificates have the correct Subject Alternative Names for the node IPs.

Test the scrape manually from the Prometheus Pod:

kubectl exec -n monitoring prometheus-0 -- wget -qO- https://192.168.1.10:2379/metrics --ca-certificate=/etc/prometheus/secrets/etcd-ca.crt | head

Expected output shows Prometheus metrics. If you get a certificate error, check the certificate validity and SANs.

Quick check 1 of 2

What is the recommended way to back up an etcd cluster according to the Kubernetes documentation?

The passage states: 'Backing up an etcd cluster can be accomplished in two ways: etcd built-in snapshot and volume snapshot.'

Verification and Diagnostics

After enabling metrics collection, verify that Prometheus is receiving etcd metrics and that they are not empty.

Query Prometheus for the up metric for etcd targets:

up{job="etcd"}

Expected result: all etcd targets should show 1.

Check for any metric sample, e.g., etcd_server_has_leader:

etcd_server_has_leader

Expected result: the value is 1 on the leader node and 0 on followers. If the metric is absent, check the scrape configuration and network policies.

Diagnose common issues:

  • Target down in Prometheus: Verify the endpoint is reachable from Prometheus Pod. Check firewall rules or Kubernetes NetworkPolicy.
  • Certificate errors: Ensure the CA and client certs in Prometheus match the etcd server certs. etcd uses mutual TLS by default in kubeadm clusters.
  • No metrics on port 2381: Confirm the --listen-metrics-urls flag was applied. Check etcd Pod logs for startup errors:
kubectl logs -n kube-system etcd-control-plane-1

You can also use etcdctl to check cluster health and member list to ensure no split brain:

kubectl exec -n kube-system etcd-control-plane-1 -- etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key member list

Expected output lists all members with their client URLs and peer URLs.

If metrics are flowing, you can now build dashboards. For Grafana, import a dashboard that uses etcd metrics. A popular choice is the etcd dashboard from the Kubernetes Mixins or the Grafana community. Key panels to include:

  • Leader changes: increase(etcd_server_leader_changes_seen_total[1h]) should be low. Frequent changes indicate network issues.
  • Raft proposal duration: histogram_quantile(0.99, sum(rate(etcd_disk_wal_fsync_duration_seconds_bucket[5m])) by (le)) should be less than 25ms.
  • Disk operation duration: histogram_quantile(0.99, sum(rate(etcd_disk_backend_commit_duration_seconds_bucket[5m])) by (le)) less than 50ms.
  • gRPC request rate: sum(rate(grpc_server_handled_total[5m])) by (grpc_method) to see request load.

Failure Modes and Recovery

Etcd failures can be subtle. Knowing the common failure modes helps you react quickly.

Leader Election Storms

If the etcd leader changes frequently (more than once per hour), it indicates network latency or packet loss between members. This can cause write timeouts for the API server.

Diagnosis: Check increase(etcd_server_leader_changes_seen_total[1h]). Also check network latency between nodes with ping or iperf.

Recovery: Investigate network performance. If caused by high disk I/O on the leader, reduce load or move etcd to dedicated disks. If network issue persists, consider adding a faster network or adjusting etcd heartbeat interval (not recommended without deep knowledge).

Disk Space Exhaustion

Etcd stores data in a directory. If the disk fills up, etcd may stop accepting writes or crash.

Diagnosis: Check disk usage on the node:

df -h /var/lib/etcd

Expected output example:

Filesystem      Size  Used Avail Use% Mounted on
/dev/sda1       100G   95G   5G  95% /var/lib/etcd

If usage is above 80%, plan for compaction and defragmentation.

Recovery: Compact the etcd history to remove old revisions:

kubectl exec -n kube-system etcd-control-plane-1 -- etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key compact $(etcdctl ... get / --prefix --keys-only | wc -l)

Better to use the current revision minus some safety margin. Then defragment:

kubectl exec -n kube-system etcd-control-plane-1 -- etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key defrag

Run defrag on each member sequentially to avoid quorum loss. Schedule regular compaction via a CronJob if your etcd version does not auto-compact.

Quorum Loss

If more than (n/2) members fail, the cluster loses quorum and cannot process writes. Read-only requests may still work depending on configuration.

Diagnosis: Check etcd_server_has_leader metric; if it is 0 on all nodes, quorum is lost. Also run etcdctl endpoint status to see leader and raft term.

Recovery: Restore the failed members as soon as possible. If the majority is permanently lost, you must perform an etcd disaster recovery from a snapshot. Always take regular snapshots:

kubectl exec -n kube-system etcd-control-plane-1 -- etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key snapshot save /var/lib/etcd/snapshot.db

Store snapshots off-node. Test restoration in a non-production environment regularly.

Upgrading Etcd with Monitoring in Place

Upgrading etcd requires careful planning. Monitor the cluster during the upgrade to catch issues early.

Pre-upgrade checklist:

  • Verify current version and cluster health (see Version and Environment Inventory).
  • Take a snapshot of the etcd data.
  • Check compatibility: Kubernetes supports etcd 3.5.x for recent versions; do not jump major versions without testing.
  • Notify stakeholders and schedule a maintenance window if necessary.

For kubeadm clusters, upgrading etcd is part of the control-plane upgrade. First, upgrade kubeadm:

sudo apt-get update && sudo apt-get install -y kubeadm=1.28.0-00

Then run the upgrade plan:

sudo kubeadm upgrade plan

This will show the recommended etcd version. Apply the upgrade on the first control-plane node:

sudo kubeadm upgrade apply v1.28.0

This upgrades etcd, kube-apiserver, kube-controller-manager, and kube-scheduler on that node. Monitor etcd metrics during the process. Watch for leader changes and quorum loss. The cluster should still function because of redundancy.

After the first node is upgraded, upgrade other control-plane nodes one by one:

sudo kubeadm upgrade node

Monitor etcd after each node. Check etcd_server_has_leader and etcd_server_leader_changes_seen_total. Ensure metrics are still scraped.

If you use an external etcd, follow the etcd documentation for in-place upgrades. Never upgrade more than one member at a time, and allow the cluster to stabilize between upgrades.

Post-upgrade verification:

  • Check etcd version: kubectl exec ... etcdctl version
  • Check cluster health: etcdctl endpoint health
  • Check metrics in Prometheus: up{job="etcd"} is 1.
  • Check dashboards for anomalies.

If an upgrade fails, you can roll back by restoring the snapshot if necessary. However, etcd upgrades are usually backward compatible within the same major version. Rolling back Kubernetes is more complex; always test in staging.

Quick check 2 of 2

What is the recommended number of members for a production etcd cluster?

The passage states: 'A five-member cluster is recommended in production.'

Configuring Alerts

Set up Prometheus alerts for etcd. Use the following alert rules as a starting point. Adjust thresholds based on your environment.

Create an alert rule file (e.g., etcd-alerts.yaml) and load it into Prometheus.

groups:
- name: etcd
  rules:
  - alert: EtcdNoLeader
    expr: etcd_server_has_leader == 0
    for: 5m
    labels:
      severity: critical
    annotations:
      summary: "etcd cluster has no leader"
      description: "etcd instance {{ $labels.instance }} has no leader for more than 5 minutes."

  - alert: EtcdHighFsyncDuration
    expr: histogram_quantile(0.99, sum(rate(etcd_disk_wal_fsync_duration_seconds_bucket[5m])) by (le)) > 0.025
    for: 10m
    labels:
      severity: warning
    annotations:
      summary: "etcd fsync latency high"
      description: "etcd instance {{ $labels.instance }} 99th percentile fsync duration is above 25ms for 10 minutes."

  - alert: EtcdHighCommitDuration
    expr: histogram_quantile(0.99, sum(rate(etcd_disk_backend_commit_duration_seconds_bucket[5m])) by (le)) > 0.05
    for: 10m
    labels:
      severity: warning
    annotations:
      summary: "etcd commit latency high"
      description: "etcd instance {{ $labels.instance }} 99th percentile commit duration is above 50ms for 10 minutes."

  - alert: EtcdLeaderChangesFrequent
    expr: increase(etcd_server_leader_changes_seen_total[1h]) > 3
    for: 5m
    labels:
      severity: warning
    annotations:
      summary: "etcd leader changes are frequent"
      description: "etcd cluster has had more than 3 leader changes in the last hour."

  - alert: EtcdMemberDown
    expr: up{job="etcd"} == 0
    for: 5m
    labels:
      severity: critical
    annotations:
      summary: "etcd member is down"
      description: "etcd instance {{ $labels.instance }} has been down for more than 5 minutes."

These alerts cover leadership, disk latency, and member availability. Integrate them with your notification system (e.g., PagerDuty, Slack). Define an escalation policy: critical alerts page on-call immediately, warnings can wait for business hours.

For each alert, assign an owner. For example:

  • EtcdNoLeader: Priya Shah, Engineering Lead (primary), Ravi Kumar, SRE (secondary). Revisit alert thresholds monthly.
  • EtcdHighFsyncDuration: Ravi Kumar, SRE. Revisit monthly.
  • EtcdHighCommitDuration: Ravi Kumar, SRE. Revisit monthly.
  • EtcdLeaderChangesFrequent: Priya Shah. Revisit weekly if triggered.
  • EtcdMemberDown: On-call engineer. Immediate response.

Document these owners in your runbook.

Common Pitfalls and How to Avoid Them

1. Monitoring the Wrong Port

Many users assume etcd metrics are on the client port 2379, but if you set a dedicated metrics port like 2381, your Prometheus scrape job must use that port. Pitfall: forget to update the scrape config after enabling --listen-metrics-urls. Always verify with a manual curl.

2. Certificate Mismatch

Etcd uses mutual TLS. If Prometheus scrapes with wrong or expired certs, the target shows as down. Pitfall: not renewing certificates before expiry. Use a certificate management tool (cert-manager) and set alerts on certificate expiry.

3. Ignoring Disk Latency

Etcd is sensitive to disk latency. Running etcd on network-attached storage or slow disks causes fsync issues. Pitfall: assuming any disk works. Use fast local SSD storage dedicated to etcd. Monitor etcd_disk_wal_fsync_duration_seconds_bucket and alert if p99 exceeds 25ms.

4. Overlooking Compaction and Defragmentation

Etcd keeps a history of all changes. Without compaction, the database grows and performance degrades. Pitfall: not scheduling compaction. Enable automatic compaction in etcd by setting --auto-compaction-mode=periodic --auto-compaction-retention=1h or create a CronJob to run etcdctl compact and defrag regularly.

5. Upgrading All Etcd Members Simultaneously

If you upgrade all etcd members at once, you risk quorum loss and data unavailability. Pitfall: not following rolling upgrade. Always upgrade one member at a time and wait for the cluster to become healthy before the next.

6. No Backup or Snapshot Strategy

If etcd data is lost, without snapshots recovery is impossible. Pitfall: not testing restores. Take snapshots regularly using a CronJob, store them off-cluster, and test the restore procedure in a staging environment quarterly.

Operations Checklist

Use this checklist before and after any etcd configuration change or upgrade.

Pre-change:

  • [ ] Record baseline metrics: leader, fsync duration, commit duration, member list.
  • [ ] Take an etcd snapshot and verify its integrity.
  • [ ] Notify team and schedule if needed.
  • [ ] Ensure all members are healthy.
  • [ ] Check available disk space.

During change:

  • [ ] Apply change to one member only (if applicable).
  • [ ] Monitor etcd metrics continuously.
  • [ ] Watch for leader changes or quorum loss.
  • [ ] Verify that Prometheus up remains 1.

Post-change:

  • [ ] Verify etcd version and health.
  • [ ] Check metrics and dashboards for anomalies.
  • [ ] Confirm alerts are not firing unexpectedly.
  • [ ] Document the change and outcome in the operations log.
  • [ ] Revisit after 24 hours to check stability.

Assign an owner for each checklist execution: the on-call engineer is responsible for pre-change checks, the SRE lead approves the change, and the on-call engineer performs post-change verification. Review this checklist quarterly for necessary updates based on incidents.

Conclusion

Etcd monitoring and alerting are not set-and-forget tasks. They require continuous attention, especially during upgrades and configuration changes. This guide provided a practical path from inventory to safe configuration, verification, alerting, and recovery.

By following the commands and examples, you can build a monitoring stack that catches etcd issues early and reduces downtime. Remember to:

  • Version-scope every recommendation.
  • Observe before changing.
  • Limit blast radius.
  • Use placeholders for secrets.
  • Verify results.
  • Document recovery procedures.

As a next step, choose one low-risk verification from this guide, such as checking metrics on port 2381 or creating a test alert. Record the current state, run the check, compare with expected output, and review dependencies like the Kube API Server and certificates.

A reliable etcd monitoring workflow makes failures visible, protects sensitive values, and defines recovery before an incident forces the decision.

Related Research

Article Quality Score

Reader usefulness 100%
  • check_circle Reader-ready guide
  • check_circle Practical examples included
  • check_circle Clean SEO article URL