E-NO
Kube API Server monitoring 7 Min Read

Kubernetes API Server Monitoring and Alerts: A Practical Implementation Guide

calendar_today Published: 2026-09-29
update Last Updated: 2026-09-29
analytics SEO Efficiency: 100%
Technical guide illustration for Kubernetes API Server Monitoring and Alerts: A Practical Implementation Guide.

Introduction

The Kubernetes API server is the central control plane component that processes all REST requests for the cluster. It validates and configures data for API objects such as pods, services, and deployments. When the API server degrades or fails, every automation, dashboard, and user command that depends on Kubernetes is affected. Monitoring it effectively means knowing which metrics matter, how to collect them, how to turn them into actionable dashboards and alerts, and how to respond when things go wrong.

This guide provides a practical, step-by-step approach to Kubernetes API server monitoring. It covers inventory and version checks, safe configuration changes, verification and diagnostics, failure modes and recovery, and an operations checklist. Every section includes concrete commands, expected outputs, and decision criteria. The focus is on operational safety: observe before changing, limit blast radius, use placeholders instead of secrets, verify results, and document recovery paths.

This article is for developers, DevOps consultants, and technical startup teams who operate Kubernetes clusters and need to monitor the API server effectively. It assumes basic familiarity with Kubernetes and the command line.

Version and Environment Inventory

Before making any changes, you must understand the cluster's version, topology, and current state. This section walks through the essential inventory commands and explains what information to capture and why it matters.

Identify the Kubernetes Version and Deployment Topology

First, determine the Kubernetes version. The API server's behavior, available metrics, and configuration flags vary by version. Run:

kubectl version --short

Expected output (example for v1.28):

Client Version: v1.28.2
Kustomize Version: v5.0.4-0.20230601165947-6ce0bf390ce3
Server Version: v1.28.2

If the server version is not shown, ensure your kubeconfig points to the correct cluster and that you have network access to the API server endpoint.

Next, identify how the API server is deployed. In managed clusters (e.g., EKS, GKE, AKS), the control plane is managed by the cloud provider, and you may not have direct access to the API server process or its metrics endpoint. In self-managed clusters, the API server typically runs as a static pod on control plane nodes or as a systemd service.

To check if the API server is a static pod:

kubectl get pods -n kube-system -l component=kube-apiserver

Expected output if static pod:

NAME                    READY   STATUS    RESTARTS   AGE
kube-apiserver-node1   1/1     Running   0          10d

If no pods are listed, check for a systemd service:

systemctl status kube-apiserver

Capture Current Observable State (Read-Only)

Collect baseline information without changing anything. Use the following commands to gather health and configuration snapshots.

API server health endpoints (available in Kubernetes 1.26+):

kubectl get --raw /readyz?verbose

This returns a list of health checks and their statuses. Example output:

[+]ping ok
[+]log ok
[+]etcd ok
[+]poststarthook/start-kube-apiserver-admission-initializer ok
...
readyz check passed

For older versions, use /healthz:

kubectl get --raw /healthz

Expected output:

ok

API server metrics endpoint (if accessible):

curl -k https://<api-server-ip>:6443/metrics

If you are using a managed cluster, you may need to port-forward or use the cloud provider's monitoring integration.

Record the following in your inventory:

  • Kubernetes version (client and server)
  • API server deployment method (static pod, systemd, managed)
  • Control plane node names and IPs
  • Health endpoint output
  • Relevant configuration flags (see next section for how to retrieve them)
  • Timestamp of when the data was collected

Prerequisites and Blast Radius for Changes

Before making any configuration change, confirm:

  • You have kubectl access with cluster-admin privileges or equivalent.
  • You can access control plane nodes (for self-managed clusters).
  • You have a rollback plan: for static pods, the kubelet will restart the API server if the manifest changes; for managed clusters, configuration changes are typically via the provider's API.
  • You are aware of the blast radius: changing API server flags affects the entire cluster. Always test in a staging environment first, if possible.

Version and Environment Inventory Example

Suppose you have a self-managed cluster on v1.28, with one control plane node. You run the inventory commands and record:

  • Server version: v1.28.2
  • API server pod: kube-apiserver-node1 in kube-system
  • /readyz?verbose output shows all checks passed except etcd which shows a warning (latency above threshold).
  • Flags from the static pod manifest include --etcd-servers=https://127.0.0.1:2379 and --audit-log-path=/var/log/kube-apiserver-audit.log.

This baseline is critical before any tuning.

Safe Configuration Path

This section describes how to safely review and adjust API server configuration, focusing on metrics and monitoring-related flags.

Review Current API Server Configuration

To view the API server's command-line flags, inspect the static pod manifest (for self-managed clusters):

sudo cat /etc/kubernetes/manifests/kube-apiserver.yaml

Look for flags related to monitoring and metrics. Common ones include:

  • --metrics-bind-address: IP address on which to serve metrics (default 127.0.0.1).
  • --metrics-port: Port for metrics (default 6443? Actually, metrics are served on the secure port by default, but there is a separate --metrics-port for the insecure metrics endpoint in older versions; in newer versions, metrics are available at /metrics on the secure port).
  • --audit-log-path and audit policy flags: for audit logging, which is valuable for security monitoring.
  • --profiling: enables profiling endpoints, which can be useful for debugging but may expose sensitive information if not secured.

For managed clusters, use the provider's command or console to view the equivalent configuration.

Change Metrics Exposure Safely

If you need to expose metrics to Prometheus or another scraper, ensure the API server's metrics are reachable. By default, in many clusters, the API server serves metrics on the same secure port (6443) at /metrics. However, if you need to change the bind address or enable a separate insecure metrics endpoint (not recommended for security), you would edit the manifest.

Example scenario: You want Prometheus to scrape metrics without using client certificates. You could enable the insecure metrics endpoint on a specific port (e.g., 8080) bound to localhost or a network interface. However, this is a security risk. A better approach is to configure Prometheus with proper TLS and authentication.

Safe change example:

Suppose the API server currently has no explicit --metrics-bind-address. To make metrics available only on a specific control plane node's internal IP for Prometheus, add the flag to the manifest:

spec:
  containers:
  - command:
    - kube-apiserver
    - --metrics-bind-address=192.168.1.10
    # other flags

Then, the kubelet automatically restarts the pod. Always verify the change by checking pod status and metrics endpoint.

Prerequisites:

  • Control plane node access.
  • Backup of the existing manifest file.

Blast radius:

  • If the IP is unreachable or incorrect, metrics scraping fails (but the API server continues to serve normal API requests).
  • If the bind address is changed to a public interface without firewall rules, metrics could be exposed.

Verification:

Check pod status:

kubectl get pods -n kube-system -l component=kube-apiserver

Ensure the pod restarts and is Running. Then test the endpoint:

curl -k https://192.168.1.10:6443/metrics | head -n 5

Expected output includes lines like:

# HELP apiserver_request_total Counter of apiserver requests broken out for each verb, dry run value, group, version, resource, scope, component, and HTTP response code.
# TYPE apiserver_request_total counter

Recovery:

If the pod fails to start, revert the manifest change and restore the backup. Check logs:

kubectl logs -n kube-system <apiserver-pod-name>

Look for errors indicating invalid flag value or address already in use.

Enable and Configure Audit Logging (Optional but Important)

Audit logs provide a detailed record of API requests, which is valuable for monitoring and incident investigation. To enable audit logging, you need to set audit policy and log path flags.

Example addition to the API server manifest:

- --audit-log-path=/var/log/kube-apiserver-audit.log
- --audit-policy-file=/etc/kubernetes/audit-policy.yaml
- --audit-log-maxage=30
- --audit-log-maxbackup=10
- --audit-log-maxsize=100

You must also create the audit policy file. A minimal policy that logs metadata for all requests at the Metadata level:

apiVersion: audit.k8s.io/v1
kind: Policy
rules:
- level: Metadata

Place this file on the control plane node at /etc/kubernetes/audit-policy.yaml and ensure the API server pod can read it (often by mounting a hostPath volume).

Verification:

After the pod restarts, check that the log file is created and receiving entries:

sudo tail -f /var/log/kube-apiserver-audit.log

You should see JSON lines with request details.

Recovery:

If the API server fails to start, review the logs for errors such as invalid audit policy or permission issues. Correct or remove the flags.

Verification and Diagnostics

Now that the API server is running and metrics are exposed, you need to verify that monitoring is working and be able to diagnose issues.

Verify Metrics Collection

If you use Prometheus, check that your scrape configuration includes the API server. Typical Prometheus job for Kubernetes API server:

- job_name: 'kubernetes-apiservers'
  kubernetes_sd_configs:
  - role: endpoints
  scheme: https
  tls_config:
    ca_file: /var/run/secrets/kubernetes.io/serviceaccount/ca.crt
  bearer_token_file: /var/run/secrets/kubernetes.io/serviceaccount/token
  relabel_configs:
  - source_labels: [__meta_kubernetes_namespace, __meta_kubernetes_service_name, __meta_kubernetes_endpoint_port_name]
    action: keep
    regex: default;kubernetes;https

This assumes the API server is exposed via the kubernetes service in the default namespace. Verify that Prometheus is scraping it:

curl -s 'http://prometheus:9090/api/v1/targets' | jq '.data.activeTargets[] | select(.labels.job=="kubernetes-apiservers") | {health, lastError, scrapeUrl}'

Expected output (healthy):

{
  "health": "up",
  "lastError": "",
  "scrapeUrl": "https://10.96.0.1:443/metrics"
}

If health is "down", check network connectivity, TLS certificates, and service account token permissions.

Key Metrics for Diagnostics

Here are critical API server metrics to monitor and what they indicate:

MetricDescriptionWhat to Watch For
apiserver_request_totalCounter of API requests by verb, resource, codeSudden spikes or drops; high error rate
apiserver_request_duration_secondsLatency of API requestsHigh p99; slow requests may indicate etcd or API server overload
apiserver_current_inflight_requestsNumber of requests currently being processedApproaching or exceeding --max-requests-inflight
apiserver_longrunning_gaugeLong-running requests (e.g., watches)Unusually high count may indicate too many watchers
etcd_request_duration_secondsLatency of etcd requests from API serverHigh latency suggests etcd performance issues
apiserver_registered_watchersNumber of watch registrationsMay grow unbounded if clients do not close connections
apiserver_request_terminations_totalRequests terminated due to client disconnect or timeoutHigh termination rate may indicate client timeouts
workqueue_adds_total, workqueue_depthAPI Priority and Fairness metricsQueue depth growth indicates request backlog

Diagnostic Commands and Expected Outputs

Use the following commands to query these metrics directly from the API server (if you have access) or via Prometheus.

Check current inflight requests:

curl -k -s https://<api-server-ip>:6443/metrics | grep apiserver_current_inflight_requests

Expected output:

# HELP apiserver_current_inflight_requests Maximal number of currently used inflight request limit of this apiserver per request kind in last second.
# TYPE apiserver_current_inflight_requests gauge
apiserver_current_inflight_requests{requestKind="mutating"} 2
apiserver_current_inflight_requests{requestKind="readOnly"} 5

If the mutating requests count is near the limit (default 200), the API server may be under heavy write load.

Check request latency percentiles:

Via Prometheus:

histogram_quantile(0.99, sum(rate(apiserver_request_duration_seconds_bucket{job="kubernetes-apiservers"}[5m])) by (le, resource, verb))

This returns the 99th percentile latency per resource and verb. If the p99 exceeds your SLO (e.g., 1 second), investigate.

Check etcd latency:

histogram_quantile(0.99, sum(rate(etcd_request_duration_seconds_bucket[5m])) by (le, operation))

If etcd p99 latency is high, the API server will be slow.

Troubleshooting Common Issues

  • High API server latency: Check etcd health and latency, API server request rate, and node resource usage (CPU, memory). Also check if the API server is configured with --max-requests-inflight too high.
  • API server memory usage growing: Could be a memory leak or too many watch streams. Check apiserver_longrunning_gauge and apiserver_registered_watchers.
  • Frequent 5xx responses: Check API server logs and metrics like apiserver_request_total{code="500"}. Also check etcd availability.

Failure Modes and Recovery

This section covers common failure scenarios for the Kubernetes API server and how to recover from them.

Failure Mode: API Server Pod CrashLoopBackOff

Symptoms:

  • kubectl get pods -n kube-system shows the API server pod restarting repeatedly.
  • Pod status is CrashLoopBackOff.
  • kubectl get --raw /readyz fails.

Diagnosis:

Check pod logs:

kubectl logs -n kube-system <apiserver-pod> --previous

Common causes:

  • Invalid configuration flag (e.g., incorrect etcd URL, bad admission plugin name).
  • Certificate expired or missing.
  • Insufficient permissions on mounted files (e.g., audit log path unwritable).
  • Resource constraints (memory limit too low).

Recovery:

For static pods, check the manifest file on the control plane node. Ensure syntax and values are correct. If you recently changed a flag, revert it. If certificates are expired, renew them using kubeadm certs renew (if using kubeadm) or the appropriate method.

Example: revert a bad flag by editing /etc/kubernetes/manifests/kube-apiserver.yaml and removing or correcting the flag. The kubelet will restart the pod. Then verify pod is Running and health endpoints pass.

Failure Mode: API Server Unreachable (No Pod Restart)

Symptoms:

  • API server pod is Running but API requests time out.
  • Health endpoint may or may not respond.

Diagnosis:

Check the API server process CPU and memory usage on the node:

top -p <pid>

Check system logs for OOM kills or other events:

journalctl -u kubelet | grep -i apiserver

Check etcd availability:

ETCDCTL_API=3 etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key endpoint health

Recovery:

If the API server is overloaded, increase resource limits in the pod manifest (after confirming node capacity). If etcd is down, restore etcd from backup or resolve etcd quorum issues.

Failure Mode: Metrics Endpoint Not Working

Symptoms:

  • Prometheus cannot scrape metrics.
  • curl to /metrics returns 404 or connection refused.

Diagnosis:

Check if the metrics path is enabled and correct. In newer versions, metrics are always available at /metrics on the secure port. If you previously changed --metrics-bind-address or port, verify the flag value.

Check network connectivity from the Prometheus server to the API server.

Recovery:

  • If using a separate metrics port, ensure the port is open in firewall rules.
  • If using TLS, ensure certificates are valid.
  • Adjust the Prometheus scrape configuration if the endpoint changed.

Failure Mode: Audit Log Filling Disk

Symptoms:

  • API server pod may be evicted or node disk pressure.
  • Audit log file size exceeds limit.

Diagnosis:

Check disk usage:

df -h /var/log

Check the size of the audit log:

ls -lh /var/log/kube-apiserver-audit.log

Recovery:

  • Rotate the log manually: sudo logrotate -f /etc/logrotate.d/kube-apiserver-audit (if configured).
  • Increase --audit-log-maxsize and --audit-log-maxbackup if needed.
  • Consider shipping logs to an external system to avoid disk filling.

Operations Checklist

This checklist summarizes the key monitoring and alerting tasks for the Kubernetes API server. It includes owners and review frequencies.

Daily Checks (Automated via Alerts)

  • API server availability: Monitor /readyz endpoint. Alert if not passing for more than 2 minutes.
  • Owner: On-call SRE (rotating).
  • Review: Incident review after alert fires.
  • API server error rate: Alert on sum(rate(apiserver_request_total{code=~"5.."}[5m])) / sum(rate(apiserver_request_total[5m])) > 0.01 (1% error rate over 5 minutes).
  • Owner: Platform Team Lead (Priya Shah).
  • Review: Weekly metrics review.
  • API server latency: Alert on p99 latency > 1 second for 10 minutes.
  • Owner: API Platform Engineer (Carlos Mendes).
  • Review: Monthly SLO review.
  • etcd latency: Alert on p99 etcd request duration > 500ms for 5 minutes.
  • Owner: Database Reliability Engineer (Aisha Khan).
  • Review: Weekly.
  • Inflight requests near limit: Alert when apiserver_current_inflight_requests > 80% of limit for 5 minutes.
  • Owner: Cluster Administrator (David Lee).
  • Review: Weekly capacity review.

Weekly Checks

  • Review API server metrics trends: request rate, latency, error rate, watch counts.
  • Ensure audit logs are being generated and archived properly.
  • Verify Prometheus scrape targets are healthy.
  • Check for any API server pod restarts in the last 7 days.

Monthly Checks

  • Review alert thresholds and adjust based on actual traffic patterns.
  • Test disaster recovery: simulate API server failure in a staging cluster and practice recovery procedures.
  • Rotate certificates if they are approaching expiration (check with kubeadm certs check-expiration).
  • Review RBAC permissions for monitoring tools and users.

Common Pitfalls and How to Avoid Them

  1. Ignoring etcd monitoring: The API server heavily depends on etcd. If etcd is slow, the API server is slow. Always monitor etcd alongside the API server.
  • Avoid: Set up etcd dashboards and alerts.
  1. Alert fatigue due to overly sensitive thresholds: Setting latency alerts too low (e.g., 100ms p99) can trigger false positives. Base thresholds on actual SLOs and baseline measurements.
  • Avoid: Measure baseline for 2-4 weeks before finalizing alert thresholds.
  1. Not securing metrics endpoints: Exposing metrics without authentication or network restrictions can leak sensitive data (e.g., request URIs). Always use TLS and authentication for metrics endpoints.
  • Avoid: Use Kubernetes RBAC and network policies.
  1. Overlooking audit logs: Audit logs are critical for security investigations. Not enabling them leaves a blind spot. Enable at least Metadata level logging.
  • Avoid: Start with minimal audit policy and refine based on compliance needs.
  1. Not having a rollback plan for configuration changes: Editing the API server manifest without a backup can lead to extended downtime if the change breaks the server.
  • Avoid: Always back up manifests before editing and test changes in staging.

Conclusion

Monitoring the Kubernetes API server is a foundational practice for cluster reliability. By systematically inventorying your environment, safely configuring metrics and audit logging, verifying data collection, and preparing for failure modes, you can detect and resolve issues before they impact users.

Start with the basics: check your API server health endpoints, ensure metrics are being collected, and set up a few high-signal alerts. Then expand to dashboards and advanced diagnostics as needed. Remember to document your configurations, test your recovery procedures, and continuously refine your monitoring based on actual incidents.

A reliable technical workflow makes failure visible, protects sensitive values, limits changes to the intended resource, and defines recovery verification before an incident forces the decision. Apply these principles to your Kubernetes API server monitoring, and you will have a robust, observable control plane.

Article Quality Score

Reader usefulness 100%
  • check_circle Reader-ready guide
  • check_circle Practical examples included
  • check_circle Clean SEO article URL