E-NO
Kubernetes Node Pressure Eviction monitoring 7 Min Read

Kubernetes Node Pressure Eviction Monitoring and Alerts: A Practical Guide

calendar_today Published: 2026-08-29
update Last Updated: 2026-08-29
analytics SEO Efficiency: 100%
Technical guide illustration for Kubernetes Node Pressure Eviction Monitoring and Alerts: A Practical Guide.

Intro

Kubernetes nodes can experience resource pressure when pods consume more CPU, memory, or disk than the node can provide. When this happens, the kubelet may evict pods to reclaim resources and keep the node stable. Node pressure eviction is a critical mechanism for maintaining cluster health, but it can also cause unexpected application downtime if not properly monitored.

This guide focuses on monitoring and alerting for node pressure eviction. We will cover the key metrics to watch, how to set up alerts that trigger before evictions occur, log signals that indicate impending pressure, dashboard examples, and incident response workflows. By the end, you will have a practical monitoring setup that helps you detect and respond to resource pressure before it impacts your workloads.

Version and Environment Inventory

Before implementing monitoring, establish the versions and topology of your cluster. The kubelet's eviction behavior can vary between Kubernetes versions, so it is important to know what you are working with.

Prerequisites:

  • A running Kubernetes cluster (version 1.20 or later recommended for stable eviction features).
  • kubectl command-line tool configured to access the cluster.
  • A monitoring stack that can scrape metrics from the kubelet and Kubernetes API server. Prometheus is commonly used, but other systems like Datadog or Grafana Cloud can also work.
  • Access to node logs, typically via journalctl or a log aggregation system.

Topology considerations:

  • Identify which nodes are most likely to experience pressure (e.g., nodes with high pod density or limited resources).
  • Determine if you are using any node auto-scaling; if so, eviction thresholds may interact with scaling decisions.

Run the following command to get the Kubernetes version:

kubectl version --short

Expected output (example):

Client Version: v1.25.0
Server Version: v1.25.0

For node resource capacity, use:

kubectl describe nodes | grep -A 5 "Capacity:"

Example output:

Capacity:
  cpu:                4
  ephemeral-storage:  100Gi
  memory:             16Gi
  pods:               110

This inventory will inform the thresholds you set for alerts. For instance, if your nodes have 16Gi of memory, you might set a warning alert when available memory drops below 2Gi.

Quick check 1 of 2

What is the purpose of the `evictionHard` setting in kubelet configuration?

The `evictionHard` setting configures the thresholds for eviction signals like memory.available. When the signal falls below the threshold, kubelet attempts to evict pods to reclaim resources.

Safe Configuration Path

When monitoring node pressure eviction, you need to collect the right metrics and configure alerts based on them. The kubelet exposes eviction-related metrics on its /metrics endpoint, which is typically scraped by Prometheus.

Key metrics to monitor:

  • kubelet_evictions: counter of evictions by reason (e.g., memory, disk). This is the primary signal that an eviction has occurred.
  • node_memory_MemAvailable_bytes: available memory on the node, useful for predicting memory pressure.
  • node_filesystem_avail_bytes: available filesystem space, for disk pressure.
  • kube_node_status_condition: condition of the node, including MemoryPressure, DiskPressure, and PIDPressure. This metric has labels for condition and status.

Example Alerting Rules: Below are sample Prometheus alert rules. Adjust thresholds based on your node capacity and workload patterns.

groups:
- name: node-pressure
  rules:
  - alert: NodeMemoryPressure
    expr: kube_node_status_condition{condition="MemoryPressure",status="true"} == 1
    for: 5m
    labels:
      severity: warning
    annotations:
      summary: "Node {{ $labels.node }} is under memory pressure"
  - alert: NodeDiskPressure
    expr: kube_node_status_condition{condition="DiskPressure",status="true"} == 1
    for: 5m
    labels:
      severity: warning
    annotations:
      summary: "Node {{ $labels.node }} is under disk pressure"
  - alert: HighEvictionRate
    expr: increase(kubelet_evictions[1h]) > 3
    labels:
      severity: critical
    annotations:
      summary: "High eviction rate on node {{ $labels.node }}"

Important: The kube_node_status_condition metric reports the node condition as a time series with value 1 when the condition is true. This is reliable for alerting on pressure conditions. However, some clusters may not expose this metric if the kube-state-metrics service is not installed. Ensure kube-state-metrics is deployed.

Scoped implementation: Start with a single node or a small set of nodes to validate your metrics and alerts. Test the alert rules in a staging environment if possible. This aligns with the principle that a narrow pilot is easier to inspect and adjust locally.

Additionally, consider creating a preemptive alert based on available memory to catch pressure before the node condition flips to true. For example:

- alert: NodeMemoryLow
  expr: node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes < 0.10
  for: 5m
  labels:
    severity: warning
  annotations:
    summary: "Node {{ $labels.node }} has less than 10% memory available"

This alert can give you earlier warning than waiting for the kubelet to set the MemoryPressure condition.

Verification and Diagnostics

Once you have configured metric collection and alerts, verify that everything works as expected. This involves checking that metrics are being scraped, alerts fire when conditions are met, and you can diagnose issues from logs.

Verify metric scraping: Check if Prometheus is scraping kubelet metrics by querying for a known metric:

kubelet_evictions

If no data appears, check the Prometheus targets page to ensure the kubelet endpoints are reachable.

Simulate memory pressure (for testing only): To test your alerts, you can temporarily create a pod that consumes a large amount of memory. Example:

apiVersion: v1
kind: Pod
metadata:
  name: memory-hog
spec:
  containers:
  - name: memory-hog
    image: polinux/stress
    command: ["stress"]
    args: ["--vm", "1", "--vm-bytes", "2G", "--vm-hang", "1"]

Apply this pod and observe if the node condition changes to MemoryPressure and if your alert fires. Be cautious and remove the pod immediately after testing with kubectl delete pod memory-hog.

Diagnostic commands:

  • Check node conditions:
kubectl get nodes -o custom-columns=NAME:.metadata.name,MemoryPressure:.status.conditions[?(@.type=="MemoryPressure")].status,DiskPressure:.status.conditions[?(@.type=="DiskPressure")].status

Expected output (example):

NAME       MemoryPressure   DiskPressure
node-1     False            False
node-2     True             False
  • View eviction events:
kubectl get events --all-namespaces | grep Evicted
  • Inspect kubelet logs for eviction decisions:
journalctl -u kubelet | grep -i evict

Look for messages like "evicting pod" or "pressure eviction" to understand the cause.

Dashboard example: A Grafana dashboard can visualize eviction metrics. Key panels:

  • Eviction rate over time (using rate(kubelet_evictions[5m])).
  • Node memory usage vs. eviction threshold.
  • Node conditions status.
  • Number of pods in Evicted state.

This dashboard helps in quickly identifying trends and potential issues.

Quick check 2 of 2

Which node condition is reported when the available memory on a node falls below the eviction threshold?

The kubelet maps the eviction signal `memory.available` to the `MemoryPressure` node condition when the threshold is satisfied.

Failure Modes and Recovery

Monitoring itself can fail, and evictions can have unintended consequences. It's important to understand potential failure modes and how to recover.

Monitoring failure modes:

  • Prometheus not scraping kubelet metrics due to network or authentication issues. Fix by checking service discovery and RBAC permissions.
  • Alert rules misconfigured (e.g., wrong metric names). Validate rules with promtool check rules.
  • Alertmanager not routing alerts. Test notification channels.

Eviction-related failure modes:

  • Aggressive eviction thresholds can cause pod churn. Review kubelet eviction settings (--eviction-hard, --eviction-soft).
  • Evicted pods may not be rescheduled if resources are still tight, leading to service disruption. Set up pod disruption budgets to protect critical workloads.
  • Disk pressure can be caused by container logs or ephemeral storage. Ensure log rotation is configured.

Recovery steps:

  • Identify the cause of pressure (memory, disk, PID).
  • If memory: check for memory leaks, adjust pod limits, or scale down.
  • If disk: clean up unused images or logs, increase node storage.
  • If PID: adjust pidLimit or identify processes leaking PIDs.

Rollback of monitoring changes: If a monitoring configuration change causes issues, revert to the previous version. Keep configuration in version control for easy rollback. For alert rules, document the baseline thresholds.

Operations Checklist

Use the following checklist for ongoing operations and periodic review:

  • Verify that kubelet metrics are being collected continuously.
  • Review alert thresholds quarterly to align with workload changes.
  • Test alert notifications monthly to ensure they reach the right channels.
  • Monitor eviction rate trends; a sudden increase may indicate capacity issues.
  • Ensure node conditions (MemoryPressure, DiskPressure, PIDPressure) are regularly checked.
  • Keep kube-state-metrics and Prometheus versions up to date.
  • Document incident response runbooks for eviction events.
  • Review pod resource requests and limits to avoid overcommitment.
  • Check disk usage on nodes and clean up when necessary.
  • Review kubelet eviction configuration after cluster upgrades.
TaskFrequencyResponsible Role
Metric collection verificationDailyDevOps Engineer
Alert threshold reviewQuarterlySRE
Notification testMonthlyOn-call
Eviction rate trend analysisWeeklyCapacity Planner
Node condition checkContinuous (automated)Monitoring System
Incident runbook reviewBi-annuallyTeam Lead

This checklist ensures that monitoring remains effective and that the team is prepared for eviction events.

Conclusion

Monitoring Kubernetes node pressure eviction is essential for maintaining cluster stability and application availability. By collecting key metrics, setting up proactive alerts, and having a clear incident response plan, you can prevent many eviction-related disruptions.

We covered the inventory of versions and environment, safe configuration of monitoring, verification methods, failure modes, and an operations checklist. The next steps are to implement these practices in your environment, starting with a small pilot, and iterate based on observed behavior.

Remember to keep your monitoring aligned with your cluster's capacity and workload patterns. Regularly review and update your alert thresholds and runbooks to ensure they remain effective. With these measures, you will be well-equipped to handle resource pressure and minimize the impact of evictions.

Related Research

Article Quality Score

Reader usefulness 100%
  • check_circle Reader-ready guide
  • check_circle Practical examples included
  • check_circle Clean SEO article URL