## Intro

Kubernetes Horizontal Pod Autoscaler (HPA) automatically scales the number of pods in a deployment, replication controller, or stateful set based on observed CPU utilization or custom metrics. For operators and developers running production workloads, understanding HPA's architecture is essential for ensuring applications remain responsive under load while avoiding unnecessary infrastructure costs.

This article is for developers, DevOps consultants, and technical startup teams who need to move beyond basic autoscaling concepts. It explains the core components of HPA, how they interact, and how to configure, verify, and troubleshoot autoscaling in real clusters. By the end, you will be able to set up HPA safely, interpret its behavior, and recover from common misconfigurations.

We emphasize operational safety: always observe before changing, limit the blast radius, use placeholders instead of secrets, verify results, and document recovery steps.

## How Horizontal Pod Autoscaler Works

The Horizontal Pod Autoscaler is implemented as a control loop that runs periodically (default every 15 seconds) and queries the metrics server for resource utilization. It then calculates the desired number of replicas based on the ratio of current metric value to target value, and updates the target workload's replica count accordingly.

The core algorithm determines the desired replica count using the formula:

```
desiredReplicas = ceil( currentReplicas * ( currentMetricValue / desiredMetricValue ) )
```

For example, if a deployment has 2 replicas, the current CPU utilization is 200m (0.2 cores), and the target utilization is 100m (0.1 cores), the ratio is 2.0, so the desired replicas become 4. The HPA controller then scales the deployment to 4 pods.

HPA works with several Kubernetes resources:

- **Metrics Server**: The default source for CPU and memory metrics. It aggregates resource usage from kubelets and exposes them via the Metrics API. HPA periodically fetches these metrics.
- **Custom Metrics API**: For scaling based on application-specific metrics (e.g., requests per second, queue length), you can integrate with tools like Prometheus Adapter or Google Cloud Monitoring.
- **External Metrics API**: Allows scaling based on metrics from outside the cluster, such as cloud provider load balancer metrics.

### Key Components

#### 1. HPA Controller

The HPA controller is part of the kube-controller-manager. It watches for HorizontalPodAutoscaler resources and performs the scaling calculations. The controller runs a reconciliation loop that:

1. Fetches the target workload (e.g., Deployment).
2. Retrieves current metrics from the Metrics Server or custom metrics API.
3. Computes the desired replica count.
4. Updates the workload's `spec.replicas` field if the desired count differs from the current count, respecting `minReplicas` and `maxReplicas`.

#### 2. Metrics Server

The Metrics Server is a cluster-wide aggregator of resource usage data. It is not deployed by default in most clusters. You can install it with:

```bash
kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
```

Verify it is running:

```bash
kubectl get deployment metrics-server -n kube-system
```

Expected output:
```
NAME             READY   UP-TO-DATE   AVAILABLE   AGE
metrics-server   1/1     1            1           5m
```

#### 3. Metrics APIs

HPA queries the Metrics API, which is served by the metrics-server. This API provides resource metrics such as CPU and memory. For custom metrics, you need to deploy an adapter that implements the custom.metrics.k8s.io API.

## Version and Environment Inventory

Before configuring HPA, you must understand your cluster version and capabilities, because HPA behavior has evolved across Kubernetes releases. Key features like stabilization windows and scaling policies are only available in newer versions.

**Check your Kubernetes version:**

```bash
kubectl version --short
```

Example output:
```
Client Version: v1.24.3
Server Version: v1.24.3
```

**Verify HPA API availability:**

```bash
kubectl api-versions | grep autoscaling
```

Expected output should include `autoscaling/v2` (or v2beta2 in older versions) if you want to use advanced features:
```
autoscaling/v1
autoscaling/v2
```

**Check if Metrics Server is installed:**

```bash
kubectl get apiservice v1beta1.metrics.k8s.io
```

If installed, the output shows `Available: True`:
```
NAME                     SERVICE                      AVAILABLE   AGE
v1beta1.metrics.k8s.io   kube-system/metrics-server   True        10m
```

If not available, install it before proceeding.

### Prerequisites

- A running Kubernetes cluster (version 1.23+ recommended for v2 API).
- Metrics Server installed and working.
- `kubectl` configured to access the cluster.
- A deployment with resource requests set (e.g., CPU and memory). HPA cannot scale without resource requests.

**Example deployment manifest with resource requests:**

```yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: php-apache
spec:
  replicas: 1
  selector:
    matchLabels:
      run: php-apache
  template:
    metadata:
      labels:
        run: php-apache
    spec:
      containers:
      - name: php-apache
        image: registry.k8s.io/hpa-example
        ports:
        - containerPort: 80
        resources:
          limits:
            cpu: 500m
          requests:
            cpu: 200m
```

Apply it and verify:

```bash
kubectl apply -f deployment.yaml
kubectl get pods -l run=php-apache
```

Expected: one pod running.

## Safe Configuration Path

Configuring HPA safely means starting with a minimal configuration, validating its impact in a test environment, and then gradually applying to production. The following steps outline a safe approach.

### Step 1: Observe Current Baseline

Before enabling autoscaling, establish the current workload behavior under normal traffic. Use these commands:

```bash
kubectl get pods -o wide
kubectl top pods
```

The `kubectl top` command shows current resource usage:
```
NAME                         CPU(cores)   MEMORY(bytes)
php-apache-<pod-id>          5m           10Mi
```

Record the current CPU usage as a baseline.

### Step 2: Define HPA Object

Create an HPA manifest using the autoscaling/v2 API for more control. Example:

```yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: php-apache
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: php-apache
  minReplicas: 1
  maxReplicas: 10
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 50
```

This sets HPA to maintain average CPU utilization of 50% across all pods. `minReplicas` and `maxReplicas` bound the scaling.

Apply it:

```bash
kubectl apply -f hpa.yaml
```

### Step 3: Verify HPA Status

Check the HPA status and events:

```bash
kubectl get hpa
kubectl describe hpa php-apache
```

Example `kubectl get hpa` output:
```
NAME         REFERENCE               TARGETS   MINPODS   MAXPODS   REPLICAS   AGE
php-apache   Deployment/php-apache   0%/50%    1         10        1          1m
```

Here, current utilization is 0% because no traffic is hitting the pod, target is 50%.

### Step 4: Generate Load to Test Scaling

To test autoscaling, you need to generate load. Create a temporary load generator pod:

```bash
kubectl run -i --tty load-generator --rm --image=busybox:1.28 --restart=Never -- /bin/sh -c "while sleep 0.01; do wget -q -O- http://php-apache; done"
```

Keep this running in a separate terminal and watch the HPA:

```bash
kubectl get hpa php-apache --watch
```

Within a minute or two, you should see the current CPU utilization rise above the target, and the replica count increase.

### Step 5: Validate Scaling Behavior

After load stops, the HPA will scale down after the default stabilization window (5 minutes for scale-down). You can observe the events:

```bash
kubectl describe hpa php-apache
```

Look for events similar to:
```
Events:
  Type    Reason             Age   From                       Message
  ----    ------             ----  ----                       -------
  Normal  SuccessfulRescale  5m    horizontal-pod-autoscaler  New size: 4; reason: cpu resource utilization (percentage of request) above target
```

### Important Configuration Considerations

- **Resource Requests**: HPA will not work if the target workload does not have resource requests set. The utilization percentage is calculated relative to the request, not the limit.
- **Multiple Metrics**: You can specify multiple metrics; HPA calculates the desired replicas for each metric and takes the maximum.
- **Stabilization Windows**: The `behavior` field in v2 API lets you control how quickly HPA scales up or down. For example, to reduce flapping, you can set a scale-down stabilization window of 300 seconds.

Example behavior configuration:

```yaml
behavior:
  scaleDown:
    stabilizationWindowSeconds: 300
    policies:
    - type: Percent
      value: 100
      periodSeconds: 15
  scaleUp:
    stabilizationWindowSeconds: 0
    policies:
    - type: Percent
      value: 100
      periodSeconds: 15
    - type: Pods
      value: 4
      periodSeconds: 15
    selectPolicy: Max
```

This ensures scale-down only happens after metrics have been below threshold for 5 minutes, and scale-up happens immediately but limited to either 100% of current replicas or 4 pods per 15 seconds, whichever results in more replicas.

## Verification and Diagnostics

Verification goes beyond simple `kubectl get hpa`. You need to ensure HPA is making correct decisions based on accurate metrics.

### Checking Metrics Availability

First, confirm that the metrics server is providing data:

```bash
kubectl get --raw "/apis/metrics.k8s.io/v1beta1/namespaces/default/pods" | jq
```

If you see entries with CPU and memory values, metrics are flowing correctly.

### Common Diagnostic Commands

1. **View HPA events:**

```bash
kubectl describe hpa <hpa-name>
```

Look for warnings or errors. Common error messages include:
- `unable to get metrics for resource cpu: no metrics returned from resource metrics API` means metrics-server is not installed or not working.
- `missing request for cpu` indicates the deployment lacks resource requests.
- `failed to get cpu utilization: unable to get metrics for resource cpu: no metrics returned from resource metrics API` similar to above.

2. **Check pods resource usage:**

```bash
kubectl top pods
```

If `kubectl top` returns an error, metrics-server may be down.

3. **Inspect metrics server logs:**

```bash
kubectl logs -n kube-system deployment/metrics-server
```

Look for errors connecting to kubelets or API server.

4. **Validate HPA calculation:**

You can manually compute the desired replicas using the formula. For instance, if you have 2 replicas, current CPU is 300m, target is 200m (i.e., 100% utilization for a 200m request), then desired replicas = ceil(2 * (300/200)) = ceil(3) = 3.

### Verifying Scaling Policies

If you have defined `behavior`, test scale-down and scale-up events:

```bash
kubectl get events --field-selector involvedObject.name=<hpa-name> --watch
```

This shows real-time events when HPA scales.

## Failure Modes and Recovery

Several common failure modes can prevent HPA from working correctly. Here, we cover symptoms, causes, and recovery steps.

### 1. HPA Shows Unknown for Metrics

**Symptom:** `kubectl get hpa` shows `<unknown>` in TARGETS column.

**Cause:** Metrics server not reachable or not returning metrics.

**Recovery:**
- Verify metrics-server is running:
```bash
kubectl get pods -n kube-system | grep metrics-server
```
- Check metrics API availability:
```bash
kubectl get apiservice v1beta1.metrics.k8s.io
```
If not available, redeploy metrics-server.

### 2. HPA Not Scaling Despite High Load

**Symptom:** CPU usage high but replica count unchanged.

**Causes:**
- No resource requests on target pods.
- HPA maxReplicas reached.
- HPA trying to scale but failing due to resource quota or insufficient cluster capacity.
- Stabilization window preventing immediate scale-up (less common).

**Recovery:**
- Check HPA events for errors:
```bash
kubectl describe hpa <name>
```
- Ensure `maxReplicas` is high enough.
- Check cluster capacity: `kubectl get nodes` and `kubectl describe nodes` to see if resources are exhausted.
- Check resource quotas:
```bash
kubectl get resourcequota -n <namespace>
```
If quota is exceeded, increase quota or reduce other workloads.

### 3. Flapping or Oscillating Replica Counts

**Symptom:** Replica count frequently changes up and down.

**Cause:** HPA thresholds set too tightly or erratic metric patterns.

**Recovery:**
- Adjust target utilization to avoid borderline conditions.
- Use stabilization windows to smooth scaling:
```yaml
behavior:
  scaleDown:
    stabilizationWindowSeconds: 300
  scaleUp:
    stabilizationWindowSeconds: 60
```
- Consider custom metrics that better reflect actual load.

### 4. HPA Deleted Accidentally

**Symptom:** Autoscaling stops.

**Recovery:**
- Recreate HPA from manifest backup. Always keep HPA manifests in version control.

## Operations Checklist

For day-2 operations, ensure the following checklist is part of your routine:

### Weekly

- [ ] Review HPA status:
```bash
kubectl get hpa --all-namespaces
```
Check that all HPAs have valid metrics (not unknown) and replica counts within min/max.

- [ ] Check for events:
```bash
kubectl get events --all-namespaces --field-selector reason=FailedGetResourceMetric
```

### Monthly

- [ ] Validate metrics server health and upgrade if needed.
- [ ] Review scaling policies and adjust based on traffic patterns.
- [ ] Test scale-down by reducing load and confirming replicas decrease after stabilization window.
- [ ] Update HPA target thresholds based on performance reviews.

### Quarterly

- [ ] Conduct load testing to verify HPA responds appropriately under peak loads.
- [ ] Review cluster capacity and node autoscaling integration.
- [ ] Audit HPA configurations for security and compliance.

**Accountability:** Assign an owner for HPA operations. For example, in a startup, the DevOps lead (e.g., Priya Shah, Engineering Lead) should be responsible for weekly checks, and the platform team should own monthly reviews. The owner revisits the checklist and updates the owner assignment quarterly.

## Common Pitfalls and How to Avoid Them

### 1. Not Setting Resource Requests

**Why it happens:** Teams forget to specify `resources.requests` on containers, thinking limits are enough. HPA uses requests for utilization calculations.

**How to avoid:** Always set requests for CPU and memory on target workloads. Automate validation with OPA/Kyverno policies.

**Recovery:** Add requests to the deployment spec and reapply. HPA will start working once requests are present.

### 2. Using CPU Utilization Percentage Without Understanding It

**Why it happens:** The target utilization percentage is relative to the request, not the limit. For example, if request is 500m and target is 50%, HPA targets 250m. If limit is 1000m, the pod can use up to 1000m but average across pods should be 250m. Misunderstanding leads to over- or under-scaling.

**How to avoid:** Set targets based on SRE principles. For CPU, aim for 50-70% of request to maintain headroom. For memory, be cautious because memory is not compressible; set target utilization lower (e.g., 70-80%) to avoid OOM kills.

### 3. Ignoring Custom Metrics When CPU Is Not the Bottleneck

**Why it happens:** Many applications are I/O-bound or have uneven load distribution, so CPU-based scaling may be insufficient.

**How to avoid:** Use custom metrics (e.g., requests per second, queue depth) via Prometheus Adapter or other tools. Design metrics that correlate with user-facing performance.

**Recovery:** Implement custom metrics and update HPA to use them.

### 4. Not Setting Max Replicas Appropriately

**Why it happens:** Teams set `maxReplicas` too high to avoid scaling limits, but this can lead to cost overruns or cluster overload.

**How to avoid:** Determine max replicas based on capacity planning and cost constraints. Use HPA with cluster autoscaler to ensure nodes can be added when needed.

**Recovery:** Adjust `maxReplicas` after observing actual peak loads and cluster capacity.

### 5. Forgetting About Scale-Down Delays

**Why it happens:** Operators expect immediate scale-down after load drops, but HPA has a default 5-minute stabilization window to prevent flapping. This can lead to unnecessary cost if not accounted for.

**How to avoid:** Understand and configure `behavior` appropriately. For aggressive cost savings, reduce stabilization window, but be aware of flapping risk.

### 6. Using HPA with Stateful Applications Without Care

**Why it happens:** StatefulSets can be scaled by HPA, but careful coordination is needed with persistent volumes and ordering. HPA might scale down pods that are in the middle of important operations.

**How to avoid:** Use custom metrics or lifecycle hooks to delay scale-down during critical operations. Consider using Vertical Pod Autoscaler for stateful workloads.

## Conclusion

Kubernetes Horizontal Pod Autoscaler is a powerful tool for maintaining application performance and cost efficiency. By understanding its architecture, you can configure it safely and troubleshoot issues effectively. In this article, we covered:

- HPA core components and algorithm.
- Environment inventory and prerequisites.
- Step-by-step safe configuration with practical examples.
- Verification and diagnostic techniques.
- Common failure modes and recovery steps.
- Operations checklist and common pitfalls.

As a next step, choose one low-risk verification from the checklist, record your current HPA state, run the documented check, and compare results against expected signals. Review dependencies such as Deployment, Resource Quota, and Pod capacities. A reliable workflow makes failures visible, protects sensitive values, limits changes to intended resources, and defines recovery verification before incidents force decisions. With these practices, you can operate HPA with confidence.
