## Intro
Helm upgrades and migrations are everyday tasks for Kubernetes operators, yet they remain a leading source of midnight pages and production incidents. A release that worked on Helm 3.9 can behave differently on Helm 3.12; a chart that looks simple can pull in CRDs that are not ready for production; a 'quick upgrade' can cascade into a multi-team fire drill.
This guide gives you a repeatable, safe path for Helm upgrade and migration work. It is written for developers, DevOps consultants, and technical startup teams who run Helm in Kubernetes and want to move from 'it worked on my machine' to a verified, recoverable rollout. You will learn how to inventory your current state before touching anything, how to stage configuration changes safely, how to verify an upgrade with concrete commands and expected output, how to diagnose failures, how to roll back or recover when things go wrong, and how to maintain an operational checklist that turns each Helm upgrade into a controlled, auditable event.
The core idea is operational safety: observe before changing, limit the blast radius, use placeholders instead of secrets, verify the result, and document how to recover if the expected state is not reached.
## Version and Environment Inventory
Before you run helm upgrade , you need an accurate inventory of what is installed, where it is running, and what version boundaries you are crossing. Skipping this step is like doing surgery without reading the chart.
### Capture the Current State
Start with read-only commands that record the current Helm client version, the Kubernetes cluster version, and the release state. Do not rely on memory or a screenshot from last month; the cluster changes.
# Helm client version
helm version --short
# Example output: v3.12.0+g1f7d5d5
# Kubernetes server version (via kubectl)
kubectl version --short
# Example output: Client Version: v1.25.0, Server Version: v1.27.3
# List all releases in the current namespace
helm list --namespace production
# Example output:
# NAME NAMESPACE REVISION UPDATED STATUS CHART APP VERSION
# payments production 12 2024-01-15 14:03:22.123456 +0000 UTC deployed payments-1.4.2 1.4.0
Record the output with a timestamp and a unique identifier for the cluster and namespace. This is your baseline.
### Identify the Deployment Topology
Now map the release to its underlying resources. Use helm get to pull the rendered manifests and see what the chart actually deployed.
# Show the rendered Kubernetes objects for the release
helm get manifest payments --namespace production | less
Look for:
- Names of Deployments, StatefulSets, DaemonSets, Services, Ingresses, ConfigMaps, and Secrets.
- Custom Resource Definitions (CRDs) that the chart may have installed. CRDs are dangerous because they are global cluster resources and are not deleted when you uninstall a release by default.
- Resource requests and limits that might prevent scheduling or cause eviction during the upgrade.
Example of a critical finding: you see a CRD named prometheusrules.monitoring.coreos.com that was installed by an old Prometheus Operator chart. If you are migrating away from that chart, that CRD may remain and block the new chart's install.
### Understand Version Boundaries
Check the Helm compatibility matrix before you upgrade. Helm guarantees compatibility between the client and the Kubernetes server within a narrow range. For example, Helm 3.9 supports Kubernetes 1.21 to 1.24; Helm 3.12 supports Kubernetes 1.24 to 1.27. Running a Helm client that is newer than the cluster API can cause unexpected behavior.
Within a chart project, read the chart's Chart.yaml and values.schema.json to see what breaking changes exist between the installed version and the target version. Look for:
- Changes in default values that alter resource names.
- New required fields in the values file.
- Removal or rename of existing fields.
- Changes to immutable fields (e.g., StatefulSet volume claim templates) that force a delete/recreate.
Document this as a short table:
| Component | Current Version | Target Version | Breaking Change? | Source of Truth |
| Helm CLI | 3.9.2 | 3.12.0 | No | Official docs |
| Kubernetes cluster | 1.24 | 1.24 | No | kubectl version |
| payments chart | 1.4.2 | 1.6.0 | Yes: service.port moved from integer to string | Chart changelog |
| postgresql subchart | 11.9.0 | 12.1.0 | Yes: requires initdb password secret | Subchart changelog |
### Define Expected Result and Failure Signal
Before running any change, write down exactly what you expect to see as a result and what would indicate failure. For example:
- Expected result: helm list shows REVISION increased by 1 and STATUS is deployed ; kubectl rollout status deployment/payments returns successfully rolled out .
- Failure signal: helm upgrade returns a non-zero exit code; kubectl get pods shows CrashLoopBackOff or ImagePullBackOff for the new pods; helm list shows STATUS failed .
This explicit contract turns a vague 'did it work?' into a pass/fail check.
## Safe Configuration Path
Helm upgrades often fail because someone changed too many things at once. The fix is to stage changes and use built-in Helm mechanisms to reduce risk.
### Use Explicit Value Overrides
Never rely on default values for anything that matters. Always provide overrides via --set or a values file. Prefer a values file for anything that is not trivial, because it can be version-controlled and reviewed.
Example minimal overrides.yaml for a payment service upgrade:
# overrides.yaml
replicaCount: 3
image:
repository: registry.example.com/payments
tag: 1.6.0
pullPolicy: IfNotPresent
service:
type: ClusterIP
port: "8080" # note: string, because target chart 1.6.0 expects string
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
cpu: 500m
memory: 256Mi
Run the upgrade with a dry run first:
# Dry-run to see what would change
helm upgrade payments ./payments-chart --namespace production --values overrides.yaml --dry-run --debug
Then apply for real:
helm upgrade payments ./payments-chart --namespace production --values overrides.yaml --atomic --timeout 5m
The --atomic flag rolls back the release automatically if the upgrade fails, which is a simple safety net.
### Limit Blast Radius with Staging Namespaces
If you are uncertain about the chart changes, test the upgrade in a dedicated staging namespace before touching production. You can use the same chart and a copy of the values file with a different image tag or resource limits.
# Create a staging namespace
kubectl create namespace payments-staging
# Install the new chart version there
helm install payments-staging ./payments-chart --namespace payments-staging --values staging-values.yaml
# Verify
kubectl --namespace payments-staging rollout status deployment/payments-staging
This does not test production data or traffic patterns, but it catches syntax errors, missing required values, and obvious template mistakes.
### Handle Secrets Correctly
Never put real secrets in a values file that will be committed to Git. Use Kubernetes Secrets or a secret management tool. If the chart requires a secret value for upgrade (e.g., a database password change), reference an existing Secret rather than passing the value on the command line.
Example of a template that expects a secret name:
# values.yaml
postgresql:
existingSecret: "payments-db-credentials"
secretKeys:
adminPasswordKey: "password"
Then ensure the Secret exists before upgrading:
kubectl get secret payments-db-credentials --namespace production
If it does not exist, create it with the correct keys. A failed upgrade because of a missing secret is a common and preventable failure.
### Preserve History and Enable Rollback
Set a reasonable history limit for the release so you can roll back if needed. The default is 10, but you may want more for debugging.
helm upgrade payments ./payments-chart --namespace production --values overrides.yaml --history-max 20
After the upgrade, you can see the revision history:
helm history payments --namespace production
# Example output:
# REVISION UPDATED STATUS CHART APP VERSION DESCRIPTION
# 12 Mon Jan 15 15:02:11 2024 deployed payments-1.6.0 1.6.0 Upgrade complete
# 11 Mon Jan 15 14:03:22 2024 superseded payments-1.4.2 1.4.0 Upgrade complete
This history is what makes rollback possible.
## Verification and Diagnostics
After an upgrade, you must verify that the release is not just 'deployed' but actually functioning. Helm's status only tells you the release object was updated; it does not guarantee the pods are healthy or the service is working.
### Immediate Verification Commands
Run these immediately after the upgrade:
# 1. Check the release status
helm status payments --namespace production
# Look for: STATUS: deployed, and NOTES section with any chart-specific post-install instructions.
# 2. Check rollout status for the main Deployment
kubectl --namespace production rollout status deployment/payments
# Expected output: deployment "payments" successfully rolled out
# 3. Check all pods in the namespace
kubectl --namespace production get pods
# Expected: all pods Running and Ready (e.g., 3/3 Running)
If any pod is not Ready, describe it and check logs:
kubectl --namespace production describe pod
# Look at Conditions, Events, and Last State.
kubectl --namespace production logs --previous
# Shows logs from the previous crashed container, helpful for crash loops.
### Functional Validation
Depending on the service, run a simple health check from within the cluster or via port-forward.
# Port-forward to the service
kubectl --namespace production port-forward svc/payments 8080:8080 &
# Hit a health endpoint
curl -s http://localhost:8080/health
# Expected: {"status":"ok"}
# Kill the port-forward when done
kill %1
For a database upgrade, run a simple query:
kubectl --namespace production exec -it -- psql -U myuser -c "SELECT 1;"
# Expected: ?column? = 1
### Monitoring and Alerts
Check your monitoring dashboards for error rate, latency, and resource usage. Compare pre-upgrade and post-upgrade metrics. If you have Prometheus, run an instant query for the five minutes after the upgrade:
rate(http_requests_total{job="payments", status=~"5.."}[5m])
If this rate spikes, you have a problem even if pods are running.
### Diagnostic Workflow for Failures
If something is wrong, follow a systematic path:
- Look at Helm's output : helm upgrade prints errors and often tells you the exact resource that failed. Read the last 20 lines.
- Check the release status : helm status --namespace might show a FAILED status and messages.
- Inspect Kubernetes resources : kubectl describe and kubectl get events --sort-by=.metadata.creationTimestamp in the namespace.
- Check container logs : kubectl logs , and if the pod is crash looping, add --previous .
- Compare with previous revision : helm get manifest --revision --namespace to see what changed in the rendered objects.
- Look at the values diff : helm get values --namespace and compare with your overrides file.
Example of a common error:
Error: UPGRADE FAILED: cannot patch "payments" with kind Deployment: Deployment.apps "payments" is invalid: spec.selector: Invalid value: ... field is immutable
This means the chart tried to change the Deployment's selector, which is immutable. The only fix is to delete and recreate the Deployment, which causes downtime. To avoid this, check the chart's changelog for changes to selector labels before upgrading.
## Failure Modes and Recovery
Even with careful planning, upgrades fail. Know how to recover quickly.
### Common Failure Modes
- Image pull failure : The new image tag does not exist in the registry, or the cluster lacks pull credentials. Pods get ImagePullBackOff . Recovery: fix the tag or pull secret, then rerun helm upgrade with the same revision (or use --force if needed).
- Configuration error : A value was set incorrectly, causing the container to exit immediately. Pods get CrashLoopBackOff . Recovery: correct the values file and run helm upgrade again. Helm will create a new revision.
- Resource exhaustion : New pods are stuck in Pending due to insufficient CPU/memory. Recovery: adjust resource requests/limits in values and upgrade, or scale down old pods manually (risky).
- Immutable field change : As above, trying to change a Deployment selector or a StatefulSet volume claim template. Recovery: delete the resource and re-create via Helm (accepting downtime), or use a migration strategy (e.g., create new Deployment with different name and switch Service).
- CRD issues : The chart installs a new CRD that conflicts with an existing one. Recovery: check CRD ownership and remove conflicting CRD if safe, or use a different installation method.
### Helm Rollback: The Fastest Recovery
If the upgrade fails or you detect a problem after a short time, the fastest recovery is to roll back to the previous revision.
# Rollback to the previous revision
helm rollback payments 1 --namespace production
# Wait for rollout
kubectl --namespace production rollout status deployment/payments
The argument 1 is the target revision number. You can get the list from helm history .
Rollback restores the previous chart version and values. Note that rollback does not undo any external side effects (e.g., database schema changes). If your chart ran a migration job that altered the database, a rollback of the chart will not reverse the database change; you need to handle that separately.
### Recovery Beyond Rollback
If rollback is not enough (e.g., the previous revision is also broken), you may need to:
- Manually fix Kubernetes resources : e.g., delete a bad pod so a new one is created with the same spec.
- Use --force to recreate resources : helm upgrade --force deletes and recreates resources that cannot be updated. This can cause downtime, so use only when necessary.
- Restore from etcd backup : If the cluster state itself is corrupted, restore from a cluster backup following your disaster recovery plan.
- Contact the chart maintainer : For chart bugs, report the issue with full logs and your values file (sanitized).
### Post-Incident Review
After any failed upgrade, hold a quick blame-free post-mortem. Document:
- What was the intended change?
- What was the observed failure?
- What was the root cause?
- What was the recovery action and time to recover?
- What should be changed in the upgrade process to prevent recurrence?
Add the outcome to your operational runbook.
## Common Pitfalls and How to Avoid Them
Here are the most common mistakes operators make with Helm upgrades and migrations, along with why they happen and how to avoid or recover from each.
### 1. Running an Upgrade Without a Dry-Run
Why it happens : Operator is in a hurry, or they think the change is trivial. How to avoid : Always run helm upgrade --dry-run --debug first. It renders the templates and shows you exactly what would change. Review the diff for unexpected changes. Recovery : If you skipped it and the upgrade failed, roll back immediately and then run the dry-run to understand what went wrong.
### 2. Not Checking Helm and Kubernetes Version Compatibility
Why it happens : Helm CLI is auto-updated, or a new cluster is provisioned with a different version. How to avoid : Before any upgrade, run helm version --short and kubectl version --short and compare with the compatibility matrix. Use the correct Helm client for the cluster version. Recovery : If you used an incompatible Helm, the upgrade may have partially applied; use helm rollback and then use the correct Helm version to redo the upgrade.
### 3. Hardcoding Secrets in Values Files
Why it happens : Convenience; developers often put a placeholder secret in the repository and forget to replace it. How to avoid : Use existing Kubernetes Secrets and reference them by name. Use a tool like sops or sealed-secrets if secrets must be stored in Git. Never pass secrets via --set on the command line (they appear in shell history). Recovery : If a secret was exposed, rotate it immediately, update the Secret, and then run a new upgrade that uses the new secret.
### 4. Ignoring CRD Lifecycle
Why it happens : CRDs are installed once and often forgotten; charts may bundle CRDs in a crds/ directory that Helm installs but does not upgrade or delete. How to avoid : Check if your chart includes CRDs. If it does, upgrade the CRDs manually before upgrading the chart, or use a chart that manages CRDs as separate releases. Never let a chart install CRDs during an upgrade without review. Recovery : If wrong CRDs were installed, you may need to delete them and reinstall the correct version. Be careful: deleting a CRD deletes all custom resources of that type.
### 5. Updating Too Many Things at Once
Why it happens : A new chart version also changes values, and the operator modifies several other parameters at the same time. How to avoid : Split the upgrade into smaller steps. First upgrade the chart version with the same values (if compatible). Then change values one at a time. Use version control to track changes. Recovery : If the combined upgrade fails, roll back and then try a staged approach.
### 6. Not Setting a Verification Timeout
Why it happens : Helm waits indefinitely for resources to become ready, which can hang the terminal and the CI pipeline. How to avoid : Always set --timeout to a reasonable value (e.g., 5m0s ). Use --atomic to auto-rollback on failure. In CI, never run Helm without a timeout. Recovery : If a previous upgrade is hung, you can interrupt it (Ctrl+C) and then check the release status; Helm will mark it as failed if it timed out, and you can roll back.
### 7. Forgetting to Update the Operations Runbook
Why it happens : The upgrade is successful, so nobody documents the new state. How to avoid : After every successful upgrade, update the version inventory, the known-good revision, and any changed operational steps in your runbook. Recovery : If undocumented, future upgrades become guesswork. Start by capturing the current state and rebuilding the inventory.
## Operations Checklist
Use this checklist before, during, and after each Helm upgrade or migration. Each item has an explicit owner (an individual, not a team) and a revisit cadence.
### Pre-Upgrade Checklist
| # | Item | Owner | Cadence | Verification |
| 1 | Confirm Helm CLI version and Kubernetes cluster version compatibility | Priya Shah, Engineering Lead | Every upgrade | helm version --short and kubectl version --short output recorded in runbook |
| 2 | Inventory current release revision, chart version, and app version | Priya Shah | Every upgrade | helm list --namespace <ns> output saved with timestamp |
| 3 | Review chart changelog for breaking changes and migration notes | Alex Chen, DevOps Consultant | Every upgrade | Changelog highlights added to upgrade ticket |
| 4 | Prepare values file with explicit overrides; never use defaults for critical settings | Alex Chen | Every upgrade | Values file reviewed and merged in Git |
| 5 | Ensure all required Kubernetes Secrets exist | Sam Rivera, Platform Engineer | Every upgrade | kubectl get secret <name> --namespace <ns> succeeds |
| 6 | Run helm upgrade --dry-run --debug and inspect diff | Priya Shah | Every upgrade | Dry-run output shows expected changes only |
### During Upgrade Checklist
| # | Item | Owner | Cadence | Verification |
| 7 | Execute upgrade with --atomic --timeout 5m | Priya Shah | Every upgrade | Command exits 0 and release status becomes deployed |
| 8 | Watch pod rollout and logs for immediate errors | Alex Chen | Every upgrade | kubectl rollout status deployment/<name> returns success |
| 9 | Monitor dashboards for error rate and latency spikes | Sam Rivera | First 30 minutes after upgrade | No alert triggered on golden metrics |
### Post-Upgrade Checklist
| # | Item | Owner | Cadence | Verification |
| 10 | Run functional smoke tests (health endpoint, DB query) | Priya Shah | Every upgrade | Expected response received |
| 11 | Record new revision number and updated chart version in runbook | Alex Chen | Every upgrade | Runbook updated within 1 hour of upgrade |
| 12 | Update post-incident review if any failure occurred | Priya Shah | After any failure | Post-mortem completed within 48 hours |
| 13 | Schedule a review of the upgrade process for continuous improvement | Sam Rivera | Quarterly | Process review meeting held and action items tracked |
Use this checklist as a template. Adapt it to your organization's naming conventions and tools, but keep the core checks.
## Conclusion
Helm upgrade and migration is a skill that rewards preparation and punishes shortcuts. The difference between a routine upgrade and a production outage often comes down to whether you checked the version compatibility, ran a dry-run, set a timeout, and had a rollback plan ready.
In this guide, you learned a full operational workflow: inventory your environment, stage configuration changes safely, verify the upgrade with commands and expected output, diagnose failures methodically, recover quickly with rollback or other means, and codify the process with a checklist that names owners and cadences.
The next step is to apply this to your environment. Choose one low-risk Helm release, follow the checklist, and perform a safe upgrade. Record the current state, run the documented checks, compare results with expected signals, and note any surprises. Then expand to more critical releases, always keeping the same discipline.
Remember: a reliable technical workflow makes failure visible, protects sensitive values, limits changes to the intended resource, and defines recovery verification before an incident forces the decision. That is how you make Helm upgrades boring, and in operations, boring is good.