Introduction
Kubernetes Virtual IPs (VIPs) are the stable network identities that clients use to reach services running in a cluster. They are not bound to a single pod or node; instead, they are implemented by the control plane and data plane components such as kube-proxy, cloud controller managers, and network plugins. Over time, clusters need to upgrade these components, migrate from one VIP implementation to another (for example, from iptables mode to IPVS mode), or move services between different load balancer technologies. Done incorrectly, these changes can cause widespread application downtime, dropped traffic, or broken network policies.
This guide provides a practical, step-by-step approach to upgrading and migrating Kubernetes Virtual IPs. It is written for platform engineers, DevOps consultants, and technical startup teams who need to perform these operations safely. We focus on real commands, expected outputs, failure signals, and recovery decisions. The goal is operational safety: observe before changing, limit the blast radius, protect sensitive data, verify every step, and document recovery paths before you need them.
Throughout the article, we will use a running example of upgrading kube-proxy from iptables mode to IPVS mode, and migrating a service from a cloud provider's legacy load balancer to a newer implementation. The principles apply to other VIP-related components, but this concrete scenario keeps the guidance grounded.
1. Version and Environment Inventory
Before any upgrade or migration, you must know exactly what is running. This includes cluster version, node operating system, kube-proxy mode and version, network plugin, and the cloud provider's load balancer controller version. A version mismatch between components can cause VIPs to behave unpredictably.
Read-Only Observation Commands
Start with read-only commands to capture the current state. These do not modify anything and can be run safely at any time.
Check cluster and node versions:
kubectl version --short
kubectl get nodes -o wide
Expected output shows the client and server versions, and for each node its internal IP, external IP (if any), and OS. For example:
Client Version: v1.28.2
Server Version: v1.27.6
NAME STATUS ROLES AGE VERSION INTERNAL-IP EXTERNAL-IP OS-IMAGE
node1 Ready control-plane 12d v1.27.6 10.0.0.4 <none> Ubuntu 22.04.3 LTS
node2 Ready <none> 12d v1.27.6 10.0.0.5 <none> Ubuntu 22.04.3 LTS
Check kube-proxy mode and version. The mode is usually set via a ConfigMap or command-line flag. On each node, you can inspect the running kube-proxy pod (if it runs as a DaemonSet) or the process.
If kube-proxy runs as a DaemonSet:
kubectl get daemonset -n kube-system kube-proxy -o wide
kubectl logs -n kube-system <kube-proxy-pod> | grep -i mode
Expected log line for iptables mode:
I1012 10:00:00.000000 1 server_others.go:212] Using iptables Proxier.
For IPVS mode:
I1012 10:00:00.000000 1 server_others.go:212] Using ipvs Proxier.
Check the cloud provider's load balancer controller version, if applicable. For example, on AWS, check the AWS Load Balancer Controller deployment:
kubectl get deployment -n kube-system aws-load-balancer-controller -o yaml | grep image
Prerequisites and Compatibility
Before changing anything, confirm that the new version or mode is compatible with your cluster version and network plugin. For example, IPVS mode requires kernel modules ip_vs, ip_vs_rr, ip_vs_wrr, ip_vs_sh, and nf_conntrack. You can check for these on a node:
lsmod | grep ip_vs
If they are not loaded, you may need to install them or load them manually.
Also review the Kubernetes changelog for any changes to kube-proxy or Service VIP behavior. For cloud load balancer migrations, check the provider's migration guide for supported annotations and CRDs.
Smallest Justified Change
Never perform a big-bang upgrade. For example, if you are moving from iptables to IPVS, do it one node at a time or use a rolling update of the kube-proxy DaemonSet. For a service migration, start with a test service that has minimal traffic, or use a canary deployment.
2. Safe Configuration Path
Once you have the inventory, plan the configuration change. The key is to make the change in a controlled, reversible way.
Back Up Current Configuration
Before modifying any component, export its current configuration. For kube-proxy, the configuration is often in a ConfigMap:
kubectl get configmap -n kube-system kube-proxy -o yaml > kube-proxy-config-backup.yaml
For a service, export its YAML:
kubectl get service my-service -o yaml > my-service-backup.yaml
Store these backups in a version-controlled location.
Apply Changes Gradually
For kube-proxy mode change, you can edit the ConfigMap and then restart kube-proxy pods. However, a rolling update of the DaemonSet is safer. If you use a configuration file, modify it and apply:
kubectl edit configmap -n kube-system kube-proxy
Change mode: "" to mode: "ipvs" (or from iptables to ipvs), then save. After saving, trigger a rolling update:
kubectl rollout restart daemonset -n kube-system kube-proxy
Monitor the rollout:
kubectl rollout status daemonset -n kube-system kube-proxy
Expected output:
Waiting for daemon set "kube-proxy" rollout to finish: 0 out of 3 new pods have been updated...
Waiting for daemon set "kube-proxy" rollout to finish: 1 out of 3 new pods have been updated...
...
daemon set "kube-proxy" successfully rolled out
For service migration, you might create a new service with the new load balancer annotations and gradually shift traffic. For example, if migrating from a legacy AWS ELB to a Network Load Balancer (NLB), you would create a new service with the appropriate annotations:
apiVersion: v1
kind: Service
metadata:
name: my-service-nlb
annotations:
service.beta.kubernetes.io/aws-load-balancer-type: "nlb"
spec:
selector:
app: my-app
ports:
- protocol: TCP
port: 80
targetPort: 8080
type: LoadBalancer
Apply it, then update DNS or ingress to point to the new VIP gradually.
Use Placeholders Instead of Secrets
When testing configuration changes, never use real secrets. Use dummy values or environment-specific placeholders (e.g., example.com for DNS, 10.0.0.0/16 for CIDR) in your test manifests. This prevents accidental leakage and makes it clear what needs to be customized per environment.
3. Verification and Diagnostics
After applying changes, you must verify that the VIP is functioning correctly. This involves checking connectivity, observing data plane behavior, and inspecting logs.
Check Service and Endpoints
First, confirm that the service has a valid VIP and that endpoints are populated:
kubectl get service my-service -o wide
kubectl get endpoints my-service
Expected output for a healthy service:
NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE
my-service LoadBalancer 10.96.0.100 203.0.113.10 80:30080/TCP 1h
NAME ENDPOINTS AGE
my-service 10.244.1.5:8080 1h
If endpoints are empty, the selector may not match any pods. Check pod labels.
Test Connectivity
From within the cluster, test access to the ClusterIP:
kubectl run test-pod --image=alpine --restart=Never --rm -it -- sh
/ # wget -qO- http://10.96.0.100
You should see the response from your application. For a LoadBalancer service, test the external IP from outside the cluster (or from a node).
Diagnose kube-proxy Issues
If traffic fails, check kube-proxy logs for errors:
kubectl logs -n kube-system <kube-proxy-pod> --tail=50
Look for messages about failing to sync rules, IPVS errors, or connectivity to the API server. For IPVS mode, you can inspect the IPVS rules on the node:
ipvsadm -Ln
Expected output shows virtual services and their real servers, for example:
IP Virtual Server version 1.2.1 (size=4096)
Prot LocalAddress:Port Scheduler Flags
-> RemoteAddress:Port Forward Weight ActiveConn InActConn
TCP 10.96.0.100:80 rr
-> 10.244.1.5:8080 Masq 1 0 0
If the IPVS rules are missing or incorrect, kube-proxy may not be running in IPVS mode, or there may be a permission issue.
Validate Traffic Shift
When migrating services, use a canary approach: send a small percentage of traffic to the new VIP and monitor error rates. For example, use a simple script with curl to hit both endpoints and compare responses. Or, if you have an ingress controller, adjust its routing rules to split traffic.
4. Failure Modes and Recovery
No matter how careful you are, things can go wrong. Plan for common failure modes and know how to recover quickly.
Common Failure Modes
- kube-proxy pods crash after mode change: This often happens when required kernel modules are missing or permissions are insufficient. The pod logs will show errors. Recovery: check node kernel modules, ensure kube-proxy has the necessary privileges (usually it runs as privileged or with CAP_NET_ADMIN), or roll back to the previous mode by editing the ConfigMap back and restarting.
- Service VIP becomes unreachable after migration: This could be due to firewall rules, network policy changes, or cloud load balancer misconfiguration. Recovery: check cloud provider console for load balancer status, verify security groups, and test connectivity step by step. If needed, revert the service YAML to the previous configuration using the backup.
- Partial rollout leaves mixed modes: If only some nodes have switched to IPVS, traffic through those nodes may behave differently (e.g., different scheduling algorithms). Recovery: complete the rollout or roll back all nodes by changing the ConfigMap back and restarting all kube-proxy pods.
- DNS resolution fails for the VIP: If you are using ExternalDNS or similar, the new VIP may not be registered in DNS. Recovery: check ExternalDNS logs, ensure the service annotations are correct, and manually verify DNS records.
Rollback Procedure
Always have a rollback plan ready before making changes. For kube-proxy mode change, rollback is straightforward:
kubectl edit configmap -n kube-system kube-proxy
# change mode back to previous value
kubectl rollout restart daemonset -n kube-system kube-proxy
For service migration, you can simply delete the new service and reapply the old one if traffic is not yet fully shifted. If traffic has been shifted and you need to roll back, update the DNS or ingress to point back to the old VIP, then remove the new service after confirming no traffic is going to it.
Monitoring and Alerting
Set up monitoring for key metrics: kube-proxy sync duration, iptables/ipvs errors, service endpoint availability, and load balancer health checks. Use Prometheus and Grafana dashboards to visualize these. Alerts should trigger if endpoints are empty for a service, if kube-proxy is not running on any node, or if external IP is not responding.
5. Operations Checklist
Use this checklist before, during, and after the upgrade or migration. Each item has an owner and a review frequency. The owner is a single accountable person, not a team.
| Step | Action | Command / Signal | Owner | Review Frequency |
|---|---|---|---|---|
| 1. Inventory | Capture cluster and component versions | kubectl version --short, kubectl get nodes -o wide, check kube-proxy mode logs | Priya Shah, Platform Lead | Before each upgrade, quarterly baseline |
| 2. Backup | Export current configs | kubectl get configmap -n kube-system kube-proxy -o yaml > backup.yaml and service YAML backups | Carlos Mendez, SRE | Every change, stored in git |
| 3. Compatibility | Verify prerequisites for new version/mode | Check kernel modules, cloud provider docs | Lin Wei, Network Engineer | Per upgrade |
| 4. Test in staging | Apply change to a test cluster first | Run through the migration steps in staging | Priya Shah | Before production change |
| 5. Production rollout | Apply change gradually | kubectl edit configmap ..., kubectl rollout restart daemonset ... | Carlos Mendez | During change window |
| 6. Verify | Test VIP connectivity | kubectl run test-pod ... wget http://<cluster-ip> | Lin Wei | Immediately after rollout |
| 7. Monitor | Watch for errors and traffic anomalies | Check kube-proxy logs, load balancer metrics | Carlos Mendez | For 24 hours post-change |
| 8. Document | Update runbooks and architecture diagrams | Write post-incident report if any issues | Priya Shah | Within 1 week after change |
6. Common Pitfalls and How to Avoid Them
In our experience, the following mistakes are common when dealing with Kubernetes VIPs.
Pitfall 1: Changing Too Much at Once
Why it happens: Teams want to minimize downtime and bundle multiple changes into one maintenance window. For example, upgrading kube-proxy and changing network plugin at the same time.
How to avoid: Make only one change at a time. If you must upgrade multiple components, sequence them with testing in between. For example, upgrade kube-proxy first, verify, then upgrade the network plugin.
Recovery: If you already made a bundled change and problems occur, roll back to the last known good configuration for all changed components, then re-apply changes one by one.
Pitfall 2: Not Checking Kernel Modules for IPVS
Why it happens: IPVS mode requires specific kernel modules that may not be loaded on minimal OS images. The kube-proxy pod may start but fail to program IPVS rules, leading to dropped traffic.
How to avoid: Before switching to IPVS, run lsmod | grep ip_vs on all nodes. If modules are missing, install ipvsadm and load the modules, or configure the OS to load them at boot. Some managed Kubernetes services handle this automatically, but self-managed clusters require manual checks.
Recovery: If you discover missing modules after the change, you can either load them on the fly and restart kube-proxy, or roll back to iptables mode until the modules are properly installed cluster-wide.
Pitfall 3: Ignoring Endpoint Slices
Why it happens: Endpoint Slices are a newer API for tracking endpoints, and older components may not be compatible. During an upgrade, if kube-proxy is updated but the EndpointSlice controller is not, VIP routing may break.
How to avoid: Ensure all components that consume endpoints are updated together or are backward compatible. Check the Kubernetes version compatibility matrix.
Recovery: If endpoints are not being populated in EndpointSlices, check the controller logs and roll back if necessary.
Pitfall 4: Not Testing with a Canary
Why it happens: Teams apply a new LoadBalancer service and immediately switch all traffic via DNS, hoping for the best.
How to avoid: Use a canary approach. Create the new service alongside the old one, test it with a small subset of traffic (e.g., using weighted DNS or an ingress controller), and gradually increase the percentage while monitoring error rates.
Recovery: If the new service fails under canary traffic, simply remove the canary routing rule and investigate without affecting most users.
Conclusion
Upgrading and migrating Kubernetes Virtual IPs is a critical operation that requires careful planning, observation, and verification. The key takeaways from this guide are:
- Always start with a thorough inventory of versions and configurations.
- Make changes in small, reversible steps.
- Use concrete commands to verify functionality at each stage.
- Know the common failure modes and have a rollback plan.
- Maintain an operations checklist with clear owners and review frequencies.
By following these practices, you can minimize downtime and ensure your services remain reachable throughout the upgrade or migration. As a next step, choose one low-risk verification from this guide, such as checking kube-proxy mode or testing service connectivity, and run it on your cluster to establish a baseline.
Remember: a reliable technical workflow makes failure visible, protects sensitive values, limits changes to the intended resource, and defines recovery verification before an incident forces the decision.