E-NO
Docker Overlay Networking upgrade 6 Min Read

Docker Overlay Networking Upgrade and Migration: A Comprehensive Guide

calendar_today Published: 2026-10-02
update Last Updated: 2026-10-02
analytics SEO Efficiency: 97%
Technical guide illustration for Docker Overlay Networking Upgrade and Migration: A Comprehensive Guide.

Introduction

Docker Overlay Networking lets containers on different hosts communicate directly over a virtual, encrypted network. As Docker releases new engine versions, the overlay implementation changes, bringing new features, bug fixes, and occasionally breaking changes. Upgrading or migrating overlay networks is a critical operation that can disrupt running services if not carefully managed.

This guide provides a practical, step-by-step approach to planning and executing Docker Overlay Networking upgrades and migrations. We cover inventorying your current environment, scoping a safe rollout, verifying the new configuration, handling common failures, and building a repeatable operational checklist.

Whether you are upgrading Docker Engine, moving from a legacy overlay network to a new one, or migrating workloads between Swarm clusters, this guide will help you reduce risk and ensure a smooth transition.

Version and Environment Inventory

Before changing anything, document your current state. Start by checking Docker versions, Swarm node status, existing overlay networks, and attached services across every node in the cluster.

On each node, check the Docker Engine version:

docker version --format '{{.Server.Version}}'

Example output: 20.10.12 or 23.0.1. Note any version mismatches. Swarm can run mixed engine versions, but large gaps can cause API compatibility issues and unstable overlay behavior.

Verify Swarm membership and node status:

docker node ls

Ensure all nodes are Ready and Active and that managers are Reachable. A Down or Unknown node can break overlay routing.

List existing overlay networks:

docker network ls --filter driver=overlay

Inspect each network to capture its configuration, including subnet, gateway, encryption, and any custom options:

docker network inspect <network-name>

The output shows the IPAM configuration, attachable status, and whether encryption is enabled. Save this output for comparison after the upgrade.

Determine which services are attached to each overlay network:

docker service ls --format 'table {{.Name}}	{{.Networks}}'

This lists every service and the networks it is connected to. Note services that attach to multiple networks, such as a reverse proxy that connects to both a frontend and backend overlay.

Before making any change, verify that VXLAN UDP port 4789 is open between all nodes that will participate in the overlay. Use nc or telnet to test:

nc -z -u <remote-node-ip> 4789

If this fails, overlay traffic will not flow, and containers on different nodes will not be able to communicate.

Create an inventory table to track details. In production, this might be a spreadsheet or a document in your configuration management system. An example:

NodeDocker VersionRoleOverlay NetworksNotes
node120.10.12Managerprod-overlay, test-overlayLeader
node220.10.12Workerprod-overlay, test-overlay
node319.03.15Workerprod-overlayNeeds upgrade

This inventory is your source of truth for the migration. Without it, you risk missing a service that depends on a network you are about to delete.

Safe Configuration Path

Upgrading overlay networking usually means upgrading Docker Engine or changing network configurations. A safe approach is to scope the change to a single low-risk network or service and use a controlled pilot.

Step 1: Plan the Pilot

Choose a low-traffic overlay network to test the new configuration. For example, if you have a test-overlay network used by internal tools, start with that instead of the production network. If possible, pick a service that can tolerate brief downtime and has a small number of replicas.

Step 2: Prepare the New Configuration

If you are upgrading Docker Engine, follow the official upgrade procedure for your operating system. For example, on Ubuntu:

sudo apt-get update
sudo apt-get install docker-ce docker-ce-cli containerd.io

However, do not upgrade all nodes at once. Upgrade one worker node first, verify overlay connectivity, then proceed to the next node.

For a network migration, create a new overlay network with the desired settings before removing the old one. Example: create a new encrypted overlay network with a specific subnet:

docker network create \
  --driver overlay \
  --opt encrypted=true \
  --subnet 10.10.0.0/24 \
  --gateway 10.10.0.1 \
  new-prod-overlay

The --opt encrypted=true enables IPsec encryption for overlay traffic. The --subnet and --gateway options give you control over the IP addressing, which is useful if you need to avoid conflicts with existing networks or meet compliance requirements. You can also attach custom labels for tracking:

docker network create \
  --driver overlay \
  --label environment=staging \
  --label owner=platform-team \
  staging-overlay

Labels help you filter networks later with docker network ls --filter label=environment=staging.

If you need to migrate an existing overlay network's configuration but keep the same name, you cannot simply edit the network in place. Docker networks are immutable for most parameters. Instead, create a new network with the desired settings, migrate services to it, and then delete the old one.

Step 3: Attach a Service to the New Network

Attach a low-risk service to the new network while keeping the old network attached. This gives you a fallback path:

docker service update \
  --network-add new-prod-overlay \
  my-service

This command triggers a rolling update of the service. Docker reschedules the service tasks with the new network attachment. The service continues to be available on the old network during the transition.

Step 4: Monitor and Verify

Check that the service can communicate with other services on the new network. The easiest way is to run an ephemeral container attached to the same network and use ping or curl:

docker run --rm -it \
  --network new-prod-overlay \
  alpine \
  ping my-service

If the service has multiple replicas, use the service name, and Docker's embedded DNS will return the virtual IP (VIP) of the service, load-balancing across replicas.

Also check the service's logs for errors:

docker service logs my-service

If the service fails to start on the new network, you will see connection refused or DNS resolution errors in the logs.

Step 5: Remove the Old Network Attachment

Once you have verified that the service works correctly on the new network, remove the old network attachment:

docker service update \
  --network-rm old-prod-overlay \
  my-service

This triggers another rolling update. Monitor the service state to ensure it remains healthy:

docker service ps my-service

Look for tasks in Running state with no recent failures. If any task fails, Docker will attempt to reschedule it. Check the logs for the root cause.

Step 6: Remove the Old Network

After all services have been migrated off the old network, delete it:

docker network rm old-prod-overlay

If you get an error that the network is still in use, use docker network inspect old-prod-overlay to see which containers or services are still attached. Detach them first.

Always keep the old network as a fallback until you are confident the new one works under production load. A common timeline is to keep the old network for one to two weeks after the migration, then delete it during a maintenance window.

Quick check 1 of 2

What is a prerequisite for attaching a container to an overlay network?

The passage states that a prerequisite for attaching a container to an overlay network is that the hosts have joined the same Swarm.

Verification and Diagnostics

After applying changes, verify that the overlay network is functioning correctly. Here are key checks to run:

1. Network Connectivity

From within a container on the overlay network, ping another container by service name:

docker exec <container-id> ping <target-service-name>

Expected: successful ICMP replies. If pings fail, check that both services are on the same overlay network and that the VXLAN port is open.

2. DNS Resolution

Check that embedded DNS resolves service names to the correct virtual IP:

docker exec <container-id> nslookup <target-service-name>

Expected: output showing the virtual IP address of the target service. Typically something like 10.10.0.2. If the name does not resolve, verify that the service is attached to the correct network and that Docker's embedded DNS is running.

3. Overlay Encryption Status

If you enabled encryption, verify that it is active:

docker network inspect new-prod-overlay | grep -i encrypt

Expected: "Encrypted": true in the network configuration. If it shows false, encryption is not active. This could mean the --opt encrypted=true flag was not applied correctly, or the Docker daemon on some node does not support encryption.

4. Swarm Routing Mesh

Check that the routing mesh is functioning by accessing a published port from outside the swarm:

curl http://<any-node-ip>:<published-port>

Expected: a response from the service, regardless of which node the service is running on. If you get no response, check that the port is published correctly and that the service is healthy.

5. Logs and Events

Inspect Docker daemon logs for errors related to networking:

journalctl -u docker | grep -i network

Look for messages about overlay network state, VXLAN errors, or encryption handshake failures. These logs often point to the root cause of connectivity problems.

Run these checks after every migration step, not just at the end. Early detection reduces the blast radius of a bad change.

Failure Modes and Recovery

Even with careful planning, failures can occur. The most common failure modes are:

Network Partition

What happens: Nodes cannot communicate with each other because of a firewall misconfiguration, a security group change, or a routing issue. Overlay traffic uses VXLAN on UDP port 4789, so that port must be open between all nodes.

How to detect: nc -z -u <remote-node-ip> 4789 fails. Containers on different nodes cannot ping each other.

Recovery: Re-open the port, verify connectivity, and confirm that containers can communicate. If using a cloud provider, check security group rules and network ACLs.

Version Incompatibility

What happens: Mixed Docker versions across Swarm nodes can cause overlay instability. For example, a node running Docker 19.03 may not fully interoperate with a node running Docker 24.0 due to changes in the overlay control plane.

How to detect: Intermittent connectivity, tasks failing to schedule on certain nodes, or errors in the Docker daemon logs about unsupported features.

Recovery: Roll back the upgraded node to the previous Docker version, or upgrade all nodes to the same version. To roll back Docker Engine on Ubuntu:

sudo apt-get install docker-ce=<previous-version> docker-ce-cli=<previous-version> containerd.io

Then restart the Docker daemon and verify the node rejoins the swarm correctly.

Service Disruption During Rolling Update

What happens: A rolling update of a service goes wrong, causing some tasks to fail. This can happen if the service's health check fails on the new network, or if the new network is misconfigured.

How to detect: docker service ps my-service shows tasks in Failed or Rejected state. The service's replica count drops below desired.

Recovery: Reattach the service to the old network immediately:

docker service update --network-add old-prod-overlay my-service

This restores connectivity while you diagnose the issue. You can also roll back the service to its previous spec with docker service rollback my-service.

Overlay Encryption Handshake Failure

What happens: When encryption is enabled, nodes exchange keys over the control plane. If a node's clock is significantly skewed or its certificates are invalid, the encryption handshake fails, and all overlay traffic is dropped.

How to detect: Containers cannot communicate even though port 4789 is open. Docker daemon logs show errors like Failed to establish IPsec security association.

Recovery: Synchronize clocks on all nodes using NTP. Check the validity of the swarm certificates. If necessary, remove the node from the swarm and rejoin it.

Recovery Checklist

When a failure occurs, follow this sequence:

  1. Identify the scope of the failure. Which services are affected? Which nodes?
  2. Roll back the most recent change. If you upgraded Docker, downgrade it. If you changed a network, reattach the old network.
  3. Verify services are reachable on the fallback network.
  4. Analyze logs to find the root cause. Check Docker daemon logs, container logs, and network traces.
  5. Plan a new attempt with adjusted configuration. Do not try the same change again without understanding why it failed.
  6. Document the incident and the fix in your operations runbook.

Operations Checklist

Use this checklist for any future overlay networking upgrades or migrations. Each item has an owner and a review cadence to ensure accountability.

Pre-Change Checklist

  • [ ] Complete inventory of nodes, versions, networks, and services. Owner: Platform Engineer. Review: before every migration.
  • [ ] Document current network configurations (subnets, encryption, options). Owner: Platform Engineer. Review: before every migration.
  • [ ] Test connectivity between all nodes on VXLAN port 4789. Owner: Network Engineer. Review: before every migration.
  • [ ] Choose a low-risk pilot network or service. Owner: Service Owner. Review: at migration planning meeting.
  • [ ] Create a new overlay network with desired settings. Owner: Platform Engineer. Review: before attaching any service.
  • [ ] Back up the Swarm state on all manager nodes. Owner: Platform Engineer. Review: before any change.

Migration Checklist

  • [ ] Attach pilot service to new network alongside old. Owner: Service Owner. Review: immediately after attachment.
  • [ ] Verify connectivity, DNS, and encryption on the pilot. Owner: Platform Engineer. Review: within 1 hour of attachment.
  • [ ] Gradually migrate all services to new network. Owner: Service Owners. Review: per service migration.
  • [ ] Remove old network attachments and then old network. Owner: Platform Engineer. Review: after all services are verified.
  • [ ] Update documentation and inventory. Owner: Platform Engineer. Review: within 1 week of migration.

Post-Change Checklist

  • [ ] Schedule regular reviews of Docker versions and network settings. Owner: Platform Engineering Manager. Review: quarterly.
  • [ ] Run a post-mortem after any failed migration. Owner: Platform Engineering Manager. Review: within 5 business days.

By assigning a single owner to each step and defining when that step is reviewed, you avoid the bystander effect and ensure the migration happens consistently.

Quick check 2 of 2

How can you enable encryption for an overlay network?

The passage specifies that the `--opt encrypted` flag enables IPsec encryption for overlay network application data.

Common Pitfalls and How to Avoid Them

Even experienced teams hit the same obstacles when working with overlay networks. Here are the most common pitfalls, why they happen, and how to avoid or recover from them.

1. Forgetting the VXLAN Port

Why it happens: In a private data center, firewall rules may be managed by a separate team. When adding new nodes to the swarm, the necessary UDP port 4789 rule is easily forgotten.

How to avoid: Include port 4789 in your standard node provisioning checklist. Use infrastructure as code to enforce security group rules.

How to recover: If you discover the port is closed, open it on all affected nodes, then test with nc -z -u. Overlay connectivity should resume immediately.

2. Not Backing Up Swarm State

Why it happens: Docker Swarm does not have a built-in backup mechanism, so many teams skip backups until they need them.

How to avoid: Regularly back up the /var/lib/docker/swarm directory on all manager nodes. Use a cron job or a configuration management tool to automate this.

How to recover: If Swarm state is lost or corrupted, restore the backup directory and restart the Docker daemon. The swarm should recover its service definitions and network configurations.

3. Deleting an Old Network Too Early

Why it happens: After a successful migration, there is a temptation to clean up immediately.

How to avoid: Keep the old network for a defined period, such as two weeks, and set a reminder to delete it later. Monitor for any stragglers that still depend on the old network.

How to recover: If you delete a network prematurely, you must recreate it and reattach affected services. This can cause downtime, so avoid it by waiting.

4. Not Testing DNS Resolution

Why it happens: Many teams test connectivity by pinging an IP address, not a service name. They miss DNS issues because the network layer works.

How to avoid: Always test with service names, not IPs. Use nslookup or a simple curl to a service name to verify DNS resolution.

How to recover: If DNS fails, check that the service is attached to the correct network and that Docker's embedded DNS is running. Restarting the service may also refresh its DNS records.

5. Upgrading All Nodes at Once

Why it happens: In a small swarm, it seems efficient to upgrade all nodes in a single maintenance window.

How to avoid: Upgrade nodes one at a time, starting with a worker. Verify overlay connectivity after each node upgrade before moving to the next.

How to recover: If you upgrade all nodes and hit a compatibility issue, you must roll back all nodes, which is more disruptive. Avoid by staging the upgrade.

6. Ignoring Encryption Overhead

Why it happens: Enabling overlay encryption adds CPU overhead for IPsec processing. Teams may enable it without considering the performance impact on network-intensive services.

How to avoid: Benchmark your services with encryption enabled in a test environment. If performance is unacceptable, consider using a dedicated encryption offload or limiting encryption to sensitive networks.

How to recover: If encryption causes performance problems, you can disable it by recreating the network without the --opt encrypted flag. Migrate services to the new unencrypted network.

Example Migration Workflow

To illustrate the process, here is a complete example of migrating a production service from an old overlay network to a new encrypted one.

Assume you have a Swarm cluster with two manager nodes and three worker nodes. The service webapp is attached to the network legacy-overlay, which has no encryption and uses the default subnet. You want to migrate to a new network secure-overlay with encryption and a custom subnet.

  1. Inventory

Run the inventory commands from the first section. Document that webapp is the only service on legacy-overlay and that it runs on all three worker nodes.

  1. Create new network
   docker network create \
     --driver overlay \
     --opt encrypted=true \
     --subnet 10.20.0.0/24 \
     --gateway 10.20.0.1 \
     secure-overlay
  1. Attach webapp to new network while keeping old
   docker service update \
     --network-add secure-overlay \
     webapp

This triggers a rolling update. Monitor progress:

   docker service ps webapp

Wait until all tasks are running and old tasks are stopped.

  1. Verify connectivity on new network

Run an ephemeral container:

   docker run --rm -it \
     --network secure-overlay \
     alpine \
     ping webapp

You should see replies from the virtual IP of webapp.

Also test DNS:

   docker run --rm -it \
     --network secure-overlay \
     alpine \
     nslookup webapp

The output should show an address in the 10.20.0.0/24 subnet.

  1. Remove old network attachment

After confirming traffic flows on secure-overlay, detach webapp from legacy-overlay:

   docker service update \
     --network-rm legacy-overlay \
     webapp

Monitor again to ensure no task failures.

  1. Delete old network

Once webapp is only on secure-overlay, delete the old network:

   docker network rm legacy-overlay
  1. Verify and update documentation

Run the verification checks from the earlier section. Update your inventory and network diagrams.

This workflow can be adapted to any service. For larger migrations, batch services by dependency group and move one group at a time.

Conclusion

Upgrading and migrating Docker Overlay Networking requires careful planning and deliberate execution. A thorough inventory, a scoped pilot, rigorous verification, and a solid rollback plan are the keys to minimizing disruption. Use the operations checklist to standardize future changes and assign clear ownership.

Start by applying this guide in a test environment. Document your specific steps, adjust the checklist to your organization's standards, and refine your process based on lessons learned. With these practices, your overlay network upgrades will be smooth, predictable, and low-risk.

Related Research

Article Quality Score

Reader usefulness 97%
  • check_circle Reader-ready guide
  • check_circle Practical examples included
  • check_circle Clean SEO article URL