ClusterTriage / Blog / Networking

Live Migration on the Wrong Network: The Silent Hyper-V Pitfall

Live Migration quietly using the wrong network is the single most common finding in ClusterTriage Hyper-V cluster assessments. It is the kind of misconfiguration that does not fail, it just runs slow, competes with cluster heartbeats during patch windows, and produces incidents that look like switch problems when they are really mis-routed VM memory transfers.

This article walks through the three configuration layers, the PowerShell to verify the actual path being used, and the remediation we apply on every engagement.

By Hans Vredevoort · 17 May 2026 · 12 minute read · Networking

1. The three configuration layers

Network selection for Live Migration runs across three independent configuration layers. We regularly find clusters where two are correct and the third silently breaks the whole thing.

Cluster network role. Each cluster network (subnet) has a Role: None, Cluster, or ClusterAndClient. The cluster only considers networks with role Cluster or ClusterAndClient for internal traffic, including Live Migration.

Cluster Live Migration priority. Inside the cluster object sits an ordered list of networks the cluster prefers for Live Migration. Set it through Set-ClusterParameter or through the GUI under Networks → Live Migration Settings.

Per-host VM Migration network list. Each Hyper-V host has its own VMMigrationNetwork list. That is what vmms actually uses to bind to a source IP when starting a migration. If the cluster says "use 10.10.20.0/24" but the host has no VMMigrationNetwork entry for that subnet, the host falls back to whichever interface responds, usually the management network.

Why this stays quiet The cluster logs no warning when fallback occurs. No event. The migration completes successfully, only over the wrong NIC. The only way to know this is verifying the actual TCP connection during the migration, or watching per-NIC byte counters during a test.

2. What "wrong network" actually looks like

The classic symptoms in production:

  • Migrations take 4 to 6 minutes where 30 to 45 seconds should be achievable. Usually because LM is going over a 1 GbE management network instead of the 25 GbE storage network.
  • Cluster heartbeat drops during CAU runs. If LM and heartbeat share a NIC, big memory transfers can saturate the link and mark heartbeats as missing.
  • Switch port utilisation on the management VLAN peaks at 90% plus during patching. Network ops notices and opens a ticket. The Hyper-V team makes no connection to LM.
  • SMB Multichannel sessions in unexpected places. If LM is on SMB (default since WS2016) but has no clean RDMA path, it falls back to plain TCP-over-SMB on whatever interface is reachable.

The cluster is healthy in all these scenarios by standard monitoring criteria. Test-Cluster passes. Failover Cluster Manager is green. The only correlation is timing. That is precisely why this is a typical Hyper-V cluster assessment finding, not something reactive monitoring catches. See Hyper-V cluster assessment: Top 10 Issues We Find in 2026 for the wider pattern.

3. Verifying the path in use

Verification requires a live test. Inspecting configuration alone is not enough because the three layers interact in non-obvious ways. The procedure ClusterTriage uses during a cluster assessment:

Step 1, snapshot all three configuration layers:

# Cluster networks and their roles
Get-ClusterNetwork |
    Select-Object Name, Address, Role, AutoMetric, Metric |
    Sort-Object Metric

# Cluster Live Migration priority
$lmNetworks = (Get-ClusterResourceType -Name 'Virtual Machine' |
    Get-ClusterParameter -Name MigrationNetworkOrder).Value
$lmExcluded = (Get-ClusterResourceType -Name 'Virtual Machine' |
    Get-ClusterParameter -Name MigrationExcludeNetworks).Value
Write-Host "LM order:    $lmNetworks"
Write-Host "LM excluded: $lmExcluded"

# Per-host VM migration network configuration
Invoke-Command -ComputerName (Get-ClusterNode).Name {
    [PSCustomObject]@{
        Node     = $env:COMPUTERNAME
        Enabled  = (Get-VMHost).VirtualMachineMigrationEnabled
        AuthType = (Get-VMHost).VirtualMachineMigrationAuthenticationType
        PerfMode = (Get-VMHost).VirtualMachineMigrationPerformanceOption
        Networks = (Get-VMMigrationNetwork | Select-Object -Expand Subnet) -join ', '
    }
}

Step 2, start an actual Live Migration of a VM with at least 8 GB of active memory:

Move-ClusterVirtualMachineRole -Name 'TEST-LM-VM' -Node HV02

Step 3, on the source host during the migration, capture the active TCP connections of the VMMS process:

$vmms = Get-Process vmms | Select-Object -First 1 -ExpandProperty Id
Get-NetTCPConnection -OwningProcess $vmms |
    Where-Object {$_.State -eq 'Established'} |
    Select-Object LocalAddress, RemoteAddress, LocalPort, RemotePort

# Live Migration uses TCP/6600 (or 445 in SMB mode); we want
# to see the expected storage/migration subnet on both sides.

If the local address comes back as a management subnet IP, or worse, the heartbeat IP, the configuration is broken regardless of what the three layers report.

4. The fix, explicit and validated configuration

The remediation pattern is the same on every engagement. We make all three layers explicit and consistent, and validate with a test migration.

1. Set cluster network roles correctly:

# Identify subnets by IP range first, names vary per deployment
Get-ClusterNetwork | Select-Object Name, Address, Role

# Storage network, typically Cluster only (not available for client traffic)
(Get-ClusterNetwork 'Storage1').Role = 1   # Cluster

# Management network, ClusterAndClient
(Get-ClusterNetwork 'Management').Role = 3 # ClusterAndClient

# Heartbeat network if separate, Cluster only
(Get-ClusterNetwork 'Heartbeat').Role = 1  # Cluster

2. Set the cluster Live Migration priority:

# Build the priority list, IDs not names
$preferred = Get-ClusterNetwork | Where-Object {
    $_.Name -in 'Storage1','Storage2'
}
$excluded = Get-ClusterNetwork | Where-Object {
    $_.Name -in 'Management','Heartbeat'
}

Get-ClusterResourceType -Name 'Virtual Machine' |
    Set-ClusterParameter -Name MigrationNetworkOrder -Value ($preferred.ID -join ';')

Get-ClusterResourceType -Name 'Virtual Machine' |
    Set-ClusterParameter -Name MigrationExcludeNetworks -Value ($excluded.ID -join ';')

3. Set per-host migration networks explicitly on every node:

Invoke-Command -ComputerName (Get-ClusterNode).Name {
    # Remove all existing entries first, start clean
    Get-VMMigrationNetwork | Remove-VMMigrationNetwork

    # Add only the subnets we want to use
    Add-VMMigrationNetwork -Subnet '10.20.0.0/24' -Priority 10
    Add-VMMigrationNetwork -Subnet '10.21.0.0/24' -Priority 20
}

Slow Live Migration is not a capacity problem, it is almost always a routing problem. The three layers above are validated as standard in every ClusterTriage Hyper-V cluster assessment, with before-and-after numbers in the report.

Schedule a Hyper-V cluster assessment intro call →

5. Performance options: TCP, Compression, SMB, RDMA

Hyper-V supports three Live Migration performance modes. Choosing the right one matters as much as choosing the right network.

  • TCP/IP. One TCP stream per migration, oldest option, lowest throughput on modern hardware. Avoid except for compatibility.
  • Compression. TCP with on-the-fly memory compression. Good when CPU is abundant and the network is the bottleneck. Default in Windows Server 2012 R2.
  • SMB. Uses SMB Multichannel, supports RDMA, scales automatically across multiple NICs. The right choice on any modern cluster with multiple NICs or RDMA-capable hardware. Default since Windows Server 2016.
Invoke-Command -ComputerName (Get-ClusterNode).Name {
    # SMB is what you want on modern hardware
    Set-VMHost -VirtualMachineMigrationPerformanceOption SMB
    # And concurrent migration limits, tuned to your bandwidth
    Set-VMHost -MaximumVirtualMachineMigrations 4
    Set-VMHost -MaximumStorageMigrations 2
}

When SMB mode is selected and NICs support RDMA with DCB/PFC switch-side, Live Migration automatically runs over RDMA. Verify with:

Get-SmbClientNetworkInterface | Select-Object FriendlyName, RdmaCapable, Speed
Get-SmbConnection                     # During an active migration
Get-SmbMultichannelConnection         # Session detail

On Windows Server 2025 you can layer dynamic CPU compatibility on top of this, enabling Live Migration between hosts with different CPU generations without VM shutdown. For heterogeneous clusters spanning multiple hardware refreshes, that is a noticeable operational win.

6. End-to-end validation

Configuration is not validation. The only way to be sure is to run an actual migration and confirm the path. The ClusterTriage validation procedure:

  1. Build a test VM with 16 GB RAM and dirty most of it (a memory stress script works).
  2. Start a packet counter on every NIC of the source host: Get-NetAdapter | Get-NetAdapterStatistics.
  3. Migrate the VM and time it:
    Measure-Command { Move-ClusterVirtualMachineRole -Name 'TEST-LM' -Node HV02 }
  4. Read NIC statistics again. The bytes-per-second delta on the intended migration NIC should account for nearly the entire VM RAM size. Other NICs should stay essentially flat.
  5. Repeat in the reverse direction to confirm symmetry.

We capture this in the report with before-and-after numbers. It is one of the most tangible measurements in a cluster assessment: the cluster goes from "migrations take 4 minutes and saturate management" to "migrations take 28 seconds and are invisible on management".

7. Preventing regression

Live Migration configuration drift usually happens during three operations: NIC replacement, virtual switch reconfiguration, and node reimaging. The ClusterTriage remediation runbook for this finding includes:

  • A documented baseline of Get-VMMigrationNetwork, Get-VMHost migration settings, and Get-ClusterParameter output, stored in source control.
  • A 10-line PowerShell script that compares current state to baseline and returns non-zero exit on drift. Hooked into the monthly maintenance window.
  • A two-line validation step in the post-patch checklist: run a test migration, confirm the path.
  • A note in the node build runbook: every new node must explicitly remove default VMMigrationNetwork entries before joining the cluster.

The fix for Live Migration on the wrong network is not configuration, it is the discipline to validate that configuration end to end after every change.

ClusterTriage delivers this as part of the standard Hyper-V cluster assessment. The configuration baseline and drift detection script ship with the remediation report.

Schedule a Hyper-V cluster assessment intro call →

Frequently asked questions

Why does Hyper-V not log a warning when Live Migration uses the wrong network?

Because from the cluster's perspective there is no fallback to see. The host had no VMMigrationNetwork entry for the intended subnet, so vmms picked an interface that worked. From the Cluster Service's view that is a successful migration. It is a design limitation, not a bug, and the only defence is verification of the actual path.

Does SMB mode always perform better than Compression?

On a cluster with multiple NICs or RDMA-capable hardware, yes. On a single-NIC cluster without RDMA, Compression can occasionally be slightly faster because it leans harder on CPU. In production we have not seen a scenario in the last three years where Compression was the right choice; SMB is the modern default.

What is the difference between MigrationNetworkOrder at cluster level and VMMigrationNetwork per host?

The cluster level influences which networks the cluster prefers when it coordinates a Live Migration. The per-host list is what vmms actually uses to bind a source IP. Both must be consistent. When they disagree, the per-host list wins in practice, because that is what actually establishes the TCP connection.

What is a reasonable MaximumVirtualMachineMigrations value?

On 25 GbE with SMB and RDMA, 4 to 8 parallel migrations is common. On 10 GbE without RDMA, 2 to 4. Test it during a maintenance window with your actual VM mix, because VMs with heavy active memory can bottleneck each other even on fast networks.

Does this affect Storage Live Migration too?

Yes, partially. Storage Live Migration uses SMB and shares the MaximumStorageMigrations setting. Network selection happens through SMB Multichannel, so the priority lives in SMB configuration rather than Live Migration-specific settings. The core principle still holds: know which path the traffic is actually taking.