ClusterTriage / Blog / Cluster assessment
Hyper-V cluster assessment: 10 issues we keep finding in 2026
After years of Hyper-V cluster assessments across Windows Server 2016, 2019, 2022 and 2025 environments, ten findings keep coming back. Across industries, scales and operations teams. None of them are exotic. Most stay quiet until a patch round, a hardware refresh or a single switch reboot turns them into a Sev-1 incident at 03:00 in the morning.
This article is the practical list. For each issue: root cause, the PowerShell to diagnose it, and the remediation we document in the customer report. Recognise three or more in your own environment? That is not an outlier, that is the average.
1. Live Migration over the wrong network
The most common finding. Live Migration falls back to the management or heartbeat VLAN because the dedicated migration network has not been enabled for cluster traffic, or has never been added to the explicit Live Migration network list per Hyper-V host. The symptom is rarely a failure. What you see is slow migrations during patch windows and unexplained latency on cluster heartbeats when multiple VMs move at the same time.
Two configuration planes matter. The Role per cluster network at cluster level, and the explicit Live Migration network list per host. They are independent. We regularly find one configured correctly and the other forgotten. Full detail in Live Migration on the wrong network: a common Hyper-V pitfall.
Diagnostic detail: since Windows Server 2016 the default Live Migration transport is SMB, so traffic uses port 445, not the older TCP port 6600. Filtering only on 6600 returns nothing on modern clusters.
Diagnosis:
# Cluster network roles
Get-ClusterNetwork | Select-Object Name, Role, Address, AutoMetric, Metric
# Live Migration configuration per host
Get-VMHost | Select-Object Name, VirtualMachineMigrationEnabled,
VirtualMachineMigrationAuthenticationType,
VirtualMachineMigrationPerformanceOption,
@{N='MigrationNetworks';E={(Get-VMMigrationNetwork).Subnet -join ', '}}
# Effective path during a test move (SMB mode: port 445, TCP mode: port 6600)
Get-NetTCPConnection -RemotePort 445,6600 |
Select-Object LocalAddress, RemoteAddress, OwningProcess
For certainty, a short pktmon capture on the source NIC during a manual Live Migration confirms which physical path is actually used.
Remediation: on each host, constrain the Live Migration network list to the subnet you actually want to use, and ensure that subnet has Role = ClusterAndClient (or Cluster only) on the cluster object. Add-VMMigrationNetwork adds a subnet; Remove-VMMigrationNetwork removes unwanted ones. Validate end-to-end with a manual Live Migration.
2. Witness misconfigured or absent
Quorum is the invisible piece that keeps a cluster alive when nodes disappear. We regularly find two-node clusters without a witness, File Share Witnesses on a single-node, single-disk file server without a UPS, and Disk Witnesses on the same SAN LUN as the CSVs they are supposed to protect.
On Dynamic Quorum: on Windows Server 2016 and later, Dynamic Quorum combines with Dynamic Witness and the cluster adjusts votes dynamically to survive single failures. This is not a replacement for a witness. In a simultaneous network partition or double failure without an external tiebreaker, split-brain is not reliably resolved. A witness remains required for any production configuration.
Get-ClusterQuorum | Format-List *
Get-ClusterResource | Where-Object {$_.OwnerGroup -eq 'Cluster Group'} |
Format-Table Name, State, OwnerNode, ResourceType
Decision matrix as a rule of thumb:
- Single-site two nodes: Cloud Witness, by definition outside your failure domain.
- Single-site three or more nodes: Cloud Witness, unless Azure connectivity is not an option, otherwise File Share Witness on infrastructure that is genuinely independent.
- Two-site stretched: Cloud Witness, or a File Share Witness in a third site, never in one of the two cluster sites.
- Disconnected or air-gapped: File Share Witness on documented independent storage and power.
Full comparison in File Share Witness vs Cloud Witness vs Disk Witness: which to choose.
3. CSV ownership imbalance
Cluster Shared Volume ownership piles up on one node after a patch round, a failover, or a node drain that nobody rebalanced afterwards. The symptom is uneven CPU on the coordinator node and read-heavy workloads experiencing latency spikes that look random until you correlate them with CSV ownership.
# Current distribution
Get-ClusterSharedVolume | Select-Object Name, OwnerNode, State |
Group-Object OwnerNode | Select-Object Name, Count
# Rebalance round-robin over active nodes, with a brief pause per move
$nodes = (Get-ClusterNode | Where-Object State -eq 'Up').Name
$i = 0
Get-ClusterSharedVolume | ForEach-Object {
Move-ClusterSharedVolume -Name $_.Name -Node $nodes[$i % $nodes.Count]
Start-Sleep -Seconds 5
$i++
}
Every move causes a short I/O pause on the CSV concerned. Schedule rebalance in a quiet window or after a patch round, not in the middle of the day. We treat this as an ongoing operational task; the post-patch checklist in our remediation reports always ends with a CSV rebalance. Deep dive in CSV ownership imbalance: causes and fixes.
4. Firmware and driver drift between nodes
Cluster nodes leave the factory identical and slowly drift apart. One host gets a NIC firmware update during an incident, another a new HBA driver during a planned maintenance window, a third neither. Six months later the cluster fails in subtle ways: Live Migration takes three times as long over an Intel X710 versus a Mellanox CX-4, or storage timeouts occur only on one host.
This is the leading cause of intermittent cluster instability we investigate. Firmware drift is invisible to most monitoring tools because the values sit in the BMC or NIC EEPROM, not in Windows.
# NIC driver version per node (firmware via Windows is vendor dependent)
Invoke-Command -ComputerName (Get-ClusterNode).Name {
Get-NetAdapter | Select-Object PSComputerName, Name, InterfaceDescription,
DriverVersion, DriverDate
} | Sort-Object PSComputerName, Name
# Storport and MPIO driver versions
Invoke-Command -ComputerName (Get-ClusterNode).Name {
Get-CimInstance Win32_PnPSignedDriver |
Where-Object DeviceClass -in 'SCSIADAPTER','SYSTEM' |
Select-Object PSComputerName, DeviceName, DriverVersion, DriverDate
}
For reliable firmware inventory of NICs, HBAs and BIOS, vendor tooling is required: Dell iDRAC (racadm), HPE iLO (ilorest or SUM), Lenovo XClarity Controller, Supermicro IPMI. Windows cmdlets such as Get-NetAdapterAdvancedProperty -RegistryKeyword '*FirmwareVersion' work for some Mellanox and Intel drivers, but not universally.
Remediation: treat firmware as configuration. Set a baseline (vendor solution profile from Dell, HPE or Lenovo), document it in your runbook, and upgrade the cluster as a whole during a planned window.
5. Mismatched Windows Server build numbers
The Cluster Service tolerates short-lived mixed builds during a rolling patch, but we regularly find clusters that have been mid-patch for half a year. The result is functional regressions: a fix in a later CU is intermittent because half the nodes lack it, and behaviour depends on which node owns a resource at the moment of failure.
Invoke-Command -ComputerName (Get-ClusterNode).Name {
[PSCustomObject]@{
Node = $env:COMPUTERNAME
OS = (Get-CimInstance Win32_OperatingSystem).Caption
Build = (Get-CimInstance Win32_OperatingSystem).BuildNumber
UBR = (Get-ItemProperty 'HKLM:\SOFTWARE\Microsoft\Windows NT\CurrentVersion').UBR
LastBoot = (Get-CimInstance Win32_OperatingSystem).LastBootUpTime
}
} | Sort-Object Node
If the UBR column shows different values across nodes, the cluster is in a state Microsoft Support will want repaired before they escalate. Patch the cluster as a whole, validate with Test-Cluster, and treat the post-patch state as the new known-good.
6. No anti-affinity for DCs and management VMs
After a node drain all domain controllers cheerfully end up on the same surviving host. One host crash then takes down all DCs simultaneously, and with them authentication for the entire environment, including the management plane you need to recover the cluster.
# Clustered VMs without anti-affinity
Get-ClusterGroup | Where-Object GroupType -eq 'VirtualMachine' |
Select-Object Name, OwnerNode, AntiAffinityClassNames |
Where-Object {-not $_.AntiAffinityClassNames}
# Set anti-affinity for all DCs
Get-ClusterGroup | Where-Object Name -like 'DC*' | ForEach-Object {
$g = Get-ClusterGroup -Name $_.Name
$g.AntiAffinityClassNames.Clear()
$g.AntiAffinityClassNames.Add('DomainControllers') | Out-Null
}
Anti-affinity is soft: it is violated if no other valid host exists. That is exactly the right behaviour. The cluster keeps the DCs running; it just does not stack them by default.
7. CSV BlockCache disabled or undersized
CSV BlockCache was opt-in for years and many older clusters are still set to zero. Enabling it for read-heavy workloads (VDI, file server clusters, OLTP databases with high cache locality) can lower read latency by an order of magnitude with no risk to data integrity. It is a read-side cache, not a write buffer.
# Current value
(Get-Cluster).BlockCacheSize
# Enable on a five-node cluster, 2 GB per node
(Get-Cluster).BlockCacheSize = 2048
# Verify per-CSV cache status
Get-ClusterSharedVolume | ForEach-Object {
$_ | Select-Object Name,
@{N='CacheEnabled';E={$_.SharedVolumeInfo.Partition.IsCsvCacheEnabled}}
}
Important: the new BlockCache value only takes effect on a CSV after it goes offline and back online, or after a cluster restart. In production this means a phased rebalance, not instant activation. Plan for it.
The trade-off is RAM. Size against your VM density and confirm in real workloads that the node has headroom for VM consolidation under N-1 conditions. Do not size BlockCache from the failover-evacuated state.
8. Cluster network on a single physical path
Cluster heartbeats running over a single switch or single uplink. We see this mostly in lift-and-shift cluster builds where the original two-switch design was stripped down in the procurement phase. Result: a switch reboot, a stack split or a failing SFP partitions the cluster, and the surviving partition arbitrates poorly because the witness is reachable over the same path.
Cluster networks require redundancy at Layer 1, not just at IP level. Two NICs in the same switch is not redundancy. Teamed NICs across two switches is the minimum for production.
On teaming technology: LBFO (Load Balancing and Failover) is not supported under the Hyper-V virtual switch on Windows Server 2022 and 2025. SET (Switch Embedded Teaming) is the only supported option for new deployments. Existing LBFO teams still attached to the Hyper-V vSwitch on these OS versions are a hard remediation item, not a recommendation.
During a Hyper-V cluster assessment we always trace cluster heartbeat traffic to the physical port, not just the logical NIC name.
9. Unsupported or unvalidated storage configurations
The catch-all. Examples we have documented:
- SMB Direct paths presented as RDMA-capable because the NIC supports RoCEv2, but DCB and PFC are not configured on the switch side. Silent fallback to TCP without warning. iWARP does not have this switch dependency; RoCEv2 does.
- SMB Multichannel over dissimilar NICs (one 25 GbE, one 10 GbE), where the cluster picks the slower path for half the sessions.
- MPIO with Round Robin on an active/passive array that strongly prefers one controller.
- S2D deployments where one node, after a hardware swap, has a deviating drive bay configuration, invisible to the cluster but visible in Storage Spaces health.
Validate with the native tools, production-safe variant without storage disruption:
Test-Cluster -Include 'Inventory','Network','System Configuration','Cluster Configuration' `
-ReportName Cluster Assessment
Get-SmbClientNetworkInterface | Select-Object FriendlyName, RdmaCapable, Speed, IpAddress
Get-PhysicalDisk | Group-Object MediaType, OperationalStatus, FirmwareVersion
A full storage validation including the Storage category requires a maintenance window or a separate test LUN that is not in use by a clustered role.
10. No documented DR validation
Stretched or multisite clusters that have never been failed over end-to-end. The runbook exists, the storage replication dashboard is green, and nobody has ever actually tried it in production. A Cluster Assessment often surfaces the gap between runbook and reality: DNS TTLs that are too long, application-layer connection strings hardcoded to one site, witness placement that does not survive the test.
We recommend at minimum an annual planned failover, treated as production impact and rehearsed in a maintenance window.
The pattern behind the list
Nine of these ten findings have nothing to do with workload performance. It is silent configuration debt that accumulates over patch cycles, hardware refreshes and staff turnover. It is also exactly the type of finding that monitoring tools miss, because in isolation each setting looks correct: a witness exists, a Live Migration network is configured, CSV ownership is valid. The drift sits in the relationship between settings, and the only reliable way to surface that is a structured external review.
Recognise three or more issues above? You are not alone, and the fix usually costs no more than a single weekly maintenance window.
Plan a Hyper-V cluster assessment introduction →
Frequently asked questions
A cluster assessment starts with a read-only measurement you run yourself from your management server, about fifteen minutes. The report and a one-hour review follow. In Assurance that repeats every quarter, so progress is measured as well.
The inventory scripts are read-only and run safely during business hours. The optional Test-Cluster validation can also be run on a live production cluster without impact when the storage tests are skipped, or when a dedicated test LUN is available that is not in use by a clustered role. A full storage validation against production CSVs requires those CSVs to be taken offline for the duration of the test, and we always schedule that outside business hours in agreement with the customer.
A Cluster Assessment is technical and focused on operational stability, performance and supportability. An audit is broader and often touches security and compliance. We do cluster assessments; for audits we partner with specialists.
Yes, with adapted scope. We replace the Storage Fabric and Storage Array sections with a Storage Spaces Direct section. See Azure Local migration readiness checklist.
Both. A Cluster Assessment can be delivered standalone, or as the starting point for a remediation engagement where we resolve the findings together with the customer.