ClusterTriage / Blog / Storage
CSV Ownership Imbalance in Hyper-V Clusters: Causes and Fixes
After every patch round, every Live Migration storm, every planned failover, Cluster Shared Volume ownership drifts apart. One node ends up owning all the CSVs, becomes coordinator for every metadata operation, and quietly caps read performance for the entire cluster. CSV ownership imbalance is a top three finding in ClusterTriage Hyper-V cluster assessments, and also one of the simplest to fix permanently.
This article explains what actually happens under the hood, how to detect it, and how to keep CSVs balanced without constant manual intervention.
1. What CSV ownership actually means
A Cluster Shared Volume is one NTFS or ReFS volume that all cluster nodes can read and write simultaneously. Behind that simple description sits a coordinator role. One node, the coordinator node, owns the metadata operations of the volume. Allocating space, extending a VHD, creating a snapshot, changing attributes, all of that runs through the coordinator. Other nodes do direct I/O for normal reads and writes, bypassing the coordinator entirely.
This works well when the cluster is balanced. The coordinator role for one CSV asks a small amount of CPU and memory on the owner, so the load spreads naturally. It works poorly when one node owns all CSVs at once. That node then handles metadata work for every workload in the cluster, and any operation that runs through redirected I/O (which happens automatically when direct I/O is impossible) goes through that node.
Two situations force redirected I/O. The local node has lost storage connectivity to the volume (the node has no direct path, so I/O routes over the cluster network to the coordinator, which does still have a path). Or the operation is one that always runs through the coordinator (most metadata operations, some snapshot work, backup software accessing the volume in certain modes).
Redirected I/O is supported and safe. It is also significantly slower than direct I/O and can saturate cluster networks if left unchecked.
2. Why ownership drifts, three field patterns
Pattern 1: the Patch Tuesday parade
CAU (Cluster Aware Updating) drains a node, patches it, brings it back, drains the next one. Each node hands its CSVs to whichever node has capacity at that moment, usually the just-patched node, which has the lowest load. At the end of the run, the last-patched node owns most or all CSVs.
Pattern 2: the midnight failover
An unplanned failover (usually a NIC firmware glitch, a flapping storage path, or a host briefly losing its quorum vote) moves all CSVs from the failing node to one other node. Nobody notices in the morning because workloads are running. The imbalance stays until the next intervention.
Pattern 3: the "I'll set it back later" debug session
An engineer pinned all CSVs to one node to rule out a node-specific issue during an incident, fixed the issue, and never undid the pin. We find PreferredOwners values on CSV resources months after the original incident, with the engineer who set them long gone from the company.
3. Performance impact, with numbers
On a four-node Hyper-V cluster running an average VDI workload, ClusterTriage recently measured the impact during a Hyper-V cluster assessment.
Workload: 220 VDI desktops, mix of boot storm and steady state. Cluster: 4 × Dell R7625, S2D, all-NVMe, 25 GbE storage.
State A: CSVs balanced (3 per node, 12 CSVs total)
Average read latency: 0.62 ms
P99 read latency: 2.1 ms
Coordinator CPU: 5-8% per node
State B: One node owns all 12 CSVs (post-patch, no rebalance)
Average read latency: 1.08 ms
P99 read latency: 9.4 ms (4.5x worse)
Coordinator CPU: 22-28% on the owner
Storage cluster network: saturated during boot storm
In State B the cluster was technically healthy. Nothing alarmed. The only user-side symptom was complaints about login times, which the ops team was investigating under the working hypothesis of a network problem in the VDI farm. The fix took eight minutes.
A note on where these numbers come from: this was an S2D, all-NVMe cluster, and S2D is where imbalance hits hardest. On S2D the owner node also holds the CSV read cache and a large share of the storage-network routing, so piling every CSV onto one node concentrates metadata, cache locality and routing in the same place at once. On a classic SAN, every node keeps its own direct path to the LUN, so steady-state reads stay direct no matter who owns the CSV, and the read-latency penalty is usually a good deal smaller. What still hurts on SAN is the metadata load on the coordinator and anything that drops into redirected I/O, a node losing a storage path being the common trigger. The imbalance is worth correcting on both architectures, just expect it to surface in different places: steady-state read latency on S2D, metadata-heavy operations and path failures on SAN.
4. Diagnosing ownership and traffic
Three views give the full picture:
# Ownership distribution
Get-ClusterSharedVolume |
Select-Object Name, OwnerNode, State |
Group-Object OwnerNode |
Sort-Object Count -Descending |
Select-Object Name, Count
# Redirected I/O, anything not 'Direct' is suspect
Get-ClusterSharedVolume | ForEach-Object {
[PSCustomObject]@{
Name = $_.Name
Owner = $_.OwnerNode
IOMode = $_.SharedVolumeInfo.Partition.IsBlockIOMode # $true = direct
Redirected = $_.SharedVolumeInfo.RedirectedAccess
}
}
# Coordinator CPU per node, run during a normal workday
Get-Counter -Counter '\Cluster Shared Volume(*)\IO Read Bytes/sec' `
-SampleInterval 5 -MaxSamples 12
The ClusterTriage Cluster Assessment report includes a one-liner per CSV: owner, mode (direct or redirected), average read latency over the inventory window, and a flag if PreferredOwners is set to anything other than default.
5. Manual rebalance, the safe procedure
A manual rebalance is safe during business hours. Move-ClusterSharedVolume transfers the coordinator role without disrupting running VMs, the volume stays available. The only thing that happens is a brief pause in metadata operations on that CSV, measured in milliseconds.
# Identify which CSV to move
Get-ClusterSharedVolume |
Select-Object Name, OwnerNode |
Group-Object OwnerNode |
Where-Object Count -gt 3
# Move one CSV to a specific node
Move-ClusterSharedVolume -Name 'CSV-VDI-Pool-03' -Node 'HV02'
# Move and verify
Move-ClusterSharedVolume -Name 'CSV-VDI-Pool-03' -Node 'HV02' -PassThru |
Select-Object Name, OwnerNode, State
If you want to drain a node, for example for maintenance, move its CSVs explicitly instead of relying on automatic balancing:
$sourceNode = 'HV01'
$targetNodes = (Get-ClusterNode |
Where-Object {$_.State -eq 'Up' -and $_.Name -ne $sourceNode}).Name
$i = 0
Get-ClusterSharedVolume | Where-Object OwnerNode -eq $sourceNode | ForEach-Object {
$target = $targetNodes[$i % $targetNodes.Count]
Write-Host "Moving $($_.Name) from $sourceNode to $target"
Move-ClusterSharedVolume -Name $_.Name -Node $target
$i++
}
If your monitoring stack has never alerted on skewed CSV ownership and your cluster has been patched in the last six months, you are almost certainly out of balance right now. The five-line script above tells you in ten seconds.
A ClusterTriage Hyper-V cluster assessment maps this alongside nine other quiet issues in two days.
Schedule a cluster assessment intro call →6. Automating rebalance
Microsoft introduced automatic CSV balancing in Windows Server 2016. It is on by default and works well in most environments. You can verify and tune it:
(Get-Cluster).CsvBalancer
# Values:
# 0 = off
# 1 = always balance (default)
# 2 = balance only when threshold exceeded
# Force immediate re-evaluation
(Get-Cluster).CsvBalancer = 1
$cluster = Get-Cluster
$cluster.CsvBalancer = $cluster.CsvBalancer
Two cases where automatic balancing does not do what you want. CSV resources with explicit PreferredOwners, the balancer respects them. And heterogeneous node sizing, the balancer treats nodes as equals, so if one node has half the RAM/CPU of the others, even balancing distributes the load poorly.
In both cases, a scheduled PowerShell rebalance after maintenance windows is the safer pattern. ClusterTriage delivers this as part of the remediation report on every engagement where imbalance is a finding.
7. Three ways to rebalance: CSV ownership, VMM, and VM alignment
"Rebalancing" gets used for three different operations that solve different problems, and it is worth being precise about which one you mean.
- Native CSV balancing (
CsvBalancerandMove-ClusterSharedVolume, above) moves the coordinator role between nodes. It spreads metadata load and, on S2D, cache locality. It moves no VMs and has no opinion on where a VM runs relative to the CSV it lives on. - VMM Dynamic Optimization moves VMs between hosts to balance CPU and memory. It is compute-aware and storage-blind: it will live-migrate a VM onto a node that does not own the VM's CSV, which on S2D quietly adds redirected reads and cross-node storage traffic. Good for compute hotspots, but it can work against storage locality at the same time.
- Aligning VMs with their storage moves VMs the other way: onto the node that owns the CSV their disks live on, so reads stay direct and the coordinator path is short. This is the axis neither CsvBalancer nor VMM touches. ClusterTriage runs it from a site-aware placement script built on Darryl van der Peijl's original align-VMs-with-storage idea, with WhatIf safety and a before/after view.
The three are complementary, not competing. A healthy pattern: let CsvBalancer keep coordinator ownership even, let VMM (if you run it) handle compute hotspots, and run an alignment pass after maintenance so VMs sit with their storage rather than wherever the last failover left them. The full write-up of our alignment approach is in Aligning VMs With Their Storage, and the original idea it builds on is Darryl van der Peijl's align VMs with storage script.
8. CSV BlockCache and why it changes the math
CSV BlockCache is a read-side memory cache on each node. When enabled, it dramatically lowers the cost of redirected reads because most are served from local RAM. It does not solve imbalance, but it softens the consequences.
(Get-Cluster).BlockCacheSize # value in MB, 0 = off
# Enable with 2 GB per node, tune to your RAM budget
(Get-Cluster).BlockCacheSize = 2048
# Verify per-CSV cache state
Get-ClusterSharedVolume | ForEach-Object {
$_ | Select-Object Name,
@{N='CacheEnabled';E={$_.SharedVolumeInfo.Partition.IsCsvCacheEnabled}}
}
Size BlockCache from your steady-state RAM headroom, not from the N-1 evacuated state. Verify that VM density still fits comfortably under failover conditions. We typically recommend 1 to 4 GB per node on Hyper-V clusters, and 4 to 16 GB on storage-heavy workloads.
Full context on BlockCache trade-offs is in Hyper-V cluster assessment: Top 10 Issues We Find in 2026, issue 7.
9. Preventing recurrence
The ClusterTriage remediation pattern for CSV ownership imbalance has four parts.
- Remove explicit
PreferredOwnerson every CSV unless there is a documented reason. Most that we find are debug session residue. - Enable CSV BlockCache at 1 to 4 GB per node, and verify enabled state on every CSV partition.
- Add a scheduled rebalance as a post-maintenance step in every runbook. CAU's post-update task list is the right place; if you do not use CAU, your patch automation has to handle it.
- Monitor the ownership distribution as a metric. We add a five-line script to ops dashboards: count of CSVs per node, alarm when one node owns more than
(total / nodes) + 1.
CSV ownership imbalance is the textbook example of a finding that monitoring tools miss because each individual component looks healthy. Surfacing it is exactly what an independent Hyper-V cluster assessment is for.
Schedule a Hyper-V cluster assessment intro call →
Frequently asked questions
After every patch round, every unplanned failover, and every node drain for maintenance. In practice that comes down to monthly or bi-monthly for most production clusters. Make it part of the standard post-maintenance checklist and nobody forgets.
No. It only transfers the coordinator role, not the volume itself. VMs on the CSV keep running without interruption. There is a brief pause in metadata operations during the handover, measured in milliseconds and invisible to workloads 99% of the time.
Limitedly. Direct writes go straight to storage and do not touch the coordinator. Metadata-heavy operations (VHD extension, snapshots, file creation) do hit the coordinator. On strongly metadata-heavy workloads (Exchange, some database engines with many small files) balancing has more impact than on read-heavy workloads.
Mode 2 balances only when imbalance exceeds an internal threshold, instead of continuously aiming for perfect distribution. Useful on clusters where VMs have PreferredOwners set and you want to avoid the balancer making constant small adjustments that PreferredOwners then undoes.
On Azure Local and Azure Stack HCI with Storage Spaces Direct, CSV BlockCache behaves differently because S2D has its own cache layer on the NVMe tier. For SAN-attached Hyper-V clusters, BlockCache is still very much worth using and is regularly forgotten.