ClusterTriage / Blog / Cluster assessment
From Cluster Assessment to Patch Night: The Method at Work
This post walks through a worked example of the five-step ClusterTriage Method applied to a production Hyper-V cluster. Twelve High Risk findings were discovered, remediated in dependency order, re-measured to verify the fixes, and then executed as a patch run so clean that the on-call administrator didn't need to get out of bed.
0. Setting the scene
This was a planned Windows Update round on a production multi-site Hyper-V cluster: six Windows Server 2022 nodes stretched across two Dutch datacentres, carrying around 135 VMs. The updates had been deliberately scheduled as the final step of a cluster assessment engagement, after weeks of remediation, and executed remotely with Cluster Aware Updating.
My working morning in Bangkok went into the final checks: the run script, a fresh DNS audit of every node, and a full Failover Cluster Validation, all clean by quarter past eleven. Then, at 06:37 Amsterdam time, while the customer's administrators were still asleep, I started the run. By 09:08 it was finished: 23 updates installed, zero failures, zero cancelled, every node back up, every VM where it should be, before the first users came in on a quiet Friday morning. One of the two administrators was on holiday. The other checked in over Teams, mid-dog-walk, when the run was already almost done. Nobody had to get out of bed.
That is the whole story, and that is the point. A patch run on a production stretch cluster should be the most boring event of the month. This one was, but only because of everything that happened in the weeks before it.
1. Intake: know what you are standing on
The first Method step. Every engagement starts with topology, storage model, management tooling, backup, monitoring and recent incidents. Not because it is a formality, but because every later judgement depends on it.
This cluster: the six nodes from the opening, SAN storage over iSCSI on dedicated 25 GbE networks, SET virtual switches for LAN and Live Migration, managed through Virtual Machine Manager, backed up with Veeam. The VMs were distributed unevenly across the two sites.
The intake also surfaces what is in motion. Here, two things were: the customer was mid-way through replacing their domain controllers, which meant the DNS landscape was changing under the cluster's feet, and one Linux VM had a known Live Migration problem. Both of these would matter at the very end. An intake that misses the moving parts produces a findings document that is stale on delivery.
2. Scripted inventory: read-only, repeatable, no agents
Step two. The data comes from PowerShell scripts tailored to the environment, executed from the customer's own management server. Nothing gets installed, nothing gets changed. The inventory script is read-only by design and says so in its header, because the person approving the run should be able to verify that claim in the source.
A few design choices have earned their keep across engagements:
Pre-flight before inventory. Section 0 tests every node for DNS resolution into the management subnet, ICMP, WinRM reachability, listener bindings and firewall state. Any node that fails pre-flight is excluded from every downstream section, with an inline banner naming it, so the report stays self-documenting instead of silently incomplete.
Resumability. A state file records which sections completed. If a run is interrupted at section 9 of 12, the next invocation picks up at section 10 and appends to the same output file. On a six-node cluster a full inventory takes a while; losing it to a dropped session is not acceptable.
One text file out. The output is a single UTF-8 text file. It diffs cleanly against the previous measurement, which is exactly what step four needs.
3. Analysis against best practice: numbers, not impressions
Step three. Every data point is compared to the Microsoft baseline for the topology, and each deviation gets a risk level with a supporting reference. The first full measurement of this cluster produced twelve High Risk findings. A sample, because the concrete numbers are what make a findings document actionable:
Clock drift of up to 86 seconds. One site was clean, within a second. The other site, plus one stray node, sat 64 to 86 seconds off. That pattern points at a failed time source on one side, not random drift. Kerberos tolerates five minutes; CSV ownership and cluster heartbeats degrade well before that. Finding, root cause direction, and a hard gate: no patch round until all nodes are within one second.
Defender exclusions: empty on all six nodes. No paths, no processes, no extensions. Every VHDX I/O was eligible for real-time scanning, the most common root cause of Event ID 9 (slow I/O on a clustered volume) and almost certainly behind the backup-window slowness the customer had lived with for a year.
Pending reboots on five of six nodes, confirmed down to PendingFileRenameOperations in the registry, with around 200 days since the last restart.
Patch level seven months behind. The newest hotfix on any node dated from the previous October.
None of these is exotic. All of them are invisible from a dashboard that only shows green VMs.
4. The findings document, and the re-measurement that keeps it accurate
Step four. Findings go into a structured document: (1) current situation, (2) best practice, (3) deviation, (4) recommendation, (5) action, (6) reference, colour-coded by risk. But a findings document is a snapshot, and remediation changes the picture. So the document gets delta versions: re-measure, compare against the previous output file, and move findings to resolved only when the data says so.
That discipline catches things a checkbox process never would. One example worth sharing: the customer rolled out the Defender exclusions through Group Policy, exactly as recommended. The verification run still showed all six nodes scanning everything. The cause sat inside the GPO editor: in the Administrative Template for exclusions, the value name and the value had been swapped. Defender takes the value name as the exclusion, so the policy had loaded twenty labels like an exclusion list and excluded nothing. Flip the columns, gpupdate, re-verify: 162 checks passed. Without a scripted re-measurement, that policy would have sat there looking compliant for years.
Over three delta cycles the High Risk list shrank from twelve to a handful: BIOS versions aligned across sites, drivers aligned, pending reboots from before the cluster assessment cleared through a controlled restart cycle (weeks before the patch run), clock drift fixed at the time source, exclusions verified active. What remained High Risk was the patch backlog itself, deliberately scheduled last, because patching a cluster with broken time sync, pending file renames and a shifting DNS layer is how you convert a backlog into an outage.
5. The action list, and the dependency chain that gates it
The fifth Method step turns recommendations into an executed plan, and the key word is order. Each action gets an owner, a dependency chain and a validation step. For the patch run, the chain looked like this: clocks within a second, pending reboots cleared, WinRM reachable on all nodes, and, because of the in-flight domain controller migration, DNS verified clean on every node.
That last gate deserves emphasis. The new domain controllers had different IP addresses than the ones they replaced. A cluster node that reboots mid-patch and comes up pointing at a decommissioned resolver is a node that may not rejoin cleanly. So on the morning of the run, a standalone read-only DNS audit walked every NIC on every node: configured DNS servers, suffixes, registration settings, with an explicit match flag against the list of old DC addresses. Zero hits. Every node pointed at two valid, new domain controllers.
Then the final gate: a full Failover Cluster Validation. Not a partial run, the whole report, and every yellow line read, not skimmed. It came back clean. That was the green light.
6. The patch night
The run itself was scripted around Invoke-CauRun in self-updating mode. No CAU cluster role, no virtual computer object, no permanent account in the cluster, because this customer patches deliberately, on their schedule, not automatically. The script does four things worth copying.
It refuses to start unless the environment matches the plan. All six nodes up, no CAU run already in progress, and one very specific check: the Linux VM with the known Live Migration problem had been shut down manually in advance, with its cluster stop action set accordingly. The script verifies that VM is actually offline and aborts if it is not, because a drain that hits an unmigratable VM stalls the whole run at three in the morning. (This could have been automated with a successful-run trigger, but manual shutdown was safer: an automated script that stops a VM and then fails mid-patch leaves you explaining to the customer why their guest is offline. Not worth the risk.)
It patches one site at a time. CAU has no native site awareness, but -NodeOrder accepts an explicit list. Site one's three nodes first, then site two's. On a stretch cluster this keeps the failover capacity story simple at every moment of the run: at most one node down, and you always know which site is carrying the drained workload.
$cauParams = @{
ClusterName = $Cluster # FQDN
CauPluginName = 'Microsoft.WindowsUpdatePlugin'
NodeOrder = $NodeOrder # site 1, then site 2
RequireAllNodesOnline = $true
Force = $true
EnableFirewallRules = $true
}
$job = Start-Job { param($p)
Import-Module ClusterAwareUpdating
Invoke-CauRun @p
} -ArgumentList $cauParams
It shows you what is happening. Invoke-CauRun blocks and prints nothing useful to a console. Launching it as a background job leaves the main thread free to poll Get-CauRun and Get-ClusterNode every twenty seconds, printing only transitions: a node moving to Suspending as it drains, Restarting as it reboots, the cluster state flipping to Down and back to Up. I watched the run over a casual lunch at my desk: the fairly boring log feed making its progress on one screen, Failover Cluster Manager open on the node being processed on the second, and the Cluster Aware Updating console on the third. All three agreed, every twenty seconds, that this was going like a charm. That triple confirmation is the difference between confidence and refreshing Failover Cluster Manager nervously for three hours.
It treats a failed launch as a failure. The first attempt died in six seconds on a wrong plugin name (Microsoft.WindowsUpdatePlugin, not Microsoft.WindowsUpdate; Get-CauPlugin lists the registered names). The lesson baked into the script afterwards: collect the job's error stream and throw loudly, because a script that prints "complete" after a run that never started is worse than no script.
The run installed the cumulative update, the .NET rollup and the definition updates on every node, each node draining, patching, rebooting and resuming in order, one site and then the other. Get-CauReport confirmed it afterwards: succeeded, twenty-three updates, nothing failed, nothing cancelled.
7. Why this worked, and what it cost
The temptation with a seven-month patch backlog is to just run the updates and hope. The backlog was the symptom that was visible. The conditions that made patching dangerous, drifting clocks, stale pending operations, an unverified policy rollout, a DNS layer mid-migration, were not visible until something measured them.
The sequence is the Method doing its job. Measure everything, rank the deviations, fix in dependency order, re-measure after every change, and only then execute the risky operation, gated behind a final validation and wrapped in a script that checks its own preconditions. The patch night was uneventful because nothing about it was left to chance, including the timing. The five to six hour difference between Bangkok and the Netherlands turns out to be a genuine operational advantage for critical maintenance: I spent a fresh morning on the final checks, launched the run in the customer's quietest hour, and watched it complete over lunch while their working day was only just beginning.
Where does your cluster stand before the next Windows Update? A ClusterTriage Hyper-V cluster assessment follows exactly this method, ending with the scripted patch run that makes the difference between an uneventful night and an incident at three in the morning.
Book a Hyper-V cluster assessment →Frequently asked questions
The five-step Method: (1) Intake,understand topology, storage, tooling, backup, monitoring. (2) Scripted inventory,read-only, repeatable PowerShell across all nodes. (3) Analysis,compare against Microsoft baselines, rank deviations by risk. (4) Findings document,structured output with delta re-measurements to verify remediation. (5) Action list,execute recommendations in dependency order, gated by validation.
Clock drift up to 86 seconds (breaks Kerberos, CSV), no Defender exclusions (causes slow I/O), pending reboots on five nodes, and seven months of missed patches. None visible on a dashboard showing only green VMs. Scripted measurement finds what dashboards hide.
Clocks within one second, pending reboots cleared, WinRM reachable on all nodes, and DNS verified,critical because domain controllers were mid-migration. Each gate had a validation step. A full Failover Cluster Validation with every line read, not skimmed, was the final green light.
From 06:37 to 09:08 Amsterdam time (2 hours, 31 minutes). Six nodes, 23 updates installed, zero failures, zero cancelled. Every node back up, every VM where it should be, before the first users arrived on a quiet Friday morning.
It had a known Live Migration problem that would stall a drain at three in the morning. The script explicitly verifies the VM is offline and aborts if not, because a stalled drain mid-patch is unacceptable. Known blockers must be eliminated before the run.