ClusterTriage / Blog / Root cause analysis
One incident, two reports: how my AI assistant knocked a node out of my Azure Local cluster, and what the RCA made of it
I asked for a test file. I got a cluster incident. Then I asked for the root cause analysis, written as if nobody knew the answer, and then I asked for it again for a reader who has never heard of a heartbeat. This is the story of one 80-second event on a two-node Azure Local lab cluster, and of why the same event deserves two different reports.
The short version: a live kernel memory dump froze one node for 52 seconds, the cluster did exactly what it was told and threw that node out, and the evidence package showed it to the second. The longer version is more useful, because it shows how an RCA should be built, what three very different readers need from it, and what it means when the person who caused the incident also writes the report.
1. The lab: why we built an Azure Local cluster to break
ClusterTriage assesses and troubleshoots Hyper-V Failover Cluster, HCI and Azure Local estates. The free entry point is ClusterDown.com (from this site: click Cluster issues? in the menu): when a cluster is in trouble, the customer runs one script, sends back one zip file, and gets the five most important issues back, with the option of a full root cause analysis afterwards.
You cannot build a tool like that on theory. You need clusters that misbehave, and you need to know the truth about every misbehaviour, so you can check whether the tool finds it. So in September I built a lab: a two-node Azure Local 2609 cluster, nested on a single rented server with an Intel Core i9-13900 and 128 GB of memory. The outer layer is Proxmox, the two cluster nodes are virtual machines running Azure Local, and inside one of those nodes runs the Arc Resource Bridge, itself another virtual machine. Three layers of virtualisation, four virtual processors per node.
Getting Azure Local to accept that setup was a project of its own (the validator wants to see ECC memory, physical-looking network adapters and unique disk identities, among other things), and that story is for another article. What matters here: the cluster was deployed, registered in Azure, and running. A perfect place to break things, as long as you know what you broke.
2. The request that turned into an incident
Most of the build work on this lab was done together with Claude, the AI assistant I use for the ClusterTriage toolchain. Claude runs the commands, I make the decisions.
On 1 October I wanted a memory dump file to test the dump detection in ClusterDown. I asked for exactly that: "make a test dump file of N01." Claude chose a live kernel dump, the kind Windows can take without crashing the server, using the storage diagnostics cmdlet:
Get-StorageDiagnosticInfo -StorageSubSystemFriendlyName 'Windows Storage on AZL-N01' `
-DestinationPath C:\Temp\LiveDumpTest -IncludeLiveDump
It reported back: 4.6 GB written in 124 seconds, node still running, nothing crashed. All true. Also incomplete.
What neither of us looked at that afternoon: during those 124 seconds, the cluster had removed N01 from its membership, paused a shared volume, moved every clustered role to N02, and N01's own Cluster service had terminated itself with a fatal error. Eighty seconds later N01 was back in the cluster, and from the outside everything looked normal again.
It only surfaced the next day, when another session asked for the "ground truth" of an event at 06:41 UTC that the analysis kept pointing at. Claude read the event logs and found its own fingerprints all over them. Its message to me opened with the admission that it had caused the incident itself.
I did not ask for a crash. I asked for a dump. The difference between those two, on a small, busy cluster node, turned out to be one cluster heartbeat setting. That is the first lesson, and we will come back to it.
3. Writing the RCA as if nobody knew
Then I did something deliberately awkward. I collected a fresh ClusterDown package from the cluster, gave it to Claude, and asked for a root cause analysis with three rules:
- You do not know the cause. Prove it from the evidence in this package, and from combinations of evidence.
- Every claim names its source.
- Say what you are missing to write a better report.
The obvious problem: the writer knew the answer. That is not a hypothetical problem; it is exactly the situation of every engineer who writes the RCA for an incident in their own team. So the method had to protect the report from its author. Claude kept its own knowledge out entirely, including the parts that would have made the report stronger, such as the exact command that was run. If the package could not prove it, it did not go in. The ground truth was written down separately, in its own file, so a reviewer could check the conclusion against it afterwards.
Every RCA should work this way, whether the analyst is a person or not. A report that leans on what the author happens to know is an opinion. A report that leans on evidence can be checked by somebody else.
4. What the evidence showed, to the second
The operator's problem statement in the package said: "Cluster node crashed unexpectedly." That was the first thing the RCA disproved. Both nodes had last booted on 30 September. There was no event 41 or 6008, no blue screen and no crash dump. No node crashed. A node lost its membership. That distinction matters, because it sends you looking in a completely different place.
Then the timeline. Three pieces of evidence, from three different sources, line up on the same seconds:
| Time (UTC) | Source | What it says |
|---|---|---|
| 06:40:49 | Kernel-LiveDump channel on N01 | "Live Dump Capture Dump Data API started" |
| ~06:40:49.7 | Cluster log on N02, heartbeat route history | Last heartbeat answer from N01 (recorded at 06:41:09 as 19.65 seconds old) |
| 06:41:41 | Kernel-LiveDump channel on N01 | Last capture phase ended |
| 06:41:41.505 | System log on N02, event 1135 | N01 removed from the cluster membership |
The heartbeats stopped in the same second the dump started. The node was removed in the same second the dump finished. Between those two moments, N01 logged nothing except the dump's own progress. And N01 itself only noticed missing heartbeats at 06:41:55, after it woke up, with its own route history showing the last heartbeat as 26 seconds old. That is what a paused server looks like. A server with a broken network notices the problem while it happens; a frozen one only notices afterwards, when the clock has jumped.
The cluster's heartbeat tolerance, read from the same package: 25 missed heartbeats at one per second. The node was silent for about 52 seconds. The cluster did nothing wrong; it removed a node that did not answer.
The report then tested the alternatives, one by one:
- Network: all three networks (management and two storage) failed at the same moment, in both directions, with zero events in the NDIS channel. A network fault does not do that.
- Hardware: zero WHEA events on either node.
- Storage: the paused volume and the file-system write failures all come after the removal. Consequence, not cause.
- Operating-system crash: disproved by the boot times and the absence of any dump or crash event.
- Resource exhaustion: possible as a contributing factor (four processors per node, plus the Arc Resource Bridge VM on N01), but not provable, because the package held no CPU history for that minute.
And then the part I like most: the question the evidence could not answer. Who started the live dump? The flags field in the LiveDump event is not rendered, and no dump file was found in the standard folders. That absence was itself a clue: a dump Windows takes on its own lands in LiveKernelReports, so a dump that is not there was probably requested on purpose, with its own destination. The report says exactly that, labelled as an inference, and lists what would settle it: process-creation auditing, PowerShell logging, and the operator's own account of what was run at 06:40.
The report could not name me. The evidence did not contain me. That is how it should be.
5. Same event, two reports
The first report is the one an engineer writes for engineers: six pages, every statement tied to an evidence ID, six hypotheses tested, heartbeat route history arithmetic, a confidence statement per conclusion. It is complete, and almost nobody outside the cluster team will read past page one.
So I asked for a second version: for a regular server administrator who is not a cluster specialist, with a management summary on top that their manager can read on its own. Same event, same evidence, same conclusion. Three pages.
Some of the differences, side by side:
| Technical RCA | Administrator version with management summary | |
|---|---|---|
| The cause | "A live kernel memory dump capture suspended the node for about 52 seconds, exceeding the 25-second heartbeat tolerance (SameSubnetThreshold 25, SameSubnetDelay 1000)." | "A full memory snapshot of that server was being taken. While the snapshot is taken, the server is paused. Here for about 50 seconds; the cluster only waits 25." |
| The proof | Evidence IDs E2304 to E2445, route history, six hypotheses | "The timing matches to the second", plus a table of what was ruled out and why |
| The open question | Rendered flags of event 1, Security 4688, PowerShell 4103/4104 | "Who or what asked for that snapshot. That is the real question." |
| The advice | Five recommendations with their reasoning | Suspend-ClusterNode -Drain before a dump, in a ready-to-run form |
Read both and compare for yourself:
- The technical RCA (PDF, 6 pages, 187 KB)
- The administrator version with management summary (PDF, 3 pages, 113 KB)
Nothing was simplified into something untrue. That was the hard part. "The cluster has run normally since" would have been a nice sentence for a manager, but nothing in the package measured it, so the report says "when the data was collected the next day, both servers were part of the cluster again." Plain language is not permission to stretch the evidence.
Cluster in trouble right now? ClusterDown.com gives you the five most important issues from one script and one zip file, free of charge. A root cause analysis like this one is the next step when you need it. On this site, click Cluster issues? in the menu.
Go to ClusterDown.com →6. Why each reader gets what they need
The manager reads six short paragraphs and can make three decisions without understanding a single cluster term: was anything lost (no server crashed; for about one and a half minutes the cluster ran on one server), is it likely to happen again (low, unless the snapshot was triggered automatically, which is why finding out who triggered it is action number one), and what does it cost to prevent (a working agreement and possibly bigger servers). No heartbeat thresholds, no event IDs, but also no reassurance that the evidence does not support.
The administrator gets what an administrator actually does with an RCA: a timeline in plain words, the reasons the obvious suspects were ruled out (so they do not spend tomorrow replacing a network switch), a short list of places to look themselves, and the exact command that prevents a repeat. The report also tells them what it still needs from them, which turns them from a reader into a participant.
The engineer gets the version they can disagree with on specifics. Every claim has a source they can open, every hypothesis has its evidence for and against, and every inference is marked as one. If they think the 52 seconds is wrong, they can redo the arithmetic from the route history. An RCA that a sceptical engineer can verify is worth more than one they have to take on trust, especially when the conclusion is uncomfortable.
The three readers want different things from the same facts. Writing one report and hoping it serves all three is how you end up with a document the manager does not understand and the engineer does not believe.
7. What we changed afterwards
The incident was small, but it changed a few things for real.
Drain before you dump. A live kernel dump is meant to be non-disruptive, and on a big, idle server it may well be. On a small cluster node it can freeze the machine past the heartbeat tolerance. If you need a dump from a cluster node, take the node out first:
Suspend-ClusterNode -Name AZL-N01 -Drain
# take the dump
Resume-ClusterNode -Name AZL-N01
Say what an action does to the cluster, not just to the server. "The node keeps running" was true and beside the point. The question that should have been asked before the command was: what does the cluster see while this runs? That applies to people and to AI assistants alike.
The ground truth is now a test case. Because we know exactly what happened at 06:41, this event is now a fixed comparison case for the ClusterDown analysis: any version of the tool should land on "a live dump on N01, started by an administrator", and any version that blames the network has a bug.
And one more thing, for the toolchain itself. Before the package in this article existed, the ClusterDown incident collector came back with zero of 27 event channels from this cluster, on two runs in a row. Measuring channel by channel showed why: on these nodes, PowerShell's Get-WinEvent needed about 70 milliseconds per event on the two SMB client channels, where reading the same events directly took less than a millisecond. Two small logs consumed the whole time budget. The collector was fixed, and the next package read every node in full. Breaking your own lab is how you find out what your tools do when it matters.
Frequently asked questions
It does not crash the server, but it pauses it while memory is captured. On this two-node cluster with four processors per node, the pause lasted about 52 seconds, twice the cluster's heartbeat tolerance. Drain the node first with Suspend-ClusterNode -Drain, and the cluster will not react when it goes quiet.
Because the cluster reacted correctly. A node that does not answer for 52 seconds should be treated as gone. Raising the threshold to hide a self-inflicted pause would also delay the reaction to a real failure.
Only if the method protects the report from its author. Here, the rule was that nothing went into the report unless the evidence package proved it, and the known answer was written down separately so a reviewer could check the conclusion against it. The report still could not name who started the dump, because the evidence did not contain that. That limitation is the proof the rule was followed.
An executive summary on top of a technical report still leaves the administrator in the middle with a document written for someone else. The administrator needs the steps, the ruled-out list and the commands, in plain words. The manager needs only the summary. The engineer needs the evidence trail. One document can carry the manager and the administrator together; the engineer deserves the full version.
At ClusterDown.com, or click Cluster issues? in the menu of this site. One script, one zip file back, and the five most important issues, free of charge; a full root cause analysis is the next step when you need one.