ClusterTriage / Blog / Quorum
File Share Witness vs Cloud Witness vs Disk Witness: Which Type Fits Your Cluster?
Quorum is the invisible thing keeping a Windows Server Failover Cluster alive when nodes go away. Witness configuration is also one of the top three areas where ClusterTriage Hyper-V cluster assessments find quiet resilience gaps. The witness exists, the cluster reports healthy, validation passes, and one fault domain still takes the cluster down because the witness sits exactly in the wrong place.
This article compares the three witness types in production terms: failure modes, topology fit, the configuration mistakes we keep finding, and the decisions you want to get right once.
1. Short primer, quorum and Dynamic Quorum
A cluster votes on its own health. Every running node gets a vote, the witness (if configured) gets a vote, and the cluster needs a majority of votes to stay alive. Loss of majority means the Cluster Service stops to prevent split-brain.
Dynamic Quorum (introduced in Windows Server 2012 R2, default on since) adjusts the vote count as nodes leave the cluster. Lose one node from a five-node cluster, the cluster reduces vote count and keeps running. In a two-node cluster with witness you can lose either node and survive on the remaining node plus the witness. Lose the witness and a node simultaneously, and Dynamic Witness has already adjusted the witness vote weight to maximise survival odds.
The witness exists to break tie votes, and to keep a majority alive in even-numbered counts where natural majorities are impossible. Without a witness in an even-node cluster, any 50/50 split takes the cluster down.
2. The three witness types at a glance
| Witness type | Where it lives | Best for | Worst for |
|---|---|---|---|
| Disk Witness | Shared SAN LUN visible to all nodes | Single-site clusters with SAN where the witness LUN is on independent storage relative to the data | Multisite/stretched clusters (cannot span sites), S2D clusters (no shared SAN), small clusters where the SAN is the data path |
| File Share Witness | Network SMB share on a server outside the cluster | Two-site clusters with the share in a third site, disconnected environments, large environments with a witness fileserver | Single site if the share sits in the same fault domain as the cluster |
| Cloud Witness | Azure Storage Account, reached over HTTPS | Any cluster with reliable Azure connectivity, including two-node, stretched, and disaster recovery clusters | Clusters with no permitted outbound Azure connectivity |
3. File Share Witness, where it fits and where it fails
The File Share Witness is the oldest non-disk witness and still the right answer in disconnected or strict-policy environments. It is also the type where we find the most configuration mistakes, because it looks like "just a file share" and gets treated that way.
What it actually is: a small file on an SMB share that the cluster periodically touches to claim its quorum vote. The share itself does not need to be highly available, just available, period. And critically: in a different fault domain than any of the cluster nodes.
What we find in ClusterTriage Hyper-V cluster assessments:
- File Share Witness on a single physical server, no UPS, in the same rack as half the cluster. The rack PDU then becomes a single point of failure for the entire cluster.
- File Share Witness on a NAS that is itself a single-node appliance. Reboot the NAS for firmware, lose the witness during a critical maintenance window.
- File Share Witness on one of the cluster nodes itself. We have seen this in production. The admin was short on fileservers and "had room left on HV03". That is not a witness, that is a permanently broken configuration.
- FSW path configured with a NetBIOS name resolving to multiple IPs (DFS namespace), where some targets sit in the wrong site.
Requirements for a usable File Share Witness:
- Hosted on Windows Server 2012 R2 or newer (older versions miss the required SMB features).
- SMB 3.0 or newer, not on a NAS that only speaks SMB 2.x.
- In a fault domain genuinely independent from the cluster: different rack, different power feed, ideally a different site.
- Cluster computer object with Full Control on the share folder (not just Read/Write).
- Share path uses a stable name, FQDN or short name not dependent on DNS aliases.
4. Cloud Witness, the modern default
Cloud Witness was introduced in Windows Server 2016 and is the witness type ClusterTriage recommends for nearly every modern deployment with permitted Azure connectivity. It is a small file in an Azure Storage Account, reached by every node over HTTPS to *.blob.core.windows.net.
Why this is the default answer:
- By definition outside the cluster's fault domain. Whatever burns down at the customer site, the witness in Azure keeps running.
- Costs essentially nothing, fractions of a cent per month for storage, no egress because the file is tiny.
- Works for two-node, multi-node, and stretched clusters identically.
- Requires no separate fileserver, share permissions, or DNS configuration.
When it is the wrong answer:
- The cluster sits in a network segment without outbound internet and security policy forbids adding it. Hard-air-gapped environments.
- The cluster runs on a hardware platform that cannot reliably reach Azure (rare, but we have seen it on industrial control systems).
- Latency from cluster to Azure exceeds the cluster heartbeat threshold, extremely rare, requires intercontinental links without peering.
Configuration is one PowerShell command (see the configuration section below). The only real planning decision is which Azure subscription and which storage account region. Pick one close to the cluster, preferably with a stable link.
5. Disk Witness, still relevant in specific cases
Disk Witness is older than the other two and remains the only option in a few specific scenarios. It is a small clustered disk resource, typically a 1 GB LUN, visible to every node through the SAN.
When this is still the right answer:
- Single-site cluster with shared SAN storage where the SAN genuinely delivers independent storage paths relative to the data LUNs (different storage controllers, ideally different physical arrays).
- The customer is heavily invested in SAN-based clustering and there is no operational appetite for Azure connectivity.
When it is the wrong answer:
- Stretched/multisite, Disk Witness cannot span sites.
- S2D / Azure Local, there is no shared SAN.
- Witness LUN on the same SAN array as the CSVs. Lose the array, lose the witness, lose the cluster. We see this surprisingly often.
- Witness LUN on a SAN volume with its own quorum logic (some active/passive arrays do unexpected things during controller failover).
6. Decision matrix per topology
| Topology | First choice | Acceptable alternative | Avoid |
|---|---|---|---|
| Two-node, single site, SAN | Cloud Witness | Disk Witness on independent SAN | FSW on the SAN-attached fileserver |
| Two-node, single site, S2D | Cloud Witness | FSW on independent infrastructure | Anything on the cluster nodes themselves |
| 3 to 8 nodes, single site, SAN | Cloud Witness | FSW or Disk Witness on independent storage | Disk Witness on the same array as the CSVs |
| 3 to 8 nodes, single site, S2D / Azure Local | Cloud Witness | FSW in a different fault domain | Anything internal to the cluster |
| Two-site stretched cluster | Cloud Witness | FSW in a third physical site | FSW in either of the two cluster sites |
| Disconnected / air-gapped | FSW on independent storage and power | Disk Witness on independent SAN (where applicable) | Witness on cluster nodes or shared fault domain |
Witness topology is a subject monitoring stays silent on. The cluster is healthy, validation passes, the report is green. Only during a real failure does it surface that the witness shared a fault domain with half the cluster.
A ClusterTriage Hyper-V cluster assessment tests witness placement against your actual topology and documents the gap with severity and remediation steps.
Schedule a cluster assessment intro call →7. Five configuration mistakes we keep running into
- Witness in the same fault domain as the cluster. FSW on a fileserver in the same rack with the same PDU. Disk Witness on the same SAN as the CSVs. The witness exists, but adds no resilience.
- Cluster computer object lacks Full Control on the file share. The witness looks configured, quorum operations fail when the cluster tries to claim it.
- Cloud Witness in an Azure region geographically far away. Witness latency spikes during regional issues. Pick a region close to the cluster.
- No witness at all on a two-node cluster. Dynamic Quorum does its best, that is not enough for any predictable outcome.
- Witness configured but Dynamic Witness off. The cluster then cannot adjust witness weight as topology degrades. We see this on older clusters carried through multiple upgrades.
8. PowerShell, configuring and verifying
Inspect current state:
# Current quorum configuration
Get-ClusterQuorum | Format-List *
# Witness resource detail
Get-ClusterResource | Where-Object {$_.OwnerGroup -eq 'Cluster Group'} |
Format-Table Name, State, OwnerNode, ResourceType
# Verify Dynamic Quorum and Dynamic Witness
(Get-Cluster).DynamicQuorum
(Get-Cluster).WitnessDynamicWeight
Configure Cloud Witness:
# Requires the Azure storage account name and one of the access keys
Set-ClusterQuorum -CloudWitness `
-AccountName 'sthcwitnessprodweu' `
-AccessKey '...' `
-Endpoint 'core.windows.net'
# Verify
Get-ClusterResource 'Cloud Witness' | Format-List *
Test-NetConnection sthcwitnessprodweu.blob.core.windows.net -Port 443
Configure File Share Witness:
# Requires Full Control on the share for the cluster computer account
Set-ClusterQuorum -FileShareWitness '\\fs01.corp.example\witness$\hv-cluster'
# Verify
Get-ClusterResource 'File Share Witness' | Format-List *
# Verify share permissions from a cluster node
Get-Acl '\\fs01.corp.example\witness$\hv-cluster' | Format-List Access
Configure Disk Witness:
# Witness LUN must already be a clustered disk resource
Get-ClusterAvailableDisk | Add-ClusterDisk
Set-ClusterQuorum -DiskWitness 'Cluster Disk Witness'
End-to-end validation:
Test-Cluster -Include 'Inventory','Quorum Configuration','Network' `
-ReportName witness-check
The bottom line
For nearly every modern cluster ClusterTriage reviews, Cloud Witness is the right answer and implementation takes one command. The most common cluster assessment finding is not "no witness", it is "witness that adds no resilience because it shares a fault domain with the cluster". Closing that gap is one of the highest-leverage remediation steps in a typical engagement.
A witness in the same fault domain as the cluster is not a witness, it is decoration.
The ClusterTriage Hyper-V cluster assessment includes witness topology analysis in standard scope. If the witness does not fit the topology, that is documented with severity scaling to blast radius, and the remediation step lands on the action list.
Schedule a Hyper-V cluster assessment intro call →
Frequently asked questions
Strictly speaking no, because odd node counts give natural majorities. But the moment you lose one node, you are in a two-node state where Dynamic Quorum substantially improves survival odds with a witness. We always configure a witness, regardless of node count.
No, a cluster has only one witness at a time. You can switch types with Set-ClusterQuorum without downtime.
Dynamic Witness adjusts the witness vote weight as soon as the cluster detects unreachability. The cluster keeps running as long as a majority of remaining votes (nodes plus any witness) holds. The witness is automatically reweighted as soon as connectivity returns.
Yes, the -Endpoint parameter supports different Azure cloud endpoints. For Azure Government use core.usgovcloudapi.net, for Azure China core.chinacloudapi.cn. Sovereign clouds with data residency requirements can use Cloud Witness as long as the Azure tenant lives in the correct cloud.
Negligibly under normal operation. The cluster only touches the witness during quorum events (node failure, reboot, witness state change). For stretched clusters with latency between sites, witness location matters for failover times, but not for steady-state performance.