ClusterTriage / Blog / Time Service

Time Drift in Hyper-V Clusters: The Silent Killer of Kerberos and CSV

A Hyper-V cluster with clock skew between nodes is a time bomb. Not figuratively, literally. When clocks on cluster nodes drift more than five minutes apart, Kerberos breaks. From fifty seconds, cluster heartbeats lose their synchronisation. And the moment a Hyper-V host loses its time reference, domain-joined VMs running on that host can no longer maintain authentication to AD.

In most Hyper-V cluster assessments ClusterTriage runs, we find time issues. Not always dramatic, often still below the threshold where something fails outright, but clearly out of balance. And it is exactly the kind of problem monitoring tools miss until it is too late, because a few seconds of skew never triggers an alert.

This article explains why time service on a Hyper-V cluster works differently from a standalone server, which configurations we treat as best practice in 2026, and how to detect drift before the first Kerberos failure.

By Hans Vredevoort · 28 May 2026 · 13 minute read · Time Service

1. Why time drift in Hyper-V clusters is specifically risky

On a standalone server, clock drift is an annoyance. On a Hyper-V cluster, it is a cascade trigger.

Kerberos. Active Directory uses Kerberos for authentication, and Kerberos has a default clock tolerance of five minutes. Difference more than that between client and KDC, and all new authentication fails with KRB_AP_ERR_SKEW. Existing sessions keep working until their ticket expires, so the first symptoms appear 8 to 10 hours after the drift and look like they have no cause. We have seen this on customer site: Monday morning nobody can log in, everything points to DNS, and the real cause is a Hyper-V host that lost its NTP source on Saturday.

Cluster heartbeats. Failover Clustering expects all nodes to sit within reasonable time tolerance to correctly track heartbeat events. With large skew, nodes report events in the wrong order, and the Cluster Service may decide a node is out of cluster while that node is running fine.

CSV coordinator arbitration. Cluster Shared Volume coordinator role transfers are logged with timestamps. Skew between nodes makes log analysis impossible, and in rare cases it can put CSV rebalancing into a loop where the balancer interprets its own actions as concurrent with the other node.

Live Migration timing. The live migration handshake uses timestamps for sequence validation. Large skew between source and target host can abort migrations without a clear error, or worse, report the migration complete while memory pages were skipped.

VM clocks that drift. Domain-joined VMs sync with their PDC emulator. If that PDC emulator in turn depends on a Hyper-V host not managing its own time well, you get drift in two layers, and the solution becomes much more complex than "fix the NTP source".

2. The Microsoft time hierarchy and where Hyper-V hosts fit

The correct time architecture on a domain environment:

  • The top level. One DC with the PDC emulator FSMO role synchronises with an external, reliable time source. Preferably time.windows.com, pool.ntp.org, or a hardware stratum-1 source. Configuration: Type = NTP, NtpServer = <source>,0x9.
  • All other DCs. Synchronise with the PDC emulator through the domain hierarchy. Configuration: Type = NT5DS.
  • All domain-joined servers, including Hyper-V hosts. Synchronise with the domain hierarchy. Configuration: Type = NT5DS.
  • All domain-joined VMs. Synchronise with the domain hierarchy through their own w32time service. Configuration: Type = NT5DS. And critically: the Hyper-V Time Synchronization Integration Service must be off for these VMs. See section 3.
  • Non-domain VMs or isolated clusters. Sync directly with an external NTP source or with an NTP source in the management network. Configuration: Type = NTP.

What we find in a typical Hyper-V cluster assessment is that this hierarchy is broken somewhere mid-stream. The PDC emulator sits on Type = NT5DS (syncs with itself, effectively with the hardware clock), the Hyper-V hosts sit on Type = NTP (each pointing to a different source), and the VMs have their Hyper-V Time Sync IS enabled while also trying to do NT5DS. Result: three time sources each pulling in a different direction.

3. The Hyper-V Time Synchronization Integration Service, and when to disable it

The Hyper-V Time Sync Integration Service syncs the VM clock with the Hyper-V host's clock. That sounds useful, and for non-domain-joined VMs it is. For domain-joined VMs it is harmful and should be disabled.

Why off for domain-joined VMs:

A domain-joined VM should get its time from AD through w32time NT5DS. If the Hyper-V Time Sync IS is also active, you get two competing time sources: the host (with its own time state) and the domain. The moment host and domain disagree (which happens more often than you think), the VM clock starts ping-ponging between both.

The Microsoft guidance:

Microsoft's documentation has said this explicitly since Windows Server 2012 R2: for domain-joined VMs, Time Sync should only be active for the "Time Synchronization" sub-service during save/restore operations, not for regular clock sync. Practically that means the integration service itself off, or the specific sync component off.

Check per VM:

# Per VM, on the Hyper-V host
Get-VMIntegrationService -VMName 'MyVM' |
    Where-Object Name -eq 'Time Synchronization' |
    Select-Object VMName, Name, Enabled

Disable for domain-joined VMs:

# For one VM
Disable-VMIntegrationService -VMName 'MyVM' -Name 'Time Synchronization'

# For all VMs on a host (filter by name, OS type, or simply all)
Get-VM | Disable-VMIntegrationService -Name 'Time Synchronization'

In a cluster assessment we report this per VM, with a flag on every domain-joined VM that still has Time Sync enabled.

4. PDC emulator configuration, the root of all time

Everything starts at the PDC emulator. If that is wrong, the error propagates through the entire domain.

Which node holds the PDC emulator role:

# From any domain-joined machine
Get-ADDomain | Select-Object PDCEmulator

Configure the PDC emulator with an external time source:

# On the PDC emulator itself
w32tm /config /manualpeerlist:"time.windows.com,0x9 pool.ntp.org,0x9" `
    /syncfromflags:manual /reliable:yes /update

# Stop and restart w32time
Restart-Service w32time

# Force a sync and check status
w32tm /resync /rediscover
w32tm /query /status
w32tm /query /source

Tolerance settings we set as standard:

# Maximum 5 minutes in both directions (default is far higher)
w32tm /config /computer:<PDC> `
    /maxposphasecorrection:300 `
    /maxnegphasecorrection:300 `
    /update

Restart-Service w32time

The MaxPosPhaseCorrection and MaxNegPhaseCorrection values determine how large a time adjustment may be before w32time rejects it. Default values in older Windows versions are absurd (54 years), which means a corrupted NTP source can yank your clock wide open without w32time complaining. 300 seconds (5 minutes) is a safe threshold that forces manual intervention on large drift.

Verify that the PDC emulator actually syncs externally:

w32tm /query /source
# Expected: <external NTP source>, not "Local CMOS Clock" or "Free-running System Clock"

If we do not find this, the PDC emulator is effectively running on its own hardware clock, and all other DCs and servers in the domain cheerfully sync with that drifting clock.

5. Cluster node configuration, w32time settings that actually matter

Cluster nodes as domain-joined servers belong on Type = NT5DS. That is usually the default after domain join. But several settings deserve specific attention on Hyper-V hosts.

Verify per cluster node:

Invoke-Command -ComputerName (Get-ClusterNode).Name {
    [PSCustomObject]@{
        Node              = $env:COMPUTERNAME
        Type              = (w32tm /query /configuration | Select-String 'Type:').Line.Trim()
        Source            = (w32tm /query /source)
        LastSync          = (w32tm /query /status | Select-String 'Last Successful').Line.Trim()
        Stratum           = (w32tm /query /status | Select-String 'Stratum').Line.Trim()
        MaxPosPhaseCorr   = (Get-ItemProperty 'HKLM:\SYSTEM\CurrentControlSet\Services\W32Time\Config').MaxPosPhaseCorrection
        MaxNegPhaseCorr   = (Get-ItemProperty 'HKLM:\SYSTEM\CurrentControlSet\Services\W32Time\Config').MaxNegPhaseCorrection
    }
} | Format-Table -AutoSize

What the output should show:

  • Type must be NT5DS on every cluster node. NTP on a node that is not a DC or PDC is a misconfiguration.
  • Source must be a domain controller in the same domain. "Local CMOS Clock" or "Free-running System Clock" is a symptom of a w32time service that has lost its source.
  • LastSync must be recent, within the last sync interval (default 64 to 1024 seconds depending on stability).
  • Stratum must be one higher than the PDC emulator's. If the PDC is stratum 2 (syncing with external stratum-1), then all other DCs are stratum 3, and cluster nodes are stratum 4.
  • MaxPosPhaseCorrection and MaxNegPhaseCorrection we set as standard to 300 seconds, same as the PDC emulator. This prevents a corrupted sync source from suddenly shifting the clock by hours.

Measure clock skew between nodes:

$nodes = (Get-ClusterNode).Name
foreach ($node in $nodes) {
    foreach ($peer in $nodes) {
        if ($node -ne $peer) {
            $offset = w32tm /stripchart /computer:$peer /dataonly /samples:1 |
                Select-String '^\d' | Select-Object -First 1
            Write-Host "$node <-> $peer : $offset"
        }
    }
}

Acceptable skew between cluster nodes: under 1 second in steady state, under 5 seconds during a patch round or node reboot. Above 30 seconds is an urgent remediation topic.

6. Detecting drift before it fails

Monitoring tools typically alert only on clock skew that has already caused damage. We add a simple drift monitor to customer environments during Cluster Assessment remediation:

# Drift detection script, run on a management host or via scheduled task
$threshold = 30  # seconds

$cluster = Get-Cluster
$nodes = (Get-ClusterNode -Cluster $cluster).Name
$baseline = $nodes[0]

foreach ($peer in $nodes | Where-Object {$_ -ne $baseline}) {
    $output = w32tm /stripchart /computer:$peer /dataonly /samples:1 2>&1
    $offsetLine = $output | Where-Object {$_ -match '^\d'} | Select-Object -First 1
    if ($offsetLine -match '([\d\.\-]+)s') {
        $offset = [math]::Abs([double]$matches[1])
        if ($offset -gt $threshold) {
            Write-Warning "Drift between $baseline and $peer: $offset seconds"
            # Hook to SIEM, Teams webhook, or monitoring stack here
        }
    }
}

Schedule this every 15 minutes on a management host. The value of a dedicated drift monitor is that it catches issues hours before the first Kerberos failure, at a point where remediation is still non-disruptive.

Time service configuration is one of those topics every administrator thinks is in order, and where during a cluster assessment we find something off in 7 out of 10 environments.

A ClusterTriage Hyper-V cluster assessment checks time hierarchy, PDC emulator configuration, per-node settings, and Hyper-V Time Sync IS state per VM, and includes a remediation script with the report.

Schedule a cluster assessment intro call →

7. Three common misconfiguration patterns

Pattern 1: PDC emulator on Type = NT5DS

Classic configuration error. The PDC emulator names itself as source through NT5DS, which effectively means: I sync with myself, I use my own hardware clock. Works until the hardware clock drifts, and then the entire domain drifts with it.

Fix: set the PDC emulator to Type = NTP with an external peer list and /reliable:yes. See section 4.

Pattern 2: Hyper-V host on Type = NTP to an external source

A Hyper-V host syncing independently to time.windows.com instead of to the domain hierarchy. Sounds logical ("closer to the source"), but breaks the entire principle of a domain hierarchy and produces inconsistent time across hosts.

Fix: set the host to Type = NT5DS. The host then syncs with a DC, which syncs with the PDC emulator, which syncs with external NTP.

Pattern 3: VMs with Hyper-V Time Sync IS enabled plus NT5DS

Domain-joined VM trying to sync from both the Hyper-V host and the domain hierarchy. The clock ping-pongs, especially when host and domain disagree by a fraction of a second.

Fix: disable Hyper-V Time Sync IS per VM for all domain-joined VMs. See section 3.

8. Remediation playbook

The order we use as standard during remediation:

Step 1, identify the PDC emulator:

Get-ADDomain | Select-Object PDCEmulator

Step 2, fix the PDC emulator:

Configure external NTP sources, set /reliable:yes, set MaxPosPhaseCorrection and MaxNegPhaseCorrection to 300 seconds, restart w32time. Verify that w32tm /query /source actually shows an external source.

Step 3, fix all other DCs:

Set them to Type = NT5DS, verify their source is a DC in the domain. No external NTP sources on non-PDC DCs.

Step 4, fix all Hyper-V hosts and domain-joined member servers:

Set them to Type = NT5DS. Adjust MaxPosPhaseCorrection and MaxNegPhaseCorrection. Restart w32time. Verify sync.

Step 5, fix VMs:

For all domain-joined VMs: disable Hyper-V Time Synchronization Integration Service. Verify the VMs show a DC in the domain as source through w32tm /query /source.

Step 6, monitor:

Schedule the drift detection script from section 6 every 15 minutes. Hook it into existing monitoring or SIEM. Set a threshold at 30 seconds for early warning.

Step 7, document:

Write down which settings live where, which external NTP sources are used, and what the procedure is on a PDC emulator failover. This is exactly the kind of knowledge that gets lost during staff changes, and where Cluster Assessments two years later find drift in the same environment again.

The bottom line

Time service in a Hyper-V cluster is not "set and forget". The interaction between w32time, AD replication, cluster heartbeats, Live Migration handshake, and Hyper-V Time Sync IS is complex enough that a wrong assumption on one layer surfaces as a symptom on another layer. When Kerberos fails, everyone looks at AD. When cluster heartbeats flap, everyone looks at the network. The actual cause often sits two layers deeper, in a w32time configuration nobody has touched in the last three years.

A ClusterTriage Hyper-V cluster assessment includes time service analysis as standard. Findings are reported with severity that scales with actual skew and blast radius. Remediation scripts are idempotent and can be applied by the customer team themselves.

Schedule a Hyper-V cluster assessment intro call →

Frequently asked questions

My cluster has worked fine for years, why should I worry about time service?

Drift accumulates slowly. Until the moment a PDC emulator fails or a Hyper-V host is patched, a misconfiguration can stay hidden for years. We regularly see that a planned maintenance action is the first time drift becomes visible, and that is an unfortunate moment for surprises.

Does this work the same on Azure Local and Azure Stack HCI?

Largely yes. Azure Local nodes are domain-joined and follow the same w32time hierarchy. However, for Arc-managed Azure Local clusters there is an additional layer, because the Arc agent also does its own time validation toward Azure. The threshold there is stricter (five minutes max), so drift becomes visible sooner.

What if I am not allowed to use an external NTP source due to security policy?

Then a hardware stratum-1 source in the management network is the solution. Devices from Meinberg, Microsemi, or a GPS-disciplined Linux server can serve as an internal NTP source. The PDC emulator then points to that hardware instead of time.windows.com. For strict air-gapped environments this is the standard approach.

How much clock skew is acceptable between cluster nodes?

Under 1 second in steady state. Up to 5 seconds during a patch round or node reboot. Above 30 seconds is an urgent remediation topic. Above 5 minutes is already an active incident, because Kerberos will refuse new authentication.

What is /reliable:yes and when do you use it?

/reliable:yes marks a time source as "reliable", which allows other w32time clients to sync with it. It is a required setting on the PDC emulator (without it, it is not a valid time source for the rest of the domain). On a DC that is not PDC, /reliable:no or no setting is correct.