ClusterTriage / Blog / Azure Local
Azure Local on Proxmox: what it took
In the last week of September we deployed a two-node Azure Local 2609 cluster, nested on Proxmox VE, through Microsoft's official deployment path: nodes registered with Azure Arc, validated and deployed from Azure, including the Arc Resource Bridge. Microsoft does not support Azure Local on Proxmox. It does work, once Microsoft's environment checker accepts a KVM guest as a Hyper-V virtual machine.
This article walks through every check that failed, how we found the cause, and the fix for each. It closes with a checklist and with our view on running Microsoft workloads on Proxmox in general.
1. Why Proxmox, and why this is a lab
The lab exists to test our own tooling. The scripts behind ClusterDown.com and our cluster assessments need real clusters that misbehave. We first ran it nested on Windows Server 2025 Hyper-V, on a rented Hetzner dedicated server in Helsinki (Intel Core i9-13900, 128 GB). That worked, but the host itself caused most of the outages, and the cluster we had built by hand was registered in Azure without the Azure Local stack behind it.
Rebuilding on Hyper-V was an option. We chose Proxmox for three reasons:
- To find out how well it works. I wanted to know how a virtual Azure Local cluster holds up on a platform Microsoft did not design it for, and more broadly how well Windows workloads are supported on Proxmox today.
- Customers ask about it. Many organisations leaving VMware are evaluating Proxmox. When a customer asks whether their Windows and Hyper-V knowledge carries over, we want to answer from our own measurements.
- A more predictable host. A Linux host with a small footprint, ZFS snapshots and clones that make rebuilding a node a matter of minutes, and no Windows roles or monthly host updates that can take the lab down.
We wiped the server, installed Proxmox VE 9.2 on Debian 13, and deployed Azure Local 2609 from Azure onto two nested nodes.
The Hyper-V lab had four nodes; this one has two. The reason is memory. The hand-built Hyper-V cluster ran four small nodes. A real Azure Local deployment needs at least 24 GB per virtual node (we ended at 32 GB), and one node also carries the Arc Resource Bridge VM with another 8 GB. On a 128 GB host that also runs a domain controller, a management server and other workloads, four of those nodes do not fit. We started with three and went to two. A two-node cluster is a supported Azure Local configuration; it needs a witness, so we used a cloud witness in Azure.
This is a lab, not a recommendation. Microsoft supports nested Azure Local for evaluation on Hyper-V and on Azure virtual machines, not on KVM. For Microsoft workloads in production, Hyper-V remains my hypervisor of choice.
The build, the teardown and most of the troubleshooting were done by Claude agents running the commands, with me making the decisions. Rebuilding a cluster from scratch in a few hours is what made it practical to test every fix in this article on a fresh installation.
2. Read the environment checker instead of guessing
The Azure Local environment checker runs on the nodes and is plain PowerShell. When a check failed, we opened the checker's own module on the node and read what it tests. That turned out to be faster than any search, for two reasons:
- The checks are literal. Most of them compare one property with one value.
- The checker's functions can be called directly on a node. That gives an answer in seconds, where a full validation from Azure takes 30 to 60 minutes.
The same applies to storage. Test-Cluster runs locally in about six seconds and shows whether the disks will pass:
Test-Cluster -Node $env:COMPUTERNAME -Include 'Storage Spaces Direct'
3. Making a KVM guest look like a Hyper-V VM
The first validation failed on all nodes, on four hardware checks.
Virtual or physical. The checker decides this with one comparison: Win32_ComputerSystem.Model -eq "Virtual Machine". That is what Hyper-V reports. A QEMU guest reports "Standard PC (Q35 + ICH9, 2009)" and is judged as physical hardware, with the physical minimums (32 GB of memory instead of 24 GB). Proxmox's smbios1 setting fixes it: manufacturer "Microsoft Corporation", product "Virtual Machine".
ECC memory. The checker counts memory as ECC when TotalWidth is larger than DataWidth in the SMBIOS memory device table (type 17). QEMU leaves both fields empty. We give QEMU a 60-byte type 17 table of our own with -smbios file=, with a total width of 72 bits and a data width of 64. From 32 GB up, the size field in that table is too small and the memory size goes in the extended size field.
TPM certificate age. The checker requires the TPM endorsement certificate to be at least one day old. A virtual TPM creates that certificate when the VM is created. Create the nodes the day before you validate. When you rebuild a node later, keep its TPM state disk: the certificate then keeps its original date.
Mounted media. Proxmox cannot hot-unplug a SATA CD drive, so the installation media count as mounted until the next full stop and start. Remove the drives, then power cycle the VM.
One more from the unattended installation. Under UEFI, the installer waits for "press any key to boot from CD", so we send Enter through the hypervisor for the first few seconds. When we sent it for 20 seconds, one Enter reached the installer's Cancel button and the installation stopped at 6% with "Are you sure you want to quit?". Eight seconds is enough.
4. The network adapter must report 802.3
The checker requires every cluster adapter to report its physical medium as 802.3, value 14. The virtio adapter with Red Hat's NetKVM driver, the default choice on KVM, reports 0 (unspecified). That is intentional: an emulated adapter is not physical Ethernet, and Microsoft documents 0 as the value for emulated devices.
First deployment: vmxnet3. QEMU can also emulate VMware's vmxnet3. Azure Local has an inbox driver for it that reports 14, and our first deployment ran on vmxnet3. The drawback is that an emulated adapter then presents itself as physical hardware.
Second deployment: virtio. NetKVM reads a *PhysicalMediaType setting, so we first set it to 14 in the registry. After an adapter restart, a reboot and a cold start, the registry still said 14 and the adapter still reported 0. Windows uses the value that the driver package's INF installs; changing it afterwards has no effect.
What works is an extension INF: a separate driver package without code that sets this one value on top of Red Hat's unchanged NetKVM package. Windows applies it after the base driver, and again after every update of the base driver. We sign its catalog with our own certificate and place that certificate in the node's trusted stores, so it installs with Secure Boot on and without test mode. All four adapters on both nodes reported 802.3, and the second deployment ran entirely on virtio.
We use this override on our own lab nodes only, and we did not ask the NetKVM maintainers to change their default. For an emulated adapter, 0 is the correct value.
5. VLAN ID, disk identity and the clock
Each of these three cost a full validation round before we found it.
The VLAN ID property needs a value. The checker reads the adapter property VlanId and fails when it has no value. On the storage adapters it must be exactly 0. Both vmxnet3 and NetKVM list the property without a value. Set it on every adapter:
Set-NetAdapterAdvancedProperty -Name Storage1 -RegistryKeyword VlanId -RegistryValue 0
Every disk needs a WWN. Test-Cluster failed on every disk with "Failed to get SCSI page 83h VPD descriptors". QEMU builds that identity from the drive's internal name (drive-scsi1), which is the same on every node. Failover Clustering needs a unique identifier. A serial number is not enough; give each virtual disk its own wwn= value, unique across the whole lab.
The hardware clock must match the guest's time zone. After a restart, a node dropped out of Azure Arc with its clock two hours ahead. Windows reads the hardware clock as local time in its own time zone. Proxmox therefore gives Windows guests a hardware clock in the host's local time by default, which is correct as long as host and guest use the same time zone. Ours do not: the host runs on Central European time, and the Azure Local nodes run in UTC. Each node read a Central European clock as UTC and started two hours ahead. For guests that run in UTC, set the hardware clock to UTC per guest, followed by a full stop and start:
qm set <vmid> --localtime 0
6. A validation that succeeds only once
This problem is not specific to Proxmox.
Our fourth validation passed. Every validation after it hung, with Azure showing "provisioning" for an hour while the node had already reported success. The deployment engine remembers which steps passed and skips them on the next run. On a fully passed state, the whole validation finishes in about a second. The step that started it polls every ten seconds for the validation to be in progress, never sees that state, and waits for its own sixty-minute timeout.
What we do now:
- Fix everything before the first validation, using the checker's own functions on a node.
- After a successful validation, start the deployment. Do not validate again.
- To start over, start from a clean state: remove the Azure resources, check that the Key Vault name is free again, and rebuild the OS disks of the nodes.
Microsoft's own deployment template describes the order in its parameter text: "First must pass Validate prior running Deploy".
Is your cluster in trouble right now? ClusterDown.com gives you the five most important issues from one script and one zip file, free of charge. On this site, click Cluster issues? in the menu.
Go to ClusterDown.com →7. Deployment times and node sizing
The first complete run took 6 hours and 24 minutes from a clean start. That includes:
- cleaning up and rebuilding the nodes, about 25 minutes of machine time;
- the Arc registration;
- a validation of about an hour;
- a deployment of 4 hours;
- about half an hour of waiting for an interactive sign-in.
The longest single step was the Arc Resource Bridge, at 1 hour and 45 minutes. The second deployment, on virtio, took 6 hours and 17 minutes for the deployment alone. The entire difference was in that same Arc Resource Bridge step. We have not found the cause.
The cluster runs, but with little headroom:
- The nodes are small. Each node has four virtual processors, pinned to the host's E-cores, because the faster P-cores are reserved for CI runners.
- One node carries the Arc Resource Bridge. That VM has four virtual processors of its own, inside a node that has four. This gives three layers of virtualisation.
When we measured, that node used all of its processors. That is enough to run the cluster and not much more. A live memory dump froze the node long enough for the cluster to remove it, which we described in One incident, two reports.
For a rebuild we would give each node at least eight virtual processors and keep the nodes off the slowest cores. That is our expectation; we have not measured it.
8. Microsoft workloads on Proxmox: what you gain and what you give up
I would not recommend Proxmox as the default platform for Microsoft workloads. For specific situations it is a reasonable choice, provided you know what you give up.
What you gain
- A free hypervisor. Proxmox VE is open source and free to use. An optional subscription, priced per socket, provides the stable update repository and vendor support. Compared with the VMware licences many organisations are trying to leave, that is a significant difference.
- A predictable host. A Linux host with a small footprint, without Windows roles or monthly host updates that can take the platform down.
- Fast rebuilds. ZFS snapshots, clones and templates reduce rebuilding a node to minutes.
- Mature Windows support. The virtio drivers, UEFI with Secure Boot, a virtual TPM and the guest agent all work, and mainstream backup products such as Veeam support Proxmox.
What you give up
- Microsoft's support statement. Microsoft supports Windows Server guests on hypervisors that passed its Server Virtualization Validation Program (SVVP). As far as I know, Proxmox VE is not on that list. Microsoft will usually still help, but may ask you to reproduce a problem on a supported platform. Check the current SVVP list before you rely on it.
- Application support. SQL Server, Exchange and many line-of-business applications refer to the same list. Check this per vendor.
- Microsoft platform features. Azure Local itself, Windows Server Azure Edition with hotpatching, free Extended Security Updates on Azure Local, Hyper-V Replica, System Center Virtual Machine Manager and Shielded VMs only exist, or work best, on Hyper-V, Azure Local or Azure.
- Existing knowledge and tooling. Hyper-V and Failover Cluster experience does not transfer one to one. Clustering, storage and networking work differently on Proxmox.
- Defaults that do not suit Windows. We ran into a hardware clock in local time, a network adapter that reports "unspecified", and a QEMU guest agent that stopped responding under load.
What a free hypervisor does not change
A free hypervisor does not make Windows free. Windows Server is licensed per physical core of the host, regardless of the hypervisor, and Datacenter edition remains the way to license an unlimited number of Windows VMs on a host. Hyper-V is included in Windows Server at no extra cost. The free standalone Hyper-V Server ended with version 2019. The saving is therefore large compared with VMware and small compared with Hyper-V.
Our advice
- Mostly Microsoft workloads: stay on Hyper-V or Azure Local. Support, tooling and your team's knowledge fit, and the hypervisor costs nothing extra there either.
- Mostly Linux, with a limited number of standard Windows Server VMs (domain controllers, file servers, simple application servers): Proxmox is defensible. Before you decide, check the SVVP status, support per application, and a complete backup and restore.
- Labs, test and training: Proxmox is a good fit. That is what we use it for.
We have not compared the performance of Proxmox and Hyper-V on the same hardware. This advice is about support, features and operations, not speed.
9. Checklist
- Set the SMBIOS manufacturer to "Microsoft Corporation" and the model to "Virtual Machine".
- Provide an SMBIOS type 17 table with a total width larger than the data width (ECC), and use the extended size field from 32 GB up.
- Create the nodes at least one day before the first validation. Keep the TPM state disk when you rebuild a node.
- Remove all installation media and do a full stop and start.
- Network adapters: vmxnet3, or virtio with an extension INF that sets
*PhysicalMediaTypeto 14. - Set
VlanIdto 0 on every adapter. - Give every virtual disk a unique WWN.
- Set
localtime 0on Windows guests that run in UTC. - Run the checker's own functions and
Test-Clusteron a node before the first validation from Azure. - Validate once, then deploy. Do not validate a passed state again.
- Size the nodes for the Arc Resource Bridge. Four virtual processors per node proved too tight.
- In unattended installations, send Enter for a few seconds only.
Frequently asked questions
No. Microsoft supports a nested (virtual) Azure Local deployment for evaluation and testing on two platforms: Hyper-V, and virtual machines in Azure. Proxmox and other KVM-based hypervisors are not among them. Everything in this article is a lab setup for learning and testing, not for customer workloads.
That works, and our first deployment used it. We moved to virtio because it is the native adapter on KVM, and because an emulated adapter that reports itself as physical hardware is the wrong default. The extension INF makes the override explicit and leaves Red Hat's driver unchanged.
No. The package contains no driver code, only one setting, so it only needs a catalog signature that the machine trusts. A certificate of your own in the machine's trusted stores is sufficient for machines you manage. Public distribution would require Microsoft's signing process.
About an hour of validation and four to six hours of deployment, most of it in the Arc Resource Bridge step. Plan a full day for a first deployment, plus the day the TPM certificate needs before the first validation.