Every time someone asks how to safely take a host offline for updates or hardware work in an XCP-ng pool, the first reply is almost always the same: enable maintenance mode and let the system evacuate the VMs. I followed that exact path on my own three-host HA pool with load balancing turned on. The result was an immediate HOST_NOT_ENOUGH_FREE_MEMORY error. The evacuation tried to shove every VM onto a single remaining host instead of spreading the load, and the free memory simply wasn’t there. I had to turn load balancing off, manually migrate VMs across the other two hosts, and only then could I finish the job. It worked, but it felt like the system was fighting me.
After digging around and testing, I found a far more reliable starting approach: make sure every VM that needs to survive has its HA restart priority set to Restart rather than Best-effort. It costs nothing extra in hardware, yet it changes the evacuation behaviour enough that maintenance mode actually succeeds without the memory error. In my experience this small configuration choice is the difference between a smooth maintenance window and an afternoon of manual VM shuffling.
The practical win that keeps paying off
The real value isn’t just “maintenance mode works.” It’s that the pool finally has a coherent plan for where every protected VM should land when a host disappears. Once the priorities are set correctly, the evacuation logic stops treating Best-effort VMs as optional luggage and starts treating them as part of the protected set. That single change opens the door to a bunch of everyday operations that used to feel risky.
I’ve used it for:
- Rolling kernel and XCP-ng updates across a three-host pool without ever dropping the critical VMs.
- Swapping a failed drive or upgrading RAM on one host while the rest of the workload stays online.
- Testing new storage repositories by evacuating a host, attaching the new SR, and bringing it back.
- Running firmware updates on the servers themselves during a short maintenance window.
- Moving a host into a different physical rack or power circuit without scheduling a full outage.
- Simply reclaiming a host temporarily for some heavy benchmarking while the production VMs stay happily distributed.
All of those tasks still feel safe months later, long after the initial setup. The configuration doesn’t become obsolete once you move past the beginner stage; it actually becomes more useful as the pool grows and the number of VMs increases.

When the workload gets heavier
Compared with leaving everything on Best-effort (the default a lot of people never touch), the Restart setting is stricter. You lose a little flexibility because the HA planner now has to guarantee restart capacity for every protected VM. On a three-host pool that tolerates one failure, that can occasionally surface an HA_OPERATION_WOULD_BREAK_FAILOVER_PLAN message if you’re already running very close to the memory edge. I’ve hit that once or twice when I was being aggressive with overcommitment. In those cases I either temporarily lowered a couple of non-critical VMs back to Best-effort or simply shut a couple of them down for the duration of the maintenance window. It’s a minor annoyance, not a deal-breaker.
I’ve also tried the alternative of just disabling HA for the maintenance window. It works, but it feels like removing the seatbelts so you can change a tire. Once the pool is back online you still have to remember to turn HA back on and re-check every priority. Setting the priorities correctly once and leaving them there has been less error-prone for me.
If you’re comparing the effort to buying a fourth host purely so you never have to think about memory during evacuations, the configuration change wins hands-down on cost. A proper extra host is a significant capital outlay; changing a few HA settings is free.
When another approach might be better?
There are situations where this isn’t the right first step. If your pool is only two hosts and is configured to tolerate zero failures, the current evacuation behaviour is still limited and an upcoming change in the upstream code may eventually help more. In that specific case, temporarily disabling HA or manually migrating VMs can still be faster. Likewise, if you have a large number of truly non-critical VMs that you never want to protect, leaving them on Best-effort and only protecting the important ones is cleaner than forcing everything to Restart. And if you’re still in pure lab mode with no real uptime requirements, the whole HA machinery may be more trouble than it’s worth.
For any production-leaning three-host (or larger) pool, though, getting the priorities right first has saved me more time than any other single habit.
In short, the next time you need to put a host into maintenance mode and the system complains about free memory, check the HA restart priorities before you start manually moving VMs or turning features off. Set the ones that matter to Restart, try the evacuation again, and you’ll usually be done in a couple of minutes instead of an hour. It’s a small configuration detail that quietly makes the whole pool more predictable.