The Physics of Two-Node HA
EXECUTIVE BRIEF
Two-node high availability has an image problem: for years the default assumption was that anything less than three nodes couldn’t reliably survive a failure, because quorum math needs a majority, and two nodes can’t produce one. That assumption is now visibly outdated. OpenShift’s Two-Node with Fencing capability reached general availability in OpenShift 4.22 in mid-2026, standing alongside Red Hat’s own arbiter-based two-node option, meaning the two-node-plus-fencing pattern Alteeve has run in production for years is now a model the platform itself ships as a first-class path, not a workaround.
What actually separates these approaches isn’t node count. It’s whether the system knows a lost node’s state, or merely assumes it.
Quorum, the mechanism behind Kubernetes’ and etcd’s consensus model, resolves cluster membership by majority vote: 5-% + 1 surviving nodes agree on who’s still in, and that agreement becomes the operating truth. It’s a proven model, and it works cleanly under the condition where a node’s absence from the vote reliably means the node has actually stopped acting. In a clean lab test, that condition holds: pull the power, the node dies instantly and completely, and the vote reflects reality.
Real failures are rarely that clean. A node can freeze and partially recover. It can suffer memory or CPU errors that corrupt computation without crashing anything. It can lock up into a state that looks, from the network, exactly like a partition, while eventually being able to recover and continue writing to local storage. In every one of these cases, “not answering cluster messages” doesn’t mean “no longer functioning”, and quorum has no mechanism to tell the difference.
It can only be assumed.
Fencing removes the assumption. Instead of inferring a node’s state from its silence, it forces a confirmed outcome. The suspect node is verified powered off or verified isolated before recovery proceeds. Not a vote about probability.
A known state, no assumptions.
TECHNICAL DEEP-DIVE
Quorum-based consensus – the model underlying etcd, and by extension the Kubernetes and OpenShift control plane – resolves cluster membership through majority agreement. Formally, a partition of N nodes needs ⌊N/2⌋+1 (rounded down) members to proceed safely, which is why odd node counts are the norm, but not required: with an odd N, a tie is mathematically impossible, and exactly one partition can ever hold majority.
Two node clusters break this mechanism. A 1-1 split has no majority on either side, which is why a bare two-node deployment can’t use quorum alone to resolve a partition – something else has to break the tie.
Red Hat’s own two-node arbiter model (Two-Node with Arbiter, running alongside Two-Node with Fencing since OpenShift 4.22’s general availability) solves the tie problem by adding a lightweight third voter – a small node or device that doesn’t run workloads, but does cast a vote, restoring a workable majority with only two “real” nodes doing service work. It provides its vote to one of the nodes, giving it a quorum, but it does not confirm what the losing node is actually doing. The core assumption holds: that a lost node’s silence means it has stopped acting.
Quorum is built on the assumption that a lost node behaves predictably.
Fencing takes a structurally different approach: it doesn’t ask which side has more votes, because it doesn’t need consensus about who’s “more correct.” It directly confirms – or forces – the state of the disputed node before any recovery begins.
It removes the ambiguity of the inquorate node’s behaviour. To understand the importance of this, consider these possible issues quorum alone wouldn’t protect against:
— Partial-write corruption: The surviving, quorate, node replays the file system journals and proceeds with normal disk activity. The lost node partially recovers, doesn’t realize yet that it is inquorate, and completes the transaction it was in the middle of, corrupting the file system.
— Freeze-then-partial-recovery: A common example is a virtual IP being claimed by the backup, the old primary partially recovers and claims the IP back, then freezes again. This “flapping”, common in poorly designed HA firewalls, can leave neither machine with the IP, both thinking the other has taken it.
In each case, the honest state of the missing node is not “gone.” It’s “unknown,” and quorum has no path to resolve unknown into known states — it can only proceed as if the two are the same thing.
Fencing resolves this with a confirmable state transition rather than an inference: a STONITH agent – IPMI power-cycle, PDU port-cycle, SAN-level fencing – forces or verifies the disputed node’s state directly. Recovery relies on that confirmation existing, not on a vote reaching a threshold.
For anyone auditing an HA architecture – Alteeve’s or otherwise – the diagnostic question is specific and answerable: when a node goes silent, what does the surviving side actually know about its state before acting? “We inferred it from a vote” is assumed state. “We forced it into a state we can verify” is known state. Node count is a secondary detail; this distinction is the one that determines what actually happens when a failure doesn’t die cleanly.
