EXECUTIVE BRIEF

A split-brain is, at its core, when a high availability service runs in two places at the same time without coordination. It is a common misconception that this is a problem in shared storage. Thus, the tools used to prevent a split-brain are thought to be required for clusters with shared storage, and not otherwise. 

This is not the case.

Consider the simplest of all cluster-managed services; A virtual IP address assigned to the node offering the highly available service. What happens in a split-brain in this case? 

Imagine the node that has the virtual IP stops responding. It’s frozen, but not dead. The peer node checks with the arbiter node, says “I have enough votes for quorum, I am going to recover the virtual IP”. It does so, and now the backup node takes the service, the virtual IP and updates the network. Traffic starts to flow and the service is back online.

Some time later, the hung node recovers, even if only for a minute. In that brief time, it hasn’t realised yet that the peer took the IP. It sends a message using the virtual IP and the switches think the IP has moved to a new switch port. 

You have a network split brain.

There is a way to prevent any split brain from occurring, which we will cover in a later article. The key lesson here is this; 

If a service can run in two or more places, at the same time and without coordination, you don’t need HA. Just run the services on both or all hosts. If, on the other hand, running the same services on two or more hosts at the same time causes problems, then you can suffer a split brain. 

TECHNICAL DEEP-DIVE

A split-brain event is a partition: two (or more) subsets of a cluster each conclude, independently, that they hold authority to operate. This can be caused by a network failure, stray firewall rules, or a node that is intermittently halting and recovering.  The node on the other side of the partition may be perfectly healthy. It just can’t be reached, and from the surviving side’s perspective, “can’t be reached” and “is dead” are indistinguishable without an explicit mechanism to tell them apart.

Storage access in a split-brain is the version most people have already been warned about. Take a synchronously replicated block device like DRBD: under normal operation, every write hits both nodes before it’s acknowledged. During a partition, if both nodes independently decide they’re now the sole writer, each accepts local writes the other never sees. The two copies diverge at the block level. When connectivity returns, there is no safe automatic merge. To recover, only one can be kept, and the other’s writes are discarded. A more dangerous outcome is shared storage, like an iSCSI LUN or cluster file system where there is only one storage unit with shared access. In this scenario, uncoordinated access results in data corruption where “recovery” is restoring your last backups. 

On the other end of the risk spectrum is a virtual IP address assigned to an active node. Split brains are dismissed as low risk and not sufficient to justify split-brain protection.  Two nodes each bring up the same address, potentially causing a split where the same IP is claimed by different devices. Downstream, switches and neighboring routers update ARP or MAC tables based on whichever advertisement arrived last at each hop, which means different clients, and even different packets from the same client, can resolve to different physical nodes depending on timing.

Generalize past both examples and the underlying principle is exclusivity of ownership: any resource that cannot tolerate concurrent, uncoordinated access from two actors needs a mechanism that guarantees only one actor holds it at a time. That covers block storage, network identity, and any other service that needs to be HA and coordinated.

Voting alone can’t provide a guarantee sufficient to prevent this risk. The absence of information, the lost node’s state, can not then infer the peer is dead. Fencing closes that gap by converting an inference into a forced, verifiable state change. Fencing converts the ambiguous state into a confirmed state, either by forcing the target to be off, or by severing connections to networks and shared resources in a way that can be confirmed. 

The practitioner-level takeaway, applicable to any two-node evaluation and not just Alteeve’s: ask what a given architecture actually knows about a silent node’s state before it acts on that silence. If the honest answer is “we assumed,” that’s an exposure, whether or not it’s ever been triggered yet.

Share