The Role of Backups in Availability.

The most common “redundancy” are simple backups. This is good, backups are required and even the very best availability solutions require a good backup strategy. From an availability perspective though, backups are not helpful. 

Availability speaks to how quickly you can get a system or a service back online. If you rely on backups only, your availability starts with a call to a supplier and asking how long it will take to get the hardware repaired or replacements to your site. 

Thus, the question of availability comes down to having the hardware in-hand before the emergency arises. 

The bottom rung: the cold spare

The simplest and most common availability is the “cold spare”. This is a machine that is sitting in storage, ready to be deployed when needed. A cold spare removes the need to call a supplier and hope they have what you need in stock. With a cold spare, the process of restoring availability begins immediately. 

This can, on its own, reduce recovery time by hours to days. 

Cold spare risks exist, and need to be understood. A machine that is turned off for a long time may develop undetected problems. If the cold spare sat on a shelf for two years, are the fans, capacitors, and storage healthy? 

Ideally, a cold spare would be turned on and checked, firmware and OS updates applied, and validated before being returned to storage. In practice, the ever-loaded IT department often must prioritize other tasks, and so cold spares sit idle.  

The next rung: the warm spare

A live spare is simply a cold spare that is left up and running. This solves the main concern of a cold spare failing silently. Having the spare up and running, monitored for faults and software updated, solves the main concerns of cold storage. 

Should a failure occur, there is no bring-up delay. Log in, run your restore from backup, change the IP address and your services are back online. 

Albeit, with the loss of any data between the last backup and the failure. In the “classic” world of redundancy, this is the best you can do before you enter the world of “High Availability”. 

Traditional HA

High availability takes the warm spare and turns it into a “hot spare”. The core difference is that the data on the spare is kept much closer to the live data than a backup dataset can provide. 

High availability is defined, fundamentally, but this replication of data. How that is implemented can vary quite significantly. The two core categories can be broadly described by the frequency of the data updates; “Asynchronous” vs “Synchronous”. 

Asynchronous HA

In asynchronous HA, the data on the spare is allowed to fall behind to some degree. This can be a stream of data that is allowed to lag production data, or it can be periodic “check points” where the production data is copied over on a tight time schedule, from hourly to every minute.

The primary reason async HA is chosen is performance. By allowing the spare to fall behind, you don’t “hold back” the performance of the production server. In the days of platter-based storage, this was particularly important and an appealing trade off. 

Synchronous HA

This is the gold standard for availability.

In synchronous HA, the data is stored in such a way that the backup has an up to date view of the live data. When the spare is pressed into service, employees and customers can pick up right where they left off.

There are different ways to achieve this, each with their own pros and cons. We will explore these options in future articles.

Intelligent Availability®

In traditional High Availability, the focus was on the speed of recovery, and the data loss accepted in that recovery. 

HA is reactive. 

Adding intelligence to availability means predicting failures before they take a service offline. An intelligently available system takes the core concepts of HA and extends them to be proactive. 

What does this look like?

One of the most common reasons a service fails is a storage fault. Drives can be made redundant via RAID, but their controllers can fail, taking the entire array offline. Whatever the cause may be, the machine loses the ability to run software. 

In an intelligently available system, a drive in prefailure, a RAID controller registering ECC failures, or other health indicators can be monitored. An IA system sees these signs of trouble and moves the services to the spare before the failure actually happens.

Whatever the cause of the outage would be, an IA system’s ability to proactively migrate services means that in many cases, the services are already moved off before a machine fails.  

With an IA based system, faults become truly transparent to your customers.

Available Infrastructure; Beyond The Server

It is natural when considering availability to focus on the server that runs your applications and stores your data. A robust availability plan needs to look beyond the server.

A spare is not able to run when the power feeding the production server has failed. A spare is not effective if the network path between your users and your services has failed. A spare is not useful if the room it is in has overheated. 

To some degree, these concerns can be allayed by moving the spare to a different location. Site redundancy is a valid and core part of a disaster recovery plan, and worth having for that merit alone. However, the geographic distance makes data synchrony harder or impossible to  implement. 

When considering the viability of small-window asynchronous replication, or the latency and bandwidth needed for synchronous replication, distance is the enemy. 

A fully available system requires “intelligence” that goes beyond the software. It requires the experience of understanding what external faults take systems offline. It requires the insights needed to know how to make the infrastructure around the servers just as resilient and redundant as the servers themselves.

Why this matters at the edge

In a data center, the cost justifications of investing if full redundancy is often an easy argument to make. This way of thinking about redundancy, effective as it is, can lead to the assumption that good redundancy at the edge is unfeasible. 

This isn’t the case. 

A machine’s internal computer, a retail check out, a remote field data collection system, they all provide critical services to your business. High availability is not beyond feasibility at the edge, but it does look different.

Physical space constraints, budget considerations, and most importantly, lack of available IT staff changes what the availability strategy will look like. 

With a proper understanding of the various layers and levels of redundancy and availability options possible, a sized-to-fit intelligent availability solution is available and affordable, even at the most distant and constrained edges.

Share