Session 1 - Reading the Plant, Five Alerts, Disaster Recovery and the Cloud

The opening session of a data centre operations track for someone new to the field and newly in the job. It builds just enough plant anatomy to make an alert readable - the power chain, hot and cold aisle airflow, and the N, N+1 and 2N redundancy vocabulary - then installs a four-question triage routine and runs it against five real scenarios: a rack thermal alert, a UPS transferred to battery, packet loss on a core uplink, a failed disk mid-rebuild, and a capacity threshold. It closes with disaster recovery built from RTO and RPO outwards, the 3-2-1 rule and recovery site strategies, and what regions, availability zones and the shared responsibility model change once the data centre is somebody else's building. Two traps recur deliberately: stabilising is not resolving, and redundancy is not backup.

Subject: Data Center Operations · 60 slides · applied lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Data Center Operations

Title

Session 1

The plant, five alerts you will actually see, disaster recovery, and what changes in the cloud

2. What This Session Buys You

Objectives

You asked for a routine: five plausible alerts every session, what each one means, and how to resolve it. This deck sets that routine up, and adds the two topics you asked for on top of it.

Nothing here assumes prior data centre experience. It does assume you will ask when a term goes past you.

3. Part 1 - What the Building Is Actually Doing

Section

Four systems, and only one of them is the computers

4. A Question Before Any Alerts

Warm-up

Have a guess before reading on. Being wrong here is the useful part.

Discussion prompt

Across the industry, what causes the largest share of serious data centre outages: hardware failure, software bugs, power problems, or human error during a planned change?

Hint: Think about when outages tend to happen, not what breaks.

Answer:

Power problems are consistently the largest single technical cause, and human error during planned maintenance is close behind and arguably ahead once you count it honestly.

That matters for how you read alerts. The alert that fires at two in the afternoon during a scheduled change is more suspicious than the one that fires at two in the morning, because somebody was touching something.

5. Four Systems, Stacked

Concept

A data centre is four systems that each exist to keep the next one alive. Learning them in this order makes every alert easier to place.

  1. Power. Utility feed, transfer switch, generator, uninterruptible power supply, distribution units, and finally the two supplies in each server.
  2. Cooling. Chillers or condensers, air handlers on the floor, and the airflow path that carries heat away from the chips.
  3. Network. Server to top-of-rack switch, top-of-rack to spine or core, core to the outside world.
  4. Compute and storage. The racks themselves, which are the only part of the building that does anything a customer pays for.

An alert almost always names a component in one of these four. Your first move is to say which layer you are in, out loud.

6. The Power Chain, End to End

Picture it

Trace it once now, and every power alert for the rest of your career has a location.

Figure (svg): Diagram of the power chain from utility feed through automatic transfer switch, UPS, power distribution unit, rack PDU and server supply, with a generator feeding the transfer switch and labels showing which element covers which outage duration.

Two different backups covering two different timescales. The UPS covers seconds, the generator covers hours.

The UPS is not there to run the building. It is there to keep it alive for the twenty or so seconds the generator needs to start and take load.

7. Put the Power Chain in Order

Ranking

From the street to the chip. Get this wrong and you will chase the wrong end of an outage.

Put in order

  1. utility feed from the street
  2. automatic transfer switch
  3. uninterruptible power supply
  4. floor power distribution unit
  5. rack power distribution unit
  6. the server's own power supply

Why: Power arrives from the street, hits the transfer switch that decides between utility and generator, passes through the UPS which smooths and bridges, then is distributed to the floor, then to the rack, and finally into the two supplies inside each server. Anything upstream of a failure loses everything downstream of it.

8. Cooling Is an Airflow Problem, Not a Temperature Problem

Concept

The building is not trying to be cold. It is trying to make sure the air arriving at the front of each server is cool, and that the hot air leaving the back never gets a chance to come round again.

Cold aisle — The side the servers breathe in from. Perforated floor tiles or overhead supply put cool air here.

Hot aisle — The side the servers exhaust into. Nothing should be drawing intake air from here.

The industry guideline for the air arriving at the intake is roughly 18 to 27 degrees Celsius. That is a target at the rack face, not a thermostat reading for the room.

ASHRAE TC 9.9, Thermal Guidelines for Data Processing Environments (recommended inlet envelope) TC 9.9 — The recommended inlet envelope.

9. Where the Heat Actually Goes

Picture it

One empty rack slot with no blanking panel is enough to break this.

Figure (svg): Diagram of a rack between a blue cold aisle and a red hot aisle, with arrows showing intake and exhaust, and a highlighted empty rack unit with no blanking panel letting hot air leak back to the front.

The leak path is short and it is free. Hot exhaust re-enters the intake through the empty slot.

This is why blanking panels matter more than they look. They cost a few dollars and they are load-bearing for the whole cooling design.

10. What Does the Gap Do?

Prediction

A rack has three empty units in the middle with no blanking panels fitted.

Predict first

Which server is most likely to alert on high inlet temperature?

  • The one at the very bottom of the rack
  • The one directly above the gap
  • The one directly below the gap
  • All of them equally, since the room is one temperature

Correct: The one directly above the gap

Why: Hot exhaust leaks forward through the gap and then rises, so it is drawn into the intake of whatever sits immediately above the opening. The bottom of the rack usually gets the coldest air. The last option is the beginner assumption this whole slide is aimed at: a data centre is not one temperature, it has a map of temperatures, and the rack face is the only place that counts.

11. Redundancy, Named Honestly

Concept

Redundancy is described in terms of how much of the plant you can lose and still carry the load. The notation is compact and worth reading precisely.

notationwhat it meanswhat it survives
Nexactly enough capacity for the loadnothing, including maintenance
N+1one spare unit beyond what the load needsone unit failing, or one being serviced
2Ntwo complete and independent systemsan entire system failing, including its distribution
2N+1two complete systems, each with a sparea failure during maintenance on the other side

The Uptime Institute tiers wrap this into four levels. The line worth memorising is between Tier III, which can be maintained without taking the load down, and Tier IV, which can take a fault without taking the load down.

Uptime Institute, Tier Classification System (Tiers I-IV) Tiers I-IV — The classification these terms come from.

12. Fill In the Redundancy Table

Comparison

Complete the missing cells from the definitions you just read.

Comparison matrix

ConfigurationSurvives one failureSurvives failure during maintenance
Nnono
N+1yesno
2Nyesno, one side is already down
2N+1yesyes

The third column is the one people get wrong on interviews. N+1 with a unit already pulled for service is just N, and N has no margin at all.

13. The Four Questions, Every Alert, Every Time

Pattern

This is the routine. It does not change between a thermal alert and a storage alert, which is exactly why it is worth drilling.

  1. What is the sensor actually measuring? Not what the alert is named, what the number is. A rack alert may be reading one intake sensor of six.
  2. What is the blast radius right now? One server, one rack, one row, one room, or the whole site. This decides how fast you move.
  3. Is this the symptom or the cause? A high temperature is a symptom. A stuck damper is a cause. You will usually be handed a symptom.
  4. What stabilises it now, and what fixes it properly later? Those are two different tickets and they should stay two different tickets.

Write the four answers down before you touch anything. The habit costs ninety seconds and it will one day stop you from making an outage worse.

14. Reading the Redundancy Right

Check

Solve it on paper before you click.

Check your understanding

A room is fed by a 2N power design. One of the two UPS systems has been taken offline for a scheduled battery replacement. What is the room's effective redundancy while that work is happening?

  • A. Still 2N, because the design is 2N
  • B. N+1, because one system is still spare
  • C. N, with no margin for a further failure (correct)
  • D. Zero, the room is down

Answer: C

Why: With one of the two systems out, the remaining system is carrying the whole load on its own. That is exactly N: enough capacity, no spare. The room is up and serving normally, but a single further failure now takes the load down, which is why maintenance windows are the highest-risk hours in the building.

Why A tempts people
The design is 2N, but the configuration right now is not. Redundancy is a property of the current state, not of the drawing.
Why B tempts people
N+1 would need a spare beyond the load. The surviving system is the load, not a spare.
Why D tempts people
The room is still up. The remaining system is carrying it. Losing margin is not the same as losing power.

15. Part 2 - Five Alerts

Section

The routine you asked for, run five times

16. The Triage Loop

Picture it

Keep this on screen while we work through the five. Every one of them goes through the same four boxes.

Figure (svg): Flowchart of four stacked triage steps: what the sensor measures, the blast radius, symptom versus cause, and stabilise now then fix later.

Four boxes, in order, every time. The discipline is in not skipping box three.

Notice that resolving the alert is the last box, not the first. Most bad incident handling is somebody starting at box four.

17. Alert 1 - High Inlet Temperature on a Rack

Concept

The page says: rack B14 inlet temperature 31 degrees Celsius, threshold 27, rising.

The useful instinct is that a thermal alert is a question about air, and air problems are almost always mechanical or geometric rather than electronic.

18. Triaging the Thermal Alert

Worked example

Working the four questions in order on rack B14.

Check whether neighbouring racks are also alerting.

Why: This separates a rack-local problem from a plant problem in about ten seconds, and it changes who you need to wake up.

If only B14 is hot, walk the rack front and back.

Why: Rack-local causes are visible: a missing blanking panel, a cable bundle blocking the rear exhaust, a floor tile moved during a recent install, or a failed fan in one chassis.

If the row or room is hot, move up to the cooling plant.

Why: Now you are looking at air handler status, chilled water supply temperature, a failed compressor, or a damper that has not opened. This is a facilities escalation, not a server one.

Stabilise before diagnosing further.

Why: Restoring airflow buys you time. Perforated tiles can be added, a temporary blanking panel fitted, or non-critical load shed. Servers throttle before they fail, so a few degrees of margin is worth a lot.

Verify: confirm the temperature is falling, not just the alert clearing.

Why: An alert can clear because the sensor crossed back under the threshold for one poll. Watch the trend for ten minutes. If the number is flat just under the threshold rather than falling, the cause is still there and it will alert again on the next warm afternoon.

Figure (svg): Decision diagram splitting a thermal alert on whether neighbouring racks are also hot, into rack-local causes on one branch and cooling plant causes on the other.

One question, ten seconds, and it decides which team owns the rest of the night.

19. Why Walk the Back of the Rack?

Socratic

You checked the front of B14 and the blanking panels are all fitted correctly.

Discussion prompt

Why is the rear of the rack still worth walking before you escalate to facilities?

Hint: Where does the heat leave from, and what else lives back there?

Answer:

The rear is where the exhaust leaves, and it is also where every power and network cable is. A dense bundle of cabling grown over years of moves and changes can block a real fraction of the exhaust path.

If the hot air cannot leave the back, it builds up and finds its way to the front regardless of how good your blanking panels are. Cable management is a cooling control, not a tidiness preference.

20. Turning the Cooling Up Is Not a Fix

Trap

The trap

Rack B14 is hot. The room set point is lowered by three degrees. The alert clears within twenty minutes. Ticket closed as resolved.

Reasoning: the temperature came down, and the temperature was the problem.

The fix

The alert cleared and nothing was fixed. The cause was a blocked exhaust path or a failed fan, and it is still there. You have paid for the fix with energy across the entire room, permanently, to mask one rack's defect.

Two things went wrong. The room was over-cooled to hide a local fault, which raises the power bill on every rack and degrades your efficiency number. And the real fault stayed in place, so it will resurface on the next hot day, at a worse hour, with less margin.

The right shape is: stabilise locally, raise a ticket for the cause, and put the set point back once the cause is fixed. Stabilising is allowed. Calling stabilisation a resolution is not.

21. Alert 2 - UPS Transferred to Battery

Concept

The page says: UPS 2 on battery, load 42 percent, estimated runtime 11 minutes.

This is the one alert in this deck where you do not have time to be thorough first. Establish whether the generator is running before you do anything else.

22. How Long Do You Actually Have?

Prediction

The battery is rated for 10 minutes at full load. The current load is 42 percent, and the alert reports 11 minutes remaining.

Predict first

If a colleague now powers up a large test cluster on that same UPS, taking the load to 85 percent, what happens to your remaining runtime?

  • It roughly halves
  • It stays about the same, because the battery is the same size
  • It falls by much more than half
  • It rises, because the load is better balanced

Correct: It falls by much more than half

Why: Battery runtime does not fall in a straight line with load. Discharging faster gets less total energy out of the same battery, so doubling the load cuts the runtime by more than half. This is why the first instruction during a transfer to battery is to add nothing, and if possible to shed load. The timer in the alert assumes today's load and stops being true the moment anyone changes it.

23. Alert 3 - Packet Loss on a Core Uplink

Concept

The page says: 0.8 percent packet loss on the uplink from switch tor-14 to core-a, sustained for six minutes.

The single most useful discriminator here is whether the interface is reporting errors or reporting discards. They point at opposite causes.

24. Separating a Broken Link From a Full One

Worked example

Two counters, two very different stories, and they sit next to each other in the same command output.

Read the interface error counters first.

Why: Cyclic redundancy check errors, alignment errors and symbol errors mean frames are arriving damaged. Damage points at the physical layer: a dirty or failing optic, a bent or over-bent fibre, a bad patch, or a marginal cable.

Now read the discard and queue-drop counters.

Why: Discards mean frames arrived intact and were thrown away because there was nowhere to put them. That is congestion, not damage, and no amount of replacing optics will help it.

Check whether the loss correlates with a traffic peak.

Why: Congestion loss tracks the traffic graph. Physical loss does not care what time it is, which is a fast way to tell them apart on a dashboard.

Stabilise by moving traffic off the suspect path.

Why: If the uplink is one member of a bonded pair, taking the member out of service administratively is safer than leaving a half-broken link carrying production traffic. A cleanly down link reroutes. A flapping link poisons everything.

Verify: re-check the counters after the traffic has moved.

Why: Clear the counters, wait, and read again. If errors stop accumulating once traffic is off the link, the link is the problem. If loss follows the traffic to the other uplink, the problem was never the link and you have just learned something far more important.

Figure (svg): Two panels comparing interface error counters, meaning damaged frames from a physical cause, against discard counters, meaning intact frames dropped through congestion.

Adjacent counters, opposite causes, different teams.

25. Sort the Causes by Layer

Discrimination

Same symptom, three different teams. Put each cause where it belongs.

Sort into buckets

Sort each cause of packet loss by the layer you would investigate it in.

Physical - something is damaged
a dirty fibre optic connector; a failing transceiver reporting low light
Congestion - the pipe is full
an uplink running at 98 percent utilisation at peak; a bonded link where one member is down, halving capacity
Configuration - somebody changed something
a mismatched MTU between two devices; a routing change that pushed two rows onto one path
phys
These show up as error counters, because frames are arriving corrupted. The fix is hands on hardware: clean, reseat or replace. Traffic levels are irrelevant to them.
cong
These show up as discards rather than errors, and they track the traffic graph. The fix is capacity or traffic engineering, never a new optic.
conf
These appear suddenly, correlate with a change window, and often affect only some traffic, such as only large frames in the MTU case. Always ask what changed before assuming hardware.

26. Alert 4 - Disk Failure in a Storage Array

Concept

The page says: array sn-03, physical disk 7 failed, array degraded, rebuild started to hot spare.

The thing to understand is the rebuild window. From the moment a drive fails until the rebuild completes, your redundancy is spent, and a rebuild reads every remaining drive from end to end, at full speed, which is the most stressful thing you will ever ask of them.

27. Reasoning About the Rebuild Window

Worked example

Why the clock on a degraded array matters more than the alert's calm wording suggests.

Establish which parity level the array uses.

Why: A single-parity array tolerates exactly one failed drive. While it is degraded it tolerates none. A double-parity array still has one failure of margin left while rebuilding, which is a completely different risk posture.

Estimate how long the rebuild will take.

Why: Rebuild time scales with drive capacity and with how busy the array is with real work. Large drives can take many hours, and a heavily loaded array will stretch that out considerably because production traffic is competing with the rebuild.

Recognise why rebuilds are when second failures appear.

Why: A rebuild reads every sector of every surviving drive. Drives in the same array were usually bought together, installed together, and have run the same hours. A latent defect that never surfaced under normal reads surfaces here.

Replace the failed drive rather than only watching the rebuild.

Why: The hot spare is now part of the array, so the array has no spare until you physically replace the dead drive and let a new spare take its place. Closing the ticket when the rebuild finishes leaves the array without a spare indefinitely.

Verify: confirm both that the array is optimal and that a spare exists.

Why: Two separate conditions. An array can report optimal, meaning full redundancy restored, while having zero hot spares available. Check for both before you consider this closed, because the next failure is the one the spare was for.

Figure (svg): Timeline from disk failure to rebuild completion, showing that a single-parity array has zero redundancy across the whole window while a double-parity array retains one failure of margin.

The same alert text. The parity level decides whether this window is survivable.

28. A Redundant Array Is Not a Backup

Trap

The trap

The storage is on a double-parity array with a hot spare, replicated to a second array in the same room. It can survive multiple drive failures. Therefore the data is backed up.

Reasoning: the data exists in more than one place, and that is what a backup is.

The fix

Redundancy protects against hardware failure. It does not protect against anything else, and almost nothing that destroys data in practice is a hardware failure.

A deletion replicates. Ransomware encryption replicates. A bad database migration replicates. A fire in that room takes both arrays. In each case the redundancy worked perfectly and faithfully preserved the damage.

A backup is defined by being separated in time, so you can go back to before the damage, and separated in place, so a single event cannot reach both. Redundancy gives you neither of those. This distinction is the whole reason the next part of this deck exists.

29. Alert 5 - Capacity Threshold Crossed

Concept

The page says: volume vol-prod-02 at 85 percent capacity, growth 1.4 percent per day.

Capacity alerts are the only ones in this deck you are allowed to schedule rather than act on, and that is exactly why they are the ones that eventually cause outages.

30. Turning a Capacity Alert Into a Date

Worked example

A percentage is not actionable. A date is. Converting one to the other is the entire job here.

Compute the remaining runway from the level and the growth rate.

Why: At 85 percent full and growing 1.4 percent per day, there are 15 percentage points left, which is a little under 11 days. That is now a calendar problem rather than a dashboard colour.

\[ \frac{100 - 85}{1.4} \approx 10.7 \text{ days} \]

Ask whether the growth is expected.

Why: Steady growth matching the business is a capacity request. A rate that changed recently is an incident wearing a capacity alert's clothing, and the usual culprits are a log level left on debug after troubleshooting, a retention policy that stopped running, or a job writing temporary files it never cleans up.

Decide urgency from the runway, not from the percentage.

Why: Eighty-five percent with eleven days is a ticket for this week. Ninety percent with a rate that just tripled is tonight's problem. The same colour on the dashboard, two completely different responses.

Verify: re-measure the rate after any change you make.

Why: If you deleted old logs and the level dropped to 60 percent, the level is fixed and the rate is not. Re-check the growth rate a day later. If it is still 1.4 percent per day you have bought 28 days, not solved anything, and you should say so on the ticket.

Figure (svg): Bar showing a volume at eighty-five percent full with the remaining headroom labelled as roughly ten point seven days of runway at the current growth rate.

The dashboard shows a colour. The ticket should show a date.

31. Tonight, or This Week?

Sorting

Same five alerts, now with context. Urgency comes from blast radius and runway, never from the alert's own severity label.

Sort into buckets

Sort each situation by how fast it needs a human.

Now - wake people up
UPS on battery, generator has not started; whole row alerting on inlet temperature, still climbing; single disk failed, single-parity array, 18-hour rebuild
This shift - but not a panic
packet loss on one member of a healthy bonded pair
Schedule it
volume at 85 percent, growing 1.4 percent per day; single disk failed, double-parity array, rebuild running
now
Each of these has a countdown attached. Battery runtime is minutes. A climbing row temperature is minutes to throttling and then to shutdown. A single-parity array mid-rebuild has no redundancy at all for eighteen hours, and that is the window second failures live in.
soon
Service is unaffected and there is real margin, but something is genuinely broken and the margin is already reduced. Handle it during the shift rather than leaving it for the next person.
sched
Redundancy is intact or the runway is measured in days. Doing these at three in the morning adds risk without reducing any, and tired people make changes they regret.

Notice that the two disk failures split across two buckets. The alert text is nearly identical. The parity level is what decides.

32. A Realistic Bad Night

Error analysis

One engineer's response to the thermal alert. Three of these steps are defensible and one caused a second incident.

Annotate

On: \( \text{alert} \to \text{lower set point} \to \text{alert clears} \to \text{close ticket} \)

  • Responding to the alert quickly was right. Speed is not the problem here.
  • Lowering the set point as a stabilising action is defensible, and buying thermal margin while you investigate is a legitimate move.
  • The failure is closing the ticket. Nothing was diagnosed, so the cause is untouched and will return, and it will return on the hottest day rather than a convenient one.
  • The second incident: the lowered set point stayed in place for months, so every rack in the room ran colder than it needed to. That is a permanent efficiency cost paid to hide one rack's blocked exhaust.

Stabilising and resolving are two different tickets. Merging them is how a building slowly accumulates workarounds nobody remembers the reason for.

33. Alert to First Move

Matching

Not the fix. The first thing you actually do.

Match the pairs

  • a1. UPS transferred to battery
  • a2. rack inlet temperature over threshold
  • a3. packet loss on an uplink
  • a4. disk failed, array degraded
  • a5. volume at 85 percent
  • m1. confirm the generator started and took load
  • m2. check whether neighbouring racks are also alerting
  • m3. read error counters against discard counters
  • m4. find the parity level and the rebuild estimate
  • m5. divide the headroom by the growth rate

Why: Every one of these first moves answers question two of the triage loop, the blast radius, before touching anything. The generator tells you whether you have minutes or hours. The neighbours tell you rack or room. The counters tell you which team owns it. The parity level tells you whether you have margin. The rate turns a percentage into a date.

34. The Runbook Shape

Pattern

Every alert you ever write a procedure for should have these six fields. If you cannot fill one in, the procedure is not finished.

fieldwhat goes in it
what firedthe exact sensor and threshold, not the friendly alert name
blast radiushow far it reaches now, and how far it reaches in an hour
first checkthe one command or observation that splits the likely causes
stabilisewhat buys time without pretending to be a fix
escalate towhich team owns the cause, with the condition that triggers handover
done meansthe observable condition that closes it, and it is never the alert clearing

That last row is the one that separates a real runbook from a wiki page. An alert clearing is evidence about a sensor. It is not evidence about the building.

35. Which One Gets the Phone Call?

Check

Solve it on paper before you click.

Check your understanding

It is 3am. Four alerts arrive within a minute of each other. Which one do you handle first?

  • A. A volume has crossed 90 percent capacity with 20 days of runway
  • B. A UPS has transferred to battery and the generator has not started (correct)
  • C. A disk has failed in a double-parity array and a rebuild has begun
  • D. One member of a bonded uplink pair is reporting errors

Answer: B

Why: The UPS alert is the only one of the four with a hard countdown measured in minutes and a blast radius covering everything downstream of that UPS. If the generator does not take load before the batteries are exhausted, the outage is total and it is not recoverable by acting faster afterwards.

Why A tempts people
Twenty days of runway is a ticket for the morning. Percentages look alarming and runway is what actually matters.
Why C tempts people
Double parity means there is still a failure of margin left during the rebuild. It deserves attention in the shift, not ahead of a countdown.
Why D tempts people
A bonded pair carries traffic on the surviving member. Degraded capacity is not an outage, and this is the definition of a during-the-shift item.

36. Part 3 - Building a Disaster Recovery Plan

Section

Two numbers first, everything else follows from them

37. RTO and RPO Are Different Questions

Concept

Almost every argument about disaster recovery is really two arguments that got tangled together. Separating them makes the rest of the design nearly mechanical.

RTO, recovery time objective — How long the business can tolerate the service being down. It is a target for how fast you can bring it back.

RPO, recovery point objective — How much data the business can tolerate losing. It is a target for how far back the recovered state is allowed to be.

They are independent. A service can require coming back within fifteen minutes while tolerating the loss of a day's data, or the reverse. Which one is tighter tells you where to spend money.

NIST SP 800-34 Rev. 1, Contingency Planning Guide for Federal Information Systems Rev. 1 — The contingency planning guide these terms come from.

38. The Two Numbers on One Timeline

Picture it

The incident sits in the middle. One number looks backwards and one looks forwards.

Figure (svg): Timeline with the incident marked in the centre, the last good backup to its left and service restored to its right, with the interval before the incident labelled RPO and the interval after labelled RTO.

RPO is bounded by backup frequency. RTO is bounded by how fast you can actually restore and cut over.

Draw this every time somebody says the words disaster recovery. Most confused conversations resolve the moment both sides can point at which half they are talking about.

39. Reading a Real Requirement

Notation

A service owner hands you this line. Pull it apart before agreeing to it.

Annotate

On: \( \text{RTO} = 4 \text{ hours}, \qquad \text{RPO} = 15 \text{ minutes} \)

  • Four hours to restore is generous. It permits restoring from backup onto rebuilt infrastructure, and does not force you to keep a second site running hot.
  • Fifteen minutes of tolerable data loss is tight. Nightly backups cannot meet it, and neither can hourly ones.
  • The combination points at a specific design: continuous or near-continuous replication of the data, feeding infrastructure that is built on demand rather than kept warm.
  • That mismatch is normal and worth naming out loud. Data protection and infrastructure readiness are bought separately, and this requirement is buying a lot of one and little of the other.

The two numbers together select the architecture. Neither one alone tells you what to build.

40. Deriving the RPO You Actually Have

Worked example

The RPO you have is set by your backup schedule, whatever the policy document claims.

Find the real interval between recovery points.

Why: A backup taken at midnight every night means the recovery points are 24 hours apart, so a failure at 11pm loses almost a full day of work.

Take the worst case, not the average.

Why: The RPO you can promise is the largest gap between recovery points, because a disaster is not obliged to happen at a convenient moment.

\[ \text{RPO}_{\text{actual}} = \text{longest interval between recovery points} = 24 \text{ hours} \]

Compare that with the requirement and name the gap plainly.

Why: Against a 15-minute requirement, nightly backups are wrong by roughly two orders of magnitude. That is not a tuning problem, it is a different architecture: log shipping, continuous replication, or snapshots at the storage layer.

Check that restore time still fits the RTO after the change.

Why: More frequent recovery points can mean more work at restore time, for example replaying a long chain of logs. Improving RPO can quietly damage RTO, and the two must be checked together.

Verify: test a restore and time it with a stopwatch.

Why: The only honest measurement of both numbers is an actual restore of an actual backup, timed. Everything else is an estimate, and estimates of restore time are almost always optimistic because they omit finding the media, the approvals, and the person who knows the procedure being asleep.

Figure (svg): Timeline of evenly spaced backup recovery points with a failure marked just before the next one, and the interval since the last backup shaded to show everything that would be lost.

A failure is not obliged to wait for your backup window.

41. How Much Does the Gap Actually Cost?

Estimation

A finance system takes a full backup at midnight and nothing else. It fails at 11pm on a working day.

Predict first

Roughly how much of that day's work is lost?

  • None, the backup is only hours old
  • About half a day
  • Almost a full working day
  • It depends on the restore speed

Correct: Almost a full working day

Why: The last recovery point is midnight, so everything entered since then is gone: nearly twenty-three hours, which is essentially the entire working day. Restore speed is a completely separate question, and answering with restore speed is the exact RTO and RPO confusion this part of the deck exists to break. Fixing this needs more frequent recovery points, not a faster restore.

42. The 3-2-1 Rule, and Why It Has Three Parts

Concept

The oldest rule in backup is still the best summary, and each of its three numbers defends against a different failure.

  1. Three copies of the data. Two is one failure away from zero, and one of your copies is the live system you are trying to protect.
  2. On two different kinds of media or systems. A defect in a product, a firmware bug, or a controller fault does not politely restrict itself to one array.
  3. One copy off site. This is the one that survives fire, flood, and the whole building losing power for a week.

A modern fourth clause is worth adding: one copy that cannot be modified or deleted, even by an administrator. That is the clause that survives ransomware, because ransomware arrives holding valid credentials.

43. Which Backup Claim Survives?

Two truths and a lie

Four statements about a backup design. Only one is defensible.

Eliminate the wrong options

Rule out the three that fail against a realistic disaster, and keep the survivor.

  • b1. Replicating to a second array in the same room protects against disaster.
  • b2. A backup you have never restored is a backup.
  • b3. A copy that cannot be deleted, held off site and periodically test-restored, is a backup.
  • b4. Snapshots on the production array are sufficient, since they let you roll back.

Survives elimination: b3

Why: Only the third has all three properties that matter: separated in place so one event cannot reach both, protected from deletion so a credentialed attacker cannot remove it, and proven by restore so you know it works before you need it. The other three each fail on at least one, and each of them is a real design somebody has shipped.

44. Recovery Sites, From Cold to Always On

Concept

Once you know your two numbers, the site strategy is close to a lookup. What you are buying is recovery speed, and the price is running something you are not using.

strategywhat is runningtypical RTO
cold sitespace, power and network onlydays
warm sitehardware in place, data replicated, systems offhours
hot sitea full standby, running and currentminutes
active-activeboth sites serving live trafficnear zero

Active-active is the only one where the failover path is exercised continuously, because it is not a failover path at all — it is just Tuesday. Everything above it has a procedure that is only ever run in an emergency, unless you deliberately practise it.

45. What Each Site Strategy Costs

Trade off

Fill the missing cells. Every row buys recovery speed with money and complexity.

Comparison matrix

StrategyOngoing costFailover is exercisedMeets a 1-hour RTO
cold sitelowestneverno
warm sitemoderateonly in testsusually
hot sitehighonly in testsyes
active-activehighestcontinuously, by designyes

The third column is the one that decides whether the plan works on the day. A procedure exercised only in tests is only as good as the tests, and a procedure never exercised is fiction.

46. The Plan That Has Never Been Run

Trap

The trap

The disaster recovery plan is complete, approved, and stored on the shared drive. Every system is listed with its RTO and RPO, and the failover procedure is documented step by step.

Conclusion: the organisation is prepared for a disaster.

The fix

An untested plan is a document, not a capability. Every plan that has never been run contains at least one instruction that does not work, and you find out which one at the worst possible moment.

The classics: the runbook is stored on the system that is down. The person who wrote it left the company. The DNS change needs an approval from someone unreachable at 3am. The standby has drifted three versions behind. The backup restores, but nobody documented the encryption key location. None of these show up in review, and all of them show up in a test.

Test in layers: a tabletop walkthrough on paper, then a restore test of real data to real hardware, then a genuine failover of one non-critical service, then a full exercise. Each layer finds different defects, and the cheap layers find plenty.

Write the date of the last successful test on the front page of the plan. If that date is old, the plan's real status is unknown regardless of how good the document looks.

47. Build the Plan in the Right Order

Ranking

Teams routinely start in the middle and then discover they were protecting the wrong thing.

Put in order

  1. identify which services the business actually cannot lose
  2. agree an RTO and RPO for each of those services
  3. choose backup and replication that meets the RPO
  4. choose a recovery site strategy that meets the RTO
  5. write the runbook, including who decides to invoke it
  6. test it, and record the date and what broke

Why: The two numbers come from the business, not from engineering, and everything technical is downstream of them. Choosing a replication product before agreeing an RPO is how organisations end up with expensive protection on the wrong systems and none on the payroll database. The test is last but it is not optional, and its output is a date plus a list of defects.

48. The Decision Nobody Documents

Real world

Every plan documents how to fail over. Most forget the step before it.

Discussion prompt

Who is allowed to declare a disaster and invoke the plan, and why does that need deciding in advance?

Hint: Think about what failing over costs if the outage turns out to be twenty minutes long.

Answer:

Invoking a DR plan is expensive and often one-way. Failing over can mean accepting the RPO's worth of data loss deliberately, and failing back afterwards is frequently harder than failing over was.

So there is a real decision: wait, in the hope the primary recovers, or invoke, and take the known loss. Made under pressure with no named owner, that decision gets made late, which is the worst of both options.

A usable plan names a role rather than a person, gives that role a decision deadline such as thirty minutes without a credible recovery estimate, and states explicitly what invoking costs. That turns a judgement call at 3am into a rule agreed in daylight.

49. Which Design Meets the Requirement?

Check

Solve it on paper before you click.

Check your understanding

A service requires an RPO of 5 minutes and an RTO of 30 minutes. Which design meets both?

  • A. Hourly backups shipped off site, restored to a cold site
  • B. Continuous replication to a warm site brought up on demand
  • C. Continuous replication to a hot standby that is already running (correct)
  • D. Nightly backups to an immutable off-site vault

Answer: C

Why: The 5-minute RPO forces continuous or near-continuous replication, which rules out any backup interval measured in hours. The 30-minute RTO then rules out anything that has to be built or booted on demand, leaving a standby that is already running and current.

Why A tempts people
Hourly backups give an RPO of up to an hour, twelve times the requirement, and a cold site cannot be stood up in thirty minutes.
Why B tempts people
The replication satisfies the RPO, but bringing a warm site up on demand is typically an hours-long operation, so the RTO fails.
Why D tempts people
An immutable off-site vault is excellent protection and belongs in the design, but a nightly interval gives an RPO of up to 24 hours.

50. Part 4 - When the Data Center Is in the Cloud

Section

The same physics, someone else's hands

51. Regions and Availability Zones

Concept

Cloud providers expose the physical world through two words, and getting them the right way round is most of what beginners need.

Availability zone — One or more discrete data centres with independent power, cooling and networking, inside a region. Zones are close enough for low-latency replication and far enough apart to fail separately.

Region — A geographic area containing several availability zones, connected to each other with high-bandwidth low-latency links.

The design rule that follows: spread across zones to survive a facility failure, and across regions to survive a geographic event, accepting the latency and cost that the second one brings.

AWS, Regions and Availability Zones — global infrastructure Global infrastructure — The region and zone definitions.

52. One Region, Three Zones

Picture it

This picture replaces a surprising amount of cloud architecture arguing.

Figure (svg): Diagram of a single cloud region containing three availability zones, each labelled with its own independent power, cooling and network.

Independence within a region is the property you are paying for when you spread across zones.

Everything you learned in Part 1 still exists here. It is behind the line, run by somebody else, and you are still paying for it.

53. Zone or Region?

Definition probe

Sort each risk by the smallest thing you must spread across to survive it.

Sort into buckets

Sort each scenario by what protects against it.

Spread across zones
a UPS failure in one facility; a cooling failure in one building; a fibre cut into one facility
Needs separate regions
a hurricane affecting an entire metropolitan area; a regulation requiring data to stay in a country
zone
These are single-facility failures, which is precisely what zone independence is designed for. Each zone has its own power, cooling and network paths, so a failure in one does not propagate.
region
These reach across an entire geography. A weather event can affect every zone in a region at once, and a data residency rule is about geography by definition, so neither is solved by spreading within a region.

54. The Shared Responsibility Model

Concept

Moving to the cloud does not remove work. It moves a line, and everything on your side of the line is still entirely yours.

the provider handlesyou handle
physical security of the facilitywho has access to your accounts
power, cooling and hardware replacementwhich zones you deploy into
the hypervisor and the physical networkyour firewall rules and network design
durability of the storage servicewhether your data is actually backed up
keeping the platform availabledesigning your application to survive a zone loss

The right-hand column is where cloud outages actually hurt people. A provider can meet every commitment it made while your service is down, because your service was in one zone.

AWS, Shared Responsibility Model Shared responsibility — Where the line sits.

55. Whose Job Is It?

Elimination

A storage service reports eleven nines of durability. A developer deletes the wrong bucket.

Eliminate the wrong options

Which statement is correct about who is responsible and why?

  • c1. The provider will restore it, since durability is their commitment.
  • c2. Durability is about hardware, not about deletion, so this is yours to have prepared for.
  • c3. Replication across zones would have prevented it.
  • c4. Nothing could have prevented it in a cloud environment.

Survives elimination: c2

Why: Durability describes the probability of the provider losing your data through their own failures. It says nothing about you deleting it, and the deletion was a valid authenticated request. This is exactly the redundancy-is-not-backup lesson from Part 2, which is worth noticing: the physics changed, the principle did not.

56. On-Premises Against Cloud

Comparison

Fill the missing cells. Neither column is the right answer in general.

Comparison matrix

ConcernOn-premisesCloud
who replaces a failed diskyou, on sitethe provider, invisibly
cost shapelarge up front, then lowongoing, scales with use
capacity lead timeweeks to monthsminutes
surviving a zone lossyou build the second sitestill yours to design for

The last row is the one worth carrying away. Everything physical moved. The architectural responsibility did not move at all.

57. Why Cloud Changed the Economics of DR

Concept

This is the one place where the cloud genuinely changes the answer rather than relocating it, and it is worth being precise about why.

A traditional warm site means buying a second set of hardware and letting it sit mostly idle. The cost is continuous, and it is a hard sell precisely because the thing you are buying is the absence of an event.

With capacity available on demand, you can keep replicated data continuously, which is cheap, and only create the compute when you actually fail over, which is the expensive part. That splits the bill so the expensive half is only paid during a disaster.

The catch is the same as always: a failover path you have never executed is a guess. On-demand infrastructure has to be defined as code and tested on a schedule, or you have simply moved your untested plan into a different building.

58. Your Own Alert Card

Connect it up

Before next session, build the artefact you will actually use on shift.

Draw it

Take one alert from your new job's monitoring system, any one. Fill in the six runbook fields for it: what fired, blast radius, first check, stabilise, escalate to, and done means. Where you cannot fill a field, write the question you need to ask someone. Bring both the card and the questions to next session, and we will run five more scenarios against the same four-question loop.

The questions you cannot answer are more valuable than the fields you can. They are the map of what you do not yet know about your own building.

59. One Question Before You Close

Exit ticket

The idea that connects Part 2 and Part 3.

Predict first

Your storage is on a double-parity array, replicated in real time to a second array in another building. Is the data backed up?

  • Yes, it survives multiple drive failures and a building loss
  • Yes, because the replication is real time
  • No - it is redundant in two places, but not separated in time
  • Only if both arrays are the same model

Correct: No - it is redundant in two places, but not separated in time

Why: The design is genuinely good at hardware failure and even at losing a building, which is real value. But every copy is current, so a deletion, a corruption or an encryption event reaches all of them within seconds. A backup must let you go back to before the damage, and nothing in this design does. Separation in place and separation in time are two different properties and you need both.

60. What You Can Do Now

Recap

Six things, and the routine we will repeat every session.

alertthe first move
UPS on batteryconfirm the generator started and took load
rack inlet temperature highcheck whether the neighbours are alerting too
packet losserrors or discards - damage or congestion
disk failedparity level, then rebuild time
capacity thresholdheadroom divided by growth rate, in days

The thread through all of it: an alert is one sensor's opinion about a symptom. Redundancy survives hardware failure, backup survives everything else, and a plan nobody has tested is a document rather than a capability.

Next session: five new alerts across power, cooling, network, storage and compute, plus the runbook card you build this week.

Sources

  1. ASHRAE TC 9.9, Thermal Guidelines for Data Processing Environments (recommended inlet envelope)
  2. Uptime Institute, Tier Classification System (Tiers I-IV)
  3. NIST SP 800-34 Rev. 1, Contingency Planning Guide for Federal Information Systems
  4. AWS, Regions and Availability Zones — global infrastructure
  5. AWS, Shared Responsibility Model

Want this taught 1-on-1? Alexander tutors Data Center Operations — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108