The opening session of a data centre operations track for someone new to the field and newly in the job. It builds just enough plant anatomy to make an alert readable - the power chain, hot and cold aisle airflow, and the N, N+1 and 2N redundancy vocabulary - then installs a four-question triage routine and runs it against five real scenarios: a rack thermal alert, a UPS transferred to battery, packet loss on a core uplink, a failed disk mid-rebuild, and a capacity threshold. It closes with disaster recovery built from RTO and RPO outwards, the 3-2-1 rule and recovery site strategies, and what regions, availability zones and the shared responsibility model change once the data centre is somebody else's building. Two traps recur deliberately: stabilising is not resolving, and redundancy is not backup.
Subject: Data Center Operations · 60 slides · applied lesson
Open the interactive version of this deck · Homework for this lesson
Title
Session 1
The plant, five alerts you will actually see, disaster recovery, and what changes in the cloud
Objectives
You asked for a routine: five plausible alerts every session, what each one means, and how to resolve it. This deck sets that routine up, and adds the two topics you asked for on top of it.
Nothing here assumes prior data centre experience. It does assume you will ask when a term goes past you.
Section
Four systems, and only one of them is the computers
Warm-up
Have a guess before reading on. Being wrong here is the useful part.
Discussion prompt
Across the industry, what causes the largest share of serious data centre outages: hardware failure, software bugs, power problems, or human error during a planned change?
Hint: Think about when outages tend to happen, not what breaks.
Answer:
Power problems are consistently the largest single technical cause, and human error during planned maintenance is close behind and arguably ahead once you count it honestly.
That matters for how you read alerts. The alert that fires at two in the afternoon during a scheduled change is more suspicious than the one that fires at two in the morning, because somebody was touching something.
Concept
A data centre is four systems that each exist to keep the next one alive. Learning them in this order makes every alert easier to place.
An alert almost always names a component in one of these four. Your first move is to say which layer you are in, out loud.
Picture it
Trace it once now, and every power alert for the rest of your career has a location.
Figure (svg): Diagram of the power chain from utility feed through automatic transfer switch, UPS, power distribution unit, rack PDU and server supply, with a generator feeding the transfer switch and labels showing which element covers which outage duration.
The UPS is not there to run the building. It is there to keep it alive for the twenty or so seconds the generator needs to start and take load.
Ranking
From the street to the chip. Get this wrong and you will chase the wrong end of an outage.
Put in order
Why: Power arrives from the street, hits the transfer switch that decides between utility and generator, passes through the UPS which smooths and bridges, then is distributed to the floor, then to the rack, and finally into the two supplies inside each server. Anything upstream of a failure loses everything downstream of it.
Concept
The building is not trying to be cold. It is trying to make sure the air arriving at the front of each server is cool, and that the hot air leaving the back never gets a chance to come round again.
Cold aisle — The side the servers breathe in from. Perforated floor tiles or overhead supply put cool air here.
Hot aisle — The side the servers exhaust into. Nothing should be drawing intake air from here.
The industry guideline for the air arriving at the intake is roughly 18 to 27 degrees Celsius. That is a target at the rack face, not a thermostat reading for the room.
ASHRAE TC 9.9, Thermal Guidelines for Data Processing Environments (recommended inlet envelope) TC 9.9 — The recommended inlet envelope.
Picture it
One empty rack slot with no blanking panel is enough to break this.
Figure (svg): Diagram of a rack between a blue cold aisle and a red hot aisle, with arrows showing intake and exhaust, and a highlighted empty rack unit with no blanking panel letting hot air leak back to the front.
This is why blanking panels matter more than they look. They cost a few dollars and they are load-bearing for the whole cooling design.
Prediction
A rack has three empty units in the middle with no blanking panels fitted.
Predict first
Which server is most likely to alert on high inlet temperature?
Correct: The one directly above the gap
Why: Hot exhaust leaks forward through the gap and then rises, so it is drawn into the intake of whatever sits immediately above the opening. The bottom of the rack usually gets the coldest air. The last option is the beginner assumption this whole slide is aimed at: a data centre is not one temperature, it has a map of temperatures, and the rack face is the only place that counts.
Concept
Redundancy is described in terms of how much of the plant you can lose and still carry the load. The notation is compact and worth reading precisely.
| notation | what it means | what it survives |
|---|---|---|
| N | exactly enough capacity for the load | nothing, including maintenance |
| N+1 | one spare unit beyond what the load needs | one unit failing, or one being serviced |
| 2N | two complete and independent systems | an entire system failing, including its distribution |
| 2N+1 | two complete systems, each with a spare | a failure during maintenance on the other side |
The Uptime Institute tiers wrap this into four levels. The line worth memorising is between Tier III, which can be maintained without taking the load down, and Tier IV, which can take a fault without taking the load down.
Uptime Institute, Tier Classification System (Tiers I-IV) Tiers I-IV — The classification these terms come from.
Comparison
Complete the missing cells from the definitions you just read.
Comparison matrix
| Configuration | Survives one failure | Survives failure during maintenance |
|---|---|---|
| N | no | no |
| N+1 | yes | no |
| 2N | yes | no, one side is already down |
| 2N+1 | yes | yes |
The third column is the one people get wrong on interviews. N+1 with a unit already pulled for service is just N, and N has no margin at all.
Pattern
This is the routine. It does not change between a thermal alert and a storage alert, which is exactly why it is worth drilling.
Write the four answers down before you touch anything. The habit costs ninety seconds and it will one day stop you from making an outage worse.
Check
Solve it on paper before you click.
Check your understanding
A room is fed by a 2N power design. One of the two UPS systems has been taken offline for a scheduled battery replacement. What is the room's effective redundancy while that work is happening?
Answer: C
Why: With one of the two systems out, the remaining system is carrying the whole load on its own. That is exactly N: enough capacity, no spare. The room is up and serving normally, but a single further failure now takes the load down, which is why maintenance windows are the highest-risk hours in the building.
Section
The routine you asked for, run five times
Picture it
Keep this on screen while we work through the five. Every one of them goes through the same four boxes.
Figure (svg): Flowchart of four stacked triage steps: what the sensor measures, the blast radius, symptom versus cause, and stabilise now then fix later.
Notice that resolving the alert is the last box, not the first. Most bad incident handling is somebody starting at box four.
Concept
The page says: rack B14 inlet temperature 31 degrees Celsius, threshold 27, rising.
The useful instinct is that a thermal alert is a question about air, and air problems are almost always mechanical or geometric rather than electronic.
Worked example
Working the four questions in order on rack B14.
Check whether neighbouring racks are also alerting.
Why: This separates a rack-local problem from a plant problem in about ten seconds, and it changes who you need to wake up.
If only B14 is hot, walk the rack front and back.
Why: Rack-local causes are visible: a missing blanking panel, a cable bundle blocking the rear exhaust, a floor tile moved during a recent install, or a failed fan in one chassis.
If the row or room is hot, move up to the cooling plant.
Why: Now you are looking at air handler status, chilled water supply temperature, a failed compressor, or a damper that has not opened. This is a facilities escalation, not a server one.
Stabilise before diagnosing further.
Why: Restoring airflow buys you time. Perforated tiles can be added, a temporary blanking panel fitted, or non-critical load shed. Servers throttle before they fail, so a few degrees of margin is worth a lot.
Verify: confirm the temperature is falling, not just the alert clearing.
Why: An alert can clear because the sensor crossed back under the threshold for one poll. Watch the trend for ten minutes. If the number is flat just under the threshold rather than falling, the cause is still there and it will alert again on the next warm afternoon.
Figure (svg): Decision diagram splitting a thermal alert on whether neighbouring racks are also hot, into rack-local causes on one branch and cooling plant causes on the other.
Socratic
You checked the front of B14 and the blanking panels are all fitted correctly.
Discussion prompt
Why is the rear of the rack still worth walking before you escalate to facilities?
Hint: Where does the heat leave from, and what else lives back there?
Answer:
The rear is where the exhaust leaves, and it is also where every power and network cable is. A dense bundle of cabling grown over years of moves and changes can block a real fraction of the exhaust path.
If the hot air cannot leave the back, it builds up and finds its way to the front regardless of how good your blanking panels are. Cable management is a cooling control, not a tidiness preference.
Trap
Rack B14 is hot. The room set point is lowered by three degrees. The alert clears within twenty minutes. Ticket closed as resolved.
Reasoning: the temperature came down, and the temperature was the problem.
The alert cleared and nothing was fixed. The cause was a blocked exhaust path or a failed fan, and it is still there. You have paid for the fix with energy across the entire room, permanently, to mask one rack's defect.
Two things went wrong. The room was over-cooled to hide a local fault, which raises the power bill on every rack and degrades your efficiency number. And the real fault stayed in place, so it will resurface on the next hot day, at a worse hour, with less margin.
The right shape is: stabilise locally, raise a ticket for the cause, and put the set point back once the cause is fixed. Stabilising is allowed. Calling stabilisation a resolution is not.
Concept
The page says: UPS 2 on battery, load 42 percent, estimated runtime 11 minutes.
This is the one alert in this deck where you do not have time to be thorough first. Establish whether the generator is running before you do anything else.
Prediction
The battery is rated for 10 minutes at full load. The current load is 42 percent, and the alert reports 11 minutes remaining.
Predict first
If a colleague now powers up a large test cluster on that same UPS, taking the load to 85 percent, what happens to your remaining runtime?
Correct: It falls by much more than half
Why: Battery runtime does not fall in a straight line with load. Discharging faster gets less total energy out of the same battery, so doubling the load cuts the runtime by more than half. This is why the first instruction during a transfer to battery is to add nothing, and if possible to shed load. The timer in the alert assumes today's load and stops being true the moment anyone changes it.
Concept
The page says: 0.8 percent packet loss on the uplink from switch tor-14 to core-a, sustained for six minutes.
The single most useful discriminator here is whether the interface is reporting errors or reporting discards. They point at opposite causes.
Worked example
Two counters, two very different stories, and they sit next to each other in the same command output.
Read the interface error counters first.
Why: Cyclic redundancy check errors, alignment errors and symbol errors mean frames are arriving damaged. Damage points at the physical layer: a dirty or failing optic, a bent or over-bent fibre, a bad patch, or a marginal cable.
Now read the discard and queue-drop counters.
Why: Discards mean frames arrived intact and were thrown away because there was nowhere to put them. That is congestion, not damage, and no amount of replacing optics will help it.
Check whether the loss correlates with a traffic peak.
Why: Congestion loss tracks the traffic graph. Physical loss does not care what time it is, which is a fast way to tell them apart on a dashboard.
Stabilise by moving traffic off the suspect path.
Why: If the uplink is one member of a bonded pair, taking the member out of service administratively is safer than leaving a half-broken link carrying production traffic. A cleanly down link reroutes. A flapping link poisons everything.
Verify: re-check the counters after the traffic has moved.
Why: Clear the counters, wait, and read again. If errors stop accumulating once traffic is off the link, the link is the problem. If loss follows the traffic to the other uplink, the problem was never the link and you have just learned something far more important.
Figure (svg): Two panels comparing interface error counters, meaning damaged frames from a physical cause, against discard counters, meaning intact frames dropped through congestion.
Discrimination
Same symptom, three different teams. Put each cause where it belongs.
Sort into buckets
Sort each cause of packet loss by the layer you would investigate it in.
Concept
The page says: array sn-03, physical disk 7 failed, array degraded, rebuild started to hot spare.
The thing to understand is the rebuild window. From the moment a drive fails until the rebuild completes, your redundancy is spent, and a rebuild reads every remaining drive from end to end, at full speed, which is the most stressful thing you will ever ask of them.
Worked example
Why the clock on a degraded array matters more than the alert's calm wording suggests.
Establish which parity level the array uses.
Why: A single-parity array tolerates exactly one failed drive. While it is degraded it tolerates none. A double-parity array still has one failure of margin left while rebuilding, which is a completely different risk posture.
Estimate how long the rebuild will take.
Why: Rebuild time scales with drive capacity and with how busy the array is with real work. Large drives can take many hours, and a heavily loaded array will stretch that out considerably because production traffic is competing with the rebuild.
Recognise why rebuilds are when second failures appear.
Why: A rebuild reads every sector of every surviving drive. Drives in the same array were usually bought together, installed together, and have run the same hours. A latent defect that never surfaced under normal reads surfaces here.
Replace the failed drive rather than only watching the rebuild.
Why: The hot spare is now part of the array, so the array has no spare until you physically replace the dead drive and let a new spare take its place. Closing the ticket when the rebuild finishes leaves the array without a spare indefinitely.
Verify: confirm both that the array is optimal and that a spare exists.
Why: Two separate conditions. An array can report optimal, meaning full redundancy restored, while having zero hot spares available. Check for both before you consider this closed, because the next failure is the one the spare was for.
Figure (svg): Timeline from disk failure to rebuild completion, showing that a single-parity array has zero redundancy across the whole window while a double-parity array retains one failure of margin.
Trap
The storage is on a double-parity array with a hot spare, replicated to a second array in the same room. It can survive multiple drive failures. Therefore the data is backed up.
Reasoning: the data exists in more than one place, and that is what a backup is.
Redundancy protects against hardware failure. It does not protect against anything else, and almost nothing that destroys data in practice is a hardware failure.
A deletion replicates. Ransomware encryption replicates. A bad database migration replicates. A fire in that room takes both arrays. In each case the redundancy worked perfectly and faithfully preserved the damage.
A backup is defined by being separated in time, so you can go back to before the damage, and separated in place, so a single event cannot reach both. Redundancy gives you neither of those. This distinction is the whole reason the next part of this deck exists.
Concept
The page says: volume vol-prod-02 at 85 percent capacity, growth 1.4 percent per day.
Capacity alerts are the only ones in this deck you are allowed to schedule rather than act on, and that is exactly why they are the ones that eventually cause outages.
Worked example
A percentage is not actionable. A date is. Converting one to the other is the entire job here.
Compute the remaining runway from the level and the growth rate.
Why: At 85 percent full and growing 1.4 percent per day, there are 15 percentage points left, which is a little under 11 days. That is now a calendar problem rather than a dashboard colour.
\[ \frac{100 - 85}{1.4} \approx 10.7 \text{ days} \]
Ask whether the growth is expected.
Why: Steady growth matching the business is a capacity request. A rate that changed recently is an incident wearing a capacity alert's clothing, and the usual culprits are a log level left on debug after troubleshooting, a retention policy that stopped running, or a job writing temporary files it never cleans up.
Decide urgency from the runway, not from the percentage.
Why: Eighty-five percent with eleven days is a ticket for this week. Ninety percent with a rate that just tripled is tonight's problem. The same colour on the dashboard, two completely different responses.
Verify: re-measure the rate after any change you make.
Why: If you deleted old logs and the level dropped to 60 percent, the level is fixed and the rate is not. Re-check the growth rate a day later. If it is still 1.4 percent per day you have bought 28 days, not solved anything, and you should say so on the ticket.
Figure (svg): Bar showing a volume at eighty-five percent full with the remaining headroom labelled as roughly ten point seven days of runway at the current growth rate.
Sorting
Same five alerts, now with context. Urgency comes from blast radius and runway, never from the alert's own severity label.
Sort into buckets
Sort each situation by how fast it needs a human.
Notice that the two disk failures split across two buckets. The alert text is nearly identical. The parity level is what decides.
Error analysis
One engineer's response to the thermal alert. Three of these steps are defensible and one caused a second incident.
Annotate
On: \( \text{alert} \to \text{lower set point} \to \text{alert clears} \to \text{close ticket} \)
Stabilising and resolving are two different tickets. Merging them is how a building slowly accumulates workarounds nobody remembers the reason for.
Matching
Not the fix. The first thing you actually do.
Match the pairs
Why: Every one of these first moves answers question two of the triage loop, the blast radius, before touching anything. The generator tells you whether you have minutes or hours. The neighbours tell you rack or room. The counters tell you which team owns it. The parity level tells you whether you have margin. The rate turns a percentage into a date.
Pattern
Every alert you ever write a procedure for should have these six fields. If you cannot fill one in, the procedure is not finished.
| field | what goes in it |
|---|---|
| what fired | the exact sensor and threshold, not the friendly alert name |
| blast radius | how far it reaches now, and how far it reaches in an hour |
| first check | the one command or observation that splits the likely causes |
| stabilise | what buys time without pretending to be a fix |
| escalate to | which team owns the cause, with the condition that triggers handover |
| done means | the observable condition that closes it, and it is never the alert clearing |
That last row is the one that separates a real runbook from a wiki page. An alert clearing is evidence about a sensor. It is not evidence about the building.
Check
Solve it on paper before you click.
Check your understanding
It is 3am. Four alerts arrive within a minute of each other. Which one do you handle first?
Answer: B
Why: The UPS alert is the only one of the four with a hard countdown measured in minutes and a blast radius covering everything downstream of that UPS. If the generator does not take load before the batteries are exhausted, the outage is total and it is not recoverable by acting faster afterwards.
Section
Two numbers first, everything else follows from them
Concept
Almost every argument about disaster recovery is really two arguments that got tangled together. Separating them makes the rest of the design nearly mechanical.
RTO, recovery time objective — How long the business can tolerate the service being down. It is a target for how fast you can bring it back.
RPO, recovery point objective — How much data the business can tolerate losing. It is a target for how far back the recovered state is allowed to be.
They are independent. A service can require coming back within fifteen minutes while tolerating the loss of a day's data, or the reverse. Which one is tighter tells you where to spend money.
NIST SP 800-34 Rev. 1, Contingency Planning Guide for Federal Information Systems Rev. 1 — The contingency planning guide these terms come from.
Picture it
The incident sits in the middle. One number looks backwards and one looks forwards.
Figure (svg): Timeline with the incident marked in the centre, the last good backup to its left and service restored to its right, with the interval before the incident labelled RPO and the interval after labelled RTO.
Draw this every time somebody says the words disaster recovery. Most confused conversations resolve the moment both sides can point at which half they are talking about.
Notation
A service owner hands you this line. Pull it apart before agreeing to it.
Annotate
On: \( \text{RTO} = 4 \text{ hours}, \qquad \text{RPO} = 15 \text{ minutes} \)
The two numbers together select the architecture. Neither one alone tells you what to build.
Worked example
The RPO you have is set by your backup schedule, whatever the policy document claims.
Find the real interval between recovery points.
Why: A backup taken at midnight every night means the recovery points are 24 hours apart, so a failure at 11pm loses almost a full day of work.
Take the worst case, not the average.
Why: The RPO you can promise is the largest gap between recovery points, because a disaster is not obliged to happen at a convenient moment.
\[ \text{RPO}_{\text{actual}} = \text{longest interval between recovery points} = 24 \text{ hours} \]
Compare that with the requirement and name the gap plainly.
Why: Against a 15-minute requirement, nightly backups are wrong by roughly two orders of magnitude. That is not a tuning problem, it is a different architecture: log shipping, continuous replication, or snapshots at the storage layer.
Check that restore time still fits the RTO after the change.
Why: More frequent recovery points can mean more work at restore time, for example replaying a long chain of logs. Improving RPO can quietly damage RTO, and the two must be checked together.
Verify: test a restore and time it with a stopwatch.
Why: The only honest measurement of both numbers is an actual restore of an actual backup, timed. Everything else is an estimate, and estimates of restore time are almost always optimistic because they omit finding the media, the approvals, and the person who knows the procedure being asleep.
Figure (svg): Timeline of evenly spaced backup recovery points with a failure marked just before the next one, and the interval since the last backup shaded to show everything that would be lost.
Estimation
A finance system takes a full backup at midnight and nothing else. It fails at 11pm on a working day.
Predict first
Roughly how much of that day's work is lost?
Correct: Almost a full working day
Why: The last recovery point is midnight, so everything entered since then is gone: nearly twenty-three hours, which is essentially the entire working day. Restore speed is a completely separate question, and answering with restore speed is the exact RTO and RPO confusion this part of the deck exists to break. Fixing this needs more frequent recovery points, not a faster restore.
Concept
The oldest rule in backup is still the best summary, and each of its three numbers defends against a different failure.
A modern fourth clause is worth adding: one copy that cannot be modified or deleted, even by an administrator. That is the clause that survives ransomware, because ransomware arrives holding valid credentials.
Two truths and a lie
Four statements about a backup design. Only one is defensible.
Eliminate the wrong options
Rule out the three that fail against a realistic disaster, and keep the survivor.
Survives elimination: b3
Why: Only the third has all three properties that matter: separated in place so one event cannot reach both, protected from deletion so a credentialed attacker cannot remove it, and proven by restore so you know it works before you need it. The other three each fail on at least one, and each of them is a real design somebody has shipped.
Concept
Once you know your two numbers, the site strategy is close to a lookup. What you are buying is recovery speed, and the price is running something you are not using.
| strategy | what is running | typical RTO |
|---|---|---|
| cold site | space, power and network only | days |
| warm site | hardware in place, data replicated, systems off | hours |
| hot site | a full standby, running and current | minutes |
| active-active | both sites serving live traffic | near zero |
Active-active is the only one where the failover path is exercised continuously, because it is not a failover path at all — it is just Tuesday. Everything above it has a procedure that is only ever run in an emergency, unless you deliberately practise it.
Trade off
Fill the missing cells. Every row buys recovery speed with money and complexity.
Comparison matrix
| Strategy | Ongoing cost | Failover is exercised | Meets a 1-hour RTO |
|---|---|---|---|
| cold site | lowest | never | no |
| warm site | moderate | only in tests | usually |
| hot site | high | only in tests | yes |
| active-active | highest | continuously, by design | yes |
The third column is the one that decides whether the plan works on the day. A procedure exercised only in tests is only as good as the tests, and a procedure never exercised is fiction.
Trap
The disaster recovery plan is complete, approved, and stored on the shared drive. Every system is listed with its RTO and RPO, and the failover procedure is documented step by step.
Conclusion: the organisation is prepared for a disaster.
An untested plan is a document, not a capability. Every plan that has never been run contains at least one instruction that does not work, and you find out which one at the worst possible moment.
The classics: the runbook is stored on the system that is down. The person who wrote it left the company. The DNS change needs an approval from someone unreachable at 3am. The standby has drifted three versions behind. The backup restores, but nobody documented the encryption key location. None of these show up in review, and all of them show up in a test.
Test in layers: a tabletop walkthrough on paper, then a restore test of real data to real hardware, then a genuine failover of one non-critical service, then a full exercise. Each layer finds different defects, and the cheap layers find plenty.
Write the date of the last successful test on the front page of the plan. If that date is old, the plan's real status is unknown regardless of how good the document looks.
Ranking
Teams routinely start in the middle and then discover they were protecting the wrong thing.
Put in order
Why: The two numbers come from the business, not from engineering, and everything technical is downstream of them. Choosing a replication product before agreeing an RPO is how organisations end up with expensive protection on the wrong systems and none on the payroll database. The test is last but it is not optional, and its output is a date plus a list of defects.
Real world
Every plan documents how to fail over. Most forget the step before it.
Discussion prompt
Who is allowed to declare a disaster and invoke the plan, and why does that need deciding in advance?
Hint: Think about what failing over costs if the outage turns out to be twenty minutes long.
Answer:
Invoking a DR plan is expensive and often one-way. Failing over can mean accepting the RPO's worth of data loss deliberately, and failing back afterwards is frequently harder than failing over was.
So there is a real decision: wait, in the hope the primary recovers, or invoke, and take the known loss. Made under pressure with no named owner, that decision gets made late, which is the worst of both options.
A usable plan names a role rather than a person, gives that role a decision deadline such as thirty minutes without a credible recovery estimate, and states explicitly what invoking costs. That turns a judgement call at 3am into a rule agreed in daylight.
Check
Solve it on paper before you click.
Check your understanding
A service requires an RPO of 5 minutes and an RTO of 30 minutes. Which design meets both?
Answer: C
Why: The 5-minute RPO forces continuous or near-continuous replication, which rules out any backup interval measured in hours. The 30-minute RTO then rules out anything that has to be built or booted on demand, leaving a standby that is already running and current.
Section
The same physics, someone else's hands
Concept
Cloud providers expose the physical world through two words, and getting them the right way round is most of what beginners need.
Availability zone — One or more discrete data centres with independent power, cooling and networking, inside a region. Zones are close enough for low-latency replication and far enough apart to fail separately.
Region — A geographic area containing several availability zones, connected to each other with high-bandwidth low-latency links.
The design rule that follows: spread across zones to survive a facility failure, and across regions to survive a geographic event, accepting the latency and cost that the second one brings.
AWS, Regions and Availability Zones — global infrastructure Global infrastructure — The region and zone definitions.
Picture it
This picture replaces a surprising amount of cloud architecture arguing.
Figure (svg): Diagram of a single cloud region containing three availability zones, each labelled with its own independent power, cooling and network.
Everything you learned in Part 1 still exists here. It is behind the line, run by somebody else, and you are still paying for it.
Definition probe
Sort each risk by the smallest thing you must spread across to survive it.
Sort into buckets
Sort each scenario by what protects against it.
Concept
Moving to the cloud does not remove work. It moves a line, and everything on your side of the line is still entirely yours.
| the provider handles | you handle |
|---|---|
| physical security of the facility | who has access to your accounts |
| power, cooling and hardware replacement | which zones you deploy into |
| the hypervisor and the physical network | your firewall rules and network design |
| durability of the storage service | whether your data is actually backed up |
| keeping the platform available | designing your application to survive a zone loss |
The right-hand column is where cloud outages actually hurt people. A provider can meet every commitment it made while your service is down, because your service was in one zone.
AWS, Shared Responsibility Model Shared responsibility — Where the line sits.
Elimination
A storage service reports eleven nines of durability. A developer deletes the wrong bucket.
Eliminate the wrong options
Which statement is correct about who is responsible and why?
Survives elimination: c2
Why: Durability describes the probability of the provider losing your data through their own failures. It says nothing about you deleting it, and the deletion was a valid authenticated request. This is exactly the redundancy-is-not-backup lesson from Part 2, which is worth noticing: the physics changed, the principle did not.
Comparison
Fill the missing cells. Neither column is the right answer in general.
Comparison matrix
| Concern | On-premises | Cloud |
|---|---|---|
| who replaces a failed disk | you, on site | the provider, invisibly |
| cost shape | large up front, then low | ongoing, scales with use |
| capacity lead time | weeks to months | minutes |
| surviving a zone loss | you build the second site | still yours to design for |
The last row is the one worth carrying away. Everything physical moved. The architectural responsibility did not move at all.
Concept
This is the one place where the cloud genuinely changes the answer rather than relocating it, and it is worth being precise about why.
A traditional warm site means buying a second set of hardware and letting it sit mostly idle. The cost is continuous, and it is a hard sell precisely because the thing you are buying is the absence of an event.
With capacity available on demand, you can keep replicated data continuously, which is cheap, and only create the compute when you actually fail over, which is the expensive part. That splits the bill so the expensive half is only paid during a disaster.
The catch is the same as always: a failover path you have never executed is a guess. On-demand infrastructure has to be defined as code and tested on a schedule, or you have simply moved your untested plan into a different building.
Connect it up
Before next session, build the artefact you will actually use on shift.
Draw it
Take one alert from your new job's monitoring system, any one. Fill in the six runbook fields for it: what fired, blast radius, first check, stabilise, escalate to, and done means. Where you cannot fill a field, write the question you need to ask someone. Bring both the card and the questions to next session, and we will run five more scenarios against the same four-question loop.
The questions you cannot answer are more valuable than the fields you can. They are the map of what you do not yet know about your own building.
Exit ticket
The idea that connects Part 2 and Part 3.
Predict first
Your storage is on a double-parity array, replicated in real time to a second array in another building. Is the data backed up?
Correct: No - it is redundant in two places, but not separated in time
Why: The design is genuinely good at hardware failure and even at losing a building, which is real value. But every copy is current, so a deletion, a corruption or an encryption event reaches all of them within seconds. A backup must let you go back to before the damage, and nothing in this design does. Separation in place and separation in time are two different properties and you need both.
Recap
Six things, and the routine we will repeat every session.
| alert | the first move |
|---|---|
| UPS on battery | confirm the generator started and took load |
| rack inlet temperature high | check whether the neighbours are alerting too |
| packet loss | errors or discards - damage or congestion |
| disk failed | parity level, then rebuild time |
| capacity threshold | headroom divided by growth rate, in days |
The thread through all of it: an alert is one sensor's opinion about a symptom. Redundancy survives hardware failure, backup survives everything else, and a plan nobody has tested is a document rather than a capability.
Next session: five new alerts across power, cooling, network, storage and compute, plus the runbook card you build this week.
Want this taught 1-on-1? Alexander tutors Data Center Operations — $55/session, free consultation.