Round two of the recurring alert routine the student asked for, deliberately drawn from a different family than session 1. Five scenarios, one per system, and the organising idea is that not one of them is an outage: a rack PDU branch circuit at 87 percent, which turns out to be a dual-feed arithmetic problem where sixty percent on each side is the exact condition for losing the rack; an air handler failure in a room that is still perfectly cool, which is a redundancy loss measured against a ride-through of minutes; a domain controller four minutes ahead of its reference, fifty seconds from the point where Kerberos begins refusing tickets across 38 servers; a storage controller cache battery failure that silently switches write-back to write-through and multiplies write latency twentyfold with every volume still optimal; and correctable memory errors rising a thousandfold on one module, which is a warning that lets you schedule an outage rather than receive one. It closes on alert storms, finding the parent rather than the loudest child, and the two ways monitoring systems train people to stop reading them. Every scenario is triaged with the same four questions and written on the same six-field card as session 1, and the memory replacement is planned with the reboot discipline from session 2.
Subject: Data Center Operations · 65 slides · applied lesson
Open the interactive version of this deck · Homework for this lesson
Title
Session 3
Power, cooling, time, storage cache and memory - and why not one of tonight's five is an outage
Objectives
This is the routine you asked for, back on schedule: five plausible alerts, what each one could mean, and how you resolve it. Round one was thermal, UPS on battery, packet loss, a failed disk and a capacity threshold. Round two deliberately avoids all five of those.
Same four triage questions as session 1 throughout. The alerts change, the routine does not, and that is the entire idea.
Section
Four questions, six fields, three classes of alert
Warm-up
From memory, without scrolling back to session 1.
Discussion prompt
There were four questions we agreed to ask of every alert, in order, before touching anything. What were they, and why is the order the way it is?
Hint: The first one is about the sensor. The last one is about two different tickets.
Answer:
One: what is the sensor actually measuring? Not what the alert is named, what the number is.
Two: what is the blast radius right now? One server, one rack, one row, one room, the whole site.
Three: is this the symptom or the cause underneath it? You will usually be handed a symptom.
Four: what stabilises it now, and what fixes it properly later? Those are two different tickets and they stay two different tickets.
The order is deliberate. Reading the sensor stops you inventing a story. Blast radius decides how fast you are allowed to move. Symptom against cause stops you fixing the wrong thing. And splitting stabilise from resolve is what stops a three in the morning workaround from quietly becoming the permanent design.
Concept
Everything tonight gets written up the same way, and this is the card you started building for your own building last time.
| field | what goes in it |
|---|---|
| what fired | the sensor, the number, the threshold it crossed, and when |
| blast radius | what is affected now, and what would be affected if it got worse |
| first check | the one thing you look at before anything else, and why that one |
| stabilise | what buys time, done knowing it is not the fix |
| escalate to | who, by name or by role, and at what point |
| done means | the condition under which you are allowed to close it |
The last field is the one people leave blank, and it is the one that stops an alert being closed because it went quiet rather than because it was fixed.
Concept
Here is the one new idea before we start, and it reframes all five of tonight's scenarios.
Four of tonight's five alerts are in the second and third categories. None of them will wake a customer up tonight. Two of them are the reason somebody gets woken up next week.
Uptime Institute, annual outage analysis - causes and cost of data centre outages Outage analysis — How often a serious outage turns out to have had a quiet precursor.
Picture it
The pattern is consistent enough to be worth stating as a rule.
Figure (svg): Diagram of three classes of alert: something is broken and everyone responds, something will break and most people respond, and the margin is gone where nothing is wrong and almost nobody responds.
Notice the shape of it: the further left you catch a problem, the cheaper it is to fix and the less anyone cares. Being the person who works the bottom row is a very fast way to become useful in a new job.
Definition probe
Here are the five alerts you are about to work, plus one from last session for calibration.
Sort into buckets
Sort each alert into its class.
Fill the middle
Three fields missing from a card for last session's UPS on battery alert. Fill them the way you would in a real ticket.
Fill in the blanks
What fired: UPS-2 transferred to battery at 02:11, runtime remaining 7 minutes. Blast radius: everything on that UPS now, and one step worse the whole room. First check: whether the generator started and took load. Stabilise: nothing to do at the rack, the design is already doing it. Escalate to: facilities, immediately. Done means: the UPS is back on utility or generator and recharging.
Why: The first check is the one that decides whether you have seven minutes or several hours, which is the only thing that matters at that moment. Escalation is immediate because nothing you can do at a keyboard affects a generator. And done is defined by the UPS being back off battery and recharging, not by the alert going quiet, because an alert that clears when the battery is flat has cleared for the worst possible reason.
Pattern
The same four questions from session 1, with one addition for the class of alert we just named.
Every one of the five that follows is worked in exactly this order, so by the fifth you should be able to run ahead of me.
Section
A branch circuit at 87 percent, and nothing is wrong
Picture it
This is the whole of what you get. Run the first question on it before reading on.
Figure (svg): An alert card reading warning, rack power distribution unit R14 branch circuit 3, load 20.9 amps of a 24 amp rating at 87 percent, threshold 80 percent, received at 21 minutes past two in the afternoon.
The sensor is measuring current on one branch circuit inside one rack power strip. It is not measuring the room, the row, or the UPS. Blast radius, so far, is whatever is plugged into that strip.
Concept
A rack power strip is fed by one or more circuits, each protected by a breaker with a rating in amps. The breaker exists to open before the wiring gets hot.
Continuous load — A load expected to run at its maximum for three hours or more. A rack of servers is the definition of a continuous load.
Electrical code requires a continuous load to be sized at no more than eighty percent of the breaker rating. That is not a suggestion or a safety margin somebody invented, it is the number the breaker was designed and tested against.
So a 24 amp circuit has 19.2 usable amps for a data centre load. At 20.9 amps this circuit is not near its limit. It is already past it.
National Electrical Code, Article 210.20 - continuous loads and the eighty percent rule Article 210.20 — Continuous loads and the eighty percent rule.
Prediction
Suppose somebody adds one more server and the circuit reaches its full 24 amp rating.
Predict first
What does the breaker do?
Correct: Holds for a while and then opens, sooner the further over it goes
Why: A thermal breaker has an inverse time curve: a small overload is tolerated for a long time, a large one for a very short time, and a genuine short circuit trips it instantly. That is exactly what makes an overloaded circuit dangerous to operate on. It does not fail when you make the mistake, it fails minutes or hours later, usually when the load rises slightly for some unrelated reason, and by then nobody connects the trip to the change that caused it.
Concept
The alert told you the partner circuit is at 20.4 amps. That is the line that turns this from a housekeeping item into something you escalate today.
Equipment in a rack like this is dual corded: every server has two power supplies, one plugged into the A strip and one into the B strip, fed from two independent paths. That is what buys you the ability to lose a feed and keep serving.
The arithmetic that makes it work is simple and constantly forgotten. Each feed must be able to carry the entire load on its own. So each feed may run at no more than half of its usable capacity in normal operation.
Usable capacity here is 19.2 amps, so the safe steady-state figure per feed is about 9.6 amps. This rack is running 20.9 on one side and 20.4 on the other.
Picture it
Two feeds, both comfortable. Watch what happens when one of them goes away.
Figure (svg): Diagram of a rack fed by two independent power feeds each carrying sixty percent of capacity, showing that losing one feed asks the survivor for one hundred and twenty percent, which opens its breaker and takes down both sides.
This failure mode has a particular cruelty to it. The redundancy does not merely fail to help, it actively converts a single feed loss, which should have been invisible, into a total rack outage.
Socratic
Sit with the shape of this before we move on, because the same pattern turns up in storage, in networking and in clusters.
Discussion prompt
A single-corded rack on one feed at 87 percent survives a maintenance day fine. A dual-corded rack with both feeds at 87 percent loses everything if one feed goes. Why does adding redundancy make the outcome worse here?
Hint: What did the second feed get used for, in practice, rather than in the design?
Answer:
Because the redundant capacity got spent on load rather than kept as margin. The design intent was two feeds each capable of the whole rack. The reality is two feeds each carrying their own half, and no spare anywhere.
So the second feed is not redundancy at all. It is capacity, and the rack has quietly become a single point of failure with twice the equipment and none of the protection anyone thinks it has.
The general pattern, and it is worth naming: redundancy consumed is redundancy absent. It is true of power feeds, of cooling units, of cluster members and of network uplinks, and the only defence is measuring the margin rather than the load.
Which is why the useful question about any redundant pair is never how loaded is it. It is what happens when one of them is gone, and can the survivor take it.
Check
Work it out properly before you look at the options.
Check your understanding
A dual-corded rack is fed by two 30 amp circuits. Applying the continuous load rule and the requirement that either feed alone can carry the rack, what is the highest steady-state current you should see on each feed?
Answer: B
Why: Two rules stack. The continuous load rule takes 30 amps down to 24 amps of usable capacity. The single-feed survival requirement then halves that, because either circuit must be able to carry the entire rack when its partner is gone. Twelve amps per feed in normal running, which is forty percent of the number printed on the breaker, and that gap between the printed number and the usable one is exactly why racks get quietly overloaded.
Worked example
Four questions, then the card. This is the shape every one of tonight's five will take.
What is the sensor measuring? Current on branch circuit 3 of the A strip in rack R14, sustained over 18 minutes, not a momentary spike.
Why: Sustained matters. A brief peak during a batch job is a different problem from a load that has genuinely moved up and stayed there.
What is the blast radius now, and one step worse? Now: nothing, everything is serving. One step worse: the breaker opens, that strip dies, and because the partner cannot carry the rack, the whole rack follows.
Why: The second half of the question is what promotes this from a ticket to an escalation. Nothing is affected now and an entire rack is one event away.
Symptom or cause? Symptom. Something was added to this rack, or something in it started drawing more.
Why: Circuits do not drift upward on their own. A change caused this, and finding the change is usually a five-minute conversation.
Stabilise: stop anything else being added to R14 today, and check the recent change record for what went in.
Why: The cheapest possible intervention is preventing the situation getting worse while you work out what to do properly.
Resolve: move load to a rack with headroom, or have the facility engineer provide additional circuits. Then reset the alert threshold to reflect the dual-feed figure rather than the eighty percent one.
Why: The threshold itself was wrong. Alerting at eighty percent of a circuit that must never exceed fifty percent means the alert only fires long after the design was violated.
Figure (svg): Diagram of a rack fed by two independent power feeds each at sixty percent, showing that losing one feed asks the survivor for one hundred and twenty percent and opens its breaker.
Verify: done means both feeds below the dual-feed figure, the threshold corrected, and the change that caused it identified.
Why: Three conditions, and the alert going quiet is not one of them. A load that drops because a batch job finished has resolved nothing at all.
Trap
A colleague looks at the same rack and says it is fine. Sixty percent on each side, well under the eighty percent rule, no action needed. On the numbers as stated, that reasoning is completely reasonable.
It is also how the rack ends up on the incident report. The eighty percent rule and the dual-feed rule are different rules, answering different questions, and passing one says nothing about the other.
Two questions, always both, for any redundant pair.
A dual-corded rack that passes the first and fails the second is the most common power finding in the industry, and it is invisible until the day somebody does maintenance on a feed.
Same shape as last session's redundancy is not backup. Two copies of something are only redundancy if either copy can do the whole job on its own.
Section
A unit has failed, and the room is perfectly cool
Picture it
Read all four lines, including the reassuring one.
Figure (svg): An alert card reading warning, computer room air handler CRAH-03 fan fault, unit offline, with return air temperature normal and the remaining three units in the room now running at increased fan speed.
This is the margin-is-gone class in its purest form. Every measurement a customer would care about is exactly where it was an hour ago.
Prediction
Three in the morning. Room temperature is normal and no rack is alerting.
Predict first
How urgent is this?
Correct: Urgent, because the room has no cooling redundancy left and the remaining units are near maximum
Why: The room is not about to overheat, which is why the third option is wrong and why this alert gets deferred so often. What has happened is that a room designed as N plus one is now running at N, with the survivors at ninety-one percent of their capability. The next unit to fail, and units in a room tend to fail for related reasons, takes the room below what it needs. The urgency is not in the current temperature, it is in the complete absence of margin combined with how little time a dense hall gives you once cooling actually goes short.
Concept
Losing cooling in a data hall is not like losing air conditioning in an office, and the difference is a matter of physics rather than of degree.
An office has furniture, walls, floors and people, all of which absorb heat, and a modest heat source. It warms up over hours.
A data hall has almost no thermal mass and an enormous, constant heat source. A rack drawing fifteen kilowatts is a fifteen kilowatt heater that never switches off. Every watt that goes in as electricity comes out as heat, immediately.
So the time between cooling stopping and equipment reaching temperatures that force a shutdown is commonly measured in single-digit minutes for a dense hall, not in hours. That is the number that makes this alert urgent while everything still reads normal.
ASHRAE TC 9.9, Thermal Guidelines for Data Processing Environments TC 9.9 — Inlet temperature envelopes, and what happens beyond them.
Picture it
Two rooms, both lose cooling at the same moment.
Figure (svg): Line graph comparing temperature rise after a cooling failure in an office, which climbs gently, against a dense data hall, which climbs steeply within minutes because it has high heat output and almost no thermal mass.
This is also why the room-level temperature reading is such a poor early warning. By the time the room average moves, the hot spots have been in trouble for a while.
Matching
Match each term to what it actually refers to. These come up constantly and are easy to nod along to.
Match the pairs
Why: The C and the H are the whole distinction between the first two: one makes its own cold, the other is handed cold water and only moves air. Delta T is the one worth internalising, because it is a measure of work done rather than of state. A delta T that collapses while fans run flat out usually means air is bypassing the racks and going straight back to the unit, which is an airflow problem no amount of extra cooling capacity will fix.
Worked example
Same four questions, same card. Notice how much of this is about the second half of question two.
What is the sensor measuring? A fan fault on one air handler, plus fan speed on the three survivors. Not temperature.
Why: The temperature reading is context, not the alert. Confusing the two is what leads people to close this because the room is cool.
Blast radius now: none. One step worse: a second unit fault takes the hall below the cooling it needs, and a dense hall gives you minutes.
Why: This is the whole justification for waking somebody up. The current state is fine and the next state is an evacuation of the hall.
Symptom or cause? The fan fault is a cause at the unit level. The ninety-one percent fan speed is a symptom of the fault.
Why: Worth separating, because the fix belongs to whoever owns the unit, and the risk belongs to you until they finish.
Stabilise: raise it to facilities as a redundancy loss, not as a temperature problem, and get a time estimate for the repair.
Why: The wording changes the response you get. Reported as room is cool, unit offline it waits until morning. Reported as the hall is at N with survivors at ninety-one percent it does not.
While waiting, reduce what you can: defer batch work that would add load in that hall, and confirm the hot-aisle containment and blanking panels have not been left open by recent work.
Why: Both are free and both buy real margin. An open containment door is a surprisingly common contributor and takes a minute to check.
Figure (svg): Line graph comparing temperature rise after cooling failure in an office against a dense data hall, showing the hall climbing steeply within minutes.
Verify: done means the failed unit is back in service and the survivors have returned to their normal fan speed.
Why: Redundancy restored, not temperature acceptable. Temperature was acceptable throughout, and closing on it would mean closing on a condition that never changed.
Socratic
The reason a redundancy loss is more urgent than it looks has a name, and it applies well beyond cooling.
Discussion prompt
The three surviving air handlers are identical, were installed on the same day, have run the same hours, and are now working harder than they ever have. Why does that combination make a second failure more likely than a first failure was?
Hint: Think about what the three survivors have in common, and about what changed for them when CRAH-03 went offline.
Answer:
Because failures in a set like this are not independent, and almost all the arithmetic people do about redundancy quietly assumes that they are.
The four units share an installation date, a duty cycle, a maintenance history, a batch of components and a supply of chilled water. Whatever wore out the bearing in the failed unit has been wearing out the same bearing in the other three, at the same rate, for the same number of hours.
And the load moved. The survivors went from sixty-two percent to ninety-one percent, so from this moment they are working harder than they ever have, which accelerates exactly the wear that just produced a failure.
The same reasoning applies to disks bought in one batch, to power supplies in one chassis, to cluster members built from one image and to controller batteries fitted on the same day. When you hear that two failures at once would be very unlikely, the question to ask is whether the two things have anything in common. Usually they have almost everything in common.
Two truths and a lie
Two are sound. One is the reasoning that keeps this alert in a queue until morning.
Eliminate the wrong options
Which statement survives?
Survives elimination: t1
Why: The survivors running at ninety-one percent is the practical instrument here, and it is worth naming even though it is not one of the options. A fan speed reading is the cheapest available measure of remaining cooling margin, and it moves long before any temperature does. Get into the habit of reading fan speed and pump speed as margin gauges rather than as trivia, and of arguing politely with both of the statements you just eliminated.
Section
Four minutes wrong, and it takes down authentication
Picture it
The least impressive-looking alert in this deck, and the one with the widest blast radius.
Figure (svg): An alert card reading warning, domain controller DC-02 clock offset is 247 seconds ahead of the reference time source, with the time service reporting no synchronisation for eleven hours.
There is no temperature here, no current, no failed part. This is a pure configuration and reachability alert, and it will express itself as a dozen unrelated-looking faults.
Anomaly
Everything is running. Nobody has complained. The clock is four minutes fast on one server.
Predict first
Which of these is the most immediate consequence?
Correct: Kerberos authentication starts refusing tickets once the skew passes five minutes
Why: Kerberos puts a timestamp inside every ticket specifically to stop an attacker replaying a captured one, and the default tolerance for clock difference is five minutes. This host is at four minutes and seven seconds and still drifting, so it is roughly fifty seconds from the point at which logins to anything that trusts it begin to fail. The failure looks nothing like a clock problem from the user's side: it looks like a wrong password, or an access denied on a file share, and it will be reported as thirty separate tickets.
Concept
Time is one of those dependencies that is invisible until it is wrong, and then it is wrong everywhere at once.
Microsoft Learn, Kerberos maximum tolerance for computer clock synchronization Clock skew — The five minute default tolerance, and what happens outside it.
Discrimination
Six reports from a morning where one domain controller was four minutes fast. Sort them.
Sort into buckets
Sort each report by whether a clock problem plausibly explains it.
Concept
Time distribution is a tree. Something authoritative sits at the top and everything else takes time from a level above it.
Stratum — How many hops a clock is from a reference. A radio or satellite reference is stratum zero, the server attached to it is stratum one, a server synchronising to that is stratum two, and so on.
In a Windows environment, one domain controller holds the role that makes it the authority for the whole domain, and every other machine chains up to it. That is why one wrong clock on the right host becomes everyone's problem.
Internet Engineering Task Force, RFC 5905 - Network Time Protocol version 4 RFC 5905 — Strata, selection among sources, and how a client decides who to believe.
Check
One domain controller, four minutes ahead, no synchronisation in eleven hours.
Check your understanding
Which first check most efficiently separates the likely causes?
Answer: B
Why: The alert already told you the source is unreachable, and one reachability test either confirms that or proves the alert is misleading you. Confirmed unreachable points immediately at a firewall or routing change, which is a five-minute conversation with the network team and explains why this started eleven hours ago. That single test splits the problem in half before you have changed anything on the server.
Concept
Most monitoring systems check that the time service is running. That check would have been green for all eleven hours of this incident, because the service was running perfectly and had nothing to synchronise against.
The measurement worth alerting on is the offset itself, in seconds, against a reference the monitoring system trusts independently.
That third point is the subtle one and it is the shape of this whole session. A machine whose clock is currently right and whose source is gone is not fine, it is drifting, and the only difference between it and a domain-wide authentication outage is how long you leave it.
Worked example
The card again, and note how different stabilise and resolve are here.
What is the sensor measuring? Offset against a reference, and time since the last successful synchronisation. Both matter.
Why: The offset tells you how close to the five minute cliff you are. The eleven hours tells you when it started, which is what you take to the network team.
Blast radius: 38 member servers chain from this host, so their clocks are wrong too, and anything authenticating against them is at risk. One step worse: the offset passes five minutes and authentication fails domain-wide.
Why: This is the widest blast radius of tonight's five, from the least dramatic-looking alert. That contrast is worth remembering.
Symptom or cause? The offset is a symptom. The unreachable source is closer to the cause, and the change that made it unreachable is the cause.
Why: Three layers, and fixing only the first guarantees you will be back here in another eleven hours.
Stabilise: correct the clock deliberately, and slew it rather than stepping it if the tooling allows.
Why: Slewing speeds the clock up or slows it down slightly until it converges, so time never repeats and never jumps backwards. Databases and log analysis both handle that far better than a jump.
Resolve: restore reachability to the reference, then confirm the host synchronises on its own and holds.
Why: Done is not the clock is right now. Done is the mechanism that keeps it right is working again, demonstrated by a successful synchronisation after you stop touching it.
Figure (svg): An alert card showing a domain controller clock offset of 247 seconds with no successful synchronisation for eleven hours and the reference source unreachable.
Verify: done means a successful synchronisation has happened without intervention, and the member servers have followed.
Why: The last clause is the one people forget. Fixing the authority does not instantly fix the 38 machines beneath it, and some of them will need a nudge.
Trap
A server is four minutes ahead. The obvious fix is to set it back four minutes. It takes one command and the number is immediately correct.
On a database server this can be genuinely destructive. Four minutes of timestamps now repeat, so two different events can carry the same time, and anything that assumed time only moves forward is now working from bad data. Replication ordering, transaction logs and any application logic that compares timestamps are all exposed.
Backups and log analysis inherit the same confusion, and unlike the database it will not be noticed for weeks.
Correct a clock the way the time service does it by default, which is to slew rather than step.
Same principle as session 2: the fast fix and the safe fix are different moves, and knowing which one you are making is the job.
Section
No disk failed, no data lost, and everything got slow
Picture it
Two lines. The second one is the entire story and it is written as if it were housekeeping.
Figure (svg): An alert card reading warning, storage array controller one cache battery has failed, write cache policy changed from write back to write through, with all volumes online and optimal.
This alert routinely gets triaged as low priority because of the third line, and the applications behind it are already suffering by the time anyone reads the fourth.
Hypothesis
No disk failed. No volume is degraded. Nothing is rebuilding. Write latency is twenty times worse than this morning.
Predict first
Which explanation fits every fact?
Correct: The array stopped acknowledging writes from its own memory and now waits for the disks
Why: The clue is that the policy change and the latency change are in the same alert. With a working battery the controller can safely say done as soon as a write is in its own memory, because the battery guarantees that memory survives a power loss long enough to be written out. With the battery dead that promise cannot be kept, so the controller stops making it and waits for the physical disks instead. That is the difference between microseconds and milliseconds, and the array is behaving correctly and conservatively throughout. Nothing is broken; the array has chosen safety over speed and told you so in a line that reads like a footnote.
Concept
This is worth understanding properly, because the same trade-off appears in databases, in filesystems and in every caching layer you will ever meet.
Write-back — The controller puts the write in its own memory, immediately tells the server it is done, and writes it to disk later. Fast, and it depends on that memory surviving a power failure.
Write-through — The controller waits until the data is actually on the disks before telling the server it is done. Slower, and it needs no promises about memory at all.
The battery or supercapacitor on the controller is what makes the first mode honest. It keeps the cache alive long enough for its contents to be written out after an unexpected power loss.
So when the battery fails, the controller cannot honour the promise it was making, and it stops making it. The performance collapse is not a fault. It is the array refusing to lie to you.
Picture it
Same write, two policies, and one component deciding which one is available.
Figure (svg): Diagram comparing write-back, where the controller acknowledges a write as soon as it is in controller cache, against write-through, where the controller waits for the disks, with a note that the cache battery is what makes write-back safe.
Note the direction of the trade. You have not lost data and you cannot lose data. You have lost speed, and the reason you lost it is that the protection against losing data went away.
Elimination
The array offers a setting to force write-back on regardless of battery state. Four options, one right answer.
Eliminate the wrong options
Which action is correct?
Survives elimination: g2
Why: The correct answer is unglamorous: accept the slower and safer mode, tell the people affected before they open tickets, and get the part replaced quickly. The judgement being tested is whether you will trade a data-integrity guarantee for latency under pressure. Someone will ask you to, and the answer is no unless a named person with the authority to accept that risk says so in writing.
Prediction
The array has two controllers. Controller 1's cache battery has just failed after four years in service. Controller 2's battery was fitted on the same day.
Predict first
What should you assume about controller 2?
Correct: It is probably close behind, so order two modules and check its charge and predicted end of life now
Why: Exactly the correlated failure argument from the cooling alert, in a different system. The two batteries are the same chemistry, the same batch and the same age, and they have spent four years in the same thermal environment doing the same duty. A battery failure at four years is a statement about the design life of that part in this room, not about one unlucky module. Ordering the pair and checking the survivor's charge and predicted replacement date costs almost nothing, and it converts a probable second incident into a second line on the same change record.
Worked example
Shorter than the others, because the diagnosis was handed to you in the alert. The work is in the response.
What is the sensor measuring? A battery charge level, a policy change, and a latency figure. Three different things in one alert.
Why: The policy change is the fact, the latency is the consequence, and the battery is the cause. Recognising that ordering immediately is what makes this alert quick to work.
Blast radius: every volume on that controller, so every application writing to it. One step worse: nothing gets worse, which is unusual and worth saying.
Why: This alert is stable. It will not deteriorate on its own, and that is precisely what makes it safe to fix properly rather than quickly.
Symptom or cause? The latency is the symptom. The policy change is the mechanism. The dead battery is the cause, and it is already named.
Why: One of the few alerts that hands you all three layers at once. Most do not.
Stabilise: notify the application owners with a number, not an adjective. Writes have gone from under half a millisecond to nine, and here is why.
Why: Getting ahead of it converts thirty confused tickets into one informed conversation, and it buys goodwill for the outage window you will need for the replacement.
Resolve: replace the battery module. Check whether it is hot-swappable on this model, and whether the array will restore write-back automatically or needs to be told.
Why: Both details are model-specific and both are worth knowing before the engineer arrives rather than while they are standing there.
Figure (svg): Diagram comparing write-back where the controller acknowledges from its own cache against write-through where it waits for the disks, with the cache battery making the fast path safe.
Verify: done means the battery reports charged, the policy has returned to write-back, and the latency figure is back where it was.
Why: Three conditions again, and the third is the one to insist on. A battery that reports healthy while the policy is still write-through means somebody has to go and change it back, and until then nothing has improved for the applications.
Section
Memory errors that were all corrected
Picture it
Everything in this alert was successfully handled by the hardware. That is what makes it easy to dismiss.
Figure (svg): An alert card reading warning, host ESX-14 correctable memory errors on DIMM A3, 1840 errors in the last 24 hours rising from 12 the previous week, with no uncorrectable errors and the host running normally.
Two numbers are doing all the work: the rate went up by a factor of about a thousand, and it is confined to one module out of many.
Concept
Server memory carries extra bits so the controller can check every read against the value that was written.
So a correctable error is not a failure. It is the hardware doing exactly what it was designed to do, and telling you about it.
The value of the alert is entirely in the rate and the location. Occasional isolated errors happen, including from cosmic rays, and mean nothing. A rate climbing on one module means that module is degrading, and field studies consistently find that a module producing correctable errors is far more likely to produce an uncorrectable one later than a module producing none.
Schroeder, Pinheiro and Weber, DRAM Errors in the Wild - a Large-Scale Field Study, SIGMETRICS 2009 DRAM errors in the wild — Correctable error rates as a predictor of uncorrectable failure.
Picture it
The alert is not about the left-hand box. It is about how close you are to the right-hand one.
Figure (svg): Diagram of correctable single-bit memory errors which are corrected in flight and counted, against uncorrectable double-bit errors which halt the machine, with an arrow showing that a rising correctable rate on one module predicts the second.
This ties straight back to last session. A predicted failure lets you plan a reboot: drained, in a window, with the console open. An uncorrectable error gives you a halted host and whatever was running on it.
Discrimination
Six memory reports from six different hosts. Sort them by whether they justify action.
Sort into buckets
Sort each report by whether you would act on it.
Real world
This is where tonight's session and last session's join up.
Discussion prompt
You are going to replace DIMM A3, which means the host must be powered off. It is a hypervisor host with 22 virtual machines on it. Write the plan.
Hint: Almost every step of this was in session 2, and only one step is about memory at all.
Answer:
Confirm the finding first: check whether the vendor tooling can retire the affected pages as an interim measure, and confirm the module location physically so the engineer replaces the right one.
Raise a normal change with the part, the window and the justification, which is the error rate trend rather than an incident.
Drain the host: put it into maintenance mode so the 22 guests migrate off, and confirm the migration completed rather than assuming it did.
Confirm out-of-band access to the host before it is powered down, since you will be watching it come back through the console.
Power it down properly, replace the module, power on, and watch the memory self test through the console. This is the blind zone, and it is where a badly seated module announces itself.
Verify up the ladder: the host boots, the memory total is what it should be, the error counters are cleared, the host rejoins the cluster, and only then take it out of maintenance mode.
The point worth noticing: exactly one step of that plan is about memory. Everything else is the reboot discipline from last session, which is why we did that one first.
Check
One module, 1840 corrected errors in a day, none uncorrectable, host running normally.
Check your understanding
What is the single most accurate statement about this alert?
Answer: C
Why: Every error was corrected, so nothing was lost and no data was affected, which is what makes the first option wrong. But the rate rose from about two a day to nearly two thousand a day on one specific module, and that trend is a well-documented predictor of an uncorrectable error, which halts the host. The whole value of the alert is the choice it offers: replace the module in a window of your choosing, or have the host halt at a time of its choosing.
Section
Find the parent, ignore the children
Concept
Sooner or later you will open the console to two hundred alerts that all arrived within ninety seconds. It is genuinely alarming and it is almost always one event.
Monitoring systems see symptoms, and a single failure produces symptoms everywhere. A switch stack restart makes every host behind it unreachable, which makes every service check on those hosts fail, which makes every cluster containing one of those hosts report itself degraded.
That last habit is worth practising deliberately, because the alternative is working two hundred tickets in parallel, which nobody has ever finished.
Picture it
The shape is always the same, and once you see it you cannot unsee it.
Figure (svg): Diagram of a single root cause, a switch stack reboot, producing four groups of symptom alerts including unreachable hosts, services down, degraded clusters and failed checks.
This is question three from the routine, at scale. Two hundred symptoms, one cause, and the entire skill is refusing to work the symptoms.
Ranking
Five alerts arrive within one minute of each other. Put them in the order they must have happened.
Put in order
Why: Causation runs downhill from the infrastructure to the customer, and the alerts arrive in roughly that order because each layer notices the one below it. The switch goes, so hosts become unreachable, so clusters notice a missing member, so health checks fail, so customers see errors. Work the top of the chain and everything below it clears on its own. Work the bottom of the chain and you will be busy for hours while the actual fault sits there untouched.
Concept
The failure mode of a monitoring system is not missing an alert. It is producing so many that people stop reading them, and it gets there by two routes.
Flapping — A measurement sitting on a threshold, crossing back and forth, generating an alert and a recovery every few minutes. The information content is near zero and the noise is constant.
Thresholds nobody agreed — A check set at a default that does not match how the system actually behaves, so it fires daily and is ignored daily.
Both are fixed the same way, and neither is fixed by muting the alert. A flapping check needs hysteresis, meaning it fires at one level and only clears at a distinctly lower one, or a requirement to be over the line for a sustained period. A wrong threshold needs the correct number, worked out from how the system behaves.
And when you are doing planned work, suppress deliberately. Putting a host into maintenance mode in the monitoring system before you start is what stops your own change generating a storm that trains everyone to ignore storms.
Trade off
Fill the blanks. This is a real decision you will be asked to make, and there is no universally right answer.
Comparison matrix
| Consideration | Threshold set tight | Threshold set loose |
|---|---|---|
| warning time before the limit | long | short, sometimes none |
| false alarms | frequent | rare |
| what people do after a month | stop reading it | trust it, and act when it fires |
| best use | a slow-moving trend with a real deadline | a fast-moving condition where any crossing is real |
The resolution is usually two thresholds rather than a compromise on one: a quiet one that opens a ticket for the trend, and a loud one that wakes somebody. Tonight's PDU alert is the example, and it needs both.
Trap
You do the right thing before a planned change: you put the host into maintenance mode in the monitoring system so your reboot does not generate a storm and train everyone to ignore storms.
The change goes well. The host comes back. You verify it up the ladder, close the ticket, and go home.
Nine days later the same host runs out of disk space and fills a database volume. No alert fires, because it is still in maintenance mode, and nobody notices until the application stops accepting writes.
Suppression is a change like any other, and it needs an end as clearly defined as its beginning.
This is the same lesson as done means from the card. A change is not finished when the work is finished, it is finished when everything you altered to do the work has been put back.
Explain it to yourself
One idea connects all five of tonight's alerts, and it is worth you putting it into your own words.
Discussion prompt
What did the PDU circuit, the failed air handler, the drifting clock, the dead cache battery and the memory errors have in common?
Hint: Ask what a customer would have noticed at the moment each alert fired.
Answer:
Not one of them was an outage. At the moment each alert fired, every service was working and no user would have noticed anything at all.
Four of the five were the system telling you that a margin had gone or was going: no spare feed, no spare cooling unit, no spare time before the skew limit, no spare corrections before a module fails.
The fifth, the cache battery, is the odd one out and it is instructive: nothing was lost and nothing was at risk, but the machine had quietly changed its behaviour to stay safe, and only the latency number gave it away.
So the practical skill is reading an alert where nothing is wrong. Ask what would happen if this got one step worse, and how much time you would have when it did. Those two answers are what turn an item in a queue into an action tonight.
Connect it up
Same artefact as last time, and by now you should have one or two from your own building.
Draw it
Take tonight's five alerts and write the six-field card for each one as it would apply in your data centre, not in mine: what fired, blast radius now and one step worse, first check, stabilise, escalate to by name or role, and done means. Then find the equivalent alert in your own monitoring system for at least two of the five, and write down the actual threshold it uses. Where a threshold does not match what we worked out tonight, that gap is your first real contribution.
Bring the cards and the thresholds. The gaps between what your system alerts on and what the arithmetic says it should alert on are the most useful thing you can walk into a team meeting with in your first month.
Exit ticket
The habit I most want to survive tonight.
Predict first
An alert fires and everything is working normally. What is the second question you ask?
Correct: What would happen if this got one step worse, and how much time would I have?
Why: It is the second question because the first is still what is the sensor actually measuring. But this is the one that separates the alerts that can genuinely wait from the ones that look identical and cannot. A room that is cool with no cooling redundancy and a room that is cool with full redundancy read the same on every temperature gauge in the building, and they are a very long way apart. Asking about the next step, and about how much time that step leaves you, is what makes the difference visible, and it is why four of tonight's five alerts were worth working immediately.
Recap
Eight things, and the routine holds for round three.
| alert | the first move |
|---|---|
| branch circuit over threshold | check the partner feed, then do the halving arithmetic |
| cooling unit failed, room cool | read the survivors' fan speed as a margin gauge |
| clock drift | test reachability of the reference on UDP 123 |
| cache battery failed | tell the application owners a number before they open tickets |
| correctable memory errors | compare the rate to last week, and check it is one module |
| fifty alerts at once | sort by time and read the first three |
The thread through all five: an alert about a margin is an alert about a future, and the only way to read one is to ask what happens next and how long you would have. Everything else tonight was arithmetic.
Next session, round three: five more, and I will hold back the category labels so you can place them yourself. Bring the cards, and bring the thresholds that do not match.
Want this taught 1-on-1? Alexander tutors Data Center Operations — $55/session, free consultation.