Session 3 - Five Alerts, Round Two: Power, Cooling, Time, Cache and Memory

Round two of the recurring alert routine the student asked for, deliberately drawn from a different family than session 1. Five scenarios, one per system, and the organising idea is that not one of them is an outage: a rack PDU branch circuit at 87 percent, which turns out to be a dual-feed arithmetic problem where sixty percent on each side is the exact condition for losing the rack; an air handler failure in a room that is still perfectly cool, which is a redundancy loss measured against a ride-through of minutes; a domain controller four minutes ahead of its reference, fifty seconds from the point where Kerberos begins refusing tickets across 38 servers; a storage controller cache battery failure that silently switches write-back to write-through and multiplies write latency twentyfold with every volume still optimal; and correctable memory errors rising a thousandfold on one module, which is a warning that lets you schedule an outage rather than receive one. It closes on alert storms, finding the parent rather than the loudest child, and the two ways monitoring systems train people to stop reading them. Every scenario is triaged with the same four questions and written on the same six-field card as session 1, and the memory replacement is planned with the reboot discipline from session 2.

Subject: Data Center Operations · 65 slides · applied lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Five Alerts, Round Two

Title

Session 3

Power, cooling, time, storage cache and memory - and why not one of tonight's five is an outage

2. The Routine, Resumed

Objectives

This is the routine you asked for, back on schedule: five plausible alerts, what each one could mean, and how you resolve it. Round one was thermal, UPS on battery, packet loss, a failed disk and a capacity threshold. Round two deliberately avoids all five of those.

Same four triage questions as session 1 throughout. The alerts change, the routine does not, and that is the entire idea.

3. Part 0 - The Routine, Refreshed

Section

Four questions, six fields, three classes of alert

4. Before We Open Anything

Warm-up

From memory, without scrolling back to session 1.

Discussion prompt

There were four questions we agreed to ask of every alert, in order, before touching anything. What were they, and why is the order the way it is?

Hint: The first one is about the sensor. The last one is about two different tickets.

Answer:

One: what is the sensor actually measuring? Not what the alert is named, what the number is.

Two: what is the blast radius right now? One server, one rack, one row, one room, the whole site.

Three: is this the symptom or the cause underneath it? You will usually be handed a symptom.

Four: what stabilises it now, and what fixes it properly later? Those are two different tickets and they stay two different tickets.

The order is deliberate. Reading the sensor stops you inventing a story. Blast radius decides how fast you are allowed to move. Symptom against cause stops you fixing the wrong thing. And splitting stabilise from resolve is what stops a three in the morning workaround from quietly becoming the permanent design.

5. The Six-Field Card

Concept

Everything tonight gets written up the same way, and this is the card you started building for your own building last time.

fieldwhat goes in it
what firedthe sensor, the number, the threshold it crossed, and when
blast radiuswhat is affected now, and what would be affected if it got worse
first checkthe one thing you look at before anything else, and why that one
stabilisewhat buys time, done knowing it is not the fix
escalate towho, by name or by role, and at what point
done meansthe condition under which you are allowed to close it

The last field is the one people leave blank, and it is the one that stops an alert being closed because it went quiet rather than because it was fixed.

6. Three Classes of Alert

Concept

Here is the one new idea before we start, and it reframes all five of tonight's scenarios.

  1. Something is broken. A service is down now. Everyone responds to these, and they are the easiest to work because the feedback is immediate.
  2. Something will break. A trend is heading for a limit. Most people respond, eventually, usually when the trend gets steeper.
  3. The margin is gone. Nothing is wrong, everything is serving normally, and there is no spare left. Almost nobody responds to these, and they are the ones that turn into the first category at four in the morning.

Four of tonight's five alerts are in the second and third categories. None of them will wake a customer up tonight. Two of them are the reason somebody gets woken up next week.

Uptime Institute, annual outage analysis - causes and cost of data centre outages Outage analysis — How often a serious outage turns out to have had a quiet precursor.

7. Which Class Gets Ignored

Picture it

The pattern is consistent enough to be worth stating as a rule.

Figure (svg): Diagram of three classes of alert: something is broken and everyone responds, something will break and most people respond, and the margin is gone where nothing is wrong and almost nobody responds.

Attention falls off exactly as the opportunity to act cheaply rises.

Notice the shape of it: the further left you catch a problem, the cheaper it is to fix and the less anyone cares. Being the person who works the bottom row is a very fast way to become useful in a new job.

8. Sort Tonight's Five

Definition probe

Here are the five alerts you are about to work, plus one from last session for calibration.

Sort into buckets

Sort each alert into its class.

Something is broken
web service returning errors to customers now
Something will break
rack PDU branch circuit at 87 percent of rating; domain controller clock is 4 minutes ahead; controller cache battery failed, write cache disabled; correctable memory errors rising on one module
The margin is gone
air handler CRAH-03 has failed, room temperature normal
broken
A user is experiencing failure right now. Everything else waits, you stabilise first, and you find the cause afterwards. Only the last item qualifies tonight.
will
A measurement is heading somewhere it must not go, or a protection has already quietly changed the machine's behaviour. Nothing has failed, and something will if nobody acts. These have a deadline attached even though nothing is on fire.
margin
The cooling failure is the pure case: the room is exactly as cool as it was an hour ago, every service is fine, and the only thing that changed is that there is now no spare. Nothing you can measure at the rack has moved, which is precisely why this class gets ignored.

9. Complete the Card

Fill the middle

Three fields missing from a card for last session's UPS on battery alert. Fill them the way you would in a real ticket.

Fill in the blanks

What fired: UPS-2 transferred to battery at 02:11, runtime remaining 7 minutes. Blast radius: everything on that UPS now, and one step worse the whole room. First check: whether the generator started and took load. Stabilise: nothing to do at the rack, the design is already doing it. Escalate to: facilities, immediately. Done means: the UPS is back on utility or generator and recharging.

Why: The first check is the one that decides whether you have seven minutes or several hours, which is the only thing that matters at that moment. Escalation is immediate because nothing you can do at a keyboard affects a generator. And done is defined by the UPS being back off battery and recharging, not by the alert going quiet, because an alert that clears when the battery is flat has cleared for the worst possible reason.

10. The Loop, Every Alert, Every Time

Pattern

The same four questions from session 1, with one addition for the class of alert we just named.

  1. What is the sensor actually measuring? Read the number, not the title.
  2. What is the blast radius right now, and what would it be if this got one step worse? The second half is the new part, and it is what makes a margin alert legible.
  3. Symptom or cause? Assume symptom until you have a reason not to.
  4. What stabilises it, and what resolves it? Two tickets, always.
  5. Then write the card. Six fields, and refuse to leave done means blank.

Every one of the five that follows is worked in exactly this order, so by the fifth you should be able to run ahead of me.

11. Alert Six - Power

Section

A branch circuit at 87 percent, and nothing is wrong

12. The Alert, As It Arrives

Picture it

This is the whole of what you get. Run the first question on it before reading on.

Figure (svg): An alert card reading warning, rack power distribution unit R14 branch circuit 3, load 20.9 amps of a 24 amp rating at 87 percent, threshold 80 percent, received at 21 minutes past two in the afternoon.

Three numbers and a threshold. The third line is the one that changes the answer completely.

The sensor is measuring current on one branch circuit inside one rack power strip. It is not measuring the room, the row, or the UPS. Blast radius, so far, is whatever is plugged into that strip.

13. What a Branch Circuit Is, and Why 80 Percent

Concept

A rack power strip is fed by one or more circuits, each protected by a breaker with a rating in amps. The breaker exists to open before the wiring gets hot.

Continuous load — A load expected to run at its maximum for three hours or more. A rack of servers is the definition of a continuous load.

Electrical code requires a continuous load to be sized at no more than eighty percent of the breaker rating. That is not a suggestion or a safety margin somebody invented, it is the number the breaker was designed and tested against.

So a 24 amp circuit has 19.2 usable amps for a data centre load. At 20.9 amps this circuit is not near its limit. It is already past it.

National Electrical Code, Article 210.20 - continuous loads and the eighty percent rule Article 210.20 — Continuous loads and the eighty percent rule.

14. What Happens at 100 Percent?

Prediction

Suppose somebody adds one more server and the circuit reaches its full 24 amp rating.

Predict first

What does the breaker do?

  • Opens immediately at 24 amps
  • Holds indefinitely, since that is its rating
  • Holds for a while and then opens, sooner the further over it goes
  • Nothing, breakers only respond to short circuits

Correct: Holds for a while and then opens, sooner the further over it goes

Why: A thermal breaker has an inverse time curve: a small overload is tolerated for a long time, a large one for a very short time, and a genuine short circuit trips it instantly. That is exactly what makes an overloaded circuit dangerous to operate on. It does not fail when you make the mistake, it fails minutes or hours later, usually when the load rises slightly for some unrelated reason, and by then nobody connects the trip to the change that caused it.

15. Now the Third Line of the Alert

Concept

The alert told you the partner circuit is at 20.4 amps. That is the line that turns this from a housekeeping item into something you escalate today.

Equipment in a rack like this is dual corded: every server has two power supplies, one plugged into the A strip and one into the B strip, fed from two independent paths. That is what buys you the ability to lose a feed and keep serving.

The arithmetic that makes it work is simple and constantly forgotten. Each feed must be able to carry the entire load on its own. So each feed may run at no more than half of its usable capacity in normal operation.

Usable capacity here is 19.2 amps, so the safe steady-state figure per feed is about 9.6 amps. This rack is running 20.9 on one side and 20.4 on the other.

16. The Dual Feed Arithmetic

Picture it

Two feeds, both comfortable. Watch what happens when one of them goes away.

Figure (svg): Diagram of a rack fed by two independent power feeds each carrying sixty percent of capacity, showing that losing one feed asks the survivor for one hundred and twenty percent, which opens its breaker and takes down both sides.

Sixty percent on each side looks safe and is the exact condition for losing everything.

This failure mode has a particular cruelty to it. The redundancy does not merely fail to help, it actively converts a single feed loss, which should have been invisible, into a total rack outage.

17. Why Does Redundancy Make It Worse?

Socratic

Sit with the shape of this before we move on, because the same pattern turns up in storage, in networking and in clusters.

Discussion prompt

A single-corded rack on one feed at 87 percent survives a maintenance day fine. A dual-corded rack with both feeds at 87 percent loses everything if one feed goes. Why does adding redundancy make the outcome worse here?

Hint: What did the second feed get used for, in practice, rather than in the design?

Answer:

Because the redundant capacity got spent on load rather than kept as margin. The design intent was two feeds each capable of the whole rack. The reality is two feeds each carrying their own half, and no spare anywhere.

So the second feed is not redundancy at all. It is capacity, and the rack has quietly become a single point of failure with twice the equipment and none of the protection anyone thinks it has.

The general pattern, and it is worth naming: redundancy consumed is redundancy absent. It is true of power feeds, of cooling units, of cluster members and of network uplinks, and the only defence is measuring the margin rather than the load.

Which is why the useful question about any redundant pair is never how loaded is it. It is what happens when one of them is gone, and can the survivor take it.

18. Do the Arithmetic

Check

Work it out properly before you look at the options.

Check your understanding

A dual-corded rack is fed by two 30 amp circuits. Applying the continuous load rule and the requirement that either feed alone can carry the rack, what is the highest steady-state current you should see on each feed?

  • A. About 24 amps on each feed
  • B. About 12 amps on each feed (correct)
  • C. About 15 amps on each feed
  • D. About 30 amps on each feed

Answer: B

Why: Two rules stack. The continuous load rule takes 30 amps down to 24 amps of usable capacity. The single-feed survival requirement then halves that, because either circuit must be able to carry the entire rack when its partner is gone. Twelve amps per feed in normal running, which is forty percent of the number printed on the breaker, and that gap between the printed number and the usable one is exactly why racks get quietly overloaded.

Why A tempts people
That is the continuous load rule applied on its own, which is right as far as it goes. It ignores the second rule: at 24 amps each, losing one feed asks the survivor for 48 amps on a 30 amp breaker.
Why C tempts people
Half of the breaker rating, which forgets the continuous load derate. It is closer than most wrong answers, and it still leaves the survivor asked for 30 amps continuous on a 30 amp breaker, which will open.
Why D tempts people
The full breaker rating is never a design target for a continuous load, and it fails the single-feed test by a factor of two on top of that.

19. Working the PDU Alert

Worked example

Four questions, then the card. This is the shape every one of tonight's five will take.

What is the sensor measuring? Current on branch circuit 3 of the A strip in rack R14, sustained over 18 minutes, not a momentary spike.

Why: Sustained matters. A brief peak during a batch job is a different problem from a load that has genuinely moved up and stayed there.

What is the blast radius now, and one step worse? Now: nothing, everything is serving. One step worse: the breaker opens, that strip dies, and because the partner cannot carry the rack, the whole rack follows.

Why: The second half of the question is what promotes this from a ticket to an escalation. Nothing is affected now and an entire rack is one event away.

Symptom or cause? Symptom. Something was added to this rack, or something in it started drawing more.

Why: Circuits do not drift upward on their own. A change caused this, and finding the change is usually a five-minute conversation.

Stabilise: stop anything else being added to R14 today, and check the recent change record for what went in.

Why: The cheapest possible intervention is preventing the situation getting worse while you work out what to do properly.

Resolve: move load to a rack with headroom, or have the facility engineer provide additional circuits. Then reset the alert threshold to reflect the dual-feed figure rather than the eighty percent one.

Why: The threshold itself was wrong. Alerting at eighty percent of a circuit that must never exceed fifty percent means the alert only fires long after the design was violated.

Figure (svg): Diagram of a rack fed by two independent power feeds each at sixty percent, showing that losing one feed asks the survivor for one hundred and twenty percent and opens its breaker.

The picture that justifies the escalation, and the one to show the person who asks why this is urgent.

Verify: done means both feeds below the dual-feed figure, the threshold corrected, and the change that caused it identified.

Why: Three conditions, and the alert going quiet is not one of them. A load that drops because a batch job finished has resolved nothing at all.

20. But We Are Only at Sixty Percent

Trap

The trap

A colleague looks at the same rack and says it is fine. Sixty percent on each side, well under the eighty percent rule, no action needed. On the numbers as stated, that reasoning is completely reasonable.

It is also how the rack ends up on the incident report. The eighty percent rule and the dual-feed rule are different rules, answering different questions, and passing one says nothing about the other.

The fix

Two questions, always both, for any redundant pair.

  1. Is each side within its own limit right now? That is the continuous load rule, and it is about the wiring.
  2. Can either side carry the whole thing alone? That is the redundancy rule, and it is about the design intent.

A dual-corded rack that passes the first and fails the second is the most common power finding in the industry, and it is invisible until the day somebody does maintenance on a feed.

Same shape as last session's redundancy is not backup. Two copies of something are only redundancy if either copy can do the whole job on its own.

21. Alert Seven - Cooling

Section

A unit has failed, and the room is perfectly cool

22. The Alert, As It Arrives

Picture it

Read all four lines, including the reassuring one.

Figure (svg): An alert card reading warning, computer room air handler CRAH-03 fan fault, unit offline, with return air temperature normal and the remaining three units in the room now running at increased fan speed.

Nothing is hot. Nothing is alerting at the rack. One line in the middle is doing all the work.

This is the margin-is-gone class in its purest form. Every measurement a customer would care about is exactly where it was an hour ago.

23. Is This Urgent?

Prediction

Three in the morning. Room temperature is normal and no rack is alerting.

Predict first

How urgent is this?

  • Not urgent, raise a ticket for the morning
  • Urgent, because the room has no cooling redundancy left and the remaining units are near maximum
  • Urgent, because the room is about to overheat
  • Not urgent as long as the remaining units stay below full speed

Correct: Urgent, because the room has no cooling redundancy left and the remaining units are near maximum

Why: The room is not about to overheat, which is why the third option is wrong and why this alert gets deferred so often. What has happened is that a room designed as N plus one is now running at N, with the survivors at ninety-one percent of their capability. The next unit to fail, and units in a room tend to fail for related reasons, takes the room below what it needs. The urgency is not in the current temperature, it is in the complete absence of margin combined with how little time a dense hall gives you once cooling actually goes short.

24. Why the Clock Matters Here

Concept

Losing cooling in a data hall is not like losing air conditioning in an office, and the difference is a matter of physics rather than of degree.

An office has furniture, walls, floors and people, all of which absorb heat, and a modest heat source. It warms up over hours.

A data hall has almost no thermal mass and an enormous, constant heat source. A rack drawing fifteen kilowatts is a fifteen kilowatt heater that never switches off. Every watt that goes in as electricity comes out as heat, immediately.

So the time between cooling stopping and equipment reaching temperatures that force a shutdown is commonly measured in single-digit minutes for a dense hall, not in hours. That is the number that makes this alert urgent while everything still reads normal.

ASHRAE TC 9.9, Thermal Guidelines for Data Processing Environments TC 9.9 — Inlet temperature envelopes, and what happens beyond them.

25. How Fast a Hall Heats

Picture it

Two rooms, both lose cooling at the same moment.

Figure (svg): Line graph comparing temperature rise after a cooling failure in an office, which climbs gently, against a dense data hall, which climbs steeply within minutes because it has high heat output and almost no thermal mass.

The ride-through in a dense hall is minutes. That is the entire argument for treating a redundancy loss as urgent.

This is also why the room-level temperature reading is such a poor early warning. By the time the room average moves, the hot spots have been in trouble for a while.

26. The Cooling Vocabulary

Matching

Match each term to what it actually refers to. These come up constantly and are easy to nod along to.

Match the pairs

  • m1. CRAC unit
  • m2. CRAH unit
  • m3. chiller
  • m4. supply air temperature
  • m5. return air temperature
  • m6. delta T
  • n1. air handler with its own refrigeration compressor
  • n2. air handler fed by chilled water from elsewhere
  • n3. the plant that makes the chilled water, usually outside
  • n4. what the unit blows into the cold aisle
  • n5. what the unit draws back from the hot aisle
  • n6. the difference between the two, which tells you how much heat is being carried

Why: The C and the H are the whole distinction between the first two: one makes its own cold, the other is handed cold water and only moves air. Delta T is the one worth internalising, because it is a measure of work done rather than of state. A delta T that collapses while fans run flat out usually means air is bypassing the racks and going straight back to the unit, which is an airflow problem no amount of extra cooling capacity will fix.

27. Working the Cooling Alert

Worked example

Same four questions, same card. Notice how much of this is about the second half of question two.

What is the sensor measuring? A fan fault on one air handler, plus fan speed on the three survivors. Not temperature.

Why: The temperature reading is context, not the alert. Confusing the two is what leads people to close this because the room is cool.

Blast radius now: none. One step worse: a second unit fault takes the hall below the cooling it needs, and a dense hall gives you minutes.

Why: This is the whole justification for waking somebody up. The current state is fine and the next state is an evacuation of the hall.

Symptom or cause? The fan fault is a cause at the unit level. The ninety-one percent fan speed is a symptom of the fault.

Why: Worth separating, because the fix belongs to whoever owns the unit, and the risk belongs to you until they finish.

Stabilise: raise it to facilities as a redundancy loss, not as a temperature problem, and get a time estimate for the repair.

Why: The wording changes the response you get. Reported as room is cool, unit offline it waits until morning. Reported as the hall is at N with survivors at ninety-one percent it does not.

While waiting, reduce what you can: defer batch work that would add load in that hall, and confirm the hot-aisle containment and blanking panels have not been left open by recent work.

Why: Both are free and both buy real margin. An open containment door is a surprisingly common contributor and takes a minute to check.

Figure (svg): Line graph comparing temperature rise after cooling failure in an office against a dense data hall, showing the hall climbing steeply within minutes.

The picture to put in the escalation, because it answers the question everyone asks: it is cool, why is this urgent?

Verify: done means the failed unit is back in service and the survivors have returned to their normal fan speed.

Why: Redundancy restored, not temperature acceptable. Temperature was acceptable throughout, and closing on it would mean closing on a condition that never changed.

28. Why Do Units Fail Together?

Socratic

The reason a redundancy loss is more urgent than it looks has a name, and it applies well beyond cooling.

Discussion prompt

The three surviving air handlers are identical, were installed on the same day, have run the same hours, and are now working harder than they ever have. Why does that combination make a second failure more likely than a first failure was?

Hint: Think about what the three survivors have in common, and about what changed for them when CRAH-03 went offline.

Answer:

Because failures in a set like this are not independent, and almost all the arithmetic people do about redundancy quietly assumes that they are.

The four units share an installation date, a duty cycle, a maintenance history, a batch of components and a supply of chilled water. Whatever wore out the bearing in the failed unit has been wearing out the same bearing in the other three, at the same rate, for the same number of hours.

And the load moved. The survivors went from sixty-two percent to ninety-one percent, so from this moment they are working harder than they ever have, which accelerates exactly the wear that just produced a failure.

The same reasoning applies to disks bought in one batch, to power supplies in one chassis, to cluster members built from one image and to controller batteries fitted on the same day. When you hear that two failures at once would be very unlikely, the question to ask is whether the two things have anything in common. Usually they have almost everything in common.

29. Three Statements About This Alert

Two truths and a lie

Two are sound. One is the reasoning that keeps this alert in a queue until morning.

Eliminate the wrong options

Which statement survives?

  • t1. The urgency comes from the loss of margin, not from the current temperature.
  • t2. The room is within its temperature range, so this can wait for business hours.
  • t3. As long as no rack is alerting on inlet temperature, the hall has nothing to worry about.

Survives elimination: t1

Why: The survivors running at ninety-one percent is the practical instrument here, and it is worth naming even though it is not one of the options. A fan speed reading is the cheapest available measure of remaining cooling margin, and it moves long before any temperature does. Get into the habit of reading fan speed and pump speed as margin gauges rather than as trivia, and of arguing politely with both of the statements you just eliminated.

30. Alert Eight - Time

Section

Four minutes wrong, and it takes down authentication

31. The Alert, As It Arrives

Picture it

The least impressive-looking alert in this deck, and the one with the widest blast radius.

Figure (svg): An alert card reading warning, domain controller DC-02 clock offset is 247 seconds ahead of the reference time source, with the time service reporting no synchronisation for eleven hours.

Four minutes and seven seconds. The last line is what makes it a room-wide problem rather than a host-level one.

There is no temperature here, no current, no failed part. This is a pure configuration and reachability alert, and it will express itself as a dozen unrelated-looking faults.

32. Four Minutes. So What?

Anomaly

Everything is running. Nobody has complained. The clock is four minutes fast on one server.

Predict first

Which of these is the most immediate consequence?

  • Log timestamps will be slightly wrong
  • Kerberos authentication starts refusing tickets once the skew passes five minutes
  • Scheduled jobs run four minutes early
  • Nothing measurable until the drift is much larger

Correct: Kerberos authentication starts refusing tickets once the skew passes five minutes

Why: Kerberos puts a timestamp inside every ticket specifically to stop an attacker replaying a captured one, and the default tolerance for clock difference is five minutes. This host is at four minutes and seven seconds and still drifting, so it is roughly fifty seconds from the point at which logins to anything that trusts it begin to fail. The failure looks nothing like a clock problem from the user's side: it looks like a wrong password, or an access denied on a file share, and it will be reported as thirty separate tickets.

33. What Actually Breaks When Clocks Disagree

Concept

Time is one of those dependencies that is invisible until it is wrong, and then it is wrong everywhere at once.

Microsoft Learn, Kerberos maximum tolerance for computer clock synchronization Clock skew — The five minute default tolerance, and what happens outside it.

34. Is Time the Cause?

Discrimination

Six reports from a morning where one domain controller was four minutes fast. Sort them.

Sort into buckets

Sort each report by whether a clock problem plausibly explains it.

Time could explain it
users on one site cannot map a file share, credentials rejected; a scheduled report ran twice, four minutes apart; an internal web application rejects a valid certificate; authenticator codes are refused for several users
Time cannot explain it
a disk in an array reported a media error; a switch port shows rising input errors
time
All four run on a timestamp somewhere. A ticket outside its validity window, a job whose schedule was evaluated against the wrong clock, a certificate compared against a wrong now, and a one-time code derived directly from the clock. They present as four unrelated faults and share one cause.
not
Media errors and port input errors are physical measurements of physical things. No clock affects whether a sector read correctly or a frame arrived with a bad checksum. When you are hunting a single cause for a morning of chaos, these two are evidence that something else is also going on.

35. Where Time Comes From, and Where It Goes Wrong

Concept

Time distribution is a tree. Something authoritative sits at the top and everything else takes time from a level above it.

Stratum — How many hops a clock is from a reference. A radio or satellite reference is stratum zero, the server attached to it is stratum one, a server synchronising to that is stratum two, and so on.

In a Windows environment, one domain controller holds the role that makes it the authority for the whole domain, and every other machine chains up to it. That is why one wrong clock on the right host becomes everyone's problem.

Internet Engineering Task Force, RFC 5905 - Network Time Protocol version 4 RFC 5905 — Strata, selection among sources, and how a client decides who to believe.

36. Where Do You Look First?

Check

One domain controller, four minutes ahead, no synchronisation in eleven hours.

Check your understanding

Which first check most efficiently separates the likely causes?

  • A. Restart the time service on the domain controller
  • B. Test whether the configured reference is actually reachable on UDP port 123 from that host (correct)
  • C. Manually set the clock to the correct time
  • D. Check whether other servers in the domain have also drifted

Answer: B

Why: The alert already told you the source is unreachable, and one reachability test either confirms that or proves the alert is misleading you. Confirmed unreachable points immediately at a firewall or routing change, which is a five-minute conversation with the network team and explains why this started eleven hours ago. That single test splits the problem in half before you have changed anything on the server.

Why A tempts people
Restarting the service is the reflex, and it will appear to help if the block has since been removed, which teaches you the wrong lesson. It changes state before you have gathered any, and it does nothing if the port is still blocked.
Why C tempts people
Setting the clock by hand treats the symptom, and on a domain controller a manual jump can cause its own problems. It also removes the evidence of how far it had drifted, which is how you would have estimated when this started.
Why D tempts people
A good second check, and worth doing, because it establishes blast radius. As a first move it costs longer and does not distinguish between the causes, since a wrong authority makes everyone wrong in exactly the same way.

37. Monitor the Offset, Not the Service

Concept

Most monitoring systems check that the time service is running. That check would have been green for all eleven hours of this incident, because the service was running perfectly and had nothing to synchronise against.

The measurement worth alerting on is the offset itself, in seconds, against a reference the monitoring system trusts independently.

That third point is the subtle one and it is the shape of this whole session. A machine whose clock is currently right and whose source is gone is not fine, it is drifting, and the only difference between it and a domain-wide authentication outage is how long you leave it.

38. Working the Time Alert

Worked example

The card again, and note how different stabilise and resolve are here.

What is the sensor measuring? Offset against a reference, and time since the last successful synchronisation. Both matter.

Why: The offset tells you how close to the five minute cliff you are. The eleven hours tells you when it started, which is what you take to the network team.

Blast radius: 38 member servers chain from this host, so their clocks are wrong too, and anything authenticating against them is at risk. One step worse: the offset passes five minutes and authentication fails domain-wide.

Why: This is the widest blast radius of tonight's five, from the least dramatic-looking alert. That contrast is worth remembering.

Symptom or cause? The offset is a symptom. The unreachable source is closer to the cause, and the change that made it unreachable is the cause.

Why: Three layers, and fixing only the first guarantees you will be back here in another eleven hours.

Stabilise: correct the clock deliberately, and slew it rather than stepping it if the tooling allows.

Why: Slewing speeds the clock up or slows it down slightly until it converges, so time never repeats and never jumps backwards. Databases and log analysis both handle that far better than a jump.

Resolve: restore reachability to the reference, then confirm the host synchronises on its own and holds.

Why: Done is not the clock is right now. Done is the mechanism that keeps it right is working again, demonstrated by a successful synchronisation after you stop touching it.

Figure (svg): An alert card showing a domain controller clock offset of 247 seconds with no successful synchronisation for eleven hours and the reference source unreachable.

What resolved actually looks like: not the number corrected, but the mechanism demonstrably working again.

Verify: done means a successful synchronisation has happened without intervention, and the member servers have followed.

Why: The last clause is the one people forget. Fixing the authority does not instantly fix the 38 machines beneath it, and some of them will need a nudge.

39. Never Push a Clock Backwards in a Hurry

Trap

The trap

A server is four minutes ahead. The obvious fix is to set it back four minutes. It takes one command and the number is immediately correct.

On a database server this can be genuinely destructive. Four minutes of timestamps now repeat, so two different events can carry the same time, and anything that assumed time only moves forward is now working from bad data. Replication ordering, transaction logs and any application logic that compares timestamps are all exposed.

Backups and log analysis inherit the same confusion, and unlike the database it will not be noticed for weeks.

The fix

Correct a clock the way the time service does it by default, which is to slew rather than step.

  • Slewing means running the clock slightly slow or slightly fast until it converges. Time never repeats and never reverses.
  • A large offset takes a while to slew out, and that is fine. Correctness beats speed here.
  • Where a step really is necessary, do it as a planned change with the affected services stopped, exactly as you would for a reboot.

Same principle as session 2: the fast fix and the safe fix are different moves, and knowing which one you are making is the job.

40. Alert Nine - Storage

Section

No disk failed, no data lost, and everything got slow

41. The Alert, As It Arrives

Picture it

Two lines. The second one is the entire story and it is written as if it were housekeeping.

Figure (svg): An alert card reading warning, storage array controller one cache battery has failed, write cache policy changed from write back to write through, with all volumes online and optimal.

Every volume is optimal. Nothing failed. Write latency went up by a factor of more than twenty.

This alert routinely gets triaged as low priority because of the third line, and the applications behind it are already suffering by the time anyone reads the fourth.

42. What Would Explain This?

Hypothesis

No disk failed. No volume is degraded. Nothing is rebuilding. Write latency is twenty times worse than this morning.

Predict first

Which explanation fits every fact?

  • A disk is failing silently and slowing the array down
  • The array stopped acknowledging writes from its own memory and now waits for the disks
  • The network path to the array became congested
  • A backup job is competing for the array

Correct: The array stopped acknowledging writes from its own memory and now waits for the disks

Why: The clue is that the policy change and the latency change are in the same alert. With a working battery the controller can safely say done as soon as a write is in its own memory, because the battery guarantees that memory survives a power loss long enough to be written out. With the battery dead that promise cannot be kept, so the controller stops making it and waits for the physical disks instead. That is the difference between microseconds and milliseconds, and the array is behaving correctly and conservatively throughout. Nothing is broken; the array has chosen safety over speed and told you so in a line that reads like a footnote.

43. Write-Back, Write-Through, and the Battery

Concept

This is worth understanding properly, because the same trade-off appears in databases, in filesystems and in every caching layer you will ever meet.

Write-back — The controller puts the write in its own memory, immediately tells the server it is done, and writes it to disk later. Fast, and it depends on that memory surviving a power failure.

Write-through — The controller waits until the data is actually on the disks before telling the server it is done. Slower, and it needs no promises about memory at all.

The battery or supercapacitor on the controller is what makes the first mode honest. It keeps the cache alive long enough for its contents to be written out after an unexpected power loss.

So when the battery fails, the controller cannot honour the promise it was making, and it stops making it. The performance collapse is not a fault. It is the array refusing to lie to you.

44. Two Paths for One Write

Picture it

Same write, two policies, and one component deciding which one is available.

Figure (svg): Diagram comparing write-back, where the controller acknowledges a write as soon as it is in controller cache, against write-through, where the controller waits for the disks, with a note that the cache battery is what makes write-back safe.

The battery is not a performance feature. It is the thing that makes the fast path safe, and losing it removes the fast path.

Note the direction of the trade. You have not lost data and you cannot lose data. You have lost speed, and the reason you lost it is that the protection against losing data went away.

45. What Do You Do About It?

Elimination

The array offers a setting to force write-back on regardless of battery state. Four options, one right answer.

Eliminate the wrong options

Which action is correct?

  • g1. Force write-back back on to restore performance, then replace the battery when convenient.
  • g2. Order the battery module, warn the application owners that writes are slow until it is replaced, and treat it as urgent.
  • g3. Fail the workload over to the second controller, since its battery is fine.
  • g4. Nothing, since all volumes report optimal and no data is at risk.

Survives elimination: g2

Why: The correct answer is unglamorous: accept the slower and safer mode, tell the people affected before they open tickets, and get the part replaced quickly. The judgement being tested is whether you will trade a data-integrity guarantee for latency under pressure. Someone will ask you to, and the answer is no unless a named person with the authority to accept that risk says so in writing.

46. And the Other Controller?

Prediction

The array has two controllers. Controller 1's cache battery has just failed after four years in service. Controller 2's battery was fitted on the same day.

Predict first

What should you assume about controller 2?

  • It is fine, since it has not alerted
  • It is probably close behind, so order two modules and check its charge and predicted end of life now
  • It will fail immediately, so fail everything over
  • Nothing can be inferred from one battery failing

Correct: It is probably close behind, so order two modules and check its charge and predicted end of life now

Why: Exactly the correlated failure argument from the cooling alert, in a different system. The two batteries are the same chemistry, the same batch and the same age, and they have spent four years in the same thermal environment doing the same duty. A battery failure at four years is a statement about the design life of that part in this room, not about one unlucky module. Ordering the pair and checking the survivor's charge and predicted replacement date costs almost nothing, and it converts a probable second incident into a second line on the same change record.

47. Working the Storage Alert

Worked example

Shorter than the others, because the diagnosis was handed to you in the alert. The work is in the response.

What is the sensor measuring? A battery charge level, a policy change, and a latency figure. Three different things in one alert.

Why: The policy change is the fact, the latency is the consequence, and the battery is the cause. Recognising that ordering immediately is what makes this alert quick to work.

Blast radius: every volume on that controller, so every application writing to it. One step worse: nothing gets worse, which is unusual and worth saying.

Why: This alert is stable. It will not deteriorate on its own, and that is precisely what makes it safe to fix properly rather than quickly.

Symptom or cause? The latency is the symptom. The policy change is the mechanism. The dead battery is the cause, and it is already named.

Why: One of the few alerts that hands you all three layers at once. Most do not.

Stabilise: notify the application owners with a number, not an adjective. Writes have gone from under half a millisecond to nine, and here is why.

Why: Getting ahead of it converts thirty confused tickets into one informed conversation, and it buys goodwill for the outage window you will need for the replacement.

Resolve: replace the battery module. Check whether it is hot-swappable on this model, and whether the array will restore write-back automatically or needs to be told.

Why: Both details are model-specific and both are worth knowing before the engineer arrives rather than while they are standing there.

Figure (svg): Diagram comparing write-back where the controller acknowledges from its own cache against write-through where it waits for the disks, with the cache battery making the fast path safe.

The picture to show an application owner who wants to know why their database got slow when nothing broke.

Verify: done means the battery reports charged, the policy has returned to write-back, and the latency figure is back where it was.

Why: Three conditions again, and the third is the one to insist on. A battery that reports healthy while the policy is still write-through means somebody has to go and change it back, and until then nothing has improved for the applications.

48. Alert Ten - Compute

Section

Memory errors that were all corrected

49. The Alert, As It Arrives

Picture it

Everything in this alert was successfully handled by the hardware. That is what makes it easy to dismiss.

Figure (svg): An alert card reading warning, host ESX-14 correctable memory errors on DIMM A3, 1840 errors in the last 24 hours rising from 12 the previous week, with no uncorrectable errors and the host running normally.

Every one of those 1840 errors was caught and corrected. Nothing was lost, and nothing was affected.

Two numbers are doing all the work: the rate went up by a factor of about a thousand, and it is confined to one module out of many.

50. What Error-Correcting Memory Does

Concept

Server memory carries extra bits so the controller can check every read against the value that was written.

So a correctable error is not a failure. It is the hardware doing exactly what it was designed to do, and telling you about it.

The value of the alert is entirely in the rate and the location. Occasional isolated errors happen, including from cosmic rays, and mean nothing. A rate climbing on one module means that module is degrading, and field studies consistently find that a module producing correctable errors is far more likely to produce an uncorrectable one later than a module producing none.

Schroeder, Pinheiro and Weber, DRAM Errors in the Wild - a Large-Scale Field Study, SIGMETRICS 2009 DRAM errors in the wild — Correctable error rates as a predictor of uncorrectable failure.

51. The Two Boxes

Picture it

The alert is not about the left-hand box. It is about how close you are to the right-hand one.

Figure (svg): Diagram of correctable single-bit memory errors which are corrected in flight and counted, against uncorrectable double-bit errors which halt the machine, with an arrow showing that a rising correctable rate on one module predicts the second.

The whole value of the alert is that it lets you choose the outage instead of receiving one.

This ties straight back to last session. A predicted failure lets you plan a reboot: drained, in a window, with the console open. An uncorrectable error gives you a halted host and whatever was running on it.

52. Which of These Would Worry You?

Discrimination

Six memory reports from six different hosts. Sort them by whether they justify action.

Sort into buckets

Sort each report by whether you would act on it.

Act on it
1840 correctable errors in a day on one module, up from 12 in a week; one uncorrectable error, host halted and restarted
Watch it
correctable errors spread evenly across all sixteen modules; a steady 40 correctable errors a day on one module for three months
No action needed
3 correctable errors on one module over six months; 0 errors, but the host has been up for 600 days
act
A rate that has climbed by orders of magnitude on a single module, or a genuine uncorrectable error, both justify replacing the module. The first lets you schedule it; the second already cost you an outage and you are replacing it so there is not a second one.
watch
A steady low rate is not the same as a rising one, and errors spread evenly across every module point at something other than a bad module, such as the controller, the board or a thermal issue. Both deserve a trend and a review date rather than a part order.
ignore
A handful of errors over months is background, and some of it is literally cosmic rays. Long uptime with no errors is not a memory finding at all, though session 2 would have something to say about the uptime itself.

53. Turning It Into a Change

Real world

This is where tonight's session and last session's join up.

Discussion prompt

You are going to replace DIMM A3, which means the host must be powered off. It is a hypervisor host with 22 virtual machines on it. Write the plan.

Hint: Almost every step of this was in session 2, and only one step is about memory at all.

Answer:

Confirm the finding first: check whether the vendor tooling can retire the affected pages as an interim measure, and confirm the module location physically so the engineer replaces the right one.

Raise a normal change with the part, the window and the justification, which is the error rate trend rather than an incident.

Drain the host: put it into maintenance mode so the 22 guests migrate off, and confirm the migration completed rather than assuming it did.

Confirm out-of-band access to the host before it is powered down, since you will be watching it come back through the console.

Power it down properly, replace the module, power on, and watch the memory self test through the console. This is the blind zone, and it is where a badly seated module announces itself.

Verify up the ladder: the host boots, the memory total is what it should be, the error counters are cleared, the host rejoins the cluster, and only then take it out of maintenance mode.

The point worth noticing: exactly one step of that plan is about memory. Everything else is the reboot discipline from last session, which is why we did that one first.

54. The Right Reading of This Alert

Check

One module, 1840 corrected errors in a day, none uncorrectable, host running normally.

Check your understanding

What is the single most accurate statement about this alert?

  • A. The host has a memory fault and is at risk of data corruption right now
  • B. Nothing has failed and nothing is at risk, because every error was corrected
  • C. Nothing has failed yet, and the alert is a warning that lets you schedule the outage rather than receive one (correct)
  • D. The errors are cosmic ray background and can be ignored

Answer: C

Why: Every error was corrected, so nothing was lost and no data was affected, which is what makes the first option wrong. But the rate rose from about two a day to nearly two thousand a day on one specific module, and that trend is a well-documented predictor of an uncorrectable error, which halts the host. The whole value of the alert is the choice it offers: replace the module in a window of your choosing, or have the host halt at a time of its choosing.

Why A tempts people
Correctable means the value was repaired before use, so there is no corruption. Overstating it costs you credibility with the people whose change window you are about to ask for.
Why B tempts people
True as far as it goes, and it is the reasoning that leaves this alert open for six weeks. It answers what happened and ignores what the rate predicts.
Why D tempts people
Background radiation genuinely does cause isolated correctable errors, which is why a few over months mean nothing. It does not produce a thousandfold rise confined to one module out of sixteen.

55. Part 6 - When Fifty Alerts Arrive at Once

Section

Find the parent, ignore the children

56. An Alert Storm Is One Alert

Concept

Sooner or later you will open the console to two hundred alerts that all arrived within ninety seconds. It is genuinely alarming and it is almost always one event.

Monitoring systems see symptoms, and a single failure produces symptoms everywhere. A switch stack restart makes every host behind it unreachable, which makes every service check on those hosts fail, which makes every cluster containing one of those hosts report itself degraded.

That last habit is worth practising deliberately, because the alternative is working two hundred tickets in parallel, which nobody has ever finished.

57. One Cause, Two Hundred Symptoms

Picture it

The shape is always the same, and once you see it you cannot unsee it.

Figure (svg): Diagram of a single root cause, a switch stack reboot, producing four groups of symptom alerts including unreachable hosts, services down, degraded clusters and failed checks.

The parent is in there, early, and usually with a lower severity than any of its children.

This is question three from the routine, at scale. Two hundred symptoms, one cause, and the entire skill is refusing to work the symptoms.

58. Order the Storm

Ranking

Five alerts arrive within one minute of each other. Put them in the order they must have happened.

Put in order

  1. core switch stack member rebooted
  2. 18 hosts unreachable
  3. 4 database clusters report a member missing
  4. 6 application health checks failing
  5. customer-facing error rate rising

Why: Causation runs downhill from the infrastructure to the customer, and the alerts arrive in roughly that order because each layer notices the one below it. The switch goes, so hosts become unreachable, so clusters notice a missing member, so health checks fail, so customers see errors. Work the top of the chain and everything below it clears on its own. Work the bottom of the chain and you will be busy for hours while the actual fault sits there untouched.

59. Alert Fatigue, and the Two Ways to Cause It

Concept

The failure mode of a monitoring system is not missing an alert. It is producing so many that people stop reading them, and it gets there by two routes.

Flapping — A measurement sitting on a threshold, crossing back and forth, generating an alert and a recovery every few minutes. The information content is near zero and the noise is constant.

Thresholds nobody agreed — A check set at a default that does not match how the system actually behaves, so it fires daily and is ignored daily.

Both are fixed the same way, and neither is fixed by muting the alert. A flapping check needs hysteresis, meaning it fires at one level and only clears at a distinctly lower one, or a requirement to be over the line for a sustained period. A wrong threshold needs the correct number, worked out from how the system behaves.

And when you are doing planned work, suppress deliberately. Putting a host into maintenance mode in the monitoring system before you start is what stops your own change generating a storm that trains everyone to ignore storms.

60. Where Do You Set the Threshold?

Trade off

Fill the blanks. This is a real decision you will be asked to make, and there is no universally right answer.

Comparison matrix

ConsiderationThreshold set tightThreshold set loose
warning time before the limitlongshort, sometimes none
false alarmsfrequentrare
what people do after a monthstop reading ittrust it, and act when it fires
best usea slow-moving trend with a real deadlinea fast-moving condition where any crossing is real

The resolution is usually two thresholds rather than a compromise on one: a quiet one that opens a ticket for the trend, and a loud one that wakes somebody. Tonight's PDU alert is the example, and it needs both.

61. The Suppression Nobody Removed

Trap

The trap

You do the right thing before a planned change: you put the host into maintenance mode in the monitoring system so your reboot does not generate a storm and train everyone to ignore storms.

The change goes well. The host comes back. You verify it up the ladder, close the ticket, and go home.

Nine days later the same host runs out of disk space and fills a database volume. No alert fires, because it is still in maintenance mode, and nobody notices until the application stops accepting writes.

The fix

Suppression is a change like any other, and it needs an end as clearly defined as its beginning.

  • Always set an expiry when you suppress. If the tool supports a duration, use it, and pick one slightly longer than the work should take rather than open-ended.
  • Make removing the suppression a step in the change plan, sitting alongside returning the node to the pool, not something you remember afterwards.
  • Review what is currently suppressed on a schedule. Anything suppressed for longer than a week is either a forgotten change or an alert that should have been fixed rather than muted.

This is the same lesson as done means from the card. A change is not finished when the work is finished, it is finished when everything you altered to do the work has been put back.

62. Say the Session Back

Explain it to yourself

One idea connects all five of tonight's alerts, and it is worth you putting it into your own words.

Discussion prompt

What did the PDU circuit, the failed air handler, the drifting clock, the dead cache battery and the memory errors have in common?

Hint: Ask what a customer would have noticed at the moment each alert fired.

Answer:

Not one of them was an outage. At the moment each alert fired, every service was working and no user would have noticed anything at all.

Four of the five were the system telling you that a margin had gone or was going: no spare feed, no spare cooling unit, no spare time before the skew limit, no spare corrections before a module fails.

The fifth, the cache battery, is the odd one out and it is instructive: nothing was lost and nothing was at risk, but the machine had quietly changed its behaviour to stay safe, and only the latency number gave it away.

So the practical skill is reading an alert where nothing is wrong. Ask what would happen if this got one step worse, and how much time you would have when it did. Those two answers are what turn an item in a queue into an action tonight.

63. Five More Cards

Connect it up

Same artefact as last time, and by now you should have one or two from your own building.

Draw it

Take tonight's five alerts and write the six-field card for each one as it would apply in your data centre, not in mine: what fired, blast radius now and one step worse, first check, stabilise, escalate to by name or role, and done means. Then find the equivalent alert in your own monitoring system for at least two of the five, and write down the actual threshold it uses. Where a threshold does not match what we worked out tonight, that gap is your first real contribution.

Bring the cards and the thresholds. The gaps between what your system alerts on and what the arithmetic says it should alert on are the most useful thing you can walk into a team meeting with in your first month.

64. One Question Before You Close

Exit ticket

The habit I most want to survive tonight.

Predict first

An alert fires and everything is working normally. What is the second question you ask?

  • Can this wait until the morning?
  • What would happen if this got one step worse, and how much time would I have?
  • Has anyone else noticed it?
  • Is the threshold set correctly?

Correct: What would happen if this got one step worse, and how much time would I have?

Why: It is the second question because the first is still what is the sensor actually measuring. But this is the one that separates the alerts that can genuinely wait from the ones that look identical and cannot. A room that is cool with no cooling redundancy and a room that is cool with full redundancy read the same on every temperature gauge in the building, and they are a very long way apart. Asking about the next step, and about how much time that step leaves you, is what makes the difference visible, and it is why four of tonight's five alerts were worth working immediately.

65. What You Can Do Now

Recap

Eight things, and the routine holds for round three.

alertthe first move
branch circuit over thresholdcheck the partner feed, then do the halving arithmetic
cooling unit failed, room coolread the survivors' fan speed as a margin gauge
clock drifttest reachability of the reference on UDP 123
cache battery failedtell the application owners a number before they open tickets
correctable memory errorscompare the rate to last week, and check it is one module
fifty alerts at oncesort by time and read the first three

The thread through all five: an alert about a margin is an alert about a future, and the only way to read one is to ask what happens next and how long you would have. Everything else tonight was arithmetic.

Next session, round three: five more, and I will hold back the category labels so you can place them yourself. Bring the cards, and bring the thresholds that do not match.

Sources

  1. National Electrical Code, Article 210.20 - continuous loads and the eighty percent rule
  2. ASHRAE TC 9.9, Thermal Guidelines for Data Processing Environments
  3. Microsoft Learn, Kerberos maximum tolerance for computer clock synchronization
  4. Internet Engineering Task Force, RFC 5905 - Network Time Protocol version 4
  5. Schroeder, Pinheiro and Weber, DRAM Errors in the Wild - a Large-Scale Field Study, SIGMETRICS 2009
  6. Uptime Institute, annual outage analysis - causes and cost of data centre outages

Want this taught 1-on-1? Alexander tutors Data Center Operations — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108