Session 2 - Rebooting Servers, Safely and On Purpose

The session the student asked for by name, to get their feet wet with rebooting. It treats a reboot as a controlled loss of state rather than a repair, and spends its time on the minutes either side of the command rather than on the command itself. Part 1 separates four verbs people use as one - restart, shutdown, reset and cold boot - and maps the boot chain, including the blind zone before the kernel that no remote tool can observe. Part 2 installs a six-gate pre-flight and the argument that the risk is not the reboot but the boot, since a long-running server is a pile of changes never tested against a restart. Part 3 covers ordering it on Linux and Windows with annotated commands, draining a node, rolling a three-node cluster without losing quorum, and the host-against-guest question on virtualised infrastructure. Part 4 replaces the ping with an eight-rung verification ladder. Part 5 works the machine that did not come back, from the console. Part 6 makes the case that a reboot destroys the evidence, and gives a ninety-second capture that keeps it. Two ideas recur deliberately: the risk lives in the boot, and a reboot that works is not the same as a problem that is fixed.

Subject: Data Center Operations · 72 slides · applied lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Rebooting Servers

Title

Session 2

What a reboot really is, what has to be true before you order one, how to prove it worked, and what to do when the machine does not come back

2. Getting Your Feet Wet Without Getting Soaked

Objectives

You asked to park the next five alert scenarios for one session and spend the hour on rebooting instead. That is what this is. The alerts come back next time, and there are five new ones waiting.

Nothing here needs prior experience. It does assume you will stop me the moment a term goes past you, which is what you did last session and it worked well.

3. Part 1 - What a Reboot Actually Is

Section

A controlled loss of state, not a repair

4. Before Any of the Mechanics

Warm-up

Answer from instinct first. The instinct is usually half right, and the half that is wrong is the useful part.

Discussion prompt

Rebooting a server fixes an enormous range of problems. In one sentence, what is the common property of every problem a reboot fixes?

Hint: Think about what a reboot removes, not what it restores.

Answer:

A reboot fixes problems that live in accumulated state: leaked memory, a wedged process, a stale handle on a file that was deleted an hour ago, a connection pool that never recovered from a network blip, a driver in a mode it cannot get out of.

It throws away everything the machine has become since it started, and returns it to what the configuration on disk says it should be. That is a real and powerful move.

The flip side follows immediately and it is the whole reason this session exists. If the problem is in the configuration on disk, in the hardware, or in another machine entirely, a reboot cannot touch it, and you will have destroyed the evidence for nothing.

5. What Actually Happens When You Type It

Concept

A reboot is two separate journeys with a gap in the middle, and people worry about the wrong one.

  1. Going down. The operating system tells every service to stop, in order, and waits. Then it flushes anything still sitting in memory to disk, unmounts the filesystems cleanly, and asks the hardware to restart.
  2. The gap. Firmware runs its power-on self test, finds a boot device, and hands control to a bootloader. The machine is not on any network at this point and nothing you run remotely can see it.
  3. Coming up. The kernel loads, mounts the filesystems, and starts services in dependency order until the machine is doing its job again.

Almost everyone watching a reboot is watching the first journey. Almost every reboot that ruins an evening fails on the third.

6. The Boot Chain, and the Part You Cannot See

Picture it

Trace this once. Every reboot question you will ever be asked lives somewhere on this line.

Figure (svg): Diagram of the boot chain from power on through firmware self test, boot device selection, bootloader, kernel and services, split into a blind zone before the kernel where no remote tool can observe the machine and an observable zone after it.

The risk in a reboot is concentrated in the half of the process you cannot watch from your desk.

This is the argument for the service processor, which we get to shortly. It is the only thing that can see the left-hand side.

7. Put the Journey in Order

Ranking

From the moment the command is accepted to the moment the application answers.

Put in order

  1. services are asked to stop, in dependency order
  2. filesystems are flushed and unmounted
  3. firmware runs its power-on self test
  4. the bootloader hands control to the kernel
  5. the kernel mounts the filesystems
  6. services start and the application answers

Why: The two halves are mirror images with the firmware in between. Going down, the operating system stops the things that use the disk and then releases the disk. Coming up, it claims the disk and then starts the things that use it. Knowing which half you are in tells you which tool can still see the machine, which is the practical value of memorising this at all.

8. Four Verbs People Use As If They Were One

Concept

Getting these confused is how a routine change becomes an incident, so it is worth being fussy about the words.

Restart, or reboot — The operating system shuts itself down cleanly and then starts again. Everything gets asked politely to stop first.

Shutdown — The same clean stop, but the machine stays off. Somebody or something has to turn it back on, and in a data centre that somebody may be a long drive away.

Reset, or power cycle — The chassis is forced off and on. Nothing is asked, nothing is flushed. This is the electrical equivalent of pulling the plug, done remotely.

Cold boot — Power removed long enough that the firmware and every device re-initialise from scratch. Occasionally the only thing that clears a wedged network card or a confused storage controller.

systemd, systemctl manual page - reboot, halt, poweroff and the shutdown targets systemctl — The reboot, halt and poweroff targets, and what each one does to running units.

9. The Four Verbs Side by Side

Picture it

The dividing line is one question: does the operating system get a say?

Figure (svg): Table style diagram comparing restart, shutdown, reset or power cycle, and cold boot, showing which of them ask running applications to stop first and which force the chassis off immediately.

Two of these are a request. Two of them are an interruption.

You will hear all four called rebooting in conversation. In a change ticket they are four different risks and they need four different words.

10. Which Verb Is This?

Discrimination

Sort each action by whether the operating system is given a chance to stop cleanly.

Sort into buckets

Sort each action into asked or forced.

The OS is asked
typing reboot at a shell prompt; using Graceful Shutdown in the service processor; sending a restart command from the cluster manager
The OS is not asked
clicking Power Cycle in the service processor; holding the physical power button for ten seconds; pulling both power cables from the back of the chassis
asked
In each of these the request reaches the operating system, which stops services, flushes writes and unmounts filesystems before anything electrical happens. Graceful Shutdown in the service processor belongs here because it sends a signal the operating system listens for, rather than cutting power.
forced
These take the power away without asking. Anything the operating system had in memory and had not yet written is gone, and filesystems come back up needing a check. Sometimes that is the right call, but it is never the default one.

11. Reboot, Shutdown, Power Cycle, Cold Boot

Comparison

Fill in the missing cells. The last column is the one that decides how you feel at three in the morning.

Comparison matrix

ActionAre apps asked to stop?Does it clear firmware and device state?Comes back on its own?
rebootyesnoyes
shutdownyesnono, someone must power it on
power cyclenopartlyyes
cold bootnoyes, fullyyes, once power returns

The shutdown row is the one that catches new people. A clean shutdown of a remote server with no service processor is a change you cannot undo without sending a human to the building.

12. Why Does a Clean Shutdown Take So Long?

Socratic

A database server takes four minutes to shut down. A web server takes eight seconds. Same hardware, same operating system.

Discussion prompt

What is the database spending those four minutes on, and why would rushing it be expensive?

Hint: Think about what lives in memory on a database server that does not live in memory on a web server.

Answer:

The database is flushing. Modern databases keep recently changed pages in memory and write them out lazily, because writing to disk on every change would be unbearably slow. A clean shutdown is where that debt gets paid.

It is also finishing or rolling back open transactions, and writing a checkpoint that says where recovery should start from next time.

Force it off in the middle and none of that is lost, exactly, because the write-ahead log is on disk. But the next start has to replay that log to work out what was and was not complete, and on a busy system that recovery can take considerably longer than the clean shutdown you were impatient about.

The general rule is worth keeping: the more state a service holds in memory, the more a clean stop is worth, and the more a forced stop costs you on the way back up.

13. The Verb That Does Not Mean What You Think

Trap

The trap

You need to restart a Windows server. You are in a hurry, and you type the shutdown command with the flag you half remember.

shutdown /s /t 0

The machine goes down cleanly. It never comes back. You are now looking at a monitoring dashboard that says the host is unreachable, and there is nothing left running on it that could bring it back.

On Linux the same trap wears a different hat. The bare shutdown command with no flags does not restart anything either, and on many systems it halts.

The fix

Say the verb out loud before you press the key. On Windows the restart flag is r for restart, and slash s is one key away on the same keyboard. On Linux it is minus r for the same reason, and minus h halts.

Better still, use a command that has no ambiguous mode at all, because it can only ever do one thing. On Windows that is Restart-Computer with the force flag. On Linux it is systemctl reboot.

And best of all, have the service processor console open in another tab before you press anything, so that a mistyped flag costs you thirty seconds instead of a drive to the building.

14. Part 2 - The Pre-Flight

Section

Six gates, any one of which can stop the change

15. The Risk Is Not the Reboot, It Is the Boot

Concept

Almost nothing goes wrong on the way down. The operating system has done that thousands of times and it is good at it.

What goes wrong is the way back up, and it goes wrong for a reason that is almost poetic: a running server is a snapshot of every change anyone has made since it last started, and none of those changes have been tested against a boot.

None of these are visible while the machine is running. All of them surface in the same five minutes.

16. Uptime Is Not a Trophy

Picture it

The number people are proudest of is the number that should worry them most.

Figure (svg): Bar chart comparing days since last boot against how much is known about whether the machine will boot successfully, showing confidence falling as uptime rises from seven days to nine hundred.

Every day of uptime is another day of changes that have never been proven to survive a restart.

This is the actual argument for patching on a schedule. It is not mainly about the patches. It is that a machine which reboots monthly is a machine you know will reboot.

17. How Long Since It Last Booted?

Estimation

You are asked to reboot a production server. You check, and it has been up for 812 days.

Predict first

What does that number tell you about tonight?

  • It is very stable, so the reboot is low risk
  • It is overdue, so the reboot is low risk
  • Nobody currently employed has seen this machine boot, so the risk is high and unmeasured
  • Nothing useful, uptime is just a counter

Correct: Nobody currently employed has seen this machine boot, so the risk is high and unmeasured

Why: Eight hundred days of uptime means eight hundred days of accumulated undocumented change, none of it tested against a restart, and it very likely means the people who built the machine have moved on. Stability while running says nothing at all about whether it comes back. Treat a long-uptime reboot as a high-risk change: do it in a window, with the console already open, with a rollback in mind, and with the person who knows the application awake and reachable.

18. The Six Gates

Concept

This is the pre-flight. It takes five minutes on a machine you know and half an hour on one you do not, and the half hour is the one that pays.

  1. What does this server actually do? Name the service, not the hostname. If you cannot, you do not yet know what you are about to interrupt.
  2. Who is using it right now? Sessions, jobs, queues, and the humans who will notice.
  3. Can I reach it if the operating system never returns? Service processor address, credentials, and a test that you can actually log in, done now rather than later.
  4. Is anything mid-flight? A storage rebuild, a backup job, a batch run, a database maintenance task, a replication catch-up.
  5. Has this machine booted successfully since it last changed? If not, assume the boot will surface something.
  6. Who approved it, and who did I tell? A change record and a notification, both before, not after.

ITIL 4, Change Enablement - standard, normal and emergency changes Change enablement — Standard, normal and emergency changes, and which one a reboot is.

19. The Pre-Flight as a Gate Sequence

Picture it

Read it as gates rather than as a checklist. A checklist gets ticked. A gate can refuse you.

Figure (svg): Flow diagram of six sequential pre-flight gates before a server reboot, covering the purpose of the server, current users, out of band access, work in flight, whether it has booted since it last changed, and approval and notification.

Any gate that cannot be answered is a reason to stop, not a reason to hurry.

The third gate is the one I would tattoo on the inside of your eyelids. Everything else on this list is recoverable from your desk. That one decides whether you can recover at all.

20. A Ticket Arrives

Missing information

The ticket reads, in full: Please reboot APP-PRD-07, it is running slow.

Discussion prompt

List everything you would need to know before you are willing to act on this, and say which single missing item would stop you outright.

Hint: Three of the six gates cannot be answered from that sentence at all.

Answer:

Missing: what the server does and who depends on it, whether it is one of a pool or the only one, whether anyone has looked at what running slow actually means, what the change window is, whether there is a service processor and whether you can reach it, and when it last booted.

The item that stops you outright is out-of-band access. Everything else can be discovered or negotiated. Without a console you are betting the service on the boot succeeding, with no move available if it does not.

The second most important omission is the diagnosis. Running slow is not a fault, it is a feeling. If nobody has looked, the reboot will probably appear to work, the problem will return next week, and by then the evidence will have been destroyed twice.

21. Reboot Now, Wait, or Refuse?

Sorting

Same request, six different contexts. Sort them by what you would actually do.

Sort into buckets

Sort each situation by the right response to a reboot request.

Go ahead now
a node in a pool of six, drained, inside the change window; a lab machine nobody has used in three weeks
Wait for the right moment
a storage array is 40 percent through rebuilding a failed disk; a database that is mid-backup, with the job due to finish in twenty minutes
Stop and escalate
the only domain controller in the site, at 2pm on a Tuesday; a remote server with no service processor and no one on site
now
Drained and in-window, or genuinely unused. The blast radius is understood and small, and the change is authorised. This is what a routine reboot is supposed to look like.
wait
Something is mid-flight that a reboot would interrupt and that will finish on its own shortly. A rebuild restarted from zero and a backup that has to be rerun are both real costs, and both are avoided by waiting twenty minutes.
stop
These are not yours to decide alone. A sole domain controller at midday is an outage for every login in the building, and a remote box with no console is a change you cannot recover from. Escalate, get a window, or get someone to the site first.

22. Out-of-Band Access, and Why It Comes First

Concept

Every serious server has a small computer inside it whose only job is to manage the big computer. Dell calls it iDRAC, HP calls it iLO, Lenovo calls it XCC, and the underlying standard is IPMI or its successor Redfish.

If you take one habit from this session, take this one: open the console before you order the reboot, not after it fails to come back.

Intel, IPMI Specification v2.0 - chassis control through the service processor IPMI v2.0 — Chassis control commands and the service processor model.

23. In Band and Out of Band

Picture it

Two paths to the same machine, and only one of them survives the operating system.

Figure (svg): Diagram of a server showing an in band path from an administrator login session to the operating system and a separate out of band path from a browser to the service processor, which runs on standby power with its own network port.

Your login session dies with the operating system. The service processor does not.

In band means through the thing you are about to restart. Out of band means around it. Once you have that distinction, the phrase remote hands starts to make a lot more sense too.

24. Say It Back

Explain it to yourself

This is the idea I most want to survive tonight, so put it in your own words rather than mine.

Discussion prompt

Why is opening the service processor console before a reboot worth more than opening it after the server fails to come back?

Hint: Think about what the console can show you that nothing else can, and when that information exists.

Answer:

Because the useful information appears in the blind zone, and it does not wait for you. Firmware errors, a bootloader prompt, a failed filesystem check, a service hanging on start: these are all shown on the console and only on the console.

If you open it afterwards you may still catch a static error message, and often you will. But if the machine is sitting in a boot loop, or if the message scrolled past, you have lost the one recording of what happened.

There is a second reason that is purely practical. Discovering that your service processor credentials do not work is a five-minute problem before the reboot and a very unpleasant one after it.

25. A Pre-Flight, Run Properly

Worked example

The request: reboot the reporting server, which has been slow since a patch went in on Sunday.

Name the service, not the host. It serves the nightly finance reports and an internal dashboard.

Why: Now you know who notices and when. The dashboard is used all day; the report run starts at 11pm.

Check who is on it. Two interactive sessions, one of them a colleague running a query.

Why: A reboot that kills a colleague's forty-minute query is a self-inflicted incident. Ask, do not assume.

Open the service processor and confirm you can see the console. Test the login now.

Why: This is the gate that has no workaround. Confirming it costs a minute.

Check for work in flight: no backup running, no batch job, replication is current.

Why: Nothing here would be damaged by an interruption, so the cost of waiting is zero and the cost of not checking could have been a rerun.

Check last boot. It last started 41 days ago, and the patch went in on Sunday.

Why: So the patch has never been through a boot. That is exactly the situation where a boot surfaces something, and it is a reason to have the console open rather than a reason not to proceed.

Raise the change, notify the finance team and your colleague, agree a window at 8pm, before the report run.

Why: Before the report run so a failure has three hours of slack in front of it, rather than landing on top of the nightly job.

Figure (svg): Flow diagram of six sequential pre-flight gates before a server reboot, from naming the service through to approval and notification.

The same six gates, this time with real answers written next to each one.

Verify: every gate has a written answer, and the one gate with a worrying answer has a mitigation.

Why: The pre-flight is not a formality that gets ticked. It is done when every gate has an answer you would be willing to read out during an incident review.

26. Three Statements About Change Windows

Two truths and a lie

Two of these are sound. One is the kind of thing that sounds responsible and is not.

Eliminate the wrong options

Which statement survives?

  • w1. A reboot in a maintenance window is safer mainly because fewer people are affected if it goes wrong.
  • w2. A reboot in a maintenance window is safe, because the window is when changes are allowed.
  • w3. The middle of the night is the safest time to reboot, because nobody is using the system.

Survives elimination: w1

Why: The window is about consent and blast radius, not about risk of failure. It grants permission for disruption; it does nothing at all about whether the machine comes back. And the classic overnight reboot cuts both ways: the moment it goes wrong you are alone, tired, and the people who could help you are asleep. Where the service allows it, an early evening change with a colleague reachable beats a 2am change most of the time.

27. Part 3 - Ordering the Reboot

Section

Linux, Windows, and the console when neither answers

28. Linux: Look First, Then Act

Concept

Four commands before, one command to act. The order matters more than the syntax.

uptime            # how long has it been up, and how loaded is it
who               # who else is logged in right now
last reboot       # has it rebooted cleanly before, and how often
systemctl list-units --failed   # anything already broken before you start

That last one is the one people skip and regret. If a service is already failed before you reboot, it will still be failed afterwards, and you will spend an hour believing the reboot broke it.

systemd, systemctl manual page - reboot, halt, poweroff and the shutdown targets systemctl — Unit states and the failed list.

29. Reading a Scheduled Reboot Command

Notation

This is the form I would like you to use by default on Linux, rather than the bare reboot command. Walk the callouts.

Annotate

  • sudo, because ordering a reboot is a privileged action. If this fails you have found out something useful before you did any damage.
  • The r is for restart. Without it you get a halt, and a halted remote server needs a console or a human.
  • Plus five means in five minutes. That delay is not politeness, it is your undo: shutdown -c cancels a pending shutdown, and you cannot cancel one that already happened.
  • The message is broadcast to every logged-in session. It is how the colleague with the forty-minute query finds out in time to say something.

The bare reboot command is fine on a lab box. On something that matters, the delay and the message are worth the extra few characters.

30. Windows: The Same Ideas, Different Spelling

Concept

Windows Server gives you the same three moves, with a reason code system bolted on that is genuinely useful once you are looking back at a year of events.

shutdown /r /t 300 /c "Monthly patching, back by 20:15" /d p:2:4

# or, more readable, and it cannot silently mean shutdown:
Restart-Computer -ComputerName APP-PRD-07 -Force -Wait

# and afterwards, prove who did it and why:
Get-WinEvent -FilterHashtable @{LogName='System'; Id=1074} -MaxEvents 5

Event 1074 records a clean, requested restart along with the account that asked for it. Event 6008 records the opposite: the last shutdown was unexpected. Learning to tell those two apart in a log is worth an hour of anybody's time.

Microsoft Learn, shutdown command reference - flags, timers and reason codes shutdown — Flags, timers and the reason code syntax.

31. Reading the Windows Version

Notation

Same walk, different platform. The reason code at the end is the part people leave off and later wish they had not.

Annotate

  • Slash r is restart. Slash s is shutdown, and it is one key away on the keyboard. This is the whole of the trap from Part 1.
  • Slash t 300 is a three hundred second delay. shutdown /a aborts it while the timer is still running.
  • Slash c is the comment shown to logged-in users and written into the event log.
  • Slash d p colon 2 colon 4 marks it planned, with a major and minor reason code. Six months later this is what tells you whether last March's restarts were planned work or a failing power supply.

Reason codes feel like bureaucracy on day one and feel like a gift the first time you have to explain a pattern of restarts to somebody senior.

32. The Same Move in Two Languages

Translation

You will work on both. Matching them up once saves a lot of second-guessing later.

Match the pairs

  • l1. uptime
  • l2. who
  • l3. shutdown -r +5
  • l4. shutdown -c
  • l5. systemctl list-units --failed
  • l6. last reboot
  • r1. shutdown /r /t 300
  • r2. shutdown /a
  • r3. Get-Service | Where-Object Status -ne 'Running'
  • r4. Get-WinEvent for Event ID 1074
  • r5. query user
  • r6. (Get-CimInstance Win32_OperatingSystem).LastBootUpTime

Why: The vocabulary is different and the concepts are identical: how long has it been up, who is on it, schedule a restart with a delay, cancel that restart, what is already broken, and what does the history say. Learn the six questions and you can look up either platform's spelling in thirty seconds.

33. Drain Before You Take It Away

Concept

A server that is one of several is rarely rebooted directly. You take it out of service first, wait for the work already on it to finish, and only then touch it.

Draining is what turns a reboot from an outage into a non-event. It is also the step most often skipped when somebody is in a hurry, which is why it is worth making it a reflex.

Broadcom VMware, placing an ESXi host into maintenance mode Maintenance mode — What a hypervisor does when you ask it to empty a host.

34. Draining a Node

Picture it

The load balancer stops sending new work first. The reboot happens later, and separately.

Figure (svg): Diagram of a load balancer distributing to three nodes, with node two marked as draining so no new requests are sent to it while the other two remain in rotation.

Draining ends when the last in-flight request finishes, not when you click the button.

The gap between clicking drain and the connection count reaching zero is where long-running requests live. On some systems that is one second. On a reporting service it can be several minutes.

35. The Reboot Pattern, Every Time

Pattern

This is the routine. It is deliberately the same shape as the four alert questions from last session, because both are attempts to stop you acting before you have looked.

  1. Pre-flight. The six gates, with written answers. Console open and tested.
  2. Drain. Take it out of service and wait for zero, or confirm there is nothing to drain.
  3. Capture. Thirty to ninety seconds of state you will want later, because it is about to be destroyed.
  4. Order it. With a delay and a message where the platform allows one.
  5. Watch the console. Not the ping, the console. Through the blind zone.
  6. Verify up the ladder. Then, and only then, put it back in service and close the ticket.

Six phases, and exactly one of them is the command. That ratio is the honest picture of the job.

36. Which Step Is Missing?

Check

Read the change plan carefully before you answer.

Check your understanding

An engineer writes the following plan for rebooting one of four web nodes: take it out of the load balancer, reboot it, wait for it to ping, put it back in the load balancer. Which omission is most likely to cause a customer-visible problem?

  • A. No maintenance window is named
  • B. No wait for connections to drain before rebooting, and no check beyond a ping before returning it to service (correct)
  • C. No service processor console is mentioned
  • D. No backup is taken first

Answer: B

Why: Two errors, both customer-visible, and they bracket the reboot. Removing a node from the pool stops new requests but does not finish the ones already running, so anybody mid-request gets an error. Returning it on a successful ping is worse: the network stack answers long before the web server is listening, so the load balancer starts sending real traffic to a node that will refuse it. The fix is to wait for zero connections going out, and to check the application answers a real request coming back in.

Why A tempts people
A window matters for consent and blast radius, but with four nodes and a correct drain this change is invisible to customers at any hour. It is a process gap rather than the thing that breaks.
Why C tempts people
A console is essential when a node might not return, and it belongs in the plan. But it is a recovery tool: its absence lengthens an outage rather than causing this one.
Why D tempts people
A reboot does not modify data, so a backup adds nothing here. Reaching for a backup before every change is a habit worth having in general and is not the gap in this plan.

37. Rolling a Three-Node Cluster

Worked example

Three database nodes, one of them the current leader. You need all three rebooted for a kernel update, with no outage.

Check cluster health first, and only start if all three members are healthy and in sync.

Why: Rebooting a member of an already-degraded cluster is how a maintenance task becomes an outage. If it is not green before you start, stop and fix that instead.

Start with a follower, never the leader.

Why: Followers can leave and rejoin with no election. Doing the leader first means an unnecessary failover, and it also means you learn nothing before touching the important one.

Reboot follower one. Wait for it to rejoin and to report itself fully caught up, not merely running.

Why: A node that is up but still replaying its replication backlog is not yet redundancy. Treating it as though it is means you take the next node down while you still have only one good copy.

Repeat on follower two, with the same wait.

Why: Same reasoning, and by now you also know exactly how long a boot and a catch-up take on this hardware, which makes the last step predictable.

Hand the leadership over deliberately, then reboot the old leader.

Why: A controlled handover happens when you choose, in a few seconds, with everyone watching. A failover forced by a reboot happens when the cluster notices, which is slower and less pleasant.

Figure (svg): Diagram of a load balancer distributing to three nodes with one marked as draining, illustrating the principle of taking a single member out of service at a time.

The same one-at-a-time discipline, whether the thing in front is a load balancer or a cluster manager.

Verify: all three members healthy, in sync, and the cluster has the leader you expect.

Why: The rolling reboot is done when the cluster is back in the state you started from, with one difference you can name: the new kernel. Anything else that changed is a finding, not a success.

38. What Must Stay True Through a Rolling Reboot

Invariant

Three nodes, one at a time. Watch the number of healthy members, and watch whether a majority still exists.

Step through it

  1. All three up. Two of three is a majority, so the cluster can lose one member and keep deciding things.
  2. One node down. Two remain, which is still a majority, but there is no margin left at all.
  3. The rebooted node has rejoined and caught up. Margin restored, and only now is the next reboot safe.
  4. Second node down, two healthy again for the same reason as before.
  5. The failure mode: taking the second node before the first had truly rejoined leaves one node of three, which is not a majority, and the cluster stops accepting writes.

The invariant is not the node count, it is the majority. That is why the wait between nodes is not optional politeness, it is the whole safety property.

39. Two at Once, to Save Time

Prediction

It is late, the rolling reboot is slow, and rebooting two of the three nodes together would halve the work.

Predict first

What happens to a three-node cluster when two members restart at the same time?

  • Nothing, one node is enough to keep serving
  • It loses majority, and typically stops accepting writes until a second node returns
  • It elects the remaining node leader and continues normally
  • It depends entirely on the application

Correct: It loses majority, and typically stops accepting writes until a second node returns

Why: Clusters that use majority voting need more than half their members to agree before they will commit anything. One node out of three is not a majority, so the survivor cannot safely act as leader, because from its point of view it cannot tell the difference between the other two being down and itself being cut off from them. That ambiguity is exactly what majority rules exist to prevent, and the cost of preventing it is that the cluster goes read-only or stops entirely. Rebooting members one at a time, with a full rejoin in between, is not caution. It is the arithmetic.

40. Host or Guest? Ask Before You Type

Concept

On virtualised infrastructure the same word means two wildly different blast radii, and the two consoles look similar enough to confuse at midnight.

you rebootwhat goes awayhow long
a virtual machineone workloada minute or two
a hypervisor host, not drainedevery virtual machine running on itseveral minutes each, all at once
a hypervisor host, in maintenance modenothing, the guests were moved firstas long as it takes, safely
a storage node in a hyperconverged clusterpossibly the storage for other hosts tooas long as the rebuild takes

Before you reboot anything virtual, say out loud whether the thing you are logged into is the guest or the host. It costs three seconds and it is the difference between a routine change and an incident with your name on it.

41. Which One Are You Rebooting?

Elimination

A virtual machine is unresponsive. You have four ways to act, and only one is right as a first move.

Eliminate the wrong options

Which is the correct first action?

  • e1. Reboot the hypervisor host it runs on.
  • e2. Send a guest operating system restart from the hypervisor console, and give it time to work.
  • e3. Power off the virtual machine immediately from the hypervisor.
  • e4. Restart the management agent on the host.

Survives elimination: e2

Why: The order is always the same: ask before you force, and act on the smallest thing that could be the problem. A guest restart sent from the hypervisor reaches the guest operating system through the virtualisation tools, which is the polite route, and it works even when the network to the guest does not. Give it a couple of minutes before escalating to a forced power off, and do not go near the host until you have a reason that involves the host.

42. Somebody Ran This at 4pm

Error analysis

This was run on a remote production server during business hours. Walk the callouts and count the mistakes.

Annotate

  • The h is halt. The machine stops and stays stopped. On a server in another building with no console, this is a taxi ride.
  • now means no delay and no chance to cancel. There is no shutdown -c to run, because the shutdown already happened.
  • No message, so nobody logged in got any warning at all.
  • Run in the middle of the afternoon, one command deep in an ssh line, which is the classic setup for running it against the wrong host.

The repair is short. Use minus r, give it a delay, write a message, and be looking at the console before you press return.

43. The Command That Worked Perfectly

Trap

The trap

You reboot a server. It goes down. It comes back in ninety seconds, answers a ping, accepts your login, and everything looks normal. You close the ticket.

Two hours later the nightly job fails, because a service that was running before the reboot is not running now. It was started by hand three months ago and was never enabled to start at boot.

The reboot did not break it. The reboot revealed it, and you were the last person to touch the machine, so it is now your problem.

The fix

This is why the failed-units check goes before the reboot, and why the verification ladder goes after it.

# before, and keep the output:
systemctl list-units --type=service --state=running > /tmp/before.txt

# after:
systemctl list-units --type=service --state=running > /tmp/after.txt
diff /tmp/before.txt /tmp/after.txt

Thirty seconds of work, and it converts an argument about causation into a list of exactly what changed. Do the same on Windows by capturing the running service list before and comparing after.

If the difference is a service that should have been enabled, you have found a real latent fault and fixing it is a genuine contribution, not an embarrassment.

44. Part 4 - Proving It Worked

Section

Eight rungs, and the ping is the second one

45. It Answers Is Not It Works

Concept

The single most common way a reboot goes wrong is that somebody declares it finished too early, and the machine then fails in front of a user instead of in front of you.

The reason is timing. A server answers a ping the moment its network stack is up, which is early. The application may need another two minutes to open its files, warm its caches and start listening.

In between those two moments the machine looks alive and is not. Anything that starts sending it real work in that window gets errors.

So verification is a ladder, and each rung proves strictly more than the one below it. You climb until you reach a rung that a user would recognise as working.

46. The Verification Ladder

Picture it

Read from the bottom. Each rung is cheap, and each one rules out a different failure.

Figure (svg): Ladder diagram of eight verification rungs after a reboot, from chassis power and ping at the bottom through login, services, listening ports, a real application request, cluster rejoin, and clean monitoring at the top.

The bottom two rungs are what most people check. The top three are what actually closes a ticket.

You do not always have to climb all eight. You do have to know which rung you stopped on, and say so in the ticket.

47. Climb It In Order

Ranking

Cheapest and least informative first, most expensive and most convincing last.

Put in order

  1. the chassis reports itself powered on
  2. it answers a ping
  3. you can log in
  4. the expected services are running
  5. the application answers a real request
  6. monitoring is green and the log has nothing new in it

Why: Each rung depends on the ones below it and proves something they cannot. Power says the hardware came back. Ping says the kernel and network are up. Login says authentication and the base services work. Services running says the application started. A real request says it actually functions, which is the first rung a user would recognise. Clean monitoring says nothing subtle broke on the way. Notice how far up the ladder you have to go before you have learned anything a customer would care about.

48. Proving It Really Rebooted

Concept

A surprisingly common ending: the reboot never happened. The command was rejected, or aimed at the wrong host, or the machine hung on the way down and somebody power cycled a different one.

uptime -s          # the exact timestamp this machine last booted
last reboot | head -3
journalctl --list-boots | tail -3   # every boot the log remembers

On Windows, the equivalent question is answered by the last boot time and by the system log, where a clean requested restart is Event 1074 and an unexpected loss of power is Event 6008.

Check this before you check anything else. There is no point climbing the ladder on a machine that never went down.

Microsoft Learn, advanced troubleshooting for Windows boot problems Windows boot events — Which event tells you a restart was requested and which tells you it was not.

49. Complete the Post-Reboot Check

Fill the middle

Fill in the three gaps in this verification note, written the way it would appear in a ticket.

Fill in the blanks

Rebooted at 20:05. Confirmed it actually restarted by checking the last boot timestamp. Climbed the ladder: ping, login, services running, and then a real application request which is the first check a user would recognise. Compared the running service list against the one captured before, and the difference was empty, or nothing.

Why: Three habits in one sentence. Prove the reboot happened before you evaluate anything else. Climb past the ping to a check that means something to a user. And compare the service list before against after, so that a service which quietly failed to start is caught by you tonight rather than by the nightly job at 2am.

50. When Can You Return It to Service?

Check

One node of four, behind a load balancer, back up after a reboot.

Check your understanding

Which single check most justifies putting the node back into the load balancer pool?

  • A. It responds to a ping from your workstation
  • B. You can log in over the management network
  • C. A request to the application's health endpoint on the real service port returns success (correct)
  • D. The uptime counter has reset, proving the reboot happened

Answer: C

Why: The load balancer is about to send it real user traffic on the real service port, so the check that justifies that decision is the one that exercises exactly that path. A health endpoint on the service port proves the application started, is listening on the right port, and can answer. It is the first check whose success means a user would succeed too.

Why A tempts people
A ping proves the network stack is up, which happens well before the application is listening. This is precisely the window where returning a node to the pool sends real customers to a server that will refuse them.
Why B tempts people
Logging in proves the operating system and authentication work. It says nothing about whether the application started, and management network access does not exercise the service path at all.
Why D tempts people
Confirming the reboot really happened is a necessary check and you should do it, but it proves only that the machine restarted, not that it came back correctly.

51. Verifying a Web Node, Rung by Rung

Worked example

The node rebooted after ninety seconds. Here is the climb, with what each rung rules out.

Confirm the last boot timestamp is a minute ago, not forty-one days ago.

Why: Rules out the reboot never happening, or having been aimed at a different host. Everything below depends on this being true.

Ping. It answers in under a millisecond.

Why: Rules out a hardware failure to return and a network configuration that did not survive the boot. Proves nothing about the application.

Log in. Confirm the disks that should be mounted are mounted.

Why: A filesystem missing from the boot configuration is one of the two classic post-reboot faults, and it is invisible until something tries to write to it.

Compare the running services against the list captured before the reboot. The difference is empty.

Why: This is the other classic fault: a service running by hand that was never enabled at boot. An empty difference rules it out in one command.

Check the service port is listening, then request the health endpoint over the real service port and read the response body.

Why: Listening proves it bound the port. The health response proves it can actually do work, which is the first rung a user would recognise.

Figure (svg): Ladder diagram of eight verification rungs after a reboot, from chassis power and ping through login, services, listening ports, a real request, cluster rejoin and clean monitoring.

Each step above corresponds to one rung, and each rung rules out a different way the boot could have gone wrong.

Return it to the pool, watch the connection count rise and the error rate stay flat for five minutes.

Why: The last rung. Returning traffic is itself a test, and watching for five minutes catches the failure that only appears under real load.

Verify: the log since boot has nothing new in it, monitoring is green, and the ticket records which rung you stopped on.

Why: A verification that is not written down did not happen, as far as the next person is concerned. Naming the rung you reached is what lets somebody else judge how much your green means.

52. What Goes in the Ticket

Real world

You will write these for the rest of your career, and a good one takes ninety seconds.

Discussion prompt

Write the closing note for a routine reboot. What are the five things it has to contain for it to be useful to the person reading it in six months?

Hint: Imagine you are the person in six months, looking at a pattern of restarts and trying to work out whether they were planned.

Answer:

What and when: the host, the exact time it went down and came back, and how long that took. The duration is the number people compare against next time.

Why: the reason, in a sentence a non-engineer could read. Patching, a memory leak, a vendor instruction.

What you checked before: the pre-flight answers, especially anything that was already broken before you started.

How far you verified: the rung you reached, named. Not the word verified on its own.

Anything that changed unexpectedly: services that did not come back, a filesystem that needed a check, a longer boot than expected. This is the most valuable line in the note and it is the one most often left blank.

53. Part 5 - When It Does Not Come Back

Section

Five minutes, then the console, and no guessing

54. The Five Minute Rule

Concept

Decide the wait before you press the key, and write it down. On known hardware with a known workload, most servers are back inside two or three minutes.

Pick a number, five minutes is a sensible default for a machine you know, and when the number is reached you stop refreshing and you open the console.

The reason for deciding in advance is not discipline for its own sake. It is that the alternative is refreshing a ping for twenty-five minutes, which feels like doing something and is not.

Two things worth knowing before you panic. A firmware update applied at boot can add several minutes with a blank screen. And a filesystem check on a large volume can take a long time and looks exactly like a hang until you read the console, which is one more reason to have it open.

55. It Did Not Come Back

Picture it

One decision, and it splits on a single question: is there anything on the screen?

Figure (svg): Decision diagram for a server that has not responded five minutes after a reboot, splitting on whether the console shows something, which points to firmware, bootloader, filesystem check or a service, or is blank, which points to power and hardware.

The console answers in five seconds a question that a ping cannot answer at all.

Notice what is not on this diagram: retrying the reboot. Doing the same thing again from your desk gives you no new information and costs another five minutes.

56. Ten Minutes, Still Nothing

Prediction

You rebooted a remote server. Ten minutes later it still does not answer a ping. You have service processor access.

Predict first

What is the correct next action?

  • Issue a power cycle from the service processor
  • Open the virtual console and read what is on the screen
  • Raise a hardware ticket with the vendor
  • Wait another ten minutes in case it is slow

Correct: Open the virtual console and read what is on the screen

Why: Look before you act, exactly as with an alert. The console costs nothing, takes five seconds, and answers the only question that matters: is the machine stuck somewhere, or is it dark? A stuck bootloader, a failed filesystem check waiting for a password, a service hanging on start and a firmware update in progress all look identical from the outside and completely different on the console. Power cycling first destroys that information and can make a filesystem check start over from the beginning.

57. The Usual Reasons It Did Not Return

Concept

Almost every failed boot is one of these, and the console tells you which within seconds.

what the console showswhat it meansthe usual cause
firmware or self test messages, then nothingit never reached a boot deviceboot order changed, or a disk not seen
a bootloader prompt or menu, waitingit found the loader but no defaulta failed update, or a missing kernel entry
a filesystem check with a progress figureit is working, leave it alonean unclean previous shutdown
a prompt asking for a password to unlock a volumeit is waiting for a humanan encrypted volume with no automated unlock
a service start hanging with a timer countingthe kernel is fine, one unit is stucka mount that no longer exists, or a dependency loop
a completely blank screen, no output at allit may not be powered at allpower supply, or the chassis never came on

Only the last of these is a hardware problem. The other five are configuration, and four of them were introduced by a person weeks before the reboot that revealed them.

58. Where in the Boot Did It Stop?

Discrimination

Sort each console symptom by which stage of the boot it belongs to. The stage decides who you call.

Sort into buckets

Sort each symptom by the stage it belongs to.

Firmware and hardware
no output at all and no fan noise; a memory self test failure message
Boot device and loader
a message about no bootable device found; a bootloader menu sitting there with no default selected
Kernel and services
the kernel is loading, then a service start hangs; waiting for a device that never appears, then a timeout
hw
The machine has not got as far as looking for an operating system. Nothing you can do in software helps, and this is a hardware call: power, memory, or the board itself.
boot
Firmware worked and handed over, or tried to. The problem is which disk it was told to use or what it found there. Boot order, a disk that has dropped out of the controller, or a loader entry broken by an update.
os
The kernel is running, so the hardware is fine and the disk is readable. Something in the configuration on disk is asking for a thing that is not there. This is the category that single user or safe mode exists for, and it is the most common of the three.

59. When Is a Forced Power Cycle Justified?

Edge cases

The forced verbs are not forbidden. They are just never first.

Discussion prompt

Give two situations where power cycling a server from the service processor is the correct action, and one where it clearly is not.

Hint: Ask what information you still stand to gain by waiting, and what you would destroy by not waiting.

Answer:

Justified: the console is completely unresponsive with no output and no progress, and you have already read whatever was on the screen. There is nothing left to learn and nothing left to protect.

Justified: a hung shutdown that has made no progress for a long time, where the graceful path has already been tried and ignored. At that point the machine is neither up nor down and is helping no one.

Not justified: a filesystem check in progress. It is working, it will finish, and cycling the power sends it back to the start, sometimes repeatedly, which is how a twenty-minute delay becomes a two-hour one.

The general test: what do I still stand to learn by waiting, and what do I destroy by not waiting? When both answers are nothing, force it. Until then, do not.

60. The Power Cycle That Cost Four Hours

Trap

The trap

A server is rebooted and does not return. Ten minutes pass. The engineer opens the console, sees a wall of text they do not recognise, and issues a power cycle.

It does not come back that time either. Another power cycle. And another.

What was on the screen was a filesystem check on a very large volume, at eleven percent. Each power cycle sent it back to zero, and each one added a little more inconsistency for the next check to find.

The fix

Read the console before you act on it, and read it slowly. A progress figure of any kind means the machine is working, and a working machine is left alone.

  • A percentage, a counter, or a growing list of names all mean progress. Wait.
  • The same frozen line for ten minutes with no counter means no progress. Now you can consider forcing it.
  • If you do not recognise what is on the screen, photograph it and ask. That is a two-minute delay, and it is much cheaper than the alternative.

This is the reboot version of the rule from last session: stabilising is not resolving, and acting before you have looked is how a small problem gets larger.

61. When the Shutdown Itself Hangs

Concept

The other direction fails too. You order a reboot and the machine sits there, half down, for a quarter of an hour.

Usually one service refuses to stop, and the operating system is politely waiting out a timeout on its behalf. On Linux the console will name the unit it is waiting for and count the timer down, which is the useful bit.

There is a middle option between waiting and pulling the power, and it is worth knowing about because it is genuinely clever.

Non-maskable interrupt — A signal the service processor can send that the operating system cannot ignore. On a properly configured system it triggers a crash dump: the machine writes the contents of memory to disk and then restarts.

That gives you the one thing a power cycle destroys, which is a record of what the machine was doing at the moment it stopped responding. If a vendor ever asks you to collect a dump from a hung server, this is the mechanism they mean.

62. Part 6 - When a Reboot Is the Wrong Answer

Section

It works, and that is exactly the problem

63. A Reboot Destroys the Evidence

Concept

The uncomfortable truth about rebooting is that it usually works, and its working is what makes it dangerous.

A machine that has run out of memory, or has a process in a strange state, or has thousands of stuck connections, is carrying the entire explanation for its own problem in memory.

The reboot discards all of it. The service recovers, the alert clears, the ticket closes, and the cause is now unknowable. Next month it happens again, and the month after that the reboot becomes a scheduled job, which is how a fault turns into a ritual.

You will still reboot. Frequently the service matters more than the diagnosis, and that is a legitimate call. But you can have both for about ninety seconds of work.

64. The Line the Reboot Draws

Picture it

Everything on the left of that line is available. Everything on the right is gone forever.

Figure (svg): Timeline diagram showing that memory contents, open file handles, the process list, socket state and unflushed log buffers exist before a reboot and are permanently gone after it, while written logs survive.

Written logs survive a reboot. The state that usually contains the actual cause does not.

Ninety seconds of capture, taken before you press the key, is what turns next month's identical incident into a fifteen-minute investigation instead of another reboot.

65. The Ninety Second Capture

Concept

This is the whole habit. It is short on purpose, because a habit you have to think about is a habit you skip at 3am.

date; uptime                       # when, and how loaded
ps aux --sort=-%mem | head -15     # who is holding the memory
free -m                            # how much is actually left
df -h; df -i                       # full disk, and full inode table
ss -s                              # socket totals, for connection exhaustion
dmesg -T | tail -40                # what the kernel has been complaining about
journalctl -p err -S -2h           # errors in the last two hours

Redirect all of it into one file with a name that includes the hostname and the timestamp, and attach it to the ticket. That file is the thing that lets somebody solve this properly later.

The inode check in the middle is worth a mention, because a disk that reports plenty of free space and still refuses to write is one of the classic reboot-does-not-fix-it problems, and this is the command that names it.

66. Reboot Now, or Capture First?

Trade off

Fill the blanks. Neither column is the answer in general, which is the point.

Comparison matrix

ConsiderationReboot immediatelyCapture, then reboot
time to service restoredfastest possibleabout ninety seconds slower
chance of finding the causeclose to zerogood
risk of the same incident recurringhigh, and it willmuch lower, once the cause is fixed
when it is the right calla live outage with users waiting and a clear causealmost every other time, and especially the second occurrence

The honest version of the rule: the first time, restore service and capture what you cheaply can. The second time, capture properly, because you now know it is a pattern rather than an event.

67. A Reboot That Fixed Nothing

Counterexample

A service is restarted nightly by a scheduled job because it slows down after about twenty hours. Everyone agrees this is fine.

Discussion prompt

What is the scheduled reboot actually hiding, and what would you have to measure to prove it?

Hint: Something is growing. The question is what, and whether it grows with time or with work done.

Answer:

It is almost certainly a leak: memory, file handles, database connections or threads that are allocated and never released, so the service degrades in proportion to how much work it has done.

To prove it, measure the resource over the day rather than at a moment. Memory used by the process, open file descriptor count, established connection count, thread count, sampled every few minutes and plotted. A leak is a line that only goes up and resets exactly when the service restarts.

Correlate it against work done rather than against time. If the line tracks requests served rather than the clock, that points at the code path doing the leaking, and the developers can act on it.

The scheduled restart is not wrong as a mitigation. It is wrong as an answer, because it converts a bug into a permanent operational cost and quietly caps how much load the service can ever take in a day.

68. Reboot or Investigate?

Check

Read the whole scenario before choosing.

Check your understanding

A production application server has become unresponsive for the third time this month, always after about three weeks of uptime. Users are affected right now. You have service processor access and a console. What is the best action?

  • A. Reboot immediately, because users are affected and the previous reboots worked
  • B. Investigate fully before touching it, since a third occurrence deserves a proper root cause
  • C. Capture memory, process, handle and log state for about ninety seconds, then reboot, then investigate from the capture (correct)
  • D. Schedule a nightly restart so the three week pattern cannot recur

Answer: C

Why: Users waiting means service restoration is genuinely urgent, so a long investigation while the service is down is the wrong trade. But this is the third identical occurrence with a three-week period, which is the signature of a leak, and rebooting without capturing anything guarantees a fourth. Ninety seconds of capture costs almost nothing against the outage and is the only thing that makes the investigation possible afterwards.

Why A tempts people
This is what produced occurrences one and two, and it will produce occurrence four. Restoring service is right; discarding every piece of evidence on the way is the part that keeps the cycle going.
Why B tempts people
Correct instinct, wrong moment. A full investigation with users down is a choice to extend an outage for information you could have captured in ninety seconds and studied afterwards.
Why D tempts people
This is a mitigation dressed as a fix. It hides the fault permanently, adds a nightly interruption forever, and quietly caps the load the service can carry in a day.

69. How Confident Are You?

Commit first

Commit to an answer and to how sure you are. Being confidently wrong here is the most useful thing that can happen to you tonight.

Predict first

A server has 8 GB of memory. Monitoring shows 7.6 GB used and only 400 MB free, and the application is fine. Is a reboot warranted?

  • Yes, it is nearly out of memory
  • No, unused memory is wasted memory and the figure to watch is available, not free
  • Only if it has been up more than 90 days
  • Only if swap is also in use

Correct: No, unused memory is wasted memory and the figure to watch is available, not free

Why: Operating systems deliberately use spare memory as a disk cache, because idle memory helps nobody. That cache is reclaimable the instant a process needs the space, so it shows as used and is not really occupied. The number that matters is available, which counts the reclaimable cache as free, and on a healthy machine it can be most of the total while free is nearly nothing. Rebooting for this empties a useful cache and makes the machine slower for an hour. The genuine warnings are a shrinking available figure over time, sustained swap activity rather than merely swap being in use, and the kernel log recording that it has had to kill something to get memory back.

70. Your Own Reboot Runbook Card

Connect it up

Same shape as the alert card from last session, and for the same reason: the artefact you write yourself is the one you actually use.

Draw it

Pick one real server from your new job, any one. Write its reboot card on a single side: what it does and who notices, the service processor address and whether you have tested the login, what has to be drained first and how you know draining finished, the exact command you would use, the highest verification rung you would need to reach before returning it to service, and who you would tell before and after. Where you cannot fill a field, write the question you need to ask somebody instead.

Bring the card and the questions next session. As last time, the questions you cannot answer are worth more than the fields you can, because they are a map of what your building has not told you yet.

71. One Question Before You Close

Exit ticket

The idea that runs through all six parts.

Predict first

A server has been up for 900 days and is behaving perfectly. Your manager asks whether it should be rebooted. What is the honest answer?

  • No, it is clearly stable and a reboot only adds risk
  • Yes, immediately, since long uptime is dangerous
  • Yes, but as a planned high-risk change, because the risk is already there and a reboot only reveals it
  • It does not matter either way as long as it is patched

Correct: Yes, but as a planned high-risk change, because the risk is already there and a reboot only reveals it

Why: This is the whole session in one question. The machine is not safe because it has not rebooted; it is untested, which is a different thing, and the risk of a failed boot has been accumulating quietly for 900 days whether or not anyone chooses to look. A power cut will collect that debt eventually, at a time nobody chose, with nobody watching. Choosing the moment yourself means it happens in a window, with the console open, with a colleague awake and with a rollback thought through. The answer is not never and it is not right now. It is soon, deliberately, and treated as the high-risk change it genuinely is.

72. What You Can Do Now

Recap

Eight things, and one habit that carries all of them.

situationthe first move
asked to reboot a server you do not knowname the service it provides, then find the console
it is one of a pooldrain it and wait for the connection count to reach zero
it did not come back after five minutesopen the console and read the screen, do not retry
the console shows a progress figureleave it alone, it is working
the shutdown itself is hungread which unit it is waiting for, then consider a dump before forcing
third identical incident this monthcapture first, reboot second, investigate from the capture

The habit underneath all of it: look before you act, and know what you are destroying when you do. That is the same instruction as last session's four alert questions, wearing different clothes.

Next session we go back to the routine you asked for: five new alert scenarios across power, cooling, network, storage and compute, plus whichever of the two cards you have managed to fill in.

Sources

  1. systemd, systemctl manual page - reboot, halt, poweroff and the shutdown targets
  2. Microsoft Learn, shutdown command reference - flags, timers and reason codes
  3. Microsoft Learn, advanced troubleshooting for Windows boot problems
  4. Intel, IPMI Specification v2.0 - chassis control through the service processor
  5. ITIL 4, Change Enablement - standard, normal and emergency changes
  6. Broadcom VMware, placing an ESXi host into maintenance mode

Want this taught 1-on-1? Alexander tutors Data Center Operations — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108