The session the student asked for by name, to get their feet wet with rebooting. It treats a reboot as a controlled loss of state rather than a repair, and spends its time on the minutes either side of the command rather than on the command itself. Part 1 separates four verbs people use as one - restart, shutdown, reset and cold boot - and maps the boot chain, including the blind zone before the kernel that no remote tool can observe. Part 2 installs a six-gate pre-flight and the argument that the risk is not the reboot but the boot, since a long-running server is a pile of changes never tested against a restart. Part 3 covers ordering it on Linux and Windows with annotated commands, draining a node, rolling a three-node cluster without losing quorum, and the host-against-guest question on virtualised infrastructure. Part 4 replaces the ping with an eight-rung verification ladder. Part 5 works the machine that did not come back, from the console. Part 6 makes the case that a reboot destroys the evidence, and gives a ninety-second capture that keeps it. Two ideas recur deliberately: the risk lives in the boot, and a reboot that works is not the same as a problem that is fixed.
Subject: Data Center Operations · 72 slides · applied lesson
Open the interactive version of this deck · Homework for this lesson
Title
Session 2
What a reboot really is, what has to be true before you order one, how to prove it worked, and what to do when the machine does not come back
Objectives
You asked to park the next five alert scenarios for one session and spend the hour on rebooting instead. That is what this is. The alerts come back next time, and there are five new ones waiting.
Nothing here needs prior experience. It does assume you will stop me the moment a term goes past you, which is what you did last session and it worked well.
Section
A controlled loss of state, not a repair
Warm-up
Answer from instinct first. The instinct is usually half right, and the half that is wrong is the useful part.
Discussion prompt
Rebooting a server fixes an enormous range of problems. In one sentence, what is the common property of every problem a reboot fixes?
Hint: Think about what a reboot removes, not what it restores.
Answer:
A reboot fixes problems that live in accumulated state: leaked memory, a wedged process, a stale handle on a file that was deleted an hour ago, a connection pool that never recovered from a network blip, a driver in a mode it cannot get out of.
It throws away everything the machine has become since it started, and returns it to what the configuration on disk says it should be. That is a real and powerful move.
The flip side follows immediately and it is the whole reason this session exists. If the problem is in the configuration on disk, in the hardware, or in another machine entirely, a reboot cannot touch it, and you will have destroyed the evidence for nothing.
Concept
A reboot is two separate journeys with a gap in the middle, and people worry about the wrong one.
Almost everyone watching a reboot is watching the first journey. Almost every reboot that ruins an evening fails on the third.
Picture it
Trace this once. Every reboot question you will ever be asked lives somewhere on this line.
Figure (svg): Diagram of the boot chain from power on through firmware self test, boot device selection, bootloader, kernel and services, split into a blind zone before the kernel where no remote tool can observe the machine and an observable zone after it.
This is the argument for the service processor, which we get to shortly. It is the only thing that can see the left-hand side.
Ranking
From the moment the command is accepted to the moment the application answers.
Put in order
Why: The two halves are mirror images with the firmware in between. Going down, the operating system stops the things that use the disk and then releases the disk. Coming up, it claims the disk and then starts the things that use it. Knowing which half you are in tells you which tool can still see the machine, which is the practical value of memorising this at all.
Concept
Getting these confused is how a routine change becomes an incident, so it is worth being fussy about the words.
Restart, or reboot — The operating system shuts itself down cleanly and then starts again. Everything gets asked politely to stop first.
Shutdown — The same clean stop, but the machine stays off. Somebody or something has to turn it back on, and in a data centre that somebody may be a long drive away.
Reset, or power cycle — The chassis is forced off and on. Nothing is asked, nothing is flushed. This is the electrical equivalent of pulling the plug, done remotely.
Cold boot — Power removed long enough that the firmware and every device re-initialise from scratch. Occasionally the only thing that clears a wedged network card or a confused storage controller.
systemd, systemctl manual page - reboot, halt, poweroff and the shutdown targets systemctl — The reboot, halt and poweroff targets, and what each one does to running units.
Picture it
The dividing line is one question: does the operating system get a say?
Figure (svg): Table style diagram comparing restart, shutdown, reset or power cycle, and cold boot, showing which of them ask running applications to stop first and which force the chassis off immediately.
You will hear all four called rebooting in conversation. In a change ticket they are four different risks and they need four different words.
Discrimination
Sort each action by whether the operating system is given a chance to stop cleanly.
Sort into buckets
Sort each action into asked or forced.
Comparison
Fill in the missing cells. The last column is the one that decides how you feel at three in the morning.
Comparison matrix
| Action | Are apps asked to stop? | Does it clear firmware and device state? | Comes back on its own? |
|---|---|---|---|
| reboot | yes | no | yes |
| shutdown | yes | no | no, someone must power it on |
| power cycle | no | partly | yes |
| cold boot | no | yes, fully | yes, once power returns |
The shutdown row is the one that catches new people. A clean shutdown of a remote server with no service processor is a change you cannot undo without sending a human to the building.
Socratic
A database server takes four minutes to shut down. A web server takes eight seconds. Same hardware, same operating system.
Discussion prompt
What is the database spending those four minutes on, and why would rushing it be expensive?
Hint: Think about what lives in memory on a database server that does not live in memory on a web server.
Answer:
The database is flushing. Modern databases keep recently changed pages in memory and write them out lazily, because writing to disk on every change would be unbearably slow. A clean shutdown is where that debt gets paid.
It is also finishing or rolling back open transactions, and writing a checkpoint that says where recovery should start from next time.
Force it off in the middle and none of that is lost, exactly, because the write-ahead log is on disk. But the next start has to replay that log to work out what was and was not complete, and on a busy system that recovery can take considerably longer than the clean shutdown you were impatient about.
The general rule is worth keeping: the more state a service holds in memory, the more a clean stop is worth, and the more a forced stop costs you on the way back up.
Trap
You need to restart a Windows server. You are in a hurry, and you type the shutdown command with the flag you half remember.
shutdown /s /t 0The machine goes down cleanly. It never comes back. You are now looking at a monitoring dashboard that says the host is unreachable, and there is nothing left running on it that could bring it back.
On Linux the same trap wears a different hat. The bare shutdown command with no flags does not restart anything either, and on many systems it halts.
Say the verb out loud before you press the key. On Windows the restart flag is r for restart, and slash s is one key away on the same keyboard. On Linux it is minus r for the same reason, and minus h halts.
Better still, use a command that has no ambiguous mode at all, because it can only ever do one thing. On Windows that is Restart-Computer with the force flag. On Linux it is systemctl reboot.
And best of all, have the service processor console open in another tab before you press anything, so that a mistyped flag costs you thirty seconds instead of a drive to the building.
Section
Six gates, any one of which can stop the change
Concept
Almost nothing goes wrong on the way down. The operating system has done that thousands of times and it is good at it.
What goes wrong is the way back up, and it goes wrong for a reason that is almost poetic: a running server is a snapshot of every change anyone has made since it last started, and none of those changes have been tested against a boot.
None of these are visible while the machine is running. All of them surface in the same five minutes.
Picture it
The number people are proudest of is the number that should worry them most.
Figure (svg): Bar chart comparing days since last boot against how much is known about whether the machine will boot successfully, showing confidence falling as uptime rises from seven days to nine hundred.
This is the actual argument for patching on a schedule. It is not mainly about the patches. It is that a machine which reboots monthly is a machine you know will reboot.
Estimation
You are asked to reboot a production server. You check, and it has been up for 812 days.
Predict first
What does that number tell you about tonight?
Correct: Nobody currently employed has seen this machine boot, so the risk is high and unmeasured
Why: Eight hundred days of uptime means eight hundred days of accumulated undocumented change, none of it tested against a restart, and it very likely means the people who built the machine have moved on. Stability while running says nothing at all about whether it comes back. Treat a long-uptime reboot as a high-risk change: do it in a window, with the console already open, with a rollback in mind, and with the person who knows the application awake and reachable.
Concept
This is the pre-flight. It takes five minutes on a machine you know and half an hour on one you do not, and the half hour is the one that pays.
ITIL 4, Change Enablement - standard, normal and emergency changes Change enablement — Standard, normal and emergency changes, and which one a reboot is.
Picture it
Read it as gates rather than as a checklist. A checklist gets ticked. A gate can refuse you.
Figure (svg): Flow diagram of six sequential pre-flight gates before a server reboot, covering the purpose of the server, current users, out of band access, work in flight, whether it has booted since it last changed, and approval and notification.
The third gate is the one I would tattoo on the inside of your eyelids. Everything else on this list is recoverable from your desk. That one decides whether you can recover at all.
Missing information
The ticket reads, in full: Please reboot APP-PRD-07, it is running slow.
Discussion prompt
List everything you would need to know before you are willing to act on this, and say which single missing item would stop you outright.
Hint: Three of the six gates cannot be answered from that sentence at all.
Answer:
Missing: what the server does and who depends on it, whether it is one of a pool or the only one, whether anyone has looked at what running slow actually means, what the change window is, whether there is a service processor and whether you can reach it, and when it last booted.
The item that stops you outright is out-of-band access. Everything else can be discovered or negotiated. Without a console you are betting the service on the boot succeeding, with no move available if it does not.
The second most important omission is the diagnosis. Running slow is not a fault, it is a feeling. If nobody has looked, the reboot will probably appear to work, the problem will return next week, and by then the evidence will have been destroyed twice.
Sorting
Same request, six different contexts. Sort them by what you would actually do.
Sort into buckets
Sort each situation by the right response to a reboot request.
Concept
Every serious server has a small computer inside it whose only job is to manage the big computer. Dell calls it iDRAC, HP calls it iLO, Lenovo calls it XCC, and the underlying standard is IPMI or its successor Redfish.
If you take one habit from this session, take this one: open the console before you order the reboot, not after it fails to come back.
Intel, IPMI Specification v2.0 - chassis control through the service processor IPMI v2.0 — Chassis control commands and the service processor model.
Picture it
Two paths to the same machine, and only one of them survives the operating system.
Figure (svg): Diagram of a server showing an in band path from an administrator login session to the operating system and a separate out of band path from a browser to the service processor, which runs on standby power with its own network port.
In band means through the thing you are about to restart. Out of band means around it. Once you have that distinction, the phrase remote hands starts to make a lot more sense too.
Explain it to yourself
This is the idea I most want to survive tonight, so put it in your own words rather than mine.
Discussion prompt
Why is opening the service processor console before a reboot worth more than opening it after the server fails to come back?
Hint: Think about what the console can show you that nothing else can, and when that information exists.
Answer:
Because the useful information appears in the blind zone, and it does not wait for you. Firmware errors, a bootloader prompt, a failed filesystem check, a service hanging on start: these are all shown on the console and only on the console.
If you open it afterwards you may still catch a static error message, and often you will. But if the machine is sitting in a boot loop, or if the message scrolled past, you have lost the one recording of what happened.
There is a second reason that is purely practical. Discovering that your service processor credentials do not work is a five-minute problem before the reboot and a very unpleasant one after it.
Worked example
The request: reboot the reporting server, which has been slow since a patch went in on Sunday.
Name the service, not the host. It serves the nightly finance reports and an internal dashboard.
Why: Now you know who notices and when. The dashboard is used all day; the report run starts at 11pm.
Check who is on it. Two interactive sessions, one of them a colleague running a query.
Why: A reboot that kills a colleague's forty-minute query is a self-inflicted incident. Ask, do not assume.
Open the service processor and confirm you can see the console. Test the login now.
Why: This is the gate that has no workaround. Confirming it costs a minute.
Check for work in flight: no backup running, no batch job, replication is current.
Why: Nothing here would be damaged by an interruption, so the cost of waiting is zero and the cost of not checking could have been a rerun.
Check last boot. It last started 41 days ago, and the patch went in on Sunday.
Why: So the patch has never been through a boot. That is exactly the situation where a boot surfaces something, and it is a reason to have the console open rather than a reason not to proceed.
Raise the change, notify the finance team and your colleague, agree a window at 8pm, before the report run.
Why: Before the report run so a failure has three hours of slack in front of it, rather than landing on top of the nightly job.
Figure (svg): Flow diagram of six sequential pre-flight gates before a server reboot, from naming the service through to approval and notification.
Verify: every gate has a written answer, and the one gate with a worrying answer has a mitigation.
Why: The pre-flight is not a formality that gets ticked. It is done when every gate has an answer you would be willing to read out during an incident review.
Two truths and a lie
Two of these are sound. One is the kind of thing that sounds responsible and is not.
Eliminate the wrong options
Which statement survives?
Survives elimination: w1
Why: The window is about consent and blast radius, not about risk of failure. It grants permission for disruption; it does nothing at all about whether the machine comes back. And the classic overnight reboot cuts both ways: the moment it goes wrong you are alone, tired, and the people who could help you are asleep. Where the service allows it, an early evening change with a colleague reachable beats a 2am change most of the time.
Section
Linux, Windows, and the console when neither answers
Concept
Four commands before, one command to act. The order matters more than the syntax.
uptime # how long has it been up, and how loaded is it
who # who else is logged in right now
last reboot # has it rebooted cleanly before, and how often
systemctl list-units --failed # anything already broken before you startThat last one is the one people skip and regret. If a service is already failed before you reboot, it will still be failed afterwards, and you will spend an hour believing the reboot broke it.
systemd, systemctl manual page - reboot, halt, poweroff and the shutdown targets systemctl — Unit states and the failed list.
Notation
This is the form I would like you to use by default on Linux, rather than the bare reboot command. Walk the callouts.
Annotate
The bare reboot command is fine on a lab box. On something that matters, the delay and the message are worth the extra few characters.
Concept
Windows Server gives you the same three moves, with a reason code system bolted on that is genuinely useful once you are looking back at a year of events.
shutdown /r /t 300 /c "Monthly patching, back by 20:15" /d p:2:4
# or, more readable, and it cannot silently mean shutdown:
Restart-Computer -ComputerName APP-PRD-07 -Force -Wait
# and afterwards, prove who did it and why:
Get-WinEvent -FilterHashtable @{LogName='System'; Id=1074} -MaxEvents 5Event 1074 records a clean, requested restart along with the account that asked for it. Event 6008 records the opposite: the last shutdown was unexpected. Learning to tell those two apart in a log is worth an hour of anybody's time.
Microsoft Learn, shutdown command reference - flags, timers and reason codes shutdown — Flags, timers and the reason code syntax.
Notation
Same walk, different platform. The reason code at the end is the part people leave off and later wish they had not.
Annotate
Reason codes feel like bureaucracy on day one and feel like a gift the first time you have to explain a pattern of restarts to somebody senior.
Translation
You will work on both. Matching them up once saves a lot of second-guessing later.
Match the pairs
Why: The vocabulary is different and the concepts are identical: how long has it been up, who is on it, schedule a restart with a delay, cancel that restart, what is already broken, and what does the history say. Learn the six questions and you can look up either platform's spelling in thirty seconds.
Concept
A server that is one of several is rarely rebooted directly. You take it out of service first, wait for the work already on it to finish, and only then touch it.
Draining is what turns a reboot from an outage into a non-event. It is also the step most often skipped when somebody is in a hurry, which is why it is worth making it a reflex.
Broadcom VMware, placing an ESXi host into maintenance mode Maintenance mode — What a hypervisor does when you ask it to empty a host.
Picture it
The load balancer stops sending new work first. The reboot happens later, and separately.
Figure (svg): Diagram of a load balancer distributing to three nodes, with node two marked as draining so no new requests are sent to it while the other two remain in rotation.
The gap between clicking drain and the connection count reaching zero is where long-running requests live. On some systems that is one second. On a reporting service it can be several minutes.
Pattern
This is the routine. It is deliberately the same shape as the four alert questions from last session, because both are attempts to stop you acting before you have looked.
Six phases, and exactly one of them is the command. That ratio is the honest picture of the job.
Check
Read the change plan carefully before you answer.
Check your understanding
An engineer writes the following plan for rebooting one of four web nodes: take it out of the load balancer, reboot it, wait for it to ping, put it back in the load balancer. Which omission is most likely to cause a customer-visible problem?
Answer: B
Why: Two errors, both customer-visible, and they bracket the reboot. Removing a node from the pool stops new requests but does not finish the ones already running, so anybody mid-request gets an error. Returning it on a successful ping is worse: the network stack answers long before the web server is listening, so the load balancer starts sending real traffic to a node that will refuse it. The fix is to wait for zero connections going out, and to check the application answers a real request coming back in.
Worked example
Three database nodes, one of them the current leader. You need all three rebooted for a kernel update, with no outage.
Check cluster health first, and only start if all three members are healthy and in sync.
Why: Rebooting a member of an already-degraded cluster is how a maintenance task becomes an outage. If it is not green before you start, stop and fix that instead.
Start with a follower, never the leader.
Why: Followers can leave and rejoin with no election. Doing the leader first means an unnecessary failover, and it also means you learn nothing before touching the important one.
Reboot follower one. Wait for it to rejoin and to report itself fully caught up, not merely running.
Why: A node that is up but still replaying its replication backlog is not yet redundancy. Treating it as though it is means you take the next node down while you still have only one good copy.
Repeat on follower two, with the same wait.
Why: Same reasoning, and by now you also know exactly how long a boot and a catch-up take on this hardware, which makes the last step predictable.
Hand the leadership over deliberately, then reboot the old leader.
Why: A controlled handover happens when you choose, in a few seconds, with everyone watching. A failover forced by a reboot happens when the cluster notices, which is slower and less pleasant.
Figure (svg): Diagram of a load balancer distributing to three nodes with one marked as draining, illustrating the principle of taking a single member out of service at a time.
Verify: all three members healthy, in sync, and the cluster has the leader you expect.
Why: The rolling reboot is done when the cluster is back in the state you started from, with one difference you can name: the new kernel. Anything else that changed is a finding, not a success.
Invariant
Three nodes, one at a time. Watch the number of healthy members, and watch whether a majority still exists.
Step through it
The invariant is not the node count, it is the majority. That is why the wait between nodes is not optional politeness, it is the whole safety property.
Prediction
It is late, the rolling reboot is slow, and rebooting two of the three nodes together would halve the work.
Predict first
What happens to a three-node cluster when two members restart at the same time?
Correct: It loses majority, and typically stops accepting writes until a second node returns
Why: Clusters that use majority voting need more than half their members to agree before they will commit anything. One node out of three is not a majority, so the survivor cannot safely act as leader, because from its point of view it cannot tell the difference between the other two being down and itself being cut off from them. That ambiguity is exactly what majority rules exist to prevent, and the cost of preventing it is that the cluster goes read-only or stops entirely. Rebooting members one at a time, with a full rejoin in between, is not caution. It is the arithmetic.
Concept
On virtualised infrastructure the same word means two wildly different blast radii, and the two consoles look similar enough to confuse at midnight.
| you reboot | what goes away | how long |
|---|---|---|
| a virtual machine | one workload | a minute or two |
| a hypervisor host, not drained | every virtual machine running on it | several minutes each, all at once |
| a hypervisor host, in maintenance mode | nothing, the guests were moved first | as long as it takes, safely |
| a storage node in a hyperconverged cluster | possibly the storage for other hosts too | as long as the rebuild takes |
Before you reboot anything virtual, say out loud whether the thing you are logged into is the guest or the host. It costs three seconds and it is the difference between a routine change and an incident with your name on it.
Elimination
A virtual machine is unresponsive. You have four ways to act, and only one is right as a first move.
Eliminate the wrong options
Which is the correct first action?
Survives elimination: e2
Why: The order is always the same: ask before you force, and act on the smallest thing that could be the problem. A guest restart sent from the hypervisor reaches the guest operating system through the virtualisation tools, which is the polite route, and it works even when the network to the guest does not. Give it a couple of minutes before escalating to a forced power off, and do not go near the host until you have a reason that involves the host.
Error analysis
This was run on a remote production server during business hours. Walk the callouts and count the mistakes.
Annotate
The repair is short. Use minus r, give it a delay, write a message, and be looking at the console before you press return.
Trap
You reboot a server. It goes down. It comes back in ninety seconds, answers a ping, accepts your login, and everything looks normal. You close the ticket.
Two hours later the nightly job fails, because a service that was running before the reboot is not running now. It was started by hand three months ago and was never enabled to start at boot.
The reboot did not break it. The reboot revealed it, and you were the last person to touch the machine, so it is now your problem.
This is why the failed-units check goes before the reboot, and why the verification ladder goes after it.
# before, and keep the output:
systemctl list-units --type=service --state=running > /tmp/before.txt
# after:
systemctl list-units --type=service --state=running > /tmp/after.txt
diff /tmp/before.txt /tmp/after.txtThirty seconds of work, and it converts an argument about causation into a list of exactly what changed. Do the same on Windows by capturing the running service list before and comparing after.
If the difference is a service that should have been enabled, you have found a real latent fault and fixing it is a genuine contribution, not an embarrassment.
Section
Eight rungs, and the ping is the second one
Concept
The single most common way a reboot goes wrong is that somebody declares it finished too early, and the machine then fails in front of a user instead of in front of you.
The reason is timing. A server answers a ping the moment its network stack is up, which is early. The application may need another two minutes to open its files, warm its caches and start listening.
In between those two moments the machine looks alive and is not. Anything that starts sending it real work in that window gets errors.
So verification is a ladder, and each rung proves strictly more than the one below it. You climb until you reach a rung that a user would recognise as working.
Picture it
Read from the bottom. Each rung is cheap, and each one rules out a different failure.
Figure (svg): Ladder diagram of eight verification rungs after a reboot, from chassis power and ping at the bottom through login, services, listening ports, a real application request, cluster rejoin, and clean monitoring at the top.
You do not always have to climb all eight. You do have to know which rung you stopped on, and say so in the ticket.
Ranking
Cheapest and least informative first, most expensive and most convincing last.
Put in order
Why: Each rung depends on the ones below it and proves something they cannot. Power says the hardware came back. Ping says the kernel and network are up. Login says authentication and the base services work. Services running says the application started. A real request says it actually functions, which is the first rung a user would recognise. Clean monitoring says nothing subtle broke on the way. Notice how far up the ladder you have to go before you have learned anything a customer would care about.
Concept
A surprisingly common ending: the reboot never happened. The command was rejected, or aimed at the wrong host, or the machine hung on the way down and somebody power cycled a different one.
uptime -s # the exact timestamp this machine last booted
last reboot | head -3
journalctl --list-boots | tail -3 # every boot the log remembersOn Windows, the equivalent question is answered by the last boot time and by the system log, where a clean requested restart is Event 1074 and an unexpected loss of power is Event 6008.
Check this before you check anything else. There is no point climbing the ladder on a machine that never went down.
Microsoft Learn, advanced troubleshooting for Windows boot problems Windows boot events — Which event tells you a restart was requested and which tells you it was not.
Fill the middle
Fill in the three gaps in this verification note, written the way it would appear in a ticket.
Fill in the blanks
Rebooted at 20:05. Confirmed it actually restarted by checking the last boot timestamp. Climbed the ladder: ping, login, services running, and then a real application request which is the first check a user would recognise. Compared the running service list against the one captured before, and the difference was empty, or nothing.
Why: Three habits in one sentence. Prove the reboot happened before you evaluate anything else. Climb past the ping to a check that means something to a user. And compare the service list before against after, so that a service which quietly failed to start is caught by you tonight rather than by the nightly job at 2am.
Check
One node of four, behind a load balancer, back up after a reboot.
Check your understanding
Which single check most justifies putting the node back into the load balancer pool?
Answer: C
Why: The load balancer is about to send it real user traffic on the real service port, so the check that justifies that decision is the one that exercises exactly that path. A health endpoint on the service port proves the application started, is listening on the right port, and can answer. It is the first check whose success means a user would succeed too.
Worked example
The node rebooted after ninety seconds. Here is the climb, with what each rung rules out.
Confirm the last boot timestamp is a minute ago, not forty-one days ago.
Why: Rules out the reboot never happening, or having been aimed at a different host. Everything below depends on this being true.
Ping. It answers in under a millisecond.
Why: Rules out a hardware failure to return and a network configuration that did not survive the boot. Proves nothing about the application.
Log in. Confirm the disks that should be mounted are mounted.
Why: A filesystem missing from the boot configuration is one of the two classic post-reboot faults, and it is invisible until something tries to write to it.
Compare the running services against the list captured before the reboot. The difference is empty.
Why: This is the other classic fault: a service running by hand that was never enabled at boot. An empty difference rules it out in one command.
Check the service port is listening, then request the health endpoint over the real service port and read the response body.
Why: Listening proves it bound the port. The health response proves it can actually do work, which is the first rung a user would recognise.
Figure (svg): Ladder diagram of eight verification rungs after a reboot, from chassis power and ping through login, services, listening ports, a real request, cluster rejoin and clean monitoring.
Return it to the pool, watch the connection count rise and the error rate stay flat for five minutes.
Why: The last rung. Returning traffic is itself a test, and watching for five minutes catches the failure that only appears under real load.
Verify: the log since boot has nothing new in it, monitoring is green, and the ticket records which rung you stopped on.
Why: A verification that is not written down did not happen, as far as the next person is concerned. Naming the rung you reached is what lets somebody else judge how much your green means.
Real world
You will write these for the rest of your career, and a good one takes ninety seconds.
Discussion prompt
Write the closing note for a routine reboot. What are the five things it has to contain for it to be useful to the person reading it in six months?
Hint: Imagine you are the person in six months, looking at a pattern of restarts and trying to work out whether they were planned.
Answer:
What and when: the host, the exact time it went down and came back, and how long that took. The duration is the number people compare against next time.
Why: the reason, in a sentence a non-engineer could read. Patching, a memory leak, a vendor instruction.
What you checked before: the pre-flight answers, especially anything that was already broken before you started.
How far you verified: the rung you reached, named. Not the word verified on its own.
Anything that changed unexpectedly: services that did not come back, a filesystem that needed a check, a longer boot than expected. This is the most valuable line in the note and it is the one most often left blank.
Section
Five minutes, then the console, and no guessing
Concept
Decide the wait before you press the key, and write it down. On known hardware with a known workload, most servers are back inside two or three minutes.
Pick a number, five minutes is a sensible default for a machine you know, and when the number is reached you stop refreshing and you open the console.
The reason for deciding in advance is not discipline for its own sake. It is that the alternative is refreshing a ping for twenty-five minutes, which feels like doing something and is not.
Two things worth knowing before you panic. A firmware update applied at boot can add several minutes with a blank screen. And a filesystem check on a large volume can take a long time and looks exactly like a hang until you read the console, which is one more reason to have it open.
Picture it
One decision, and it splits on a single question: is there anything on the screen?
Figure (svg): Decision diagram for a server that has not responded five minutes after a reboot, splitting on whether the console shows something, which points to firmware, bootloader, filesystem check or a service, or is blank, which points to power and hardware.
Notice what is not on this diagram: retrying the reboot. Doing the same thing again from your desk gives you no new information and costs another five minutes.
Prediction
You rebooted a remote server. Ten minutes later it still does not answer a ping. You have service processor access.
Predict first
What is the correct next action?
Correct: Open the virtual console and read what is on the screen
Why: Look before you act, exactly as with an alert. The console costs nothing, takes five seconds, and answers the only question that matters: is the machine stuck somewhere, or is it dark? A stuck bootloader, a failed filesystem check waiting for a password, a service hanging on start and a firmware update in progress all look identical from the outside and completely different on the console. Power cycling first destroys that information and can make a filesystem check start over from the beginning.
Concept
Almost every failed boot is one of these, and the console tells you which within seconds.
| what the console shows | what it means | the usual cause |
|---|---|---|
| firmware or self test messages, then nothing | it never reached a boot device | boot order changed, or a disk not seen |
| a bootloader prompt or menu, waiting | it found the loader but no default | a failed update, or a missing kernel entry |
| a filesystem check with a progress figure | it is working, leave it alone | an unclean previous shutdown |
| a prompt asking for a password to unlock a volume | it is waiting for a human | an encrypted volume with no automated unlock |
| a service start hanging with a timer counting | the kernel is fine, one unit is stuck | a mount that no longer exists, or a dependency loop |
| a completely blank screen, no output at all | it may not be powered at all | power supply, or the chassis never came on |
Only the last of these is a hardware problem. The other five are configuration, and four of them were introduced by a person weeks before the reboot that revealed them.
Discrimination
Sort each console symptom by which stage of the boot it belongs to. The stage decides who you call.
Sort into buckets
Sort each symptom by the stage it belongs to.
Edge cases
The forced verbs are not forbidden. They are just never first.
Discussion prompt
Give two situations where power cycling a server from the service processor is the correct action, and one where it clearly is not.
Hint: Ask what information you still stand to gain by waiting, and what you would destroy by not waiting.
Answer:
Justified: the console is completely unresponsive with no output and no progress, and you have already read whatever was on the screen. There is nothing left to learn and nothing left to protect.
Justified: a hung shutdown that has made no progress for a long time, where the graceful path has already been tried and ignored. At that point the machine is neither up nor down and is helping no one.
Not justified: a filesystem check in progress. It is working, it will finish, and cycling the power sends it back to the start, sometimes repeatedly, which is how a twenty-minute delay becomes a two-hour one.
The general test: what do I still stand to learn by waiting, and what do I destroy by not waiting? When both answers are nothing, force it. Until then, do not.
Trap
A server is rebooted and does not return. Ten minutes pass. The engineer opens the console, sees a wall of text they do not recognise, and issues a power cycle.
It does not come back that time either. Another power cycle. And another.
What was on the screen was a filesystem check on a very large volume, at eleven percent. Each power cycle sent it back to zero, and each one added a little more inconsistency for the next check to find.
Read the console before you act on it, and read it slowly. A progress figure of any kind means the machine is working, and a working machine is left alone.
This is the reboot version of the rule from last session: stabilising is not resolving, and acting before you have looked is how a small problem gets larger.
Concept
The other direction fails too. You order a reboot and the machine sits there, half down, for a quarter of an hour.
Usually one service refuses to stop, and the operating system is politely waiting out a timeout on its behalf. On Linux the console will name the unit it is waiting for and count the timer down, which is the useful bit.
There is a middle option between waiting and pulling the power, and it is worth knowing about because it is genuinely clever.
Non-maskable interrupt — A signal the service processor can send that the operating system cannot ignore. On a properly configured system it triggers a crash dump: the machine writes the contents of memory to disk and then restarts.
That gives you the one thing a power cycle destroys, which is a record of what the machine was doing at the moment it stopped responding. If a vendor ever asks you to collect a dump from a hung server, this is the mechanism they mean.
Section
It works, and that is exactly the problem
Concept
The uncomfortable truth about rebooting is that it usually works, and its working is what makes it dangerous.
A machine that has run out of memory, or has a process in a strange state, or has thousands of stuck connections, is carrying the entire explanation for its own problem in memory.
The reboot discards all of it. The service recovers, the alert clears, the ticket closes, and the cause is now unknowable. Next month it happens again, and the month after that the reboot becomes a scheduled job, which is how a fault turns into a ritual.
You will still reboot. Frequently the service matters more than the diagnosis, and that is a legitimate call. But you can have both for about ninety seconds of work.
Picture it
Everything on the left of that line is available. Everything on the right is gone forever.
Figure (svg): Timeline diagram showing that memory contents, open file handles, the process list, socket state and unflushed log buffers exist before a reboot and are permanently gone after it, while written logs survive.
Ninety seconds of capture, taken before you press the key, is what turns next month's identical incident into a fifteen-minute investigation instead of another reboot.
Concept
This is the whole habit. It is short on purpose, because a habit you have to think about is a habit you skip at 3am.
date; uptime # when, and how loaded
ps aux --sort=-%mem | head -15 # who is holding the memory
free -m # how much is actually left
df -h; df -i # full disk, and full inode table
ss -s # socket totals, for connection exhaustion
dmesg -T | tail -40 # what the kernel has been complaining about
journalctl -p err -S -2h # errors in the last two hoursRedirect all of it into one file with a name that includes the hostname and the timestamp, and attach it to the ticket. That file is the thing that lets somebody solve this properly later.
The inode check in the middle is worth a mention, because a disk that reports plenty of free space and still refuses to write is one of the classic reboot-does-not-fix-it problems, and this is the command that names it.
Trade off
Fill the blanks. Neither column is the answer in general, which is the point.
Comparison matrix
| Consideration | Reboot immediately | Capture, then reboot |
|---|---|---|
| time to service restored | fastest possible | about ninety seconds slower |
| chance of finding the cause | close to zero | good |
| risk of the same incident recurring | high, and it will | much lower, once the cause is fixed |
| when it is the right call | a live outage with users waiting and a clear cause | almost every other time, and especially the second occurrence |
The honest version of the rule: the first time, restore service and capture what you cheaply can. The second time, capture properly, because you now know it is a pattern rather than an event.
Counterexample
A service is restarted nightly by a scheduled job because it slows down after about twenty hours. Everyone agrees this is fine.
Discussion prompt
What is the scheduled reboot actually hiding, and what would you have to measure to prove it?
Hint: Something is growing. The question is what, and whether it grows with time or with work done.
Answer:
It is almost certainly a leak: memory, file handles, database connections or threads that are allocated and never released, so the service degrades in proportion to how much work it has done.
To prove it, measure the resource over the day rather than at a moment. Memory used by the process, open file descriptor count, established connection count, thread count, sampled every few minutes and plotted. A leak is a line that only goes up and resets exactly when the service restarts.
Correlate it against work done rather than against time. If the line tracks requests served rather than the clock, that points at the code path doing the leaking, and the developers can act on it.
The scheduled restart is not wrong as a mitigation. It is wrong as an answer, because it converts a bug into a permanent operational cost and quietly caps how much load the service can ever take in a day.
Check
Read the whole scenario before choosing.
Check your understanding
A production application server has become unresponsive for the third time this month, always after about three weeks of uptime. Users are affected right now. You have service processor access and a console. What is the best action?
Answer: C
Why: Users waiting means service restoration is genuinely urgent, so a long investigation while the service is down is the wrong trade. But this is the third identical occurrence with a three-week period, which is the signature of a leak, and rebooting without capturing anything guarantees a fourth. Ninety seconds of capture costs almost nothing against the outage and is the only thing that makes the investigation possible afterwards.
Commit first
Commit to an answer and to how sure you are. Being confidently wrong here is the most useful thing that can happen to you tonight.
Predict first
A server has 8 GB of memory. Monitoring shows 7.6 GB used and only 400 MB free, and the application is fine. Is a reboot warranted?
Correct: No, unused memory is wasted memory and the figure to watch is available, not free
Why: Operating systems deliberately use spare memory as a disk cache, because idle memory helps nobody. That cache is reclaimable the instant a process needs the space, so it shows as used and is not really occupied. The number that matters is available, which counts the reclaimable cache as free, and on a healthy machine it can be most of the total while free is nearly nothing. Rebooting for this empties a useful cache and makes the machine slower for an hour. The genuine warnings are a shrinking available figure over time, sustained swap activity rather than merely swap being in use, and the kernel log recording that it has had to kill something to get memory back.
Connect it up
Same shape as the alert card from last session, and for the same reason: the artefact you write yourself is the one you actually use.
Draw it
Pick one real server from your new job, any one. Write its reboot card on a single side: what it does and who notices, the service processor address and whether you have tested the login, what has to be drained first and how you know draining finished, the exact command you would use, the highest verification rung you would need to reach before returning it to service, and who you would tell before and after. Where you cannot fill a field, write the question you need to ask somebody instead.
Bring the card and the questions next session. As last time, the questions you cannot answer are worth more than the fields you can, because they are a map of what your building has not told you yet.
Exit ticket
The idea that runs through all six parts.
Predict first
A server has been up for 900 days and is behaving perfectly. Your manager asks whether it should be rebooted. What is the honest answer?
Correct: Yes, but as a planned high-risk change, because the risk is already there and a reboot only reveals it
Why: This is the whole session in one question. The machine is not safe because it has not rebooted; it is untested, which is a different thing, and the risk of a failed boot has been accumulating quietly for 900 days whether or not anyone chooses to look. A power cut will collect that debt eventually, at a time nobody chose, with nobody watching. Choosing the moment yourself means it happens in a window, with the console open, with a colleague awake and with a rollback thought through. The answer is not never and it is not right now. It is soon, deliberately, and treated as the high-risk change it genuinely is.
Recap
Eight things, and one habit that carries all of them.
| situation | the first move |
|---|---|
| asked to reboot a server you do not know | name the service it provides, then find the console |
| it is one of a pool | drain it and wait for the connection count to reach zero |
| it did not come back after five minutes | open the console and read the screen, do not retry |
| the console shows a progress figure | leave it alone, it is working |
| the shutdown itself is hung | read which unit it is waiting for, then consider a dump before forcing |
| third identical incident this month | capture first, reboot second, investigate from the capture |
The habit underneath all of it: look before you act, and know what you are destroying when you do. That is the same instruction as last session's four alert questions, wearing different clothes.
Next session we go back to the routine you asked for: five new alert scenarios across power, cooling, network, storage and compute, plus whichever of the two cards you have managed to fill in.
Want this taught 1-on-1? Alexander tutors Data Center Operations — $55/session, free consultation.