Every lesson in the Data Center Operations slide course, in full text: 3 decks, 197 slides.
Session 1 - Reading the Plant, Five Alerts, Disaster Recovery and the CloudThe opening session of a data centre operations track for someone new to the field and newly in the job. It builds just enough plant anatomy to make an alert readable - the power chain, hot and cold aisle airflow, and the N, N+1 and 2N redundancy vocabulary - then installs a four-question triage routine and runs it against five real scenarios: a rack thermal alert, a UPS transferred to battery, packet loss on a core uplink, a failed disk mid-rebuild, and a capacity threshold. It closes with disaster recovery built from RTO and RPO outwards, the 3-2-1 rule and recovery site strategies, and what regions, availability zones and the shared responsibility model change once the data centre is somebody else's building. Two traps recur deliberately: stabilising is not resolving, and redundancy is not backup.
Session 2 - Rebooting Servers, Safely and On PurposeThe session the student asked for by name, to get their feet wet with rebooting. It treats a reboot as a controlled loss of state rather than a repair, and spends its time on the minutes either side of the command rather than on the command itself. Part 1 separates four verbs people use as one - restart, shutdown, reset and cold boot - and maps the boot chain, including the blind zone before the kernel that no remote tool can observe. Part 2 installs a six-gate pre-flight and the argument that the risk is not the reboot but the boot, since a long-running server is a pile of changes never tested against a restart. Part 3 covers ordering it on Linux and Windows with annotated commands, draining a node, rolling a three-node cluster without losing quorum, and the host-against-guest question on virtualised infrastructure. Part 4 replaces the ping with an eight-rung verification ladder. Part 5 works the machine that did not come back, from the console. Part 6 makes the case that a reboot destroys the evidence, and gives a ninety-second capture that keeps it. Two ideas recur deliberately: the risk lives in the boot, and a reboot that works is not the same as a problem that is fixed.
Session 3 - Five Alerts, Round Two: Power, Cooling, Time, Cache and MemoryRound two of the recurring alert routine the student asked for, deliberately drawn from a different family than session 1. Five scenarios, one per system, and the organising idea is that not one of them is an outage: a rack PDU branch circuit at 87 percent, which turns out to be a dual-feed arithmetic problem where sixty percent on each side is the exact condition for losing the rack; an air handler failure in a room that is still perfectly cool, which is a redundancy loss measured against a ride-through of minutes; a domain controller four minutes ahead of its reference, fifty seconds from the point where Kerberos begins refusing tickets across 38 servers; a storage controller cache battery failure that silently switches write-back to write-through and multiplies write latency twentyfold with every volume still optimal; and correctable memory errors rising a thousandfold on one module, which is a warning that lets you schedule an outage rather than receive one. It closes on alert storms, finding the parent rather than the loudest child, and the two ways monitoring systems train people to stop reading them. Every scenario is triaged with the same four questions and written on the same six-field card as session 1, and the memory replacement is planned with the reboot discipline from session 2.
Want this taught 1-on-1? Alexander tutors Data Center Operations — $55/session, free consultation.