The most important distribution in the book, and the foundation of everything from chapter 7 onward. A normal distribution is bell shaped and symmetric about its mean, and it is fixed by exactly two numbers: the mean, which slides the curve left or right, and the standard deviation, which makes it fatter or skinnier. Since those two may be anything, there are infinitely many normal distributions, and one of special interest is the standard normal, whose mean is zero and standard deviation one. The z-score converts any normal value into a standard one by measuring its distance from the mean in units of the standard deviation, which both reduces every normal distribution to a single one and makes values on completely different scales directly comparable. The Empirical Rule then reads off the areas that matter most: about 68, 95 and 99.7 percent of values lie within one, two and three standard deviations of the mean.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 6 — The Normal Distribution
The Standard Normal Distribution
Objectives
Five outcomes, and the third is the one that makes the normal usable at all.
OpenStax Introductory Statistics 2e, §6.1 The Standard Normal Distribution §6.1, pp. 336-339 — the section these objectives are drawn from
Warm-up
Section 2.7 introduced the z-score for a data set, and chapter 5 established that probability is area under a density.
Discussion prompt
A statistics exam has mean 63 and standard deviation 5; a history exam has mean 78 and standard deviation 12. You scored 68 on statistics and 90 on history. Which was the better performance?
Hint: The raw scores are on different scales, so comparing 68 with 90 answers nothing.
Answer:
On statistics you are 5 points above a mean of 63, which is exactly one standard deviation up. On history you are 12 points above a mean of 78, which is also exactly one standard deviation up.
So the two performances are identical in the only sense that can be compared: both sit one standard deviation above their own means, even though the raw gaps are 5 points and 12 points. The number that captured this is the z-score of section 2.7.
This section does two things with that idea. It applies it to a whole distribution rather than a data set, and it notes that standardising turns every normal distribution into the same one — which is what makes a single table, or a single calculator command, able to answer questions about all of them.
Concept
The normal distribution has two numerical descriptive measures: the mean and the standard deviation, and we write X follows N(mu, sigma). The curve is symmetric about a vertical line drawn through the mean, so in theory the mean is the same as the median. Since there are infinitely many such distributions, one of special interest is the standard normal, whose mean is zero and standard deviation one.
the standard normal distribution — A normal distribution of standardized values called z-scores, with mean zero and standard deviation one. Written Z follows N(0, 1). Any normal variable becomes one by the transformation z equals x minus mu, over sigma.
\[ X \sim N(\mu, \sigma) \;\xrightarrow{\;z = (x-\mu)/\sigma\;}\; Z \sim N(0, 1) \]
The book is unusually direct about the density function itself: it is a rather complicated mathematical function, and do not memorize it — it is not necessary. That is section 5.1's point about calculus made concrete. The cumulative distribution function is calculated by a calculator or a computer, or looked up in a table, and the book adds that technology has made the tables virtually obsolete. What you need is the two parameters and the z-score, not the formula.
Figure (svg): Two panels, the left showing two bell curves shifted apart by their means and the right showing two of different widths sharing a mean
OpenStax Introductory Statistics 2e, §6.1 The Standard Normal Distribution §6.1, p. 336
Section
Section 1
Concept
A change in the mean causes the graph to shift to the left or right. A change in the standard deviation causes a change in the shape of the curve, which becomes fatter or skinnier. Since the area under the curve must equal one, those two effects are all that can happen.
N(mu, sigma) — The book's notation, in which the second entry is the standard DEVIATION rather than the variance. Some other texts write the variance there, so the convention has to be checked whenever a normal is quoted from elsewhere.
\[ \text{mean} \to \text{location}, \qquad \text{standard deviation} \to \text{spread} \]
The constraint that the total area stays one is what links the two effects: a curve squeezed narrower has to grow taller to keep its area, which is exactly the trade-off met for the uniform in section 5.2. It also means the height of a normal density is not a free choice, so the two parameters really do determine everything about the curve.
Figure (svg): Two panels, the left showing two bell curves shifted apart by their means and the right showing two of different widths sharing a mean
OpenStax Introductory Statistics 2e, §6.1 The Standard Normal Distribution §6.1, p. 336 — the chapter introduction, on the same page as the section heading
Picture it
The same family, with one parameter changed at a time.
Figure (svg): Two panels, the left showing two bell curves shifted apart by their means and the right showing two of different widths sharing a mean
It is worth noticing that neither change makes the curve asymmetric. Every normal distribution is symmetric about its mean, so the mean and median always coincide — which is why the normal is the reference point against which section 2.6's skewness was described.
Worked example
Example 6.1's opening line.
\[ X \sim N(5, 6) \]
Read the first entry
Why: The mean.
\[ \mu = 5 \]
Read the second
Why: The standard deviation.
\[ \sigma = 6 \]
Locate the centre
Why: The curve is symmetric about it.
\[ \text{centred at } 5 \]
Note the median
Why: Symmetry forces it.
\[ \text{also } 5 \]
Figure (svg): The solution to Worked example reading the notation shown as a ladder of expressions, one row per legal move
\[ \mu = 5, \quad \sigma = 6, \quad \text{median} = 5 \]
Verify: confirm the second entry is a standard deviation, not a variance
Why: Here the spread of 6 is larger than the mean of 5, which is perfectly possible and produces a curve with substantial area below zero. Some texts write N(mu, sigma squared) with the variance second, so N(5, 6) would mean a standard deviation of about 2.45 for them. Always check which convention a source uses before computing a z-score from it — this book's second entry is always the standard deviation.
OpenStax Introductory Statistics 2e, §6.1 The Standard Normal Distribution §6.1, p. 336
Fill the middle
The book's notation for a normal distribution.
Fill in the blanks
X \sim N(\mu, \sigma), \textstandard deviation ___
Why: The standard deviation. The book writes the two numerical descriptive measures as the mean and the standard deviation, so N(5, 6) has a spread of 6 rather than of the square root of 6.
Worked example
Comparing two eras of Chilean heights.
\[ X \sim N(170, 6.28) \text{ against } Y \sim N(172.36, 6.34) \]
Compare the means
Why: 172.36 against 170.
\[ Y\text{ sits } 2.36 \text{cm}\text{ right} \]
Compare the spreads
Why: 6.34 against 6.28.
Say what the curves look like
Why: Nearly the same shape.
Note what is preserved
Why: Both symmetric.
Figure (svg): The solution to Worked example what changes when a parameter changes shown as a ladder of expressions, one row per legal move
\[ \mu_Y - \mu_X = 2.36 \text{ cm}, \quad \sigma_Y - \sigma_X = 0.06 \text{ cm} \]
Verify: confirm the shift is small relative to the spread
Why: A difference of 2.36 cm against standard deviations of about 6.3 cm is well under half a standard deviation, so the two distributions overlap heavily and most individual heights are uninformative about which era they came from. Comparing a difference in means against the spread — rather than looking at the difference alone — is the habit that chapters 9 and 10 formalise into a hypothesis test.
OpenStax Introductory Statistics 2e, §6.1 The Standard Normal Distribution §6.1, pp. 338-339
Trap
\[ X \sim N(5, 6) \;\Rightarrow\; \sigma = \sqrt{6} \approx 2.45 \]
Take the second entry to be the variance
Why: Many textbooks and software packages do write it that way.
\[ z = \frac{17 - 5}{2.45} \approx 4.9 \quad \text{instead of } 2 \]
Every z-score, probability and percentile computed afterwards is wrong, and none of them looks obviously wrong.
\[ X \sim N(5, 6) \;\Rightarrow\; \sigma = 6, \quad z = \frac{17-5}{6} = 2 \]
In this book the second entry is always the standard deviation
Why: The book states it in the notation itself.
This is a genuine ambiguity in the wider literature rather than a mistake anyone should feel bad about, and the defence is to check the convention whenever a distribution is quoted from outside the book. A quick tell: if the two entries are described as mean and standard deviation in the surrounding sentence, as they always are here, the second is a standard deviation.
Prediction
Commit before reasoning.
Predict first
N(50, 5) is replaced by N(50, 15). What happens to the curve?
Correct: It becomes fatter and shorter, still centred at 50.
Why: The standard deviation controls spread, so tripling it spreads the values wider — and since the total area must stay one, a wider curve must be shorter. The centre is unmoved because the mean is unchanged. Getting the height direction right matters: fatter always means shorter for a density.
Two truths and a lie
All three concern the parameters.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and the book says so outright: the density is a rather complicated mathematical function, and do not memorize it — it is not necessary. Areas come from a calculator, a computer or a table, so the two parameters and the z-score are all that is required.
Sorting
Compare each pair against N(50, 10).
Sort into buckets
Sort each distribution by what changed.
Item (e) is a deliberate no-change, sorted with the width column because nothing moved. It is worth a moment: two normals are the same distribution only when BOTH parameters match, which is the fact chapters 10 and 13 lean on when comparing groups.
Section
Section 2
Concept
If X is normally distributed with mean mu and standard deviation sigma, the z-score is x minus mu, all over sigma. It tells you how many standard deviations the value x is above or below the mean. Values larger than the mean have positive z-scores and values smaller have negative ones, and a value equal to the mean has a z-score of zero.
z-score — A value measured in units of the standard deviation. Its magnitude gives the distance from the mean and its sign gives the direction: positive is to the right of the mean, negative to the left.
\[ z = \frac{x - \mu}{\sigma}, \qquad x = \mu + z\sigma \]
The two forms are the same statement solved for different unknowns, and both get used constantly from here on. The book demonstrates the check explicitly after Example 6.1: having found that x equals 17 has a z-score of 2, it verifies that 5 plus 2 times 6 gives 17 back. Doing that substitution the other way round is the cheapest possible guard against a sign error or a division that went the wrong way.
Figure (svg): A card giving the z-score formula with its direction and units explained
OpenStax Introductory Statistics 2e, §6.1 The Standard Normal Distribution §6.1, pp. 336-337 — the z-score definition and Example 6.1
Picture it
A numerator that measures distance, a denominator that sets the units, and a sign that gives direction.
Figure (svg): A card giving the z-score formula with its direction and units explained
Reading the sign as direction rather than as an arithmetic accident is worth insisting on. A z-score of minus 1.5 is not a smaller number than 1.5 in any useful sense — it is the same distance from the mean, on the other side, which is exactly what Example 6.4 needs.
Worked example
Example 6.1, with a value above the mean and one below.
\[ X \sim N(5, 6); \text{ find } z \text{ for } x = 17 \text{ and } x = 1 \]
Subtract the mean from 17
Why: Seventeen minus five.
\[ 12 \]
Divide by sigma
Why: Twelve over six.
\[ z = 2 \]
Subtract the mean from 1
Why: One minus five.
\[ -4 \]
Divide by sigma
Why: Minus four over six.
\[ z = -0.67 \]
Figure (svg): The solution to Worked example two z-scores from one distribution shown as a ladder of expressions, one row per legal move
\[ z = \frac{17-5}{6} = 2, \qquad z = \frac{1-5}{6} \approx -0.67 \]
Verify: confirm both by rebuilding the original values
Why: Going back the other way, 5 plus 2 times 6 gives 17, and 5 plus minus 0.67 times 6 gives about 0.98, which rounds to the 1 we started from — the book performs both of these checks. The small discrepancy in the second is purely the rounding of minus 0.67 to two decimals, and noticing that it is rounding rather than error is itself part of reading the check correctly.
OpenStax Introductory Statistics 2e, §6.1 The Standard Normal Distribution §6.1, pp. 336-337
Faded example
Jerome averages 16 points a game with a standard deviation of four. He scores 10.
Fill in the blanks
z = \frac4-1.5} = ___
Why: Minus six over four is minus 1.5, so ten points is 1.5 standard deviations to the LEFT of his mean of 16 — a below-average game.
Worked example
Example 6.3(b), where the z-score is given and the value is wanted.
\[ X \sim N(170, 6.28); \text{ a height has } z = 1.27 \]
Choose the right form
Why: The value is unknown.
\[ x = \mu + z \sigma \]
Substitute
Why: 1.27 standard deviations up.
\[ 170 + (1.27) (6.28) \]
Multiply
Why: The distance in centimetres.
\[ 7.98 \]
Add
Why: To the mean.
\[ 177.98 \text{cm} \]
Figure (svg): The solution to Worked example recovering a height from a z-score shown as a ladder of expressions, one row per legal move
\[ x = 170 + (1.27)(6.28) \approx 177.98 \text{ cm} \]
Verify: confirm the direction against the sign of z
Why: The z-score is positive, so the height must exceed the mean of 170 — and 177.98 does. A negative z-score would have required a height below 170, and getting a result on the wrong side of the mean is the clearest possible signal that the sign was mishandled. Checking direction before checking arithmetic catches this class of error immediately.
OpenStax Introductory Statistics 2e, §6.1 The Standard Normal Distribution §6.1, p. 338
Error analysis
The book's answer is minus 0.32.
Annotate
On: \( \begin{aligned} &(1)\; \frac{170 - 168}{6.28} = +0.32 \\ &(2)\; \frac{168 - 170}{170} \approx -0.012 \\ &(3)\; 168 - 170 = -2 \\ &(4)\; \frac{168 - 170}{6.28} \approx -0.32 \end{aligned} \)
Error (1) is the dangerous one because its magnitude is right and only the sign is wrong, so it survives any check that ignores direction. Asking whether x is above or below the mean BEFORE computing, and then confirming the sign matches, is the guard.
Faded example
For X ~ N(170, 6.28), a height has z = -2.
Fill in the blanks
x = 170 + (-2)(6.28) = 157.44 \text___
Why: Two standard deviations below 170 is 170 minus 12.56, which is 157.44 cm. The negative z-score puts the value below the mean, as it must.
Two truths and a lie
All three concern z-scores.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. A z-score is measured in units of the standard deviation and is itself unitless, which is precisely what allows centimetres to be compared with exam points. Reporting a z-score with units attached signals that the division by sigma was skipped.
Estimation
A value has a z-score of 3.2.
Predict first
How should that be described?
Correct: Very unusual.
Why: The Empirical Rule puts 99.7 percent of values within three standard deviations, so a z-score beyond 3 is in the outer tenth of one percent. It is not impossible — a normal distribution assigns positive probability to every value — but it is rare enough to be worth remarking on, and section 2.7's rule of thumb for outliers rests on exactly this.
Section
Section 3
Concept
The z-score allows us to compare data that are scaled differently. Two values from different normal distributions that share a z-score sit the same number of their own standard deviations from their own means, and so represent the same standardised position.
standardised comparison — Judging values from different distributions by their z-scores rather than their raw sizes. Two values with equal z-scores are equally unusual relative to their own distributions, whatever their raw magnitudes.
\[ z_X = z_Y \;\Longrightarrow\; x \text{ and } y \text{ occupy the same relative position} \]
The book's illustration is deliberately artificial and useful for that reason: X follows N(5, 6) measures weight gain for one group and Y follows N(2, 1) for a second, and x equals 17 and y equals 4 are each two standard deviations to the right of their means. The raw numbers differ by a factor of four, yet they represent the same standardized weight gain relative to their means.
Figure (svg): Two bell curves with different means and standard deviations, each with a value marked that is one and a half standard deviations below its own mean
OpenStax Introductory Statistics 2e, §6.1 The Standard Normal Distribution §6.1, pp. 337-339 — Example 6.2(c) and Example 6.4
Picture it
Example 6.4: heights from 2009-2010 and from 1984-1985.
Figure (svg): Two bell curves with different means and standard deviations, each with a value marked that is one and a half standard deviations below its own mean
The two heights differ by more than two centimetres, yet both are 1.5 standard deviations below their own era's mean — so a boy of 160.58 cm in 2009 was exactly as short, relative to his peers, as one of 162.85 cm had been in 1984. Comparing the raw heights would have suggested the second boy was taller in a meaningful sense, and relative to his own generation he was not.
Worked example
Example 6.4.
\[ x = 160.58 \text{ in } N(170, 6.28); \quad y = 162.85 \text{ in } N(172.36, 6.34) \]
z-score for x
Why: 160.58 minus 170, over 6.28.
\[ -1.5 \]
z-score for y
Why: 162.85 minus 172.36, over 6.34.
\[ -1.5 \]
Compare
Why: The same number.
Interpret
Why: Both below their own means.
Figure (svg): The solution to Worked example matching z-scores across two eras shown as a ladder of expressions, one row per legal move
\[ z_x = z_y = -1.5 \]
Verify: confirm that the raw comparison would have misled
Why: The second height is 2.27 cm greater than the first, so on raw numbers the 1984 boy looks taller. But his era's mean was also higher, by 2.36 cm, so relative to his peers he was fractionally shorter. Standardising is what exposes that reversal, and it is the reason exam scores, test results and growth measurements are almost always reported as standardised values rather than raw ones.
OpenStax Introductory Statistics 2e, §6.1 The Standard Normal Distribution §6.1, p. 339
Prediction
Commit before reasoning.
Predict first
Statistics: 68 on N(63, 5). History: 90 on N(78, 12). Which is the stronger performance relative to the class?
Correct: They tie.
Why: Sixty-eight is 5 above a mean of 63 with sigma 5, giving z = 1; ninety is 12 above 78 with sigma 12, also z = 1. The raw gap is 5 points against 12, but each is exactly one standard deviation, so the two performances are identical in relative terms. Comparing the raw scores would wrongly favour history.
Worked example
Example 6.2(c), where the raw values differ fourfold.
\[ x = 17 \text{ in } N(5, 6); \quad y = 4 \text{ in } N(2, 1) \]
z-score for x
Why: Seventeen minus five, over six.
\[ 2 \]
z-score for y
Why: Four minus two, over one.
\[ 2 \]
Compare the raw values
Why: 17 against 4.
Compare the z-scores
Why: Both equal.
Figure (svg): The solution to Worked example two groups on very different scales shown as a ladder of expressions, one row per legal move
\[ z = \frac{17-5}{6} = 2 = \frac{4-2}{1} \]
Verify: confirm which comparison answers the question being asked
Why: If the question is who gained more weight, the answer is plainly the first group's member, at 17 against 4. If it is who gained more relative to their group, the answer is that they tie. Both questions are legitimate, and the z-score answers only the second — so it is worth being explicit about which one is wanted before standardising, since standardising discards the raw scale entirely.
OpenStax Introductory Statistics 2e, §6.1 The Standard Normal Distribution §6.1, p. 337
Trap
\[ 162.85 > 160.58 \;\Rightarrow\; \text{the second boy was taller for his era} \]
Compare the two heights directly
Why: One number is plainly larger than the other.
\[ \text{but their eras had different means} \]
Relative to his own generation the taller boy was in fact fractionally shorter, since his era's mean was higher still.
\[ z_x = z_y = -1.5 \;\Rightarrow\; \text{identical relative positions} \]
Standardise both before comparing
Why: The z-score removes the difference in scale.
The rule worth carrying is that raw values may be compared only when they come from the same distribution. Across different means or different spreads — different exams, different eras, different instruments — the comparison has to be made on z-scores, and this is why standardised test reporting exists at all.
Faded example
SAT verbal scores in 2012 had mean 496 and standard deviation 114. A student scored 325.
Fill in the blanks
z = \frac-1.5below = ___, \text___ ___ \text___
Why: Minus 171 over 114 is exactly minus 1.5, so the score is 1.5 standard deviations below the mean. The same z-score as Example 6.4's heights, on a completely different scale.
Two truths and a lie
All three concern comparing across scales.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false across different distributions, and Example 6.4 is the counterexample: 162.85 exceeds 160.58 in raw terms while both have the same z-score of minus 1.5. It is true only within a single distribution, where the mean and standard deviation are shared.
Explain it
A classmate insists the 1984 boy at 162.85 cm was taller than the 2009 boy at 160.58 cm, so he cannot have been equally short.
Discussion prompt
In three sentences or fewer, reconcile the two claims.
Hint: Ask them taller compared with whom.
Answer:
Agree that in absolute terms he was taller — 162.85 is more than 160.58, and nothing changes that.
But his generation was taller too: its mean was 172.36 against 170, so he had further to climb to reach his own era's average.
Measured against their own peers both boys sit 1.5 standard deviations below the mean, so they were equally short for their time even though one was physically taller.
Section
Section 4
Concept
The standard normal distribution is a normal distribution of standardized values called z-scores. Its mean is zero and its standard deviation is one, and the transformation z equals x minus mu over sigma produces Z following N(0, 1).
Z ~ N(0, 1) — The standard normal. Because its mean is zero and its standard deviation one, a value and its z-score are the same number, so areas computed once for this distribution serve every other normal.
\[ Z = \frac{X - \mu}{\sigma} \sim N(0, 1) \]
This is the practical pay-off of the whole section. There are infinitely many normal distributions and it would be impossible to tabulate them all, but standardising collapses them onto one — so a single table of areas, or a single calculator routine, answers questions about every one. The book notes the historical shape of this: before technology, the z-score was looked up in a standard normal probability table because the math involved is too cumbersome.
Figure (svg): The standard normal curve centred at zero with ticks one unit apart
OpenStax Introductory Statistics 2e, §6.1 The Standard Normal Distribution §6.1, p. 336 — the transformation and Z ~ N(0, 1)
Picture it
Mean zero, standard deviation one, ticks a standard deviation apart.
Figure (svg): The standard normal curve centred at zero with ticks one unit apart
The horizontal axis here is labelled z rather than x, and that relabelling is the entire content of standardising. Any normal picture can be redrawn this way by rescaling its axis, which is why the shape of every normal curve in this book looks the same.
Worked example
Converting a statement about x into one about z.
\[ X \sim N(63, 5); \text{ express } 58 < x < 68 \text{ in } z \]
Standardise the lower end
Why: 58 minus 63, over 5.
\[ z = -1 \]
Standardise the upper end
Why: 68 minus 63, over 5.
\[ z = +1 \]
Rewrite the statement
Why: In terms of z.
\[ -1 < z < 1 \]
Read what it says
Why: Within one standard deviation.
\[ \text{about } 68 \% \]
Figure (svg): The solution to Worked example standardising an interval shown as a ladder of expressions, one row per legal move
\[ 58 < x < 68 \;\Longleftrightarrow\; -1 < z < 1 \]
Verify: confirm that standardising preserves the probability
Why: The two statements describe exactly the same event — a score between 58 and 68 IS a z-score between minus 1 and 1 — so they must have the same probability. That is why standardising is legitimate rather than an approximation: it relabels the axis without moving any area. Section 6.2 relies on this at every step.
OpenStax Introductory Statistics 2e, §6.1 The Standard Normal Distribution §6.1, pp. 336-338
Fill the middle
What makes a normal distribution standard.
Fill in the blanks
Z \sim N(0, 1): \textone ___
Why: One. With a mean of zero and a standard deviation of one, a value and its z-score coincide, which is what makes this the reference distribution for all the others.
Worked example
Three different distributions, one lookup.
\[ N(63, 5), \; N(170, 6.28), \; N(36.9, 13.9) \]
Standardise each
Why: One sigma above the mean.
\[ z = 1\text{ in all three} \]
Look up the area
Why: Below z = 1.
\[ 0.8413 \]
Apply to each
Why: The same number.
\[ 0.8413\text{ for all} \]
Say why
Why: They are the same event in z.
Figure (svg): The solution to Worked example why one table suffices shown as a ladder of expressions, one row per legal move
\[ P(Z < 1) = 0.8413 \]
Verify: confirm this against the Empirical Rule
Why: The rule says 68 percent lies within one standard deviation, so 32 percent lies outside, split evenly into 16 percent in each tail by symmetry. The area below z = 1 is therefore 68 plus 16, or 84 percent — matching 0.8413 to the accuracy the rule offers. Two independent routes agreeing is a genuine check on both.
OpenStax Introductory Statistics 2e, §6.1 The Standard Normal Distribution §6.1, pp. 336-338
Trap
\[ y = 162.85 \text{ standardised with } N(170, 6.28) \]
Use the parameters already written down
Why: Both variables are heights in centimetres.
\[ z = \frac{162.85 - 170}{6.28} \approx -1.14 \quad \text{instead of } -1.5 \]
The value belongs to the 1984 distribution, so it must be standardised against that era's mean and standard deviation.
\[ z = \frac{162.85 - 172.36}{6.34} = -1.5 \]
Match each value to ITS own distribution before standardising
Why: A z-score is always relative to the distribution the value came from.
This error is easy to make precisely when the two distributions are similar, as they are here, because the wrong answer looks reasonable. Writing the distribution beside each value before computing anything — y is from N(172.36, 6.34) — is the habit that prevents it, and it becomes essential in chapter 10, where two samples are compared and each has its own parameters.
Two truths and a lie
All three concern standardising.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and is a common misreading. Subtracting the mean and dividing by the standard deviation shifts and rescales a distribution but cannot change its shape — a skewed variable standardises to a skewed variable with mean zero and spread one. Only a normal variable standardises to a standard NORMAL.
Discrimination
Each value is paired with its own distribution.
Sort into buckets
Sort each by whether its z-score is 2.
Prediction
Commit before reasoning.
Predict first
For any normal distribution, what is the z-score of the mean itself?
Correct: 0.
Why: The numerator x minus mu is zero, so the z-score is zero whatever sigma may be. This is why the standard normal is centred at zero: standardising sends the mean to the origin, and every distance is then measured from there in units of the standard deviation.
Section
Section 5
Concept
For a normal distribution, about 68 percent of the values lie within one standard deviation of the mean, about 95 percent within two, and about 99.7 percent within three. It is also known as the 68-95-99.7 rule.
the Empirical Rule — The three areas within one, two and three standard deviations of the mean of a normal distribution. Because they are stated in standard deviations, the same three percentages apply to every normal distribution.
\[ P(-1 < Z < 1) \approx 0.68, \; P(-2 < Z < 2) \approx 0.95, \; P(-3 < Z < 3) \approx 0.997 \]
The rule is exactly the standard normal's areas quoted once and reused, which is why the percentages never change. The book adds the observation that matters most in practice: almost all the x values lie within three standard deviations of the mean. That is what makes a z-score beyond 3 worth investigating, and it is the basis of the outlier rule of thumb from section 2.7.
Figure (svg): A bell curve with three nested shaded bands at one, two and three standard deviations, labelled sixty-eight, ninety-five and ninety-nine point seven percent
OpenStax Introductory Statistics 2e, §6.1 The Standard Normal Distribution §6.1, pp. 338-339 — the Empirical Rule and Example 6.5
Picture it
The same curve, with one, two and three standard deviations marked.
Figure (svg): A bell curve with three nested shaded bands at one, two and three standard deviations, labelled sixty-eight, ninety-five and ninety-nine point seven percent
The bands are nested rather than separate, so the second band's 95 percent includes the first band's 68. Reading them as disjoint slices is a frequent slip: the region between one and two standard deviations holds about 27 percent, which is 95 minus 68, and not 95 itself.
Worked example
Example 6.5: a mean of 50 and a standard deviation of 6.
\[ \mu = 50, \; \sigma = 6 \]
One standard deviation
Why: Fifty give or take six.
\[ 44\text{ to } 56 \]
Two standard deviations
Why: Give or take twelve.
\[ 38\text{ to } 62 \]
Three standard deviations
Why: Give or take eighteen.
\[ 32\text{ to } 68 \]
Note the z-scores
Why: Always the same.
\[ -1\text{ and } 1, -2\text{ and } 2, -3\text{ and } 3 \]
Figure (svg): The solution to Worked example the rule in original units shown as a ladder of expressions, one row per legal move
\[ 50 \pm 6, \quad 50 \pm 12, \quad 50 \pm 18 \]
Verify: confirm the bands are nested and widening
Why: Each band contains the one before it — 38 to 62 contains 44 to 56 — and each is six units wider at each end than its predecessor, because each step adds one more standard deviation. If a computed band ever failed to contain the previous one, an arithmetic slip would be certain. The z-scores in the last column never change, which is the point of expressing the rule in standard deviations.
OpenStax Introductory Statistics 2e, §6.1 The Standard Normal Distribution §6.1, p. 339
Faded example
X has a normal distribution with mean 25 and standard deviation 5.
Fill in the blanks
68 \text20 30 \text___ ___
Why: One standard deviation either side of 25 gives 20 and 30. The z-scores at those points are minus 1 and plus 1, as they always are for the 68 percent band.
Worked example
Turning a two-sided rule into a one-sided answer.
\[ \text{about what fraction lies ABOVE } \mu + 2\sigma? \]
Start from the band
Why: Within two standard deviations.
\[ 95 \% \]
Find what is outside
Why: One hundred minus 95.
\[ 5 \% \]
Split by symmetry
Why: Two equal tails.
\[ 2.5 \%\text{ each} \]
Answer the question
Why: The upper tail alone.
\[ \text{about } 2.5 \% \]
Figure (svg): The solution to Worked example reading a tail from the rule shown as a ladder of expressions, one row per legal move
\[ \frac{100 - 95}{2} = 2.5\text{ percent} \]
Verify: confirm the split is licensed by symmetry
Why: The normal curve is symmetric about its mean, so the two tails outside a symmetric band must have equal areas — halving is exact rather than an approximation. This move is used constantly from chapter 8 onward, where confidence levels of 95 percent produce tails of 2.5 percent on each side, and it is worth recognising here where the arithmetic is easy.
OpenStax Introductory Statistics 2e, §6.1 The Standard Normal Distribution §6.1, pp. 338-339
Trap
\[ \text{between } 1\sigma \text{ and } 2\sigma \text{ lies about } 95 \text{ percent} \]
Attach each percentage to the region at that distance
Why: The rule lists 68, 95 and 99.7 against one, two and three.
\[ 68 + 95 + 99.7 > 100 \]
The three percentages cannot be slices of one distribution, since they already total more than everything.
\[ \text{between } 1\sigma \text{ and } 2\sigma: \; 95 - 68 = 27 \text{ percent} \]
Read the bands as NESTED, each containing the last
Why: Ninety-five percent lies within two standard deviations, including the 68 within one.
The impossibility check is instant and worth using: any reading of the rule that makes the percentages add to more than 100 has misread nesting as partition. Subtracting consecutive bands gives the slices — about 68 percent in the middle, 27 percent in the next ring, and 4.7 percent in the one after that.
Estimation
The rule puts 99.7 percent within three standard deviations.
Predict first
About what fraction lies more than three standard deviations ABOVE the mean?
Correct: About 0.15 percent.
Why: 0.3 percent lies outside the band altogether, and symmetry splits that evenly into about 0.15 percent in each tail — roughly one value in 700. That is why exceeding three standard deviations on one side is treated as genuinely remarkable rather than merely uncommon.
Two truths and a lie
All three concern the rule.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The exact areas are about 0.6827, 0.9545 and 0.9973, so the rule's 68, 95 and 99.7 are rounded — good enough for a sanity check but not for a reported answer. Section 6.2 uses a calculator precisely because the rule is not exact.
Matching
Match each region to the area it holds.
Match the pairs
Why: The third row is the one that has to be computed rather than recalled: 95 minus 68 gives the ring between one and two standard deviations. Rows one and two come straight from the rule, and row four is 100 minus 99.7.
Comparison
Fill the blanks. Standardising is a relabelling, not a transformation of the data's character.
Comparison matrix
| Quantity | Before standardising | After |
|---|---|---|
| Mean | mu | 0 |
| Standard deviation | sigma | 1 |
| Shape | bell shaped | bell shaped, unchanged |
| Probability of an event | unchanged | unchanged |
The last two rows carry the important content. Standardising moves the centre and rescales the spread, but it cannot alter the shape or any probability — which is exactly why it is a legitimate step rather than an approximation.
Pattern
Five steps, and the first one prevents most of the errors.
When comparing values from different distributions, standardise each against its OWN mean and standard deviation before comparing anything.
OpenStax Introductory Business Statistics 2e, §6.1 The Standard Normal Distribution §6.1 The Standard Normal Distribution
Check
Computing a z-score.
Check your understanding
For X ~ N(5, 2), what is the z-score of x = 10?
Answer: A
Why: Ten minus five is 5, divided by the standard deviation of 2 gives 2.5 — two and a half standard deviations above the mean.
Check
Recovering a value.
Check your understanding
For X ~ N(170, 6.28), a height has z = 1.27. What is the height?
Answer: A
Why: x equals mu plus z sigma, which is 170 plus 1.27 times 6.28, about 177.98 cm.
Check
The Empirical Rule.
Check your understanding
X is normal with mean 50 and standard deviation 6. About 95 percent of values lie between which two numbers?
Answer: A
Why: Ninety-five percent lies within two standard deviations, which is 50 give or take 12, so between 38 and 62.
Real world
A paediatric clinic records a two-year-old's height as 79 cm and reports it as being on the 5th percentile, flagging the child for follow-up. A parent points out that a friend's child, measured at a different clinic on a different growth chart, was 81 cm and was reported as perfectly normal, and asks how 79 can be concerning when 81 was not.
Discussion prompt
Explain how both reports can be correct, and say what information you would need to check them.
Hint: Ask what each height was being compared against.
Answer:
Both can be correct because a percentile is always relative to a reference distribution. Growth charts differ by the population they were built from, by sex, and — critically for a two-year-old — by exact age in months, since children of that age grow several centimetres a year. A height of 79 cm at 24 months and 81 cm at 30 months are being standardised against quite different means.
The comparison the parent is making is the raw one, and it is the one Example 6.4 warns against. What matters is each child's z-score against their own reference distribution, not the two heights against each other. If the 24-month reference has mean 87 cm with standard deviation 3.2 cm, then 79 cm gives a z-score of about minus 2.5, which is genuinely in the lowest few percent.
\[ z = \frac{79 - 87}{3.2} = -2.5 \quad\text{against}\quad z = \frac{81 - 90}{3.5} \approx -2.6 \]
To check the reports you would need three things for each child: the exact age in months, the sex, and which growth reference the clinic used. With those you can compute both z-scores directly and see whether the two children really are in different positions or — as the illustrative numbers above suggest — in almost the same one, reported differently because the clinics used different cut-offs for follow-up.
Two cautions belong in a careful answer. Percentiles from a single measurement are noisy, which is why paediatricians track a child's growth curve over time rather than acting on one point — a child steadily on the 5th percentile is usually less concerning than one who has fallen from the 50th. And the Empirical Rule gives the scale of what is being claimed: a z-score of minus 2.5 puts the child below about the 1st percentile of the reference, so the flag is not raised lightly.
Commit first
Answer, then rate your confidence honestly.
Predict first
Why does standardising let a single table answer questions about every normal distribution?
Correct: Because standardising sends every normal to N(0, 1) without changing any probability.
\[ P(a < X < b) = P\left(\frac{a-\mu}{\sigma} < Z < \frac{b-\mu}{\sigma}\right) \]
Why: Subtracting the mean and dividing by the standard deviation relabels the axis so that every normal becomes the standard normal, and since the relabelling describes the same events it moves no area. So one set of areas, computed once for N(0, 1), serves them all. The normal density is in fact notoriously awkward to integrate, which is precisely why tabulating one distribution rather than infinitely many mattered so much.
Explain it
They computed the z-score for x = 168 in N(170, 6.28) as positive 0.32 and concluded the height is above average.
Discussion prompt
In two sentences or fewer, locate the error.
Hint: Ask whether 168 is bigger or smaller than 170.
Answer:
Ask them whether 168 is above or below 170 — it is below, so the z-score has to come out negative before any arithmetic is checked.
They subtracted in the wrong order: the formula is x minus mu, so it is 168 minus 170, giving minus 2 and a z-score of minus 0.32.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, the second entry is always the standard deviation in this book. For the second, subtract the mean FROM x, and check the sign against which side of the mean x lies. For the third, use x equals mu plus z sigma and check the side. For the fourth, remember the bands are nested and the tails are found by symmetry. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top, draw two pairs of bell curves: one pair with the same standard deviation and different means, one pair with the same mean and different standard deviations, and label what each parameter does. Below them, write the z-score formula and its rearrangement x equals mu plus z sigma, marking on a sketch which side of the mean a positive and a negative z-score fall on. In the middle of the page, work Example 6.1 completely — X follows N(5, 6), find the z-scores for x = 17 and x = 1, then rebuild both original values from their z-scores as a check. Beside it, draw the two curves of Example 6.4, N(170, 6.28) and N(172.36, 6.34), mark 160.58 on the first and 162.85 on the second, compute both z-scores, and write one sentence on why the raw comparison misleads. At the bottom, draw a single standard normal curve, shade three nested bands at one, two and three standard deviations, and label them 68, 95 and 99.7 percent. Beneath it, convert those bands into actual values for a distribution with mean 50 and standard deviation 6, and finally compute what fraction lies ABOVE two standard deviations by subtracting from one hundred and halving.
Check the two z-scores in Example 6.4 against each other: both must come out to exactly minus 1.5, and if one does not, the likely cause is standardising a value against the other era's parameters. Check the bottom row by confirming your three bands are nested, each containing the one before it.
Recap
Five things, and the middle three are one idea used three ways.
| If you see | Then |
|---|---|
| N(mu, sigma) in this book | The second entry is the standard deviation |
| A value and a distribution | z is x minus mu, over sigma |
| A z-score and a distribution | x is mu plus z sigma |
| A negative z-score | The value lies below the mean |
| Two values on different scales | Standardise each against its OWN parameters |
| A z-score beyond 3 | Rare: about 0.3 percent lie outside that band |
| Percentages totalling over 100 | The nested bands have been read as slices |
Section 6.2 puts the standard normal to work. With a calculator supplying areas, the same three questions asked of every earlier distribution return: the probability of an interval, the probability of a tail, and the value at a given percentile — and the main difficulty becomes translating an English phrase such as at least, the bottom quartile, or the middle twenty percent into an area to the left.
OpenStax Introductory Statistics 2e, §6.1 The Standard Normal Distribution §6.1, pp. 336-339 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.