Section 6.1's machinery put to work. With a calculator or computer returning areas under the normal curve, three questions can be answered about any normally distributed quantity: the probability of a tail, the probability of an interval, and the value at a given percentile. The technology reports only the area to the left, so a right tail is found by subtracting from one and an interval by subtracting one left area from another, while a percentile reverses the process by taking an area and returning the boundary. The genuine difficulty in this section is neither the arithmetic nor the technology but the translation of English phrases into areas to the left: at most, at least, the bottom quartile, the top ten percent and the middle twenty percent all specify an area, and two of them specify the complement of the number they mention.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 6 — The Normal Distribution
Using the Normal Distribution
Objectives
Five outcomes, and the middle two are where the marks are actually lost.
OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 340-346 — the section these objectives are drawn from
Warm-up
Section 6.1 gave the z-score and the standard normal; section 5.1 gave the complement rule for areas.
Discussion prompt
Exam scores are normal with mean 63 and standard deviation 5. Roughly what fraction of students score above 65, and can you answer it without a calculator?
Hint: How many standard deviations above the mean is 65?
Answer:
A score of 65 is two points above the mean of 63, which is 0.4 standard deviations — so the answer must be a bit less than half, since a score exactly at the mean would leave half above it and 65 is slightly higher.
The Empirical Rule cannot do better than that, because 0.4 is not a whole number of standard deviations. It brackets the answer between roughly 0.32 and 0.50, which is a real constraint but not an answer.
That gap is what this section fills. A calculator gives 0.3446 exactly, and the estimate's role changes from producing the answer to checking it — which is exactly how it should be used from here on, since a mistyped command produces a number with no warning at all.
Concept
The cumulative distribution function P(X < x) gives the area to the left, and it is what a calculator, a computer or a table reports. A right tail is one minus it, and an interval is the difference of two of them. A percentile runs the process backwards: the area is given and the boundary is wanted.
area to the left — The single quantity the technology returns. Every other normal question — right tails, intervals, percentiles, middle regions, quartiles — is built from it by subtraction or by reversal.
\[ P(X > x) = 1 - P(X < x), \qquad P(a < X < b) = P(X < b) - P(X < a) \]
The book is explicit that the tables are largely historical: technology has made the tables virtually obsolete, and for that reason, as well as the fact that there are various table formats, we are not including table instructions. What survives from the table era is the convention that left areas are what get reported — so the two subtractions above are as necessary now as they were then.
Figure (svg): Two identical bell curves, the left one shaded below a cut-off and the right one shaded above it, with the two areas summing to one
OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 341-342
Section
Section 1
Concept
A right tail is one minus the area to the left of the cut-off. An interval is the area to the left of the upper endpoint minus the area to the left of the lower one. Neither needs anything the technology does not already give.
the two subtractions — One minus a left area gives a right tail; one left area minus another gives an interval. Together they answer every probability question about a normal distribution.
\[ P(a < X < b) = P(X < b) - P(X < a) \]
The book shows the historical route for Example 6.8(a) alongside the calculator one, and it is worth following once. The z-score is 65 minus 63 over 5, which is 0.4; the area to the left of 0.4 is 0.6554; so the right tail is 1 minus 0.6554, or 0.3446. That is the same answer the calculator gives directly from the unstandardised values, which confirms that standardising changes nothing — it was simply necessary when only one distribution could be tabulated.
Figure (svg): A normal curve for exam scores with the region above sixty-five shaded
OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 341-344 — Examples 6.7, 6.8(a)-(b) and 6.9(a)
Picture it
Example 6.8(a): the chance a student scores above 65.
Figure (svg): A normal curve for exam scores with the region above sixty-five shaded
The shaded region is a little under half the curve, which matches the warm-up's estimate — 65 is only just above the mean, so the tail above it should be only just below a half. Making that comparison before accepting 0.3446 costs nothing and catches a mistyped mean or standard deviation immediately.
Worked example
Example 6.8(a), by both routes.
\[ X \sim N(63, 5); \text{ find } P(x > 65) \]
Standardise the cut-off
Why: 65 minus 63, over 5.
\[ z = 0.4 \]
Find the area to the left
Why: Below z = 0.4.
\[ 0.6554 \]
Subtract from one
Why: The right tail.
\[ 0.3446 \]
Check against technology
Why: Directly from N(63, 5).
\[ 0.3446 \]
Figure (svg): The solution to Worked example more than 65 on the exam shown as a ladder of expressions, one row per legal move
\[ P(x > 65) = 1 - 0.6554 = 0.3446 \]
Verify: confirm the answer is on the right side of a half
Why: Since 65 exceeds the mean of 63, less than half the distribution lies above it — so the answer had to come out below 0.50, and 0.3446 does. Had the subtraction from one been forgotten, the answer would have been 0.6554, which is above a half and therefore impossible for a cut-off above the mean. That single comparison catches the section's commonest error.
OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 341-342
Faded example
Mandarin diameters are N(5.85, 0.24). The left area at 6.0 is 0.7340.
Fill in the blanks
P(x > 6.0) = 1 - 0.7340 = 0.2660
Why: One minus 0.7340 gives 0.2660, the book's answer for Example 6.12(a). Since 6.0 cm is above the mean of 5.85, the tail above it must be below a half — and it is.
Worked example
Example 6.9(a), an interval.
\[ X \sim N(2, 0.5); \text{ find } P(1.8 < x < 2.75) \]
Left area at 2.75
Why: The upper endpoint.
\[ 0.9332 \]
Left area at 1.8
Why: The lower endpoint.
\[ 0.3446 \]
Subtract
Why: Upper minus lower.
\[ 0.5886 \]
State it in words
Why: Hours of entertainment use.
\[ \text{about } 59 \% \]
Figure (svg): The solution to Worked example between 1.8 and 2.75 hours shown as a ladder of expressions, one row per legal move
\[ P(1.8 < x < 2.75) = 0.9332 - 0.3446 = 0.5886 \]
Verify: confirm the interval's answer is plausible for its width
Why: The interval runs from 0.4 standard deviations below the mean to 1.5 above, so it should capture well over half the distribution but nothing like all of it — and 0.5886 sits comfortably in that range. It is worth noticing that the left area at 1.8 is 0.3446, the same number as the previous example's answer, which is a coincidence of the two problems' z-scores both being 0.4 in magnitude.
OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 343-344
Error analysis
The book's answer is 0.3446.
Annotate
On: \( \begin{aligned} &(1)\; 0.6554 \\ &(2)\; 1 - P(X < 63) = 0.5 \\ &(3)\; P(X < 65) - P(X < 63) = 0.1554 \\ &(4)\; 1 - 0.6554 = 0.3446 \end{aligned} \)
Errors (1) and (3) both produce numbers that look like probabilities, and only (1) is caught by the above-or-below-a-half test. Shading the region on a sketch before computing is what separates (3) from (4), since the two describe visibly different areas.
Discrimination
Decide what each question needs.
Sort into buckets
Sort each question about a normal X.
Two truths and a lie
All three concern tails and intervals.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and reverses the relationship. A cut-off above the mean leaves LESS than half the distribution above it, so the right tail must be below 0.5. This is the fastest available check on any tail computation, and it catches a forgotten subtraction immediately.
Estimation
Smartphone users' ages are N(36.9, 13.9), and 50.8 is exactly one standard deviation above the mean.
Predict first
Roughly what is P(x < 50.8)?
Correct: About 0.84.
Why: Sixty-eight percent lies within one standard deviation, leaving 32 percent split into two tails of 16 percent each — so the area below one standard deviation above the mean is 68 plus 16, about 84 percent. The calculator gives 0.8413 for Example 6.10(b), which confirms it. Recognising that 50.8 is exactly mu plus sigma is what makes this check available.
Section
Section 2
Concept
A percentile question hands you the area and asks for the value. The 90th percentile is the score k that has 90 percent of the scores below it and 10 percent above, and it is found by reversing the cumulative function rather than evaluating it.
critical value — The book's other name for the percentile k. The term becomes standard from chapter 8, where the same reversal supplies the boundaries of a confidence interval.
\[ P(X < k) = p \;\Longrightarrow\; k = \text{invNorm}(p, \mu, \sigma) \]
The relationship to section 6.1's two formulas is worth seeing. A percentile could be found by hand in two steps: look up the z-score with the given area to its left, then convert with x equals mu plus z sigma. For the 90th percentile that is z equal to about 1.28, giving 63 plus 1.28 times 5, or about 69.4. The calculator's single command does both steps at once, which is convenient but hides the standardising — worth doing by hand once to see that nothing new is happening.
Figure (svg): A normal curve for exam scores with ninety percent of the area shaded and the cut-off marked at sixty-nine point four
OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 342-344 — Examples 6.8(c)-(d) and 6.9(b)
Picture it
Example 6.8(c): the 90th percentile of exam scores.
Figure (svg): A normal curve for exam scores with ninety percent of the area shaded and the cut-off marked at sixty-nine point four
Compare this picture with the first idea's. There the dashed line was given and the shading was the answer; here the shading is given and the dashed line is the answer. Recognising which of the two you are looking at is the whole of the difficulty, and the phrase to watch for is one naming a proportion rather than a value.
Worked example
Example 6.8(c) and (d).
\[ X \sim N(63, 5); \text{ find the } 90\text{th and } 70\text{th percentiles} \]
Set up the 90th
Why: Ninety percent below k.
\[ P(x < k) = 0.90 \]
Reverse the cumulative
Why: With mean 63 and sigma 5.
\[ k = 69.4 \]
Set up the 70th
Why: Seventy percent below.
\[ P(x < k) = 0.70 \]
Reverse again
Why: Same distribution.
\[ k = 65.6 \]
Figure (svg): The solution to Worked example the 90th and 70th percentiles shown as a ladder of expressions, one row per legal move
\[ k_{90} = 69.4, \qquad k_{70} = 65.6 \]
Verify: confirm both sit above the mean, and in the right order
Why: Both percentiles exceed 50, so both boundaries must lie above the mean of 63 — and 65.6 and 69.4 do. The 90th must also exceed the 70th, since a larger area to the left needs a boundary further right. Two percentiles that came out in the wrong order, or on the wrong side of the mean, would signal that the areas had been entered as their complements.
OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 342-343
Faded example
Smartphone users' ages are N(36.9, 13.9). Find the 80th percentile.
Fill in the blanks
P(x < k) = 0.80 \;\Rightarrow\; k \approx 48.6 \text___
Why: Eighty percent below gives k = 48.6 years, so 80 percent of users in the age range are 48.6 years old or less. The answer exceeds the mean of 36.9, as any percentile above the 50th must.
Worked example
Example 6.9(b), where the wording names a quarter rather than a percentile.
\[ X \sim N(2, 0.5); \text{ the bottom quartile uses at most how many hours?} \]
Translate the phrase
Why: The bottom quarter.
Write the equation
Why: Twenty-five percent below.
\[ P(x < k) = 0.25 \]
Reverse the cumulative
Why: With mean 2 and sigma 0.5.
\[ k = 1.66 \]
Interpret
Why: In the problem's units.
\[ 1.66\text{ hours} \]
Figure (svg): The solution to Worked example the bottom quartile shown as a ladder of expressions, one row per legal move
\[ k = \text{invNorm}(0.25,\; 2,\; 0.5) = 1.66 \]
Verify: confirm the answer falls below the mean, as a bottom quartile must
Why: Twenty-five percent is below a half, so the boundary must lie below the mean of 2 hours — and 1.66 does. It is also about 0.67 standard deviations below, which matches the standard normal's first quartile of about minus 0.674; that constant is worth recognising, since the first and third quartiles of ANY normal sit about two thirds of a standard deviation either side of the mean.
OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, p. 344
Trap
\[ \text{the } 90\text{th percentile of } N(63,5) \;\Rightarrow\; P(X < 90) \approx 1 \]
Feed the number 90 in as a score
Why: The question mentions 90, and scores go into the cumulative function.
\[ \text{an answer of } 1, \text{ which is a probability and not a score} \]
Ninety was a percentage, not a mark; the question supplied an area and wanted a boundary.
\[ P(X < k) = 0.90 \;\Rightarrow\; k = 69.4 \]
Ask what the question GIVES and what it WANTS before touching the calculator
Why: A proportion given means a percentile problem.
The units test settles it instantly: a percentile answer must come out in the problem's units — marks, hours, centimetres — while a probability is always between zero and one. An answer of 1 to a question asking for a score is self-evidently the wrong kind of object, and noticing that is faster than re-reading the question.
Two truths and a lie
All three concern percentiles.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and confuses the two kinds of question. A percentile is a value in the problem's own units — 69.4 marks, 1.66 hours, 48.6 years — while it is the AREA that lies between 0 and 1. Checking the units of the answer is the quickest way to confirm the right question was answered.
Prediction
Commit before reasoning.
Predict first
For any normal distribution, where does the 30th percentile lie?
Correct: Below the mean.
Why: Only 30 percent of the area lies to its left, which is less than half, so the boundary must sit left of centre. The standard deviation affects how far below it falls but never which side. Checking the side before computing rules out half the possible errors at no cost.
Matching
Match each percentile to where it falls.
Match the pairs
Why: These four z-scores — 0, minus 0.67, 1.28 and 0.67 — are the same for every normal distribution, which is what makes them worth recognising. The 1.28 in particular recurs constantly from chapter 8 onward, where it supplies the boundary of a 90 percent one-sided interval.
Section
Section 3
Concept
Every calculator command reports area to the left, so a worded question has to be converted into one before anything can be computed. Some phrases name that area directly and others name what remains, and the book's own worked solution spells the conversion out.
at least — The book translates it as greater than or equal to, so the given proportion is the area to the RIGHT. The area to the left is one minus it, and that is what goes into the calculation.
\[ P(x \ge k) = 0.40 \;\Rightarrow\; P(x < k) = 0.60 \]
Example 6.11(b) is worth reading in the book's own words because it shows the steps a careful solution makes explicit: find k where P(x at least k) equals 0.40; at least translates to greater than or equal to; 0.40 is the area to the right; area to the left equals one minus 0.40, which is 0.60; the area to the left of k is 0.60. Four short lines, none of them arithmetic, and they are where the question is actually answered.
Figure (svg): A translation table turning English phrases into areas to the left
OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 345-346 — Example 6.11(b) and Example 6.12(b)
Picture it
Two of them name the complement of the number they mention.
Figure (svg): A translation table turning English phrases into areas to the left
The third and fourth rows are the ones that get reversed, and they look almost identical in print — the bottom quartile uses 0.25 while the top ten percent uses 0.90. Reading the direction word before the number, rather than after it, is the habit that keeps them apart.
Worked example
Example 6.11(b), following the book's own steps.
\[ X \sim N(36.9, 13.9); \text{ forty percent are at least what age?} \]
Write the requirement
Why: Forty percent at or above k.
\[ P(x \ge k) = 0.40 \]
Identify which side
Why: At least means to the right.
\[ 0.40\text{ is the RIGHT area} \]
Convert to a left area
Why: One minus 0.40.
\[ 0.60 \]
Reverse the cumulative
Why: At 0.60.
\[ k = 40.4 \]
Figure (svg): The solution to Worked example at least what age shown as a ladder of expressions, one row per legal move
\[ P(x < k) = 0.60 \;\Rightarrow\; k \approx 40.4 \text{ years} \]
Verify: confirm the answer lies above the mean, and say why that is right
Why: Forty percent lie above k, which is less than half, so k must sit above the mean of 36.9 — and 40.4 does. Had 0.40 been used as the left area by mistake, k would have come out at about 33.4, below the mean, which contradicts the requirement that a minority lie above it. That side check catches the reversal every time it occurs.
OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 345-346
Discrimination
Translate each phrase for a normal distribution.
Sort into buckets
Sort each phrase by the area to the left it specifies.
Worked example
Showing that the same cut-off can be described from either side.
\[ \text{the top } 10 \text{ percent, against the } 90\text{th percentile} \]
Read the first phrase
Why: Ten percent above k.
\[ \text{right area } 0.10 \]
Convert
Why: One minus 0.10.
\[ \text{left area } 0.90 \]
Read the second phrase
Why: Ninety percent below k.
\[ \text{left area } 0.90 \]
Compare
Why: The same left area.
Figure (svg): The solution to Worked example two phrasings of one boundary shown as a ladder of expressions, one row per legal move
\[ P(x \ge k) = 0.10 \;\Longleftrightarrow\; P(x < k) = 0.90 \]
Verify: confirm why the two phrasings cannot disagree
Why: The regions above and below k make up the whole distribution, so specifying either one specifies the other — there is nothing left over for them to disagree about. The practical consequence is that any phrase can be converted to a left area, and the only question is whether the number quoted needs subtracting from one first. Deciding that is the entire skill.
OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 342-346
Trap
\[ \text{forty percent are at least } k \;\Rightarrow\; P(x < k) = 0.40 \]
Put the stated 0.40 straight into the calculator
Why: The question says forty percent, so 0.40 is the area.
\[ k \approx 33.4, \text{ which is BELOW the mean of } 36.9 \]
If only forty percent are at least k, then k must sit above the middle of the distribution, not below it.
\[ P(x \ge k) = 0.40 \;\Rightarrow\; P(x < k) = 0.60 \;\Rightarrow\; k \approx 40.4 \]
Ask which SIDE the quoted proportion describes before using it
Why: At least means the area to the right.
Both answers are plausible ages and neither is absurd on its face, which is what makes this the section's most costly error — nothing except the side check exposes it. Writing the phrase, then the side, then the left area on three separate lines before computing is exactly what the book's own solution does, and it is worth copying.
Faded example
Sixty percent of a normal population is at least k.
Fill in the blanks
P(x \ge k) = 0.60 \;\Rightarrow\; P(x < k) = 0.40, \textbelow k \text___ ___ \text___
Why: One minus 0.60 gives a left area of 0.40, which is below a half, so k sits below the mean. That makes sense: if a majority are at or above k, then k must be low.
Two truths and a lie
All three concern translating wording.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and believing it is what produces the section's worst errors. Phrases such as at least, the top, or more than name the area to the RIGHT, and the quoted proportion has to be subtracted from one before it can be used.
Explain it
A classmate answered forty percent are at least what age with 33.4 years for N(36.9, 13.9).
Discussion prompt
In two sentences or fewer, show them the error without doing the whole problem.
Hint: Ask whether a minority or a majority should be above their answer.
Answer:
Ask what fraction of the distribution lies above 33.4: more than half, since 33.4 is below the mean of 36.9 — but the question says only forty percent should be.
So 0.40 was the area to the RIGHT, not the left, and the calculator needed 0.60 instead, which gives 40.4 years.
Section
Section 4
Concept
A middle region is specified by the area it contains, so the remaining area is split evenly between the two tails by symmetry. Each boundary is then an ordinary percentile, and the quartiles are the special case where the middle region is fifty percent.
a middle region — The central portion holding a stated proportion. Subtract that proportion from one, halve the remainder to get each tail, and the two boundaries are the percentiles at that tail area and at one minus it.
\[ \text{middle } p \;\Rightarrow\; \text{tails of } \frac{1-p}{2} \text{ each} \]
The book works this for the middle 20 percent of mandarin diameters and spells out the arithmetic: one minus 0.20 is 0.80; the tails of the graph each have an area of 0.40; find the 40th percentile and the 60th, since 0.40 plus 0.20 is 0.60. The last step is worth pausing on — the upper boundary's left area is the lower tail plus the middle region, which is why it is 0.60 rather than 0.80.
Figure (svg): A normal curve for orange diameters with the middle twenty percent shaded between two cut-offs and forty percent left in each tail
OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 345-346 — Example 6.11(a) and Example 6.12(b)
Picture it
Example 6.12(b): mandarin diameters, with 40 percent left in each tail.
Figure (svg): A normal curve for orange diameters with the middle twenty percent shaded between two cut-offs and forty percent left in each tail
The shaded band is narrow because it holds only a fifth of the distribution, and it hugs the mean because that is where a normal's density is highest — the middle 20 percent of any normal spans only about half a standard deviation. Comparing the band's width against the curve's spread is a quick check that the tails were halved rather than used whole.
Worked example
Example 6.12(b).
\[ X \sim N(5.85, 0.24); \text{ the middle } 20 \text{ percent lie between which values?} \]
Find what is left over
Why: One minus 0.20.
\[ 0.80 \]
Split between two tails
Why: Half of 0.80.
\[ 0.40\text{ each} \]
Lower boundary
Why: The 40th percentile.
\[ 5.79 \text{cm} \]
Upper boundary
Why: 0.40 plus 0.20, the 60th.
\[ 5.91 \text{cm} \]
Figure (svg): The solution to Worked example the middle 20 percent of diameters shown as a ladder of expressions, one row per legal move
\[ k_1 = 5.79, \qquad k_2 = 5.91 \]
Verify: confirm the two boundaries are symmetric about the mean
Why: The mean is 5.85, and 5.79 and 5.91 sit 0.06 below and above it — equal distances, as symmetry requires. Any middle region of a normal distribution must come out symmetric about the mean, so boundaries at unequal distances signal that the tails were not halved correctly. This check needs no calculator at all.
OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, p. 346
Faded example
You want the middle 95 percent of a normal distribution.
Fill in the blanks
\text0.025 \frac2.5___ = ___, \text___ ___\text___ 97.5\text___
Why: Five percent is left over, so each tail holds 0.025 and the boundaries are the 2.5th and 97.5th percentiles. This exact arithmetic reappears in every 95 percent confidence interval from chapter 8 onward.
Worked example
Example 6.11(a), the middle 50 percent.
\[ X \sim N(36.9, 13.9); \text{ find the IQR} \]
Find Q3
Why: The 75th percentile.
\[ 46.2754 \]
Find Q1
Why: The 25th percentile.
\[ 27.5246 \]
Subtract
Why: Q3 minus Q1.
\[ 18.7508 \]
Round
Why: As the book does.
\[ 18.8\text{ years} \]
Figure (svg): The solution to Worked example the interquartile range shown as a ladder of expressions, one row per legal move
\[ \text{IQR} = 46.2754 - 27.5246 \approx 18.8 \]
Verify: confirm the IQR against a general rule for normal distributions
Why: The quartiles of any normal sit about 0.674 standard deviations either side of the mean, so the IQR is about 1.35 standard deviations — here 1.35 times 13.9, which is about 18.8. That constant is worth knowing because it converts an IQR back into a standard deviation, and it is how a boxplot can be read as evidence for or against normality.
OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, p. 345
Trap
\[ \text{middle } 20\text{ percent} \;\Rightarrow\; \text{tails of } 0.80 \text{ each} \]
Take the leftover 0.80 as each tail's area
Why: One minus 0.20 is 0.80, and that is what is outside.
\[ 0.80 + 0.20 + 0.80 = 1.80 > 1 \]
The three regions have to make up exactly one, so tails of 0.80 apiece are impossible before any percentile is computed.
\[ \text{tails of } \frac{1 - 0.20}{2} = 0.40 \text{ each} \]
Halve the leftover, since symmetry splits it evenly between two tails
Why: 0.40 plus 0.20 plus 0.40 gives exactly 1.
The totalling check is the reliable one and takes a second: the two tails plus the middle must sum to exactly one. It also generalises — for a middle 95 percent the tails are 0.025 each, which is the arithmetic behind every confidence interval in chapter 8, so getting the habit now pays off there.
Two truths and a lie
All three concern middle regions.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false: the leftover is SPLIT between two tails, so each holds half of it. Using the whole leftover for each tail makes the three regions total more than one, which is impossible and instantly checkable.
Estimation
A normal distribution has standard deviation 10.
Predict first
Roughly how wide is its interquartile range?
Correct: About 13.5.
Why: The quartiles sit about 0.674 standard deviations either side of the mean, so the IQR is about 1.35 standard deviations, which is 13.5 here. That ratio holds for every normal distribution, which makes it useful in reverse: dividing an observed IQR by 1.35 estimates the standard deviation.
Prediction
Commit before reasoning.
Predict first
For the middle 20 percent of a normal distribution, how far apart are the two boundaries?
Correct: About half a standard deviation.
Why: The 40th and 60th percentiles sit at z-scores of about minus 0.253 and plus 0.253, so they are about 0.51 standard deviations apart. For Example 6.12 that is 0.51 times 0.24, about 0.12 cm — matching the gap between 5.79 and 5.91. The mean shifts both boundaries equally and so cannot affect the gap.
Section
Section 5
Concept
The book asks repeatedly for the answer to be interpreted in a complete sentence, and every worked solution ends with one. A number alone does not say which quantity it refers to, in what units, or on which side of a boundary.
interpreting in context — Restating the numeric answer as a claim about the original quantity: not 48.6 but eighty percent of the smartphone users in this age range are 48.6 years old or less.
\[ 48.6 \;\longrightarrow\; \text{“80 percent are 48.6 years old or less”} \]
The interpretation is not decoration; it is where a reversed area or a misread question shows itself. Writing the sentence for the wrong answer to Example 6.11(b) would give forty percent of users are at least 33.4 years, and stating it that plainly invites the question of whether that can be right when the average is 36.9 — which it cannot. Forcing yourself to say what a number means is a genuine error-detection step.
Figure (svg): A normal curve for smartphone users' ages with the interquartile range shaded between the first and third quartiles
OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 343-346 — the book's repeated instruction to interpret in a complete sentence
Picture it
Example 6.11(a): quartiles and the IQR for smartphone users.
Figure (svg): A normal curve for smartphone users' ages with the interquartile range shaded between the first and third quartiles
Interpreted, this says the middle half of smartphone users in the 13 to 55-plus range are between about 27.5 and 46.3 years old — a span of 18.8 years. Stated that way it can be checked against what anyone knows about smartphone users, which a bare 18.8 could not be.
Worked example
Example 6.10, interpreted rather than merely computed.
\[ X \sim N(36.9, 13.9) \]
Between 23 and 64.7
Why: The probability.
\[ 0.8186 \]
Say it in words
Why: About 82 percent.
At most 50.8
Why: The left area.
\[ 0.8413 \]
The 80th percentile
Why: A value, not a probability.
\[ 48.6\text{ years} \]
Figure (svg): The solution to Worked example three answers, three sentences shown as a ladder of expressions, one row per legal move
\[ 0.8186, \quad 0.8413, \quad 48.6 \text{ years} \]
Verify: confirm the third answer is a different kind of object from the first two
Why: The first two are probabilities and must lie between zero and one; the third is an age and must lie in the plausible range of the data. Mixing them up is the error the first idea's trap describes, and simply writing the units beside each answer makes it impossible to confuse them. The book's own solution ends with a full sentence for exactly this reason.
OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 344-345
Two truths and a lie
All three concern interpreting.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and the distinction matters. Such a check confirms the ARITHMETIC given the assumed normal model; if the quantity is not really normally distributed, the check will pass while the answer remains wrong. Testing the model itself needs the goodness-of-fit methods of chapter 11.
Worked example
Example 6.10(b), where the cut-off is exactly one standard deviation up.
\[ \text{find } P(x \le 50.8) \text{ for } N(36.9, 13.9) \]
Locate the cut-off
Why: 36.9 plus 13.9.
Recall the rule
Why: 68 percent within one sigma.
\[ 32 \%\text{ outside} \]
Halve the outside
Why: One upper tail.
\[ 16 \% \]
Add
Why: 68 plus 16.
\[ \text{about } 84 \% \]
Figure (svg): The solution to Worked example checking an answer against the rule shown as a ladder of expressions, one row per legal move
\[ 0.68 + \frac{1 - 0.68}{2} = 0.84 \quad\text{against}\quad 0.8413 \]
Verify: confirm what the agreement does and does not establish
Why: It confirms the calculation was set up correctly — the right distribution, the right cut-off, the right side. It does not confirm the model: if ages are not really normal, both numbers are equally wrong. Sanity checks of this kind test the arithmetic against the assumptions, never the assumptions themselves, and keeping that distinction clear matters from chapter 9 onward.
OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 344-345
Trap
\[ \text{answer: } 48.6 \]
Stop once the calculator has produced a value
Why: The arithmetic is finished.
\[ 48.6 \text{ what, and above or below?} \]
Without units and a direction the number cannot be checked by anyone, including the person who computed it.
\[ \text{80 percent of users are } 48.6 \text{ years old or less} \]
State the quantity, the units and the direction in one sentence
Why: The book asks for this on every worked example.
The sentence is a test, not a formality. Writing eighty percent of users are 48.6 years old or less invites an immediate check against the mean of 36.9 — the boundary is above the mean, as a percentile above the 50th must be. A bare 48.6 offers nothing to check against, which is why errors survive it.
Sorting
Ask what kind of object each answer must be.
Sort into buckets
Sort each question by the kind of answer it needs.
Doing this sort before touching a calculator is worth the few seconds, because it settles which command is needed and gives an immediate test of the answer's plausibility. An answer of the wrong kind is always wrong, whatever its value.
Explain it
A classmate has correctly computed the 80th percentile of smartphone users' ages as 48.6 and stopped there.
Discussion prompt
In one or two sentences, state what the number means and one check it passes.
Hint: Say which group, which direction, and compare with the mean.
Answer:
Eighty percent of smartphone users in the 13 to 55-plus range are 48.6 years old or less, and the remaining twenty percent are older than that.
It passes the side check: 48.6 sits above the mean of 36.9, as any percentile above the 50th must.
Prediction
Commit before reasoning.
Predict first
An answer agrees with the Empirical Rule's estimate. What has been established?
Correct: That the calculation was set up correctly, assuming the model.
Why: The check compares two routes through the same assumed distribution, so it catches a wrong cut-off, a wrong side or a mistyped parameter — but both routes rest on normality, so neither can test it. The rule is also only approximate, so agreement to two decimal places is all it can offer.
Comparison
Fill the blanks. Telling them apart is most of this section.
Comparison matrix
| Probability question | Percentile question | |
|---|---|---|
| What is given | one or two boundaries | an area |
| What is wanted | an area | a boundary |
| The answer's units | none: it is between 0 and 1 | the problem's own units |
| Signal phrase | find the probability that... | the 90th percentile; at least; the top 10 percent |
The units row is the most useful check in the whole section. An answer between zero and one to a question asking for an age, or an answer in years to a question asking for a probability, is wrong before any of the arithmetic is examined.
Pattern
Six steps, and the third is where this section's marks are won and lost.
The Empirical Rule is the sanity check throughout: a cut-off one, two or three standard deviations out should give areas near 0.84, 0.977 and 0.9987 to its left.
OpenStax Introductory Business Statistics 2e, §6.2 Using the Normal Distribution §6.2 Using the Normal Distribution
Check
A right tail.
Check your understanding
Exam scores are N(63, 5). The area to the left of 65 is 0.6554. What is P(x > 65)?
Answer: A
Why: The right tail is one minus the left area: 1 minus 0.6554 gives 0.3446.
Check
Translating a phrase.
Check your understanding
For X ~ N(36.9, 13.9), forty percent of values are at least k. What left area should be used to find k?
Answer: A
Why: At least means the area to the right is 0.40, so the area to the left is one minus 0.40, which is 0.60.
Check
A middle region.
Check your understanding
You want the middle 20 percent of a normal distribution. What area lies in each tail?
Answer: A
Why: One minus 0.20 leaves 0.80, split evenly between two tails, so each holds 0.40.
Real world
A manufacturer machines bolts whose diameters are normally distributed with mean 10.00 mm and standard deviation 0.04 mm. The specification accepts bolts between 9.92 and 10.08 mm. A production manager proposes tightening the specification to 9.96 to 10.04 mm to improve quality, and expects the scrap rate to rise only slightly.
Discussion prompt
Compute the scrap rate under each specification, assess the manager's expectation, and say what would actually reduce scrap.
Hint: Express each specification limit in standard deviations before computing anything.
Answer:
Express the limits in standard deviations first. The current limits sit at 10.00 give or take 0.08, which is exactly two standard deviations, so about 95 percent of bolts conform and roughly 4.6 percent are scrapped. The proposed limits sit at one standard deviation, so about 68 percent conform and about 32 percent are scrapped.
\[ \text{current: } 1 - 0.9545 \approx 0.046, \qquad \text{proposed: } 1 - 0.6827 \approx 0.317 \]
The manager's expectation is badly wrong: scrap would rise sevenfold, from about 1 bolt in 22 to nearly 1 in 3. The intuition that halving the tolerance costs only a little fails because the normal curve is densest near the mean, so the region being given up — between one and two standard deviations — contains about 27 percent of all output, far more than the 4.6 percent already outside.
Tightening the specification does not improve quality; it only reclassifies more output as scrap. The bolts being produced are unchanged. What would genuinely reduce scrap is reducing the standard deviation of the process itself: at a sigma of 0.02 mm the proposed limits would again sit two standard deviations out, and conformance would return to about 95 percent.
Two further points belong in a full answer. The calculation assumes the process stays centred at 10.00 mm, and a drift in the mean of even half a standard deviation raises scrap sharply on one side while barely reducing it on the other — so monitoring the centre matters as much as the spread. And this is exactly the logic behind process capability indices in manufacturing, which express the specification width as a number of standard deviations rather than in millimetres, for precisely the reason this problem illustrates.
Commit first
Answer, then rate your confidence honestly.
Predict first
A question says forty percent of values are at least k. What area do you give the calculator to find k?
Correct: 0.60.
\[ P(x \ge k) = 0.40 \;\Longrightarrow\; P(x < k) = 0.60 \]
Why: At least translates to greater than or equal to, so 0.40 is the area to the right of k. The technology reports and reverses left areas only, so it needs one minus 0.40. The side check confirms it: with only forty percent above k, the boundary must lie above the mean, and 0.60 puts it there while 0.40 would put it below.
Explain it
Asked for the 90th percentile of N(63, 5), they computed P(X < 90) and got approximately 1.
Discussion prompt
In two sentences or fewer, locate the error.
Hint: Ask what kind of object the answer should be.
Answer:
Point out that the answer should be an exam mark, and 1 is a probability — so the wrong kind of question was answered before any arithmetic is even checked.
The 90 was a percentage rather than a score, so it is the AREA that equals 0.90, and reversing that gives a mark of 69.4.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, a right tail is one minus the left area and an interval is a difference of two. For the second, ask what the question gives and what it wants, then check the answer's units. For the third, decide which SIDE the quoted proportion describes before using it. For the fourth, subtract from one and halve, then check the two boundaries are symmetric about the mean. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Twenty minutes.
Draw it
Draw six normal curves down the page, all for the same distribution where possible, and label the mean on each. On the first, shade above 65 for exam scores from N(63, 5) and compute the right tail as one minus the left area, checking against 0.3446. On the second, shade between 1.8 and 2.75 for N(2, 0.5) and compute it as a difference of two left areas, checking against 0.5886. On the third, shade ninety percent of the area for N(63, 5) and mark the boundary, checking against 69.4 — and write beside it which quantity was given and which was wanted, to contrast it with the first two. On the fourth, take N(36.9, 13.9) and answer forty percent are at least what age: write the phrase, then the side it names, then the left area, then the answer, on four separate lines, and check 40.4 lies above the mean. On the fifth, shade the middle twenty percent of N(5.85, 0.24), writing the leftover, the halved tails, and the two percentiles, and confirm the boundaries 5.79 and 5.91 are equally far from 5.85. On the sixth, shade the interquartile range of N(36.9, 13.9), compute Q1 and Q3, and check the IQR against 1.35 standard deviations. Finish by writing one full interpreting sentence for each of the six answers.
Check the fifth curve by adding your three areas: 0.40 plus 0.20 plus 0.40 must give exactly 1. Check the sixth by dividing your IQR by the standard deviation of 13.9 — it should come out near 1.35 for any normal distribution, and a very different ratio means a quartile was computed from the wrong area.
Recap
Five things, and the third is the one worth practising most.
| If you see | Then |
|---|---|
| A right tail wanted | One minus the left area |
| An interval wanted | Upper left area minus lower |
| A proportion given, a value wanted | A percentile: reverse the cumulative |
| At least, or the top so-many percent | The quoted proportion is the RIGHT area |
| The middle so-many percent | Subtract from one, halve, then two percentiles |
| An answer between 0 and 1 to a value question | The wrong kind of question was answered |
| A cut-off one sigma above the mean | The left area should be about 0.84 |
That closes chapter 6. Chapter 7 supplies the result that makes the normal distribution matter far beyond the quantities that happen to be bell shaped: the central limit theorem, which says that sample means are approximately normal whatever the population looks like — so the machinery of this chapter applies to almost any inference, and chapters 8 onward are built on it.
OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 340-346 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.