6.2 Using the Normal Distribution

Section 6.1's machinery put to work. With a calculator or computer returning areas under the normal curve, three questions can be answered about any normally distributed quantity: the probability of a tail, the probability of an interval, and the value at a given percentile. The technology reports only the area to the left, so a right tail is found by subtracting from one and an interval by subtracting one left area from another, while a percentile reverses the process by taking an area and returning the boundary. The genuine difficulty in this section is neither the arithmetic nor the technology but the translation of English phrases into areas to the left: at most, at least, the bottom quartile, the top ten percent and the middle twenty percent all specify an area, and two of them specify the complement of the number they mention.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 6.2 Using the Normal Distribution

Title

Statistics · Chapter 6 — The Normal Distribution

Using the Normal Distribution

2. By the end of this lesson you can

Objectives

Five outcomes, and the middle two are where the marks are actually lost.

OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 340-346 — the section these objectives are drawn from

3. What you already have

Warm-up

Section 6.1 gave the z-score and the standard normal; section 5.1 gave the complement rule for areas.

Discussion prompt

Exam scores are normal with mean 63 and standard deviation 5. Roughly what fraction of students score above 65, and can you answer it without a calculator?

Hint: How many standard deviations above the mean is 65?

Answer:

A score of 65 is two points above the mean of 63, which is 0.4 standard deviations — so the answer must be a bit less than half, since a score exactly at the mean would leave half above it and 65 is slightly higher.

The Empirical Rule cannot do better than that, because 0.4 is not a whole number of standard deviations. It brackets the answer between roughly 0.32 and 0.50, which is a real constraint but not an answer.

That gap is what this section fills. A calculator gives 0.3446 exactly, and the estimate's role changes from producing the answer to checking it — which is exactly how it should be used from here on, since a mistyped command produces a number with no warning at all.

4. Everything is built from the area to the left

Concept

The cumulative distribution function P(X < x) gives the area to the left, and it is what a calculator, a computer or a table reports. A right tail is one minus it, and an interval is the difference of two of them. A percentile runs the process backwards: the area is given and the boundary is wanted.

area to the left — The single quantity the technology returns. Every other normal question — right tails, intervals, percentiles, middle regions, quartiles — is built from it by subtraction or by reversal.

\[ P(X > x) = 1 - P(X < x), \qquad P(a < X < b) = P(X < b) - P(X < a) \]

The book is explicit that the tables are largely historical: technology has made the tables virtually obsolete, and for that reason, as well as the fact that there are various table formats, we are not including table instructions. What survives from the table era is the convention that left areas are what get reported — so the two subtractions above are as necessary now as they were then.

Figure (svg): Two identical bell curves, the left one shaded below a cut-off and the right one shaded above it, with the two areas summing to one

The same complement rule from section 3.2 and section 5.1, now applied to normal areas.

OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 341-342

5. Tails and intervals

Section

Section 1

6. One subtraction, or two

Concept

A right tail is one minus the area to the left of the cut-off. An interval is the area to the left of the upper endpoint minus the area to the left of the lower one. Neither needs anything the technology does not already give.

the two subtractions — One minus a left area gives a right tail; one left area minus another gives an interval. Together they answer every probability question about a normal distribution.

\[ P(a < X < b) = P(X < b) - P(X < a) \]

The book shows the historical route for Example 6.8(a) alongside the calculator one, and it is worth following once. The z-score is 65 minus 63 over 5, which is 0.4; the area to the left of 0.4 is 0.6554; so the right tail is 1 minus 0.6554, or 0.3446. That is the same answer the calculator gives directly from the unstandardised values, which confirms that standardising changes nothing — it was simply necessary when only one distribution could be tabulated.

Figure (svg): A normal curve for exam scores with the region above sixty-five shaded

The technology reports area to the left, so a right tail is one minus that: 1 - 0.6554 = 0.3446.

OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 341-344 — Examples 6.7, 6.8(a)-(b) and 6.9(a)

7. A right tail from a left area

Picture it

Example 6.8(a): the chance a student scores above 65.

Figure (svg): A normal curve for exam scores with the region above sixty-five shaded

The technology reports area to the left, so a right tail is one minus that: 1 - 0.6554 = 0.3446.

The shaded region is a little under half the curve, which matches the warm-up's estimate — 65 is only just above the mean, so the tail above it should be only just below a half. Making that comparison before accepting 0.3446 costs nothing and catches a mistyped mean or standard deviation immediately.

8. Worked example: more than 65 on the exam

Worked example

Example 6.8(a), by both routes.

\[ X \sim N(63, 5); \text{ find } P(x > 65) \]

Standardise the cut-off

Why: 65 minus 63, over 5.

\[ z = 0.4 \]

Find the area to the left

Why: Below z = 0.4.

\[ 0.6554 \]

Subtract from one

Why: The right tail.

\[ 0.3446 \]

Check against technology

Why: Directly from N(63, 5).

\[ 0.3446 \]

Figure (svg): The solution to Worked example more than 65 on the exam shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(x > 65) = 1 - 0.6554 = 0.3446 \]

Verify: confirm the answer is on the right side of a half

Why: Since 65 exceeds the mean of 63, less than half the distribution lies above it — so the answer had to come out below 0.50, and 0.3446 does. Had the subtraction from one been forgotten, the answer would have been 0.6554, which is above a half and therefore impossible for a cut-off above the mean. That single comparison catches the section's commonest error.

OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 341-342

9. A right tail

Faded example

Mandarin diameters are N(5.85, 0.24). The left area at 6.0 is 0.7340.

Fill in the blanks

P(x > 6.0) = 1 - 0.7340 = 0.2660

Why: One minus 0.7340 gives 0.2660, the book's answer for Example 6.12(a). Since 6.0 cm is above the mean of 5.85, the tail above it must be below a half — and it is.

10. Worked example: between 1.8 and 2.75 hours

Worked example

Example 6.9(a), an interval.

\[ X \sim N(2, 0.5); \text{ find } P(1.8 < x < 2.75) \]

Left area at 2.75

Why: The upper endpoint.

\[ 0.9332 \]

Left area at 1.8

Why: The lower endpoint.

\[ 0.3446 \]

Subtract

Why: Upper minus lower.

\[ 0.5886 \]

State it in words

Why: Hours of entertainment use.

\[ \text{about } 59 \% \]

Figure (svg): The solution to Worked example between 1.8 and 2.75 hours shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(1.8 < x < 2.75) = 0.9332 - 0.3446 = 0.5886 \]

Verify: confirm the interval's answer is plausible for its width

Why: The interval runs from 0.4 standard deviations below the mean to 1.5 above, so it should capture well over half the distribution but nothing like all of it — and 0.5886 sits comfortably in that range. It is worth noticing that the left area at 1.8 is 0.3446, the same number as the previous example's answer, which is a coincidence of the two problems' z-scores both being 0.4 in magnitude.

OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 343-344

11. Error analysis: four attempts at P(x > 65) for N(63, 5)

Error analysis

The book's answer is 0.3446.

Annotate

On: \( \begin{aligned} &(1)\; 0.6554 \\ &(2)\; 1 - P(X < 63) = 0.5 \\ &(3)\; P(X < 65) - P(X < 63) = 0.1554 \\ &(4)\; 1 - 0.6554 = 0.3446 \end{aligned} \)

  • (1) reports the left area without subtracting. It is above a half, which is impossible for a cut-off above the mean.
  • (2) uses the mean as the cut-off instead of 65, answering a question nobody asked.
  • (3) computes the interval between the mean and 65 rather than the tail beyond 65.
  • (4) is correct: one minus the left area at 65.

Errors (1) and (3) both produce numbers that look like probabilities, and only (1) is caught by the above-or-below-a-half test. Shading the region on a sketch before computing is what separates (3) from (4), since the two describe visibly different areas.

12. Which subtraction?

Discrimination

Decide what each question needs.

Sort into buckets

Sort each question about a normal X.

One minus a left area
P(x > 65); P(x at least 40)
A left area minus another, or read directly
P(1.8 < x < 2.75); P(x < 85); P(23 < x < 64.7)
one
A right tail, so subtract the left area from one.
two
Either an interval needing two left areas, or a left area that can be read straight off.

13. One of these is false

Two truths and a lie

All three concern tails and intervals.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. A right tail is one minus the left area
  • C. An interval is the difference of two left areas
  • B. A cut-off above the mean gives a right tail above 0.5

Survives elimination: B

Why: The survivor is false and reverses the relationship. A cut-off above the mean leaves LESS than half the distribution above it, so the right tail must be below 0.5. This is the fastest available check on any tail computation, and it catches a forgotten subtraction immediately.

14. Check against the Empirical Rule

Estimation

Smartphone users' ages are N(36.9, 13.9), and 50.8 is exactly one standard deviation above the mean.

Predict first

Roughly what is P(x < 50.8)?

  • About 0.84
  • About 0.68
  • About 0.95
  • About 0.50

Correct: About 0.84.

Why: Sixty-eight percent lies within one standard deviation, leaving 32 percent split into two tails of 16 percent each — so the area below one standard deviation above the mean is 68 plus 16, about 84 percent. The calculator gives 0.8413 for Example 6.10(b), which confirms it. Recognising that 50.8 is exactly mu plus sigma is what makes this check available.

15. Percentiles

Section

Section 2

16. Given the area, find the boundary

Concept

A percentile question hands you the area and asks for the value. The 90th percentile is the score k that has 90 percent of the scores below it and 10 percent above, and it is found by reversing the cumulative function rather than evaluating it.

critical value — The book's other name for the percentile k. The term becomes standard from chapter 8, where the same reversal supplies the boundaries of a confidence interval.

\[ P(X < k) = p \;\Longrightarrow\; k = \text{invNorm}(p, \mu, \sigma) \]

The relationship to section 6.1's two formulas is worth seeing. A percentile could be found by hand in two steps: look up the z-score with the given area to its left, then convert with x equals mu plus z sigma. For the 90th percentile that is z equal to about 1.28, giving 63 plus 1.28 times 5, or about 69.4. The calculator's single command does both steps at once, which is convenient but hides the standardising — worth doing by hand once to see that nothing new is happening.

Figure (svg): A normal curve for exam scores with ninety percent of the area shaded and the cut-off marked at sixty-nine point four

The book notes that k is often called a critical value, a name that becomes standard from chapter 8 onward.

OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 342-344 — Examples 6.8(c)-(d) and 6.9(b)

17. The area is given; the boundary is not

Picture it

Example 6.8(c): the 90th percentile of exam scores.

Figure (svg): A normal curve for exam scores with ninety percent of the area shaded and the cut-off marked at sixty-nine point four

The book notes that k is often called a critical value, a name that becomes standard from chapter 8 onward.

Compare this picture with the first idea's. There the dashed line was given and the shading was the answer; here the shading is given and the dashed line is the answer. Recognising which of the two you are looking at is the whole of the difficulty, and the phrase to watch for is one naming a proportion rather than a value.

18. Worked example: the 90th and 70th percentiles

Worked example

Example 6.8(c) and (d).

\[ X \sim N(63, 5); \text{ find the } 90\text{th and } 70\text{th percentiles} \]

Set up the 90th

Why: Ninety percent below k.

\[ P(x < k) = 0.90 \]

Reverse the cumulative

Why: With mean 63 and sigma 5.

\[ k = 69.4 \]

Set up the 70th

Why: Seventy percent below.

\[ P(x < k) = 0.70 \]

Reverse again

Why: Same distribution.

\[ k = 65.6 \]

Figure (svg): The solution to Worked example the 90th and 70th percentiles shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ k_{90} = 69.4, \qquad k_{70} = 65.6 \]

Verify: confirm both sit above the mean, and in the right order

Why: Both percentiles exceed 50, so both boundaries must lie above the mean of 63 — and 65.6 and 69.4 do. The 90th must also exceed the 70th, since a larger area to the left needs a boundary further right. Two percentiles that came out in the wrong order, or on the wrong side of the mean, would signal that the areas had been entered as their complements.

OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 342-343

19. Find a percentile

Faded example

Smartphone users' ages are N(36.9, 13.9). Find the 80th percentile.

Fill in the blanks

P(x < k) = 0.80 \;\Rightarrow\; k \approx 48.6 \text___

Why: Eighty percent below gives k = 48.6 years, so 80 percent of users in the age range are 48.6 years old or less. The answer exceeds the mean of 36.9, as any percentile above the 50th must.

20. Worked example: the bottom quartile

Worked example

Example 6.9(b), where the wording names a quarter rather than a percentile.

\[ X \sim N(2, 0.5); \text{ the bottom quartile uses at most how many hours?} \]

Translate the phrase

Why: The bottom quarter.

Write the equation

Why: Twenty-five percent below.

\[ P(x < k) = 0.25 \]

Reverse the cumulative

Why: With mean 2 and sigma 0.5.

\[ k = 1.66 \]

Interpret

Why: In the problem's units.

\[ 1.66\text{ hours} \]

Figure (svg): The solution to Worked example the bottom quartile shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ k = \text{invNorm}(0.25,\; 2,\; 0.5) = 1.66 \]

Verify: confirm the answer falls below the mean, as a bottom quartile must

Why: Twenty-five percent is below a half, so the boundary must lie below the mean of 2 hours — and 1.66 does. It is also about 0.67 standard deviations below, which matches the standard normal's first quartile of about minus 0.674; that constant is worth recognising, since the first and third quartiles of ANY normal sit about two thirds of a standard deviation either side of the mean.

OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, p. 344

21. Trap: computing a probability when a percentile was asked

Trap

The trap

\[ \text{the } 90\text{th percentile of } N(63,5) \;\Rightarrow\; P(X < 90) \approx 1 \]

Feed the number 90 in as a score

Why: The question mentions 90, and scores go into the cumulative function.

\[ \text{an answer of } 1, \text{ which is a probability and not a score} \]

Ninety was a percentage, not a mark; the question supplied an area and wanted a boundary.

The fix

\[ P(X < k) = 0.90 \;\Rightarrow\; k = 69.4 \]

Ask what the question GIVES and what it WANTS before touching the calculator

Why: A proportion given means a percentile problem.

The units test settles it instantly: a percentile answer must come out in the problem's units — marks, hours, centimetres — while a probability is always between zero and one. An answer of 1 to a question asking for a score is self-evidently the wrong kind of object, and noticing that is faster than re-reading the question.

22. One of these is false

Two truths and a lie

All three concern percentiles.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. A percentile question gives an area and wants a value
  • C. A percentile above the 50th lies above the mean
  • B. The answer to a percentile question lies between 0 and 1

Survives elimination: B

Why: The survivor is false and confuses the two kinds of question. A percentile is a value in the problem's own units — 69.4 marks, 1.66 hours, 48.6 years — while it is the AREA that lies between 0 and 1. Checking the units of the answer is the quickest way to confirm the right question was answered.

23. Which side of the mean?

Prediction

Commit before reasoning.

Predict first

For any normal distribution, where does the 30th percentile lie?

  • Below the mean
  • Above the mean
  • At the mean
  • It depends on sigma

Correct: Below the mean.

Why: Only 30 percent of the area lies to its left, which is less than half, so the boundary must sit left of centre. The standard deviation affects how far below it falls but never which side. Checking the side before computing rules out half the possible errors at no cost.

24. Percentile to position

Matching

Match each percentile to where it falls.

Match the pairs

  • l1. the 50th percentile
  • l2. the 25th percentile
  • l3. the 90th percentile
  • l4. the 75th percentile
  • r1. exactly at the mean
  • r2. about 0.67 standard deviations below the mean
  • r3. about 1.28 standard deviations above the mean
  • r4. about 0.67 standard deviations above the mean

Why: These four z-scores — 0, minus 0.67, 1.28 and 0.67 — are the same for every normal distribution, which is what makes them worth recognising. The 1.28 in particular recurs constantly from chapter 8 onward, where it supplies the boundary of a 90 percent one-sided interval.

25. Translating the wording

Section

Section 3

26. The phrase names an area, sometimes its complement

Concept

Every calculator command reports area to the left, so a worded question has to be converted into one before anything can be computed. Some phrases name that area directly and others name what remains, and the book's own worked solution spells the conversion out.

at least — The book translates it as greater than or equal to, so the given proportion is the area to the RIGHT. The area to the left is one minus it, and that is what goes into the calculation.

\[ P(x \ge k) = 0.40 \;\Rightarrow\; P(x < k) = 0.60 \]

Example 6.11(b) is worth reading in the book's own words because it shows the steps a careful solution makes explicit: find k where P(x at least k) equals 0.40; at least translates to greater than or equal to; 0.40 is the area to the right; area to the left equals one minus 0.40, which is 0.60; the area to the left of k is 0.60. Four short lines, none of them arithmetic, and they are where the question is actually answered.

Figure (svg): A translation table turning English phrases into areas to the left

The book asks the middle-percent question and the at-least question in consecutive examples, because they are the two that get reversed.

OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 345-346 — Example 6.11(b) and Example 6.12(b)

27. Five phrases, five left areas

Picture it

Two of them name the complement of the number they mention.

Figure (svg): A translation table turning English phrases into areas to the left

The book asks the middle-percent question and the at-least question in consecutive examples, because they are the two that get reversed.

The third and fourth rows are the ones that get reversed, and they look almost identical in print — the bottom quartile uses 0.25 while the top ten percent uses 0.90. Reading the direction word before the number, rather than after it, is the habit that keeps them apart.

28. Worked example: at least what age?

Worked example

Example 6.11(b), following the book's own steps.

\[ X \sim N(36.9, 13.9); \text{ forty percent are at least what age?} \]

Write the requirement

Why: Forty percent at or above k.

\[ P(x \ge k) = 0.40 \]

Identify which side

Why: At least means to the right.

\[ 0.40\text{ is the RIGHT area} \]

Convert to a left area

Why: One minus 0.40.

\[ 0.60 \]

Reverse the cumulative

Why: At 0.60.

\[ k = 40.4 \]

Figure (svg): The solution to Worked example at least what age shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(x < k) = 0.60 \;\Rightarrow\; k \approx 40.4 \text{ years} \]

Verify: confirm the answer lies above the mean, and say why that is right

Why: Forty percent lie above k, which is less than half, so k must sit above the mean of 36.9 — and 40.4 does. Had 0.40 been used as the left area by mistake, k would have come out at about 33.4, below the mean, which contradicts the requirement that a minority lie above it. That side check catches the reversal every time it occurs.

OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 345-346

29. Which left area?

Discrimination

Translate each phrase for a normal distribution.

Sort into buckets

Sort each phrase by the area to the left it specifies.

Left area 0.30
the 30th percentile; seventy percent are at least this; at most this value, for 30 percent; the top 70 percent start here
Left area 0.70
seventy percent fall below this
p30
Either the phrase names 30 percent below, or it names 70 percent above — which is the same boundary.
p70
The phrase names seventy percent below the value directly.

30. Worked example: two phrasings of one boundary

Worked example

Showing that the same cut-off can be described from either side.

\[ \text{the top } 10 \text{ percent, against the } 90\text{th percentile} \]

Read the first phrase

Why: Ten percent above k.

\[ \text{right area } 0.10 \]

Convert

Why: One minus 0.10.

\[ \text{left area } 0.90 \]

Read the second phrase

Why: Ninety percent below k.

\[ \text{left area } 0.90 \]

Compare

Why: The same left area.

Figure (svg): The solution to Worked example two phrasings of one boundary shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(x \ge k) = 0.10 \;\Longleftrightarrow\; P(x < k) = 0.90 \]

Verify: confirm why the two phrasings cannot disagree

Why: The regions above and below k make up the whole distribution, so specifying either one specifies the other — there is nothing left over for them to disagree about. The practical consequence is that any phrase can be converted to a left area, and the only question is whether the number quoted needs subtracting from one first. Deciding that is the entire skill.

OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 342-346

31. Trap: using the quoted number as the left area

Trap

The trap

\[ \text{forty percent are at least } k \;\Rightarrow\; P(x < k) = 0.40 \]

Put the stated 0.40 straight into the calculator

Why: The question says forty percent, so 0.40 is the area.

\[ k \approx 33.4, \text{ which is BELOW the mean of } 36.9 \]

If only forty percent are at least k, then k must sit above the middle of the distribution, not below it.

The fix

\[ P(x \ge k) = 0.40 \;\Rightarrow\; P(x < k) = 0.60 \;\Rightarrow\; k \approx 40.4 \]

Ask which SIDE the quoted proportion describes before using it

Why: At least means the area to the right.

Both answers are plausible ages and neither is absurd on its face, which is what makes this the section's most costly error — nothing except the side check exposes it. Writing the phrase, then the side, then the left area on three separate lines before computing is exactly what the book's own solution does, and it is worth copying.

32. Convert the phrase

Faded example

Sixty percent of a normal population is at least k.

Fill in the blanks

P(x \ge k) = 0.60 \;\Rightarrow\; P(x < k) = 0.40, \textbelow k \text___ ___ \text___

Why: One minus 0.60 gives a left area of 0.40, which is below a half, so k sits below the mean. That makes sense: if a majority are at or above k, then k must be low.

33. One of these is false

Two truths and a lie

All three concern translating wording.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. At least translates to the area on the right
  • C. The top 10 percent and the 90th percentile name the same boundary
  • B. The proportion quoted in the question is always the left area

Survives elimination: B

Why: The survivor is false, and believing it is what produces the section's worst errors. Phrases such as at least, the top, or more than name the area to the RIGHT, and the quoted proportion has to be subtracted from one before it can be used.

34. Explain the reversal

Explain it

A classmate answered forty percent are at least what age with 33.4 years for N(36.9, 13.9).

Discussion prompt

In two sentences or fewer, show them the error without doing the whole problem.

Hint: Ask whether a minority or a majority should be above their answer.

Answer:

Ask what fraction of the distribution lies above 33.4: more than half, since 33.4 is below the mean of 36.9 — but the question says only forty percent should be.

So 0.40 was the area to the RIGHT, not the left, and the calculator needed 0.60 instead, which gives 40.4 years.

35. Middle regions and quartiles

Section

Section 4

36. Split what is left into two equal tails

Concept

A middle region is specified by the area it contains, so the remaining area is split evenly between the two tails by symmetry. Each boundary is then an ordinary percentile, and the quartiles are the special case where the middle region is fifty percent.

a middle region — The central portion holding a stated proportion. Subtract that proportion from one, halve the remainder to get each tail, and the two boundaries are the percentiles at that tail area and at one minus it.

\[ \text{middle } p \;\Rightarrow\; \text{tails of } \frac{1-p}{2} \text{ each} \]

The book works this for the middle 20 percent of mandarin diameters and spells out the arithmetic: one minus 0.20 is 0.80; the tails of the graph each have an area of 0.40; find the 40th percentile and the 60th, since 0.40 plus 0.20 is 0.60. The last step is worth pausing on — the upper boundary's left area is the lower tail plus the middle region, which is why it is 0.60 rather than 0.80.

Figure (svg): A normal curve for orange diameters with the middle twenty percent shaded between two cut-offs and forty percent left in each tail

A middle region is two percentile problems, and the arithmetic that finds them is the only part that can go wrong.

OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 345-346 — Example 6.11(a) and Example 6.12(b)

37. The middle twenty percent

Picture it

Example 6.12(b): mandarin diameters, with 40 percent left in each tail.

Figure (svg): A normal curve for orange diameters with the middle twenty percent shaded between two cut-offs and forty percent left in each tail

A middle region is two percentile problems, and the arithmetic that finds them is the only part that can go wrong.

The shaded band is narrow because it holds only a fifth of the distribution, and it hugs the mean because that is where a normal's density is highest — the middle 20 percent of any normal spans only about half a standard deviation. Comparing the band's width against the curve's spread is a quick check that the tails were halved rather than used whole.

38. Worked example: the middle 20 percent of diameters

Worked example

Example 6.12(b).

\[ X \sim N(5.85, 0.24); \text{ the middle } 20 \text{ percent lie between which values?} \]

Find what is left over

Why: One minus 0.20.

\[ 0.80 \]

Split between two tails

Why: Half of 0.80.

\[ 0.40\text{ each} \]

Lower boundary

Why: The 40th percentile.

\[ 5.79 \text{cm} \]

Upper boundary

Why: 0.40 plus 0.20, the 60th.

\[ 5.91 \text{cm} \]

Figure (svg): The solution to Worked example the middle 20 percent of diameters shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ k_1 = 5.79, \qquad k_2 = 5.91 \]

Verify: confirm the two boundaries are symmetric about the mean

Why: The mean is 5.85, and 5.79 and 5.91 sit 0.06 below and above it — equal distances, as symmetry requires. Any middle region of a normal distribution must come out symmetric about the mean, so boundaries at unequal distances signal that the tails were not halved correctly. This check needs no calculator at all.

OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, p. 346

39. Split the tails

Faded example

You want the middle 95 percent of a normal distribution.

Fill in the blanks

\text0.025 \frac2.5___ = ___, \text___ ___\text___ 97.5\text___

Why: Five percent is left over, so each tail holds 0.025 and the boundaries are the 2.5th and 97.5th percentiles. This exact arithmetic reappears in every 95 percent confidence interval from chapter 8 onward.

40. Worked example: the interquartile range

Worked example

Example 6.11(a), the middle 50 percent.

\[ X \sim N(36.9, 13.9); \text{ find the IQR} \]

Find Q3

Why: The 75th percentile.

\[ 46.2754 \]

Find Q1

Why: The 25th percentile.

\[ 27.5246 \]

Subtract

Why: Q3 minus Q1.

\[ 18.7508 \]

Round

Why: As the book does.

\[ 18.8\text{ years} \]

Figure (svg): The solution to Worked example the interquartile range shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{IQR} = 46.2754 - 27.5246 \approx 18.8 \]

Verify: confirm the IQR against a general rule for normal distributions

Why: The quartiles of any normal sit about 0.674 standard deviations either side of the mean, so the IQR is about 1.35 standard deviations — here 1.35 times 13.9, which is about 18.8. That constant is worth knowing because it converts an IQR back into a standard deviation, and it is how a boxplot can be read as evidence for or against normality.

OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, p. 345

41. Trap: using the leftover area without halving it

Trap

The trap

\[ \text{middle } 20\text{ percent} \;\Rightarrow\; \text{tails of } 0.80 \text{ each} \]

Take the leftover 0.80 as each tail's area

Why: One minus 0.20 is 0.80, and that is what is outside.

\[ 0.80 + 0.20 + 0.80 = 1.80 > 1 \]

The three regions have to make up exactly one, so tails of 0.80 apiece are impossible before any percentile is computed.

The fix

\[ \text{tails of } \frac{1 - 0.20}{2} = 0.40 \text{ each} \]

Halve the leftover, since symmetry splits it evenly between two tails

Why: 0.40 plus 0.20 plus 0.40 gives exactly 1.

The totalling check is the reliable one and takes a second: the two tails plus the middle must sum to exactly one. It also generalises — for a middle 95 percent the tails are 0.025 each, which is the arithmetic behind every confidence interval in chapter 8, so getting the habit now pays off there.

42. One of these is false

Two truths and a lie

All three concern middle regions.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. A middle region's boundaries are symmetric about the mean
  • C. The upper boundary's left area is the lower tail plus the middle
  • B. Each tail holds the whole of the leftover area

Survives elimination: B

Why: The survivor is false: the leftover is SPLIT between two tails, so each holds half of it. Using the whole leftover for each tail makes the three regions total more than one, which is impossible and instantly checkable.

43. How wide is the middle half?

Estimation

A normal distribution has standard deviation 10.

Predict first

Roughly how wide is its interquartile range?

  • About 13.5
  • About 10
  • About 20
  • About 5

Correct: About 13.5.

Why: The quartiles sit about 0.674 standard deviations either side of the mean, so the IQR is about 1.35 standard deviations, which is 13.5 here. That ratio holds for every normal distribution, which makes it useful in reverse: dividing an observed IQR by 1.35 estimates the standard deviation.

44. Where do the boundaries sit?

Prediction

Commit before reasoning.

Predict first

For the middle 20 percent of a normal distribution, how far apart are the two boundaries?

  • About half a standard deviation
  • About two standard deviations
  • About one standard deviation
  • It depends on the mean

Correct: About half a standard deviation.

Why: The 40th and 60th percentiles sit at z-scores of about minus 0.253 and plus 0.253, so they are about 0.51 standard deviations apart. For Example 6.12 that is 0.51 times 0.24, about 0.12 cm — matching the gap between 5.79 and 5.91. The mean shifts both boundaries equally and so cannot affect the gap.

45. Interpreting and checking

Section

Section 5

46. A sentence, and a sanity check

Concept

The book asks repeatedly for the answer to be interpreted in a complete sentence, and every worked solution ends with one. A number alone does not say which quantity it refers to, in what units, or on which side of a boundary.

interpreting in context — Restating the numeric answer as a claim about the original quantity: not 48.6 but eighty percent of the smartphone users in this age range are 48.6 years old or less.

\[ 48.6 \;\longrightarrow\; \text{“80 percent are 48.6 years old or less”} \]

The interpretation is not decoration; it is where a reversed area or a misread question shows itself. Writing the sentence for the wrong answer to Example 6.11(b) would give forty percent of users are at least 33.4 years, and stating it that plainly invites the question of whether that can be right when the average is 36.9 — which it cannot. Forcing yourself to say what a number means is a genuine error-detection step.

Figure (svg): A normal curve for smartphone users' ages with the interquartile range shaded between the first and third quartiles

Section 2.4's quartiles, computed from a model rather than from data: each is just a percentile.

OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 343-346 — the book's repeated instruction to interpret in a complete sentence

47. The middle half of an age distribution

Picture it

Example 6.11(a): quartiles and the IQR for smartphone users.

Figure (svg): A normal curve for smartphone users' ages with the interquartile range shaded between the first and third quartiles

Section 2.4's quartiles, computed from a model rather than from data: each is just a percentile.

Interpreted, this says the middle half of smartphone users in the 13 to 55-plus range are between about 27.5 and 46.3 years old — a span of 18.8 years. Stated that way it can be checked against what anyone knows about smartphone users, which a bare 18.8 could not be.

48. Worked example: three answers, three sentences

Worked example

Example 6.10, interpreted rather than merely computed.

\[ X \sim N(36.9, 13.9) \]

Between 23 and 64.7

Why: The probability.

\[ 0.8186 \]

Say it in words

Why: About 82 percent.

At most 50.8

Why: The left area.

\[ 0.8413 \]

The 80th percentile

Why: A value, not a probability.

\[ 48.6\text{ years} \]

Figure (svg): The solution to Worked example three answers, three sentences shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 0.8186, \quad 0.8413, \quad 48.6 \text{ years} \]

Verify: confirm the third answer is a different kind of object from the first two

Why: The first two are probabilities and must lie between zero and one; the third is an age and must lie in the plausible range of the data. Mixing them up is the error the first idea's trap describes, and simply writing the units beside each answer makes it impossible to confuse them. The book's own solution ends with a full sentence for exactly this reason.

OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 344-345

49. One of these is false

Two truths and a lie

All three concern interpreting.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. A percentile answer carries the problem's units
  • C. A probability answer lies between zero and one
  • B. A sanity check against the Empirical Rule confirms the model is correct

Survives elimination: B

Why: The survivor is false, and the distinction matters. Such a check confirms the ARITHMETIC given the assumed normal model; if the quantity is not really normally distributed, the check will pass while the answer remains wrong. Testing the model itself needs the goodness-of-fit methods of chapter 11.

50. Worked example: checking an answer against the rule

Worked example

Example 6.10(b), where the cut-off is exactly one standard deviation up.

\[ \text{find } P(x \le 50.8) \text{ for } N(36.9, 13.9) \]

Locate the cut-off

Why: 36.9 plus 13.9.

Recall the rule

Why: 68 percent within one sigma.

\[ 32 \%\text{ outside} \]

Halve the outside

Why: One upper tail.

\[ 16 \% \]

Add

Why: 68 plus 16.

\[ \text{about } 84 \% \]

Figure (svg): The solution to Worked example checking an answer against the rule shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 0.68 + \frac{1 - 0.68}{2} = 0.84 \quad\text{against}\quad 0.8413 \]

Verify: confirm what the agreement does and does not establish

Why: It confirms the calculation was set up correctly — the right distribution, the right cut-off, the right side. It does not confirm the model: if ages are not really normal, both numbers are equally wrong. Sanity checks of this kind test the arithmetic against the assumptions, never the assumptions themselves, and keeping that distinction clear matters from chapter 9 onward.

OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 344-345

51. Trap: reporting a bare number

Trap

The trap

\[ \text{answer: } 48.6 \]

Stop once the calculator has produced a value

Why: The arithmetic is finished.

\[ 48.6 \text{ what, and above or below?} \]

Without units and a direction the number cannot be checked by anyone, including the person who computed it.

The fix

\[ \text{80 percent of users are } 48.6 \text{ years old or less} \]

State the quantity, the units and the direction in one sentence

Why: The book asks for this on every worked example.

The sentence is a test, not a formality. Writing eighty percent of users are 48.6 years old or less invites an immediate check against the mean of 36.9 — the boundary is above the mean, as a percentile above the 50th must be. A bare 48.6 offers nothing to check against, which is why errors survive it.

52. Probability or value?

Sorting

Ask what kind of object each answer must be.

Sort into buckets

Sort each question by the kind of answer it needs.

A probability, between 0 and 1
P(x > 65); P(1.8 < x < 2.75)
A value, in the problem's units
the 90th percentile; the bottom quartile's maximum; forty percent are at least what age
prob
The question supplies the boundaries and asks how much area they enclose.
val
The question supplies the area and asks where the boundary is.

Doing this sort before touching a calculator is worth the few seconds, because it settles which command is needed and gives an immediate test of the answer's plausibility. An answer of the wrong kind is always wrong, whatever its value.

53. Write the sentence

Explain it

A classmate has correctly computed the 80th percentile of smartphone users' ages as 48.6 and stopped there.

Discussion prompt

In one or two sentences, state what the number means and one check it passes.

Hint: Say which group, which direction, and compare with the mean.

Answer:

Eighty percent of smartphone users in the 13 to 55-plus range are 48.6 years old or less, and the remaining twenty percent are older than that.

It passes the side check: 48.6 sits above the mean of 36.9, as any percentile above the 50th must.

54. What does a check confirm?

Prediction

Commit before reasoning.

Predict first

An answer agrees with the Empirical Rule's estimate. What has been established?

  • That the calculation was set up correctly, assuming the normal model
  • That the data really are normally distributed
  • That the answer is exactly right to four decimal places
  • Nothing at all

Correct: That the calculation was set up correctly, assuming the model.

Why: The check compares two routes through the same assumed distribution, so it catches a wrong cut-off, a wrong side or a mistyped parameter — but both routes rest on normality, so neither can test it. The rule is also only approximate, so agreement to two decimal places is all it can offer.

55. The two kinds of question

Comparison

Fill the blanks. Telling them apart is most of this section.

Comparison matrix

Probability questionPercentile question
What is givenone or two boundariesan area
What is wantedan areaa boundary
The answer's unitsnone: it is between 0 and 1the problem's own units
Signal phrasefind the probability that...the 90th percentile; at least; the top 10 percent

The units row is the most useful check in the whole section. An answer between zero and one to a question asking for an age, or an answer in years to a question asking for a probability, is wrong before any of the arithmetic is examined.

56. Answering a normal question, in order

Pattern

Six steps, and the third is where this section's marks are won and lost.

  1. Write the distribution as N(mu, sigma) and sketch the curve, marking the mean.
  2. Decide which kind of question it is: boundaries given and area wanted, or area given and boundary wanted.
  3. Translate the wording into an area to the LEFT, subtracting from one whenever the phrase names the right-hand side.
  4. For a middle region, subtract from one and halve, then treat each boundary as an ordinary percentile.
  5. Compute: one minus a left area for a right tail, a difference of left areas for an interval, or the reverse command for a percentile.
  6. Check the side against the mean, check the units, and state the answer in a complete sentence.

The Empirical Rule is the sanity check throughout: a cut-off one, two or three standard deviations out should give areas near 0.84, 0.977 and 0.9987 to its left.

OpenStax Introductory Business Statistics 2e, §6.2 Using the Normal Distribution §6.2 Using the Normal Distribution

57. Check yourself 1 of 3

Check

A right tail.

Check your understanding

Exam scores are N(63, 5). The area to the left of 65 is 0.6554. What is P(x > 65)?

  • A. 0.3446 (correct)
  • B. 0.6554
  • C. 0.5
  • D. 1.6554

Answer: A

Why: The right tail is one minus the left area: 1 minus 0.6554 gives 0.3446.

Why B tempts people
That is the left area itself, and it is above a half, which is impossible for a cut-off above the mean.
Why C tempts people
That would be the answer if the cut-off were exactly at the mean of 63.
Why D tempts people
That adds instead of subtracting and exceeds one, so it cannot be a probability.

58. Check yourself 2 of 3

Check

Translating a phrase.

Check your understanding

For X ~ N(36.9, 13.9), forty percent of values are at least k. What left area should be used to find k?

  • A. 0.60 (correct)
  • B. 0.40
  • C. 0.20
  • D. 0.80

Answer: A

Why: At least means the area to the right is 0.40, so the area to the left is one minus 0.40, which is 0.60.

Why B tempts people
That uses the quoted proportion as the left area, which puts k below the mean when it must be above.
Why C tempts people
That halves the quoted proportion, which applies to middle regions rather than one-sided ones.
Why D tempts people
That doubles the quoted proportion, which corresponds to no step in the translation.

59. Check yourself 3 of 3

Check

A middle region.

Check your understanding

You want the middle 20 percent of a normal distribution. What area lies in each tail?

  • A. 0.40 (correct)
  • B. 0.80
  • C. 0.10
  • D. 0.20

Answer: A

Why: One minus 0.20 leaves 0.80, split evenly between two tails, so each holds 0.40.

Why B tempts people
That is the whole leftover area, not each tail; using it would make the regions total 1.8.
Why C tempts people
That halves the middle region rather than the leftover.
Why D tempts people
That is the middle region itself, not a tail.

60. Where this shows up outside the textbook

Real world

A manufacturer machines bolts whose diameters are normally distributed with mean 10.00 mm and standard deviation 0.04 mm. The specification accepts bolts between 9.92 and 10.08 mm. A production manager proposes tightening the specification to 9.96 to 10.04 mm to improve quality, and expects the scrap rate to rise only slightly.

Discussion prompt

Compute the scrap rate under each specification, assess the manager's expectation, and say what would actually reduce scrap.

Hint: Express each specification limit in standard deviations before computing anything.

Answer:

Express the limits in standard deviations first. The current limits sit at 10.00 give or take 0.08, which is exactly two standard deviations, so about 95 percent of bolts conform and roughly 4.6 percent are scrapped. The proposed limits sit at one standard deviation, so about 68 percent conform and about 32 percent are scrapped.

\[ \text{current: } 1 - 0.9545 \approx 0.046, \qquad \text{proposed: } 1 - 0.6827 \approx 0.317 \]

The manager's expectation is badly wrong: scrap would rise sevenfold, from about 1 bolt in 22 to nearly 1 in 3. The intuition that halving the tolerance costs only a little fails because the normal curve is densest near the mean, so the region being given up — between one and two standard deviations — contains about 27 percent of all output, far more than the 4.6 percent already outside.

Tightening the specification does not improve quality; it only reclassifies more output as scrap. The bolts being produced are unchanged. What would genuinely reduce scrap is reducing the standard deviation of the process itself: at a sigma of 0.02 mm the proposed limits would again sit two standard deviations out, and conformance would return to about 95 percent.

Two further points belong in a full answer. The calculation assumes the process stays centred at 10.00 mm, and a drift in the mean of even half a standard deviation raises scrap sharply on one side while barely reducing it on the other — so monitoring the centre matters as much as the spread. And this is exactly the logic behind process capability indices in manufacturing, which express the specification width as a number of standard deviations rather than in millimetres, for precisely the reason this problem illustrates.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

A question says forty percent of values are at least k. What area do you give the calculator to find k?

  • 0.40, the proportion stated
  • 0.60, because at least names the area to the right and the calculator needs the left
  • 0.20, half of the stated proportion
  • 0.80, twice the stated proportion

Correct: 0.60.

\[ P(x \ge k) = 0.40 \;\Longrightarrow\; P(x < k) = 0.60 \]

Why: At least translates to greater than or equal to, so 0.40 is the area to the right of k. The technology reports and reverses left areas only, so it needs one minus 0.40. The side check confirms it: with only forty percent above k, the boundary must lie above the mean, and 0.60 puts it there while 0.40 would put it below.

62. Explain it to someone a year behind you

Explain it

Asked for the 90th percentile of N(63, 5), they computed P(X < 90) and got approximately 1.

Discussion prompt

In two sentences or fewer, locate the error.

Hint: Ask what kind of object the answer should be.

Answer:

Point out that the answer should be an exam mark, and 1 is a probability — so the wrong kind of question was answered before any arithmetic is even checked.

The 90 was a percentage rather than a score, so it is the AREA that equals 0.90, and reversing that gives a mark of 69.4.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Turning a right tail or an interval into left areas
  • Recognising a percentile question and reversing the cumulative
  • Translating at least, at most or the top ten percent into a left area
  • Handling a middle region or an interquartile range

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the first, a right tail is one minus the left area and an interval is a difference of two. For the second, ask what the question gives and what it wants, then check the answer's units. For the third, decide which SIDE the quoted proportion describes before using it. For the fourth, subtract from one and halve, then check the two boundaries are symmetric about the mean. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Twenty minutes.

Draw it

Draw six normal curves down the page, all for the same distribution where possible, and label the mean on each. On the first, shade above 65 for exam scores from N(63, 5) and compute the right tail as one minus the left area, checking against 0.3446. On the second, shade between 1.8 and 2.75 for N(2, 0.5) and compute it as a difference of two left areas, checking against 0.5886. On the third, shade ninety percent of the area for N(63, 5) and mark the boundary, checking against 69.4 — and write beside it which quantity was given and which was wanted, to contrast it with the first two. On the fourth, take N(36.9, 13.9) and answer forty percent are at least what age: write the phrase, then the side it names, then the left area, then the answer, on four separate lines, and check 40.4 lies above the mean. On the fifth, shade the middle twenty percent of N(5.85, 0.24), writing the leftover, the halved tails, and the two percentiles, and confirm the boundaries 5.79 and 5.91 are equally far from 5.85. On the sixth, shade the interquartile range of N(36.9, 13.9), compute Q1 and Q3, and check the IQR against 1.35 standard deviations. Finish by writing one full interpreting sentence for each of the six answers.

Check the fifth curve by adding your three areas: 0.40 plus 0.20 plus 0.40 must give exactly 1. Check the sixth by dividing your IQR by the standard deviation of 13.9 — it should come out near 1.35 for any normal distribution, and a very different ratio means a quartile was computed from the wrong area.

65. What you can do now

Recap

Five things, and the third is the one worth practising most.

If you seeThen
A right tail wantedOne minus the left area
An interval wantedUpper left area minus lower
A proportion given, a value wantedA percentile: reverse the cumulative
At least, or the top so-many percentThe quoted proportion is the RIGHT area
The middle so-many percentSubtract from one, halve, then two percentiles
An answer between 0 and 1 to a value questionThe wrong kind of question was answered
A cut-off one sigma above the meanThe left area should be about 0.84

That closes chapter 6. Chapter 7 supplies the result that makes the normal distribution matter far beyond the quantities that happen to be bell shaped: the central limit theorem, which says that sample means are approximately normal whatever the population looks like — so the machinery of this chapter applies to almost any inference, and chapters 8 onward are built on it.

OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution §6.2, pp. 340-346 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §6.2 Using the Normal Distribution — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 340-346
  2. OpenStax Introductory Business Statistics 2e, §6.2 Using the Normal Distribution — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0
  3. OpenStax Introductory Business Statistics 2e, §6.3 Estimating the Binomial with the Normal Distribution — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108