The result that makes the rest of the course possible. Chapter 6's normal machinery applies only to quantities that happen to be bell shaped, but the central limit theorem says that if random samples of size n are drawn from any population at all, then as n increases the distribution of the sample means tends toward a normal distribution. That distribution has the same mean as the original, and a standard deviation equal to the original standard deviation divided by the square root of the sample size, a quantity called the standard error of the mean. Two consequences follow. Sample means are far less variable than individual values, and they become less variable in proportion to the square root of the sample size rather than to the sample size itself. The sample size n counts the values averaged together to form one mean, not the number of times the sampling is repeated.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 7 — The Central Limit Theorem
The Central Limit Theorem for Sample Means (Averages)
Objectives
Five outcomes, and the second is the one that governs every later chapter.
OpenStax Introductory Statistics 2e, §7.1 The Central Limit Theorem for Sample Means (Averages) §7.1, pp. 366-370 — the section these objectives are drawn from
Warm-up
Chapter 6 handled normally distributed quantities; chapters 4 and 5 supplied several that are plainly not normal.
Discussion prompt
A fair die has a flat distribution over 1 to 6 — nothing bell shaped about it. Roll ten dice and average the faces. What shape would you expect that average to have if you repeated the experiment many times?
Hint: How many ways are there to average 3.5, and how many ways to average 1?
Answer:
An average of 1 requires all ten dice to show 1, which happens one time in sixty million. An average of 3.5 can arise in an enormous number of ways, since the highs and lows can offset each other in countless combinations.
So the averages pile up in the middle and thin out at the ends, even though each individual die is perfectly flat. The book's classroom demonstration is exactly this: graph the means for one die, then two, then five, then ten, and watch the shape emerge.
What emerges is a normal distribution, and this section states it as a theorem. That is the striking part — the population's own shape becomes irrelevant, and the sample means turn out normal whatever it was.
Concept
Suppose X is a random variable with a distribution that may be known or unknown — it can be any distribution — with mean mu and standard deviation sigma. If you draw random samples of size n, then as n increases the random variable consisting of sample means tends to be normally distributed, with the same mean and a standard deviation of sigma divided by the square root of n.
the sampling distribution of the mean — The distribution of the sample means themselves, as opposed to the distribution of individual values. It approaches a normal distribution as the sample size increases, whatever the population's own shape.
\[ \bar{X} \sim N\left(\mu, \frac{\sigma}{\sqrt{n}}\right) \]
The book's demonstration records three things happening as the number of dice rises from one to two to five to ten: the mean of the sample means remains approximately the same, the spread of the sample means gets smaller, and the graph appears steeper and thinner. Those three observations are the theorem's three claims — unchanged centre, shrinking spread, normal shape — arrived at by rolling dice rather than by proof.
Figure (svg): Four bell curves sharing a centre at three point five, growing progressively narrower and taller as the number of dice averaged rises from one to ten
OpenStax Introductory Statistics 2e, §7.1 The Central Limit Theorem for Sample Means (Averages) §7.1, p. 366
Section
Section 1
Concept
The sampling distribution of the mean has the same mean as the original distribution, and a variance equal to the original variance divided by the sample size. Since the standard deviation is the square root of the variance, the standard deviation of the sampling distribution is the original standard deviation divided by the square root of n.
the central limit theorem for sample means — If you repeatedly draw samples of a given size and calculate their means, those means tend to follow a normal distribution, and the approximation improves as the sample size increases.
\[ \mu_{\bar{X}} = \mu, \qquad \sigma^2_{\bar{X}} = \frac{\sigma^2}{n}, \qquad \sigma_{\bar{X}} = \frac{\sigma}{\sqrt{n}} \]
The route through the variance is worth following, because it explains the square root. It is the VARIANCE that divides cleanly by n; the standard deviation inherits a square root from that division. Anyone who expects the spread to fall as one over n rather than one over the square root of n has skipped that step, and will badly overestimate how much precision a large sample buys.
Figure (svg): A card giving the central limit theorem's distribution for the sample mean and naming the standard error
OpenStax Introductory Statistics 2e, §7.1 The Central Limit Theorem for Sample Means (Averages) §7.1, pp. 366-367 — the statement of the theorem
Picture it
The book's demonstration, drawn as four densities on one axis.
Figure (svg): Four bell curves sharing a centre at three point five, growing progressively narrower and taller as the number of dice averaged rises from one to ten
Two features of the picture matter. All four curves share a centre at 3.5, so averaging does not bias the result. And the narrowing is visibly slowing: the step from one die to two is far larger than the step from five to ten, which is the square root at work.
Worked example
Example 7.1's setup. The population's own shape is never given.
\[ \text{an unknown distribution with } \mu = 90, \; \sigma = 15; \; n = 25 \]
Note what is unknown
Why: The population's shape.
Keep the mean
Why: The centre does not move.
\[ 90 \]
Divide sigma by root n
Why: Fifteen over five.
\[ 3 \]
Write the distribution
Why: By the theorem.
Figure (svg): The solution to Worked example naming the sampling distribution shown as a ladder of expressions, one row per legal move
\[ \bar{X} \sim N\left(90, \frac{15}{\sqrt{25}}\right) = N(90, 3) \]
Verify: confirm that the unknown population shape really does not matter
Why: The problem says the distribution is unknown and never returns to it, which would be impossible for any question about an individual value — there we would need the actual distribution. It is only the theorem that lets the question be answered at all, and recognising that is the difference between using the result and merely applying a formula.
OpenStax Introductory Statistics 2e, §7.1 The Central Limit Theorem for Sample Means (Averages) §7.1, p. 367
Faded example
A population has standard deviation 15 and samples of size 100 are drawn.
Fill in the blanks
\sigma_100} = \frac1.5___}}} = ___
Why: The square root of 100 is 10, so the standard error is 1.5 — Example 7.3's value. The sample mean is ten times less variable than a single observation.
Worked example
Example 7.2, where the population happens to be normal already.
\[ \text{soccer match times} \sim N(2, 0.5); \; n = 50 \]
Note the population
Why: Already normal.
\[ N(2, 0.5) \]
Keep the mean
Why: Two hours.
\[ 2 \]
Divide sigma by root n
Why: 0.5 over root 50.
\[ 0.0707 \]
Write the distribution
Why: Much narrower.
Figure (svg): The solution to Worked example the same theorem on a normal population shown as a ladder of expressions, one row per legal move
\[ \bar{X} \sim N\left(2, \frac{0.5}{\sqrt{50}}\right) \approx N(2, 0.0707) \]
Verify: confirm what changes when the population is already normal
Why: When the population is normal, the sample mean is EXACTLY normal for every n, not merely approximately so for large n — the theorem's approximation is not needed. What the theorem adds is the case where the population is not normal, which is Example 7.1's. Both problems use the same formula, and only their justifications differ.
OpenStax Introductory Statistics 2e, §7.1 The Central Limit Theorem for Sample Means (Averages) §7.1, p. 368
Trap
\[ \sigma_{\bar{X}} = \frac{15}{25} = 0.6 \]
Divide the standard deviation by the sample size
Why: The theorem says the sample size shrinks the spread.
\[ \text{but it is the VARIANCE that divides by } n \]
A spread of 0.6 would make a sample mean of 92 more than three standard errors out, which is nothing like the truth.
\[ \sigma_{\bar{X}} = \frac{15}{\sqrt{25}} = 3 \]
Divide by the square ROOT of the sample size
Why: The variance divides by n, so the standard deviation divides by root n.
This error always makes a sample look more precise than it is, and by a large factor: at n equal to 100 it understates the standard error tenfold. The practical consequence is that quadrupling a sample only halves its standard error — which is why large studies are expensive and why the square root appears in every sample-size calculation from chapter 8 onward.
Two truths and a lie
All three concern the theorem.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The sample means have a SMALLER standard deviation, namely sigma over the square root of n — which is the whole point. If means were as variable as individual values there would be no advantage in taking a sample at all.
Prediction
Commit before reasoning.
Predict first
A sample of 25 gives a standard error of 3. What does a sample of 100 give?
Correct: 1.5.
Why: Quadrupling the sample size doubles the square root, so it halves the standard error: 3 becomes 1.5. Getting 0.75 would mean dividing by four rather than by two, which is the one-over-n error. This diminishing return is why sample-size decisions are always a cost calculation.
Estimation
The theorem is a statement about a limit as n increases.
Predict first
For a population that is already normal, how large must n be for the sample mean to be normal?
Correct: Any n at all.
Why: When the population is normal, the sample mean is exactly normal for every sample size, so no approximation is involved. The familiar rule of thumb about 30 concerns non-normal populations, where the approximation needs a reasonable sample to take hold — and the more skewed the population, the larger n has to be.
Section
Section 2
Concept
The standard deviation of the sample mean is called the standard error of the mean. The book gives its meaning directly: it is a description of how far, on average, the sample mean will be from the population mean in repeated simple random samples of size n.
standard error of the mean — Sigma divided by the square root of n. It measures the typical distance between a sample mean and the population mean, and it is what every margin of error in chapter 8 is built from.
\[ \text{SE} = \frac{\sigma}{\sqrt{n}} \]
The name is worth taking seriously: it is a standard deviation, but of a statistic rather than of the data. Calling it an error emphasises what it is used for — quantifying how wrong a single sample's mean is likely to be. Every confidence interval in chapter 8 is a sample mean plus or minus a multiple of this quantity, so the whole of inference rests on it.
Figure (svg): A decreasing curve showing the standard error falling from fifteen at a sample size of one to one point five at a sample size of one hundred
OpenStax Introductory Statistics 2e, §7.1 The Central Limit Theorem for Sample Means (Averages) §7.1, pp. 367-368 — the standard error and Example 7.1(b)
Picture it
The standard error against sample size, for a population standard deviation of 15.
Figure (svg): A decreasing curve showing the standard error falling from fifteen at a sample size of one to one point five at a sample size of one hundred
The curve is steep at the left and nearly flat at the right, and that shape has a practical reading: the first few observations buy a great deal of precision and the later ones buy very little. Moving from 25 to 100 costs three times as much data as moving from 1 to 4 and delivers exactly the same halving.
Worked example
Example 7.1(b), using the book's own formula.
\[ \bar{X} \sim N(90, 3); \text{ find the value two standard deviations above } 90 \]
Recall the standard error
Why: Fifteen over root 25.
\[ 3 \]
Use the value formula
Why: Mean plus a number of spreads.
\[ 90 + 2(3) \]
Evaluate
Why: Ninety plus six.
\[ 96 \]
Say what it describes
Why: A sample MEAN, not a value.
\[ 96 \]
Figure (svg): The solution to Worked example two standard errors above the mean shown as a ladder of expressions, one row per legal move
\[ \text{value} = \mu + (2)\left(\frac{\sigma}{\sqrt{n}}\right) = 90 + 6 = 96 \]
Verify: confirm how different the answer would be for an individual value
Why: Two standard deviations above the mean for an INDIVIDUAL would be 90 plus 2 times 15, or 120 — a completely different number. The same phrase, two standard deviations above the mean, means 96 for a sample mean of 25 and 120 for a single value, and only the context distinguishes them. Naming which variable is being described before computing is what keeps them apart.
OpenStax Introductory Statistics 2e, §7.1 The Central Limit Theorem for Sample Means (Averages) §7.1, p. 368
Faded example
Mean 8.2 minutes, standard deviation 1 minute, samples of 60.
Fill in the blanks
\text0.129 = \fracsmaller___} \approx ___, \text___ ___ \text___
Why: One divided by about 7.75 is roughly 0.129 minutes — much smaller than the population's one minute, as a standard error always is for n above 1.
Worked example
Example 7.3(a) and (b).
\[ \mu = 34, \; \sigma = 15, \; n = 100 \]
The mean of the sample mean
Why: Targets the population mean.
\[ 34 \]
The standard error
Why: Fifteen over root 100.
\[ 1.5 \]
The shape
Why: By the theorem, for large n.
Write the distribution
Why: All three together.
Figure (svg): The solution to Worked example the standard error for iPad ages shown as a ladder of expressions, one row per legal move
\[ \bar{X} \sim N(34, 1.5) \]
Verify: confirm the standard error is much smaller than the population's spread
Why: Individual ages vary with a standard deviation of 15 years, while means of 100 vary with only 1.5 — a tenfold reduction. That is what makes a sample of 100 informative about the population mean even though any single user's age tells you almost nothing. The whole logic of sampling is contained in that comparison.
OpenStax Introductory Statistics 2e, §7.1 The Central Limit Theorem for Sample Means (Averages) §7.1, p. 369
Error analysis
The correct value is 3.
Annotate
On: \( \begin{aligned} &(1)\; \frac{15}{25} = 0.6 \\ &(2)\; \frac{15}{\sqrt{15}} \approx 3.87 \\ &(3)\; 15 \\ &(4)\; \frac{15}{\sqrt{25}} = 3 \end{aligned} \)
Errors (1) and (3) fail in opposite directions and both are caught by the same check: the standard error must always be SMALLER than sigma but not dramatically so, since dividing by root n is a gentle reduction. Error (2) is caught by noticing that the units would be wrong.
Two truths and a lie
All three concern the standard error.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and describes sigma itself. The standard error is about the variability of a STATISTIC across repeated samples, not about the variability of the data. Confusing the two makes sample means look as unreliable as single observations, which would defeat the purpose of sampling.
Estimation
A study with n = 400 wants to halve its standard error.
Predict first
What sample size does it need?
Correct: 1,600.
Why: Halving the standard error requires doubling the square root of n, which means quadrupling n. Going from 400 to 1,600 is a fourfold increase in cost for a twofold gain in precision, and that trade-off is a permanent feature of sampling rather than a limitation of any particular method.
Prediction
Commit before reasoning.
Predict first
If the sample size is 1, what is the standard error?
Correct: Sigma.
Why: The square root of 1 is 1, so the standard error equals sigma — and correctly so, because the mean of a single observation is that observation. This is a useful boundary check on the formula: at n equal to 1 the sampling distribution must coincide with the population's, and it does.
Section
Section 3
Concept
The book states it plainly: the variable n is the number of values that are averaged together, not the number of times the experiment is done. In the dice demonstration, rolling ten dice a hundred times gives n equal to ten.
the sample size n — The count of observations combined into one sample mean. It appears in the standard error and nowhere else, so misidentifying it corrupts every probability computed afterwards.
\[ \bar{x} = \frac{x_1 + x_2 + \cdots + x_n}{n} \]
The confusion is easy to fall into because both numbers are counts and both appear in the setup. The reliable test is to look at the formula for a single sample mean and count how many values are being added in the numerator — that count is n. How many such means were computed affects how clearly the sampling distribution's shape shows up in a histogram, but it does not appear anywhere in the theorem.
Figure (svg): A checklist for identifying the sample size n correctly in a central limit theorem problem
OpenStax Introductory Statistics 2e, §7.1 The Central Limit Theorem for Sample Means (Averages) §7.1, p. 366 — the explicit warning about what n counts
Picture it
The distinction the book flags explicitly.
Figure (svg): A checklist for identifying the sample size n correctly in a central limit theorem problem
The last line is worth separate attention. If a problem asks about a single measurement rather than an average of several, the theorem simply does not apply — and section 7.3 makes that the first question to ask of any problem in this chapter.
Worked example
The book's classroom activity, where both numbers appear.
\[ \text{roll five dice, ten times, recording the mean each roll} \]
Count the values in one mean
Why: Five faces added and divided.
\[ 5 \]
Count the repetitions
Why: Ten means recorded.
\[ 10 \]
Choose n
Why: The values averaged together.
\[ n = 5 \]
Say what the ten does
Why: Shows the shape.
Figure (svg): The solution to Worked example reading n from the dice demonstration shown as a ladder of expressions, one row per legal move
\[ n = 5, \quad \sigma_{\bar{X}} = \frac{\sigma}{\sqrt{5}} \]
Verify: confirm by asking what would change if the activity were repeated a thousand times
Why: The histogram would look far smoother and its shape would be much clearer, but the sampling distribution itself — its centre and its spread — would be identical, because those depend only on n equal to 5. More repetitions reveal the distribution better; they do not change it. That is the cleanest way to see which number belongs in the formula.
OpenStax Introductory Statistics 2e, §7.1 The Central Limit Theorem for Sample Means (Averages) §7.1, p. 366
Sorting
Count the values combined into one mean.
Sort into buckets
Sort each description by whether the bolded number is n.
In items (b) and (d) the correct n is 5 and 12 respectively — in both cases the smaller number. That is not a rule, but it is a useful warning: the eye is drawn to the larger figure, and the larger figure is often the wrong one.
Worked example
A case where only one number is offered and it is easy to misuse.
\[ \text{a pollster surveys } 1\,200 \text{ people and reports their mean income} \]
Count the values averaged
Why: All 1,200 incomes.
\[ n = 1200 \]
Count the repetitions
Why: One survey was run.
\[ 1 \]
Choose n
Why: The values in the mean.
\[ 1, 200 \]
Note what the 1 means
Why: One sample mean observed.
Figure (svg): The solution to Worked example n in a survey shown as a ladder of expressions, one row per legal move
\[ \sigma_{\bar{X}} = \frac{\sigma}{\sqrt{1200}} \]
Verify: confirm the sampling distribution is meaningful with only one sample observed
Why: The sampling distribution describes what WOULD happen across many repeated samples, and it exists whether or not those repetitions are ever performed. That is what licenses a statement about how far this one survey's mean is likely to be from the truth — a hypothetical distribution supporting a real conclusion, which is the logical structure of every inference in the rest of the book.
OpenStax Introductory Statistics 2e, §7.1 The Central Limit Theorem for Sample Means (Averages) §7.1, pp. 366-367
Trap
\[ \text{ten dice rolled } 100 \text{ times} \;\Rightarrow\; n = 100 \]
Take the larger of the two counts
Why: A hundred sounds more like a sample size.
\[ \sigma_{\bar{X}} = \frac{\sigma}{10} \text{ instead of } \frac{\sigma}{\sqrt{10}} \]
The standard error comes out about three times too small, and every probability computed from it is wrong.
\[ n = 10, \quad \sigma_{\bar{X}} = \frac{\sigma}{\sqrt{10}} \]
Count the values combined into ONE mean
Why: Ten dice make one average.
The test that always works is to write out the formula for a single sample mean and count the terms in the numerator. Repetitions never appear there. This error is particularly costly because it always makes results look more precise than they are, and nothing in the answer signals it.
Fill the middle
The book's own warning.
Fill in the blanks
n \textaveraged ___ \text___
Why: Averaged together. The number of times the experiment is repeated does not appear anywhere in the theorem, which the book states explicitly because the confusion is so common.
Two truths and a lie
All three concern the sample size.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. Only a larger n — more values in each mean — reduces the standard error. Taking a thousand samples of ten does not make any one of those means more precise than taking a single sample of ten; it merely lets you see the distribution they come from.
Explain it
A classmate says that rolling ten dice a hundred times must give a smaller standard error than rolling them ten times.
Discussion prompt
In two sentences or fewer, correct them.
Hint: Ask what the standard error describes.
Answer:
Point out that the standard error describes how much ONE sample mean varies, and each of their sample means is still an average of just ten dice either way.
A hundred repetitions gives a clearer picture of that variation — a smoother histogram — but does not reduce it; only rolling more dice per sample would do that.
Section
Section 4
Concept
Once the sampling distribution is written down, every question about a sample mean is an ordinary normal probability question. The only change from chapter 6 is that the spread is the standard error rather than the population standard deviation.
the z-score for a sample mean — X-bar minus mu, divided by sigma over the square root of n. It differs from an individual value's z-score only in its denominator, and that difference is the whole of the chapter.
\[ z = \frac{\bar{x} - \mu}{\sigma/\sqrt{n}} \]
The book notes that the random variable X-bar has a different z-score associated with it from that of the random variable X, and it is worth dwelling on how large the difference is. In Example 7.1 a value of 92 sits 0.13 standard deviations out for an individual and 0.67 standard errors out for a mean of 25 — the same number, five times further out on the sampling scale, because means cluster so much more tightly.
Figure (svg): A normal curve for the sample mean centred at ninety with the region between eighty-five and ninety-two shaded
OpenStax Introductory Statistics 2e, §7.1 The Central Limit Theorem for Sample Means (Averages) §7.1, pp. 366-369 — the z-score for X-bar, and Examples 7.1 to 7.3
Picture it
Example 7.1(a): the chance a sample mean of 25 lands between 85 and 92.
Figure (svg): A normal curve for the sample mean centred at ninety with the region between eighty-five and ninety-two shaded
The interval from 85 to 92 covers about seven tenths of this curve, and it would cover only about a fifth of the population's own curve, which has a spread of 15 rather than 3. Reading which curve is drawn — individuals or means — before interpreting any shaded area is essential, and the axis label is the only clue.
Worked example
Example 7.1(a).
\[ \bar{X} \sim N(90, 3); \text{ find } P(85 < \bar{x} < 92) \]
Write the sampling distribution
Why: Mean 90, standard error 3.
\[ N(90, 3) \]
Left area at 92
Why: Just above the mean.
\[ 0.7475 \]
Left area at 85
Why: Below the mean.
\[ 0.0478 \]
Subtract
Why: The interval.
\[ 0.6997 \]
Figure (svg): The solution to Worked example a sample mean between 85 and 92 shown as a ladder of expressions, one row per legal move
\[ P(85 < \bar{x} < 92) = 0.6997 \]
Verify: confirm the answer against the interval's width in standard errors
Why: The interval runs from 1.67 standard errors below the mean to 0.67 above, so it should capture rather more than half the distribution but well short of all of it — and 0.70 fits. Attempting the same check with sigma equal to 15 would suggest the interval spans only about half a standard deviation, giving a much smaller answer, which is exactly the error using the wrong spread produces.
OpenStax Introductory Statistics 2e, §7.1 The Central Limit Theorem for Sample Means (Averages) §7.1, p. 367
Faded example
Soccer times are N(2, 0.5) and n = 50, so the standard error is 0.0707.
Fill in the blanks
P(1.8 < \bar0.0707 < 2.3) \text0.9977 N(2, ___) = ___
Why: The standard error is 0.5 over the square root of 50, about 0.0707, and the interval covers almost the whole sampling distribution — hence 0.9977. Using 0.5 instead would give only about 0.38.
Worked example
Example 7.3(c).
\[ \bar{X} \sim N(34, 1.5); \text{ find } P(\bar{x} > 30) \]
Find the standard error
Why: Fifteen over root 100.
\[ 1.5 \]
Locate 30
Why: Four years below 34.
\[ 2.67\text{ SEs below} \]
Take the right tail
Why: One minus the left area.
\[ 0.9962 \]
Interpret
Why: Samples of 100 iPad users.
Figure (svg): The solution to Worked example a sample mean age above 30 shown as a ladder of expressions, one row per legal move
\[ P(\bar{x} > 30) = 0.9962 \]
Verify: confirm why this is so much larger than the individual probability
Why: For an individual user, 30 is only 4 years below a mean of 34 with a spread of 15 — about a quarter of a standard deviation, giving a probability near 0.60. For a mean of 100 the same 4 years is 2.67 standard errors, which is nearly the whole distribution. The gap between 0.60 and 0.9962 is entirely the effect of averaging, and section 7.3 makes this contrast its central point.
OpenStax Introductory Statistics 2e, §7.1 The Central Limit Theorem for Sample Means (Averages) §7.1, p. 369
Trap
\[ P(85 < \bar{x} < 92) \text{ computed on } N(90, 15) \]
Use the population's standard deviation
Why: It is the number the problem supplied.
\[ \approx 0.18 \text{ instead of } 0.6997 \]
That answers a question about a single observation, not about the mean of twenty-five of them.
\[ P(85 < \bar{x} < 92) \text{ on } N(90, 3), \text{ giving } 0.6997 \]
Divide sigma by the square root of n first
Why: The sample mean has its own, smaller spread.
This is the section's defining error, and it always understates how concentrated sample means are. The guard is to write the sampling distribution on its own line — X-bar follows N(90, 3) — before any probability is attempted, exactly as the book's solutions do. A distribution written down cannot be silently ignored.
Discrimination
Decide whether each question is about an individual or a mean.
Sort into buckets
Sort each question for a population with mu = 34 and sigma = 15.
Two truths and a lie
All three concern probabilities for means.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. In Example 7.1 the interval from 85 to 92 holds 0.70 of the sampling distribution but only about 0.18 of the population's. Identical intervals give very different probabilities on the two scales, which is exactly why the distinction matters.
Estimation
A population has mean 34 and standard deviation 15.
Predict first
Which is larger: the chance a single value exceeds 36, or the chance a mean of 100 exceeds 36?
Correct: The single value's chance.
Why: For an individual, 36 is only 0.13 standard deviations above the mean, so the chance is about 0.45. For a mean of 100 it is 1.33 standard errors above, giving about 0.09. Averaging concentrates values near mu, so it makes any departure from mu LESS likely — which cuts both ways, above and below.
Section
Section 5
Concept
A percentile for a sample mean is found exactly as in section 6.2: the area is given and the boundary is wanted. The only difference is that the reversal is performed on the sampling distribution, using the standard error as the spread.
a percentile of the sample mean — The value k such that a stated proportion of all sample means of size n fall below it. It is a statement about samples, not about individuals, and its interpretation must say so.
\[ P(\bar{X} < k) = p \;\Longrightarrow\; k = \text{invNorm}\left(p, \mu, \frac{\sigma}{\sqrt{n}}\right) \]
Because the sampling distribution is so much narrower than the population's, its percentiles are bunched tightly around the mean — Example 7.3's 95th percentile is 36.5 against a population mean of 34, a gap of only 2.5 years, while the 95th percentile of individual ages would be around 58.7. Reporting one as though it were the other overstates or understates dramatically, so the interpreting sentence must name which.
Figure (svg): A decreasing curve showing the standard error falling from fifteen at a sample size of one to one point five at a sample size of one hundred
OpenStax Introductory Statistics 2e, §7.1 The Central Limit Theorem for Sample Means (Averages) §7.1, pp. 369-370 — Examples 7.3(d) and 7.4(c)
Picture it
The standard error against n, for a population spread of 15.
Figure (svg): A decreasing curve showing the standard error falling from fifteen at a sample size of one to one point five at a sample size of one hundred
At n equal to 100 the standard error is a tenth of the population's spread, so every percentile of the sampling distribution sits ten times closer to the mean than the corresponding percentile of individuals. That concentration is what makes a sample mean a useful estimate, and it is the quantitative content of the law of large numbers that section 7.3 states.
Worked example
Example 7.3(d).
\[ \bar{X} \sim N(34, 1.5); \text{ find the } 95\text{th percentile} \]
Write the sampling distribution
Why: Mean 34, standard error 1.5.
\[ N(34, 1.5) \]
Set the area
Why: Ninety-five percent below.
\[ 0.95 \]
Reverse the cumulative
Why: On that distribution.
\[ 36.47 \]
Round
Why: To one decimal place.
\[ 36.5\text{ years} \]
Figure (svg): The solution to Worked example the 95th percentile of sample mean ages shown as a ladder of expressions, one row per legal move
\[ k = \text{invNorm}(0.95,\; 34,\; 1.5) \approx 36.5 \]
Verify: confirm the interpretation is about samples rather than users
Why: The sentence must say that 95 percent of SAMPLE MEANS fall below 36.5, not that 95 percent of users are younger than 36.5 — the latter is wildly false, since individual ages have a spread of 15 years and their 95th percentile is nearly 59. Writing the interpretation carefully is the only thing that catches this confusion, which is why the book asks for a complete sentence.
OpenStax Introductory Statistics 2e, §7.1 The Central Limit Theorem for Sample Means (Averages) §7.1, p. 369
Faded example
For app engagement, mu = 8.2, sigma = 1 and n = 60, so the standard error is 0.129.
Fill in the blanks
P(\bar8.37 < k) = 0.90 \;\Rightarrow\; k \approx 1.28 \text___ ___ \text___
Why: The 90th percentile of any normal distribution is about 1.28 standard deviations above its mean, and here that is 1.28 times 0.129, which added to 8.2 gives 8.37.
Worked example
Example 7.4(c), with the book's own interpretation.
\[ \mu = 8.2, \; \sigma = 1, \; n = 60; \text{ find the } 90\text{th percentile} \]
Find the standard error
Why: One over root 60.
\[ 0.1291 \]
Set the area
Why: Ninety percent below.
\[ 0.90 \]
Reverse the cumulative
Why: On N(8.2, 0.1291).
\[ 8.365 \]
Round
Why: As the book does.
\[ 8.37\text{ minutes} \]
Figure (svg): The solution to Worked example the 90th percentile of app engagement means shown as a ladder of expressions, one row per legal move
\[ k = \text{invNorm}(0.90,\; 8.2,\; 0.1291) \approx 8.37 \]
Verify: confirm the answer is close to the mean, and why
Why: The 90th percentile sits only 0.17 minutes above the mean of 8.2, which looks surprisingly tight until you notice the standard error is only 0.129 minutes — so 8.37 is about 1.28 standard errors out, exactly where a 90th percentile belongs. Checking a percentile's distance in standard errors rather than in raw units is what makes it recognisable as reasonable.
OpenStax Introductory Statistics 2e, §7.1 The Central Limit Theorem for Sample Means (Averages) §7.1, p. 370
Trap
\[ k = 36.5 \;\Rightarrow\; \text{95 percent of iPad users are under } 36.5 \]
Read the percentile as a statement about individuals
Why: It is an age in years, so it sounds like a statement about people.
\[ \text{but individual ages have a spread of } 15 \text{ years} \]
The 95th percentile of individual ages is nearly 59, not 36.5, so the claim is wrong by more than twenty years.
\[ \text{95 percent of SAMPLE MEANS of } 100 \text{ users fall below } 36.5 \]
Name the sampling distribution in the interpretation
Why: The percentile describes averages, not people.
Both statements are about ages in years, so nothing in the units distinguishes them — only the words do. This is why the book asks for a complete sentence on every percentile in the chapter: the sentence is where the distinction lives, and an answer given as a bare number cannot record it at all.
Two truths and a lie
All three concern percentiles for means.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. For Example 7.3 they are 36.5 and about 58.7 — over twenty years apart. They coincide only when n equals 1, since only then is the sampling distribution the population's own.
Prediction
Commit before reasoning.
Predict first
As the sample size increases, what happens to the 95th percentile of the sample mean?
Correct: It moves closer to mu.
Why: The standard error shrinks as n grows, so every percentile is pulled toward the mean. In the limit the whole sampling distribution collapses onto mu, which is the law of large numbers stated in terms of percentiles — and section 7.3 makes that connection explicitly.
Explain it
A classmate reports the 95th percentile of sample mean ages as 36.5 and writes: 95 percent of iPad users are younger than 36.5.
Discussion prompt
In two sentences or fewer, fix the sentence and say why it matters.
Hint: Ask what the 1.5 in their calculation was.
Answer:
Their 1.5 was the standard error, which describes how MEANS of 100 users vary, so the sentence should read that 95 percent of samples of 100 users have a mean age below 36.5.
It matters because individual ages vary with a spread of 15 years rather than 1.5, so the 95th percentile of actual users is nearly 59 — their sentence is wrong by more than two decades.
Comparison
Fill the blanks. One population, two very different distributions.
Comparison matrix
| One value, X | A mean of n, X-bar | |
|---|---|---|
| Centre | mu | mu, the same |
| Spread | sigma | sigma over the square root of n |
| Shape | whatever the population's is | approximately normal, for large n |
| Example 7.1's spread | 15 | 3 |
The third row is the theorem itself, and the fourth is what it buys: for Example 7.1 the sampling distribution is five times narrower than the population, so a sample of 25 pins down the mean far better than any single observation could.
Pattern
Six steps, and the second and third are the ones that go wrong.
A sanity check that always applies: the standard error must be smaller than sigma but larger than sigma over n, since it lies between them at sigma over root n.
OpenStax Introductory Business Statistics 2e, §7.1 The Central Limit Theorem for Sample Means §7.1 The Central Limit Theorem for Sample Means
Check
The standard error.
Check your understanding
A population has sigma = 15. Samples of size 25 are drawn. What is the standard deviation of the sample mean?
Answer: A
Why: Fifteen divided by the square root of 25 is 15 over 5, which is 3.
Check
Identifying n.
Check your understanding
A class rolls ten dice and records the mean, repeating the experiment 40 times. What is n?
Answer: A
Why: n is the number of values averaged together to make one mean, which is the ten dice. The 40 repetitions never enter the formula.
Check
Which distribution applies.
Check your understanding
A population has mu = 34 and sigma = 15. Which spread should be used to find the probability that a sample of 100 averages more than 36?
Answer: A
Why: The question is about a sample mean, so the spread is the standard error: 15 over the square root of 100, which is 1.5.
Real world
A bottling line fills containers whose contents have mean 500 ml and standard deviation 8 ml. Quality control takes a sample of 16 bottles each hour and stops the line if the sample mean falls outside 495 to 505 ml. A supervisor argues that since individual bottles routinely fall outside that range, the line will be stopped constantly.
Discussion prompt
Compute how often the line stops when the process is on target, assess the supervisor's argument, and say what the rule is actually designed to detect.
Hint: The rule is applied to a mean of 16, not to individual bottles.
Answer:
The supervisor is comparing the wrong distributions. Individual bottles do routinely fall outside 495 to 505: those limits are only 0.625 standard deviations from the mean, so about 47 percent of single bottles lie outside them. But the rule is applied to a MEAN of 16 bottles, whose standard error is 8 over the square root of 16, which is 2 ml.
\[ \text{SE} = \frac{8}{\sqrt{16}} = 2, \qquad P(495 < \bar{x} < 505) = P(-2.5 < z < 2.5) \approx 0.9876 \]
So the line stops about 1.2 percent of the time when it is running correctly — roughly once every 80 hours, not constantly. The limits sit 2.5 standard errors from the target, which is why they are rarely crossed by chance.
What the rule is designed to detect is a shift in the process mean, not variation among bottles. Individual bottles varying by several millilitres is expected and harmless; the mean of 16 drifting to 505 would signal that the filling head has moved, and the rule catches that quickly precisely because sample means are so much less variable than individual bottles.
Two extensions are worth noting. The rule's sensitivity depends on the sample size: taking 4 bottles instead of 16 would double the standard error to 4 ml, making the limits only 1.25 standard errors out and the false-stop rate about 21 percent. And a rule of this kind cannot detect an increase in variability with the mean unchanged — that needs a separate chart on the sample's spread, which is why real control charts always run a pair.
Commit first
Answer, then rate your confidence honestly.
Predict first
Why is the standard error sigma over the square root of n rather than sigma over n?
Correct: Because the variance divides by n, and the standard deviation is its square root.
\[ \sigma^2_{\bar{X}} = \frac{\sigma^2}{n} \;\Longrightarrow\; \sigma_{\bar{X}} = \frac{\sigma}{\sqrt{n}} \]
Why: The theorem's clean statement is about the variance: the variance of the sample mean is the population variance divided by n. Taking square roots of both sides introduces the root n in the denominator. Getting this wrong always overstates a sample's precision, and by a large factor once n is big.
Explain it
They computed P(85 < x-bar < 92) for Example 7.1 using N(90, 15) and got about 0.18.
Discussion prompt
In two sentences or fewer, locate the error.
Hint: Ask what the 15 describes.
Answer:
Their 15 is how much a SINGLE observation varies, but the question asks about the mean of twenty-five of them, which varies far less.
Dividing 15 by the square root of 25 gives a standard error of 3, and on N(90, 3) the same interval holds 0.6997 rather than 0.18.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, the population may be any distribution at all. For the second, n counts the values averaged, and divide sigma by its square root. For the third, write the sampling distribution on its own line before computing anything. For the fourth, make the sentence say samples of size n. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Twenty minutes.
Draw it
At the top, draw four bell curves on one axis sharing a centre, of decreasing width, and label them as the means of one, two, five and ten dice — then write the three things the book observes happening as the number of dice rises. Below that, write the theorem's statement with its three parts: same mean, standard deviation sigma over root n, and normal shape whatever the population. Beside it write the definition of n in the book's own terms, and give one example where the wrong number is tempting. In the middle of the page, work Example 7.1 completely: write the sampling distribution N(90, 3) on its own line, shade the region from 85 to 92 and compute 0.6997, then find the value two standard errors above the mean and compare it with the value two population standard deviations above the mean. Below that, plot the standard error against n for sigma equal to 15, marking the four points at n equal to 1, 4, 25 and 100, and write one sentence on what quadrupling n does. At the bottom, work Example 7.3: compute the standard error of 1.5, the probability that a sample mean exceeds 30, and the 95th percentile — then write two sentences for that percentile, one correct and one making the individuals-versus-means error, and mark which is which.
Check your standard errors by confirming each is smaller than sigma but larger than sigma over n. And check the final pair of sentences against each other: the 95th percentile of sample means is 36.5 while that of individual ages is nearly 59, so if your two sentences do not differ dramatically, the error has not been made vividly enough to learn from.
Recap
Five things, and the second underlies every remaining chapter.
| If you see | Then |
|---|---|
| A question about a mean or average | Use the sampling distribution |
| A question about one value | Use the population's own distribution |
| An unknown population shape | The theorem still applies, for large n |
| Two counts in the setup | n is the one being averaged |
| A spread that looks too small | Check for dividing by n instead of root n |
| A percentile of the sample mean | Say samples of size n in the interpretation |
| A wish to halve the standard error | Quadruple the sample size |
Section 7.2 states the same theorem for sums rather than means. Because a sum is a mean multiplied by n, its mean is n times the population mean and its standard deviation is the square root of n times sigma — the multiplication running the opposite way from the division here, which is worth watching carefully.
OpenStax Introductory Statistics 2e, §7.1 The Central Limit Theorem for Sample Means (Averages) §7.1, pp. 366-370 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.