The first act of statistical inference in the course. Chapter 7 asked how far a sample mean is likely to fall from a known population mean; this section reverses the question and asks which values of an unknown population mean are consistent with an observed sample mean. The answer is not a single number but an interval, formed as the point estimate plus and minus a margin of error called the error bound, which equals a z-score chosen by the confidence level multiplied by the standard error of the mean. The confidence level and its complement alpha determine that z-score, since alpha is split equally between the two tails. Raising the confidence level or shrinking the sample widens the interval, and the error bound formula can be solved backwards to find the sample size a study needs. The interpretation is the section's real difficulty: the interval is the random thing and the parameter is fixed, so the confidence level describes the procedure across repeated samples rather than any one interval.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 8 — Confidence Intervals
A Single Population Mean using the Normal Distribution
Objectives
Five outcomes, and the last two sentences of the first one are where most marks are lost.
OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 406-416 — the section these objectives are drawn from
Warm-up
Chapter 7 established that a sample mean of n values has standard error sigma over the square root of n.
Discussion prompt
Songs downloaded per month have a known population standard deviation of 1, and a sample of 100 gives a mean of 2. What can be said about the unknown population mean?
Hint: How far is a sample mean typically from mu, and can that statement be turned around?
Answer:
The standard error is 1 over the square root of 100, which is 0.1. By the Empirical Rule, about 95 percent of samples give a mean within two standard errors — 0.2 units — of mu.
Now turn it round. If the sample mean is within 0.2 of mu, then mu is within 0.2 of the sample mean, since being within a distance is symmetric. So mu lies between 1.8 and 2.2 for 95 percent of samples.
That interval is the book's example, and the reversal in the second step is the whole idea of this chapter. Nothing new has been computed — the standard error came from chapter 7 and the two came from chapter 6 — but the question being answered has changed direction entirely.
Concept
A confidence interval is an estimate that is an interval of numbers rather than just one. It provides a range of reasonable values in which we expect the population parameter to fall. There is no guarantee that a given confidence interval does capture the parameter, but there is a predictable probability of success.
point estimate and interval estimate — The sample mean is the point estimate of the population mean, and the sample standard deviation is the point estimate of the population standard deviation. A confidence interval surrounds a point estimate with a margin of error to give a range instead.
\[ (\text{point estimate} - \text{EBM}, \;\text{point estimate} + \text{EBM}) \]
The book's warning about interpretation belongs here rather than later, because everything else in the chapter depends on it: the confidence interval is a random variable, and it is the population parameter that is fixed. Each new sample produces a different interval; mu never moves. That is why a confidence level is a statement about how the procedure behaves over many samples and not about the one interval in front of you.
Figure (svg): A number line showing a point estimate at sixty-eight with an error bound either side, giving an interval from sixty-seven point one eight to sixty-eight point eight two
OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, p. 407
Section
Section 1
Concept
The sample mean is the point estimate for the population mean. A confidence interval takes that point estimate and extends it by a margin of error, called here the error bound for a population mean, in both directions.
error bound for a population mean — The margin of error, abbreviated EBM. The interval is the point estimate minus EBM to the point estimate plus EBM, and because the same amount is added and subtracted these are called symmetrical confidence intervals.
\[ \left(\bar{x} - \text{EBM},\; \bar{x} + \text{EBM}\right) \]
The book notes that some reports use the phrase margin of error while others give an interval, and that these are two ways of expressing the same concept. It also adds a caveat worth remembering: although the text only covers symmetrical confidence intervals, there are non-symmetrical ones — a confidence interval for a standard deviation is its example, and section 11.6 will produce one.
Figure (svg): A number line showing a point estimate at sixty-eight with an error bound either side, giving an interval from sixty-seven point one eight to sixty-eight point eight two
OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 407-408 — the form of a confidence interval and Example 8.1
Picture it
Example 8.2: a sample mean of 68 with an error bound of 0.8225.
Figure (svg): A number line showing a point estimate at sixty-eight with an error bound either side, giving an interval from sixty-seven point one eight to sixty-eight point eight two
Sections 8.2 and 8.3 change what goes into the error bound but not this shape. A point estimate in the middle with an equal margin either side is what a confidence interval looks like throughout the rest of the book, so recognising the shape is worth more than memorising any one formula.
Worked example
Example 8.1, which supplies both pieces directly.
\[ \bar{x} = 7, \; \text{EBM} = 2.5, \; \text{CL} = 95\text{ percent} \]
Subtract the error bound
Why: Seven minus 2.5.
\[ 4.5 \]
Add the error bound
Why: Seven plus 2.5.
\[ 9.5 \]
Write the interval
Why: Lower then upper.
\[ (4.5, 9.5) \]
Write the sentence
Why: In the book's template.
\[ 95 \%\text{ confidence} \]
Figure (svg): The solution to Worked example building an interval from its parts shown as a ladder of expressions, one row per legal move
\[ (7 - 2.5,\; 7 + 2.5) = (4.5, 9.5) \]
Verify: confirm the point estimate sits at the centre
Why: The midpoint of 4.5 and 9.5 is 7, which is the sample mean — as it must be for a symmetrical interval. That check works in both directions and is used in the last idea to recover a sample mean from a published interval. An interval whose midpoint is not the point estimate has been assembled wrongly.
OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, p. 407
Faded example
A sample mean is 15 and the error bound is 3.2.
Fill in the blanks
\left(15 - 3.2,\; 15 + 3.2\right) = (11.8, 18.2)
Why: The interval runs from 11.8 to 18.2, and its midpoint is 15 — the point estimate, as it must be for a symmetrical interval.
Worked example
The chapter's opening illustration, built from the Empirical Rule.
\[ \sigma = 1, \; n = 100, \; \bar{x} = 2 \]
Find the standard error
Why: One over root 100.
\[ 0.1 \]
Take two of them
Why: The Empirical Rule's 95 percent.
\[ 0.2 \]
Subtract and add
Why: Two minus and plus 0.2.
\[ 1.8\text{ and } 2.2 \]
State the interval
Why: With its level.
\[ (1.8, 2.2) \]
Figure (svg): The solution to Worked example the Apple Music interval shown as a ladder of expressions, one row per legal move
\[ 2 \pm (2)(0.1) \;\Rightarrow\; (1.8, 2.2) \]
Verify: confirm the two possibilities the book names
Why: The book spells out what the interval implies: either the interval contains the true mean, or the sample produced a mean that is not within 0.2 units of it. The first happens for 95 percent of well-chosen samples, and the second happens for 5 percent even though correct procedures are followed. Recognising that the second case is not an error but an expected outcome is central to reading any confidence interval.
OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 406-407
Trap
\[ (1.8, 2.2) \;\Rightarrow\; P(1.8 < \mu < 2.2) = 0.95 \]
Treat the interval as fixed and mu as random
Why: The interval is the thing written down, so mu seems to be the unknown that moves.
\[ \text{but } \mu \text{ is a fixed number} \]
Either mu is in that interval or it is not; there is no randomness left once the sample has been taken.
\[ \text{95 percent of intervals built this way contain } \mu \]
Locate the randomness in the procedure, not in the parameter
Why: The book: the confidence interval is a random variable, and it is the population parameter that is fixed.
The book's own phrasing is worth quoting because it is careful: the confidence level is the percent of confidence intervals that contain the true population parameter when repeated samples are taken. That sentence is longer than the wrong one and says something different, and the next idea's figure is what makes the difference visible.
Two truths and a lie
All three concern what an interval is.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and the book contradicts it directly: there is no guarantee that a given confidence interval does capture the parameter, but there is a predictable probability of success. An interval that missed was not produced by a mistake — missing is what the other 5 percent of correct procedure looks like.
Prediction
Commit before reasoning.
Predict first
After a sample has been taken and an interval computed, what is random?
Correct: Nothing.
Why: Once the sample is drawn, the interval is a pair of specific numbers and mu was never random. The randomness lives in the procedure — in which sample you happened to get — which is why the confidence level describes what happens across repeated samples rather than making a probability statement about the interval in front of you.
Fill the middle
How every confidence interval in this chapter is written.
Fill in the blanks
\left(\textEBM - \text___,\; \text___ + ___\right)
Why: The same error bound is added and subtracted, which is what makes these symmetrical confidence intervals and what puts the point estimate exactly at the midpoint.
Section
Section 2
Concept
The confidence level is the area in the middle of the standard normal distribution. Alpha is the probability that the interval does not contain the parameter, and alpha plus the confidence level equals one. Alpha is split equally between the two tails, so each holds alpha over two.
z sub alpha over two — The z-score with an area of alpha over two to its right. For a 95 percent level, alpha is 0.05 and alpha over two is 0.025, so the z-score wanted is the one with 0.975 to its left, namely 1.96.
\[ \alpha + \text{CL} = 1, \qquad \text{each tail holds } \frac{\alpha}{2} \]
This is exactly section 6.2's middle-region arithmetic: subtract the level from one, halve the remainder, and read off the boundaries. The book's own note is a practical one — remember to use the area to the LEFT when reversing a cumulative, so a 95 percent interval needs 0.975 rather than 0.025 typed into a calculator.
Figure (svg): A bell curve with the central ninety percent shaded green and two red tails of five percent each, marked at plus and minus one point six four five
OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 408-409 — the relationship between CL and alpha, and finding the z-score
Picture it
The area the confidence level names, and the two tails alpha leaves behind.
Figure (svg): A bell curve with the central ninety percent shaded green and two red tails of five percent each, marked at plus and minus one point six four five
The 1.645 marked on this picture is the same number section 7.2's percentile check used, and the same one that will appear in every 90 percent interval and every one-tailed 5 percent test from chapter 9 onward. It is worth recognising rather than recomputing.
Worked example
Example 8.2's second step.
\[ \text{CL} = 0.90; \text{ find } z_{\alpha/2} \]
Find alpha
Why: One minus 0.90.
\[ 0.10 \]
Halve it
Why: Each tail's area.
\[ 0.05 \]
Find the area to the left
Why: One minus 0.05.
\[ 0.95 \]
Reverse the cumulative
Why: On the standard normal.
\[ 1.645 \]
Figure (svg): The solution to Worked example the z-score for 90 percent shown as a ladder of expressions, one row per legal move
\[ z_{0.05} = 1.645 \]
Verify: confirm the z-score is consistent with the level's size
Why: A 90 percent level should need a z between 1 and 2, since the Empirical Rule puts 68 percent within one standard deviation and 95 percent within two — and 1.645 sits between them, closer to the upper end as 90 is closer to 95 than to 68. Any z-score outside that bracket for a level between 68 and 95 percent signals that alpha was halved wrongly or not at all.
OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 409-410
Faded example
A 99 percent confidence level.
Fill in the blanks
\alpha = 0.01, \quad \frac2.576___ = 0.005, \quad z = \text___(0.995) \approx ___
Why: Alpha is 0.01, each tail holds 0.005, and the area to the left of the boundary is 0.995 — giving about 2.576, the largest of the four z-scores in common use.
Worked example
Example 8.3's step, with a smaller alpha.
\[ \text{CL} = 0.98; \text{ find } z_{\alpha/2} \]
Find alpha
Why: One minus 0.98.
\[ 0.02 \]
Halve it
Why: Each tail.
\[ 0.01 \]
Find the area to the left
Why: One minus 0.01.
\[ 0.99 \]
Reverse the cumulative
Why: On the standard normal.
\[ 2.326 \]
Figure (svg): The solution to Worked example the z-score for 98 percent shown as a ladder of expressions, one row per legal move
\[ z_{0.01} = 2.326 \]
Verify: confirm it exceeds the 95 percent value, and why it must
Why: A 98 percent level leaves less in the tails than a 95 percent one, so its boundary must sit further out: 2.326 against 1.96. Confidence levels and z-scores move together, always, which makes the ordering of these four numbers — 1.645, 1.96, 2.326, 2.576 — a check on any of them. A higher level with a smaller z is impossible.
OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, p. 412
Error analysis
The correct value is 1.96.
Annotate
On: \( \begin{aligned} &(1)\; \text{invNorm}(0.95) = 1.645 \\ &(2)\; \text{invNorm}(0.05) = -1.645 \\ &(3)\; \text{invNorm}(0.025) = -1.96 \\ &(4)\; \text{invNorm}(0.975) = 1.96 \end{aligned} \)
Error (1) is the dangerous one, because 1.645 is a perfectly legitimate z-score for a different level, so nothing about it looks wrong. The defence is the ordering check: a 95 percent interval must be wider than a 90 percent one, so its z cannot be the smaller of the two.
Matching
Match each confidence level to its z.
Match the pairs
Why: These four recur constantly from here to the end of the book. Note that they increase with the level, which is the ordering check: a more confident interval always needs a larger multiplier and so comes out wider.
Two truths and a lie
All three concern the level and alpha.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and is the section's commonest slip. The left area wanted is one minus alpha over two, which is 0.975 rather than 0.95 — and invNorm(0.95) returns 1.645, the z-score for a 90 percent interval.
Prediction
Commit before reasoning.
Predict first
Moving from a 90 percent to a 99 percent confidence level, what happens to the z-score?
Correct: It rises, from 1.645 to 2.576.
Why: A higher level leaves less area in the tails, so the boundary must move further out. The second option confuses a smaller alpha with a smaller boundary — they move in opposite directions. The sample size does not enter the z-score at all; it enters the standard error.
Section
Section 3
Concept
The error bound is the z-score for the confidence level multiplied by the standard error of the mean. The book stresses that the standard deviation used must be appropriate for the parameter being estimated, so it is sigma over the square root of n rather than sigma itself.
EBM — The error bound for a population mean, equal to z sub alpha over two times sigma over the square root of n. The interval is then the sample mean plus and minus that quantity.
\[ \text{EBM} = z_{\alpha/2}\left(\frac{\sigma}{\sqrt{n}}\right) \]
The whole construction depends on chapter 7. It is because the sample mean is normally distributed with that standard error that a z-score can be used at all — and the book says so, summarising that as a result of the central limit theorem, X-bar is normally distributed, and when sigma is known we use a normal distribution to calculate the error bound. Section 8.2 changes exactly that second clause.
Figure (svg): The five steps for constructing and interpreting a confidence interval
OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 408-411 — the steps and Example 8.2
Picture it
The book's own procedure.
Figure (svg): The five steps for constructing and interpreting a confidence interval
The last step carries a template worth using verbatim at first: we estimate with a given percent confidence that the true population mean, described in the words of the problem, is between two values with their units. It forces the level, the parameter and both endpoints into one sentence, which is exactly what a reader needs.
Worked example
Example 8.2, the chapter's central worked case.
\[ \sigma = 3, \; n = 36, \; \bar{x} = 68, \; \text{CL} = 0.90 \]
Find the standard error
Why: Three over root 36.
\[ 0.5 \]
Find the z-score
Why: For a 90 percent level.
\[ 1.645 \]
Compute the error bound
Why: 1.645 times 0.5.
\[ 0.8225 \]
Build the interval
Why: 68 minus and plus it.
\[ (67.18, 68.82) \]
Figure (svg): The solution to Worked example a 90 percent interval for exam scores shown as a ladder of expressions, one row per legal move
\[ 68 \pm (1.645)\left(\frac{3}{\sqrt{36}}\right) = (67.18, 68.82) \]
Verify: confirm the standard error rather than sigma was used
Why: Using sigma of 3 in place of the standard error of 0.5 would give an error bound of 4.935 and an interval from 63.07 to 72.94 — six times too wide. The book flags this explicitly, saying the standard deviation used must be appropriate for the parameter being estimated. Since the parameter is a mean, the relevant spread is the mean's, which is the standard error.
OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 409-411
Faded example
Pizza delivery times have sigma = 6 minutes; a sample of 36 gives a mean of 36 minutes. Use 90 percent.
Fill in the blanks
\text1.645 = (1.645)\left(\frac3.29___}\right) = ___, \quad \text___ = (36 - ___,\; 36 + ___) \text___ ___
Why: The standard error is 6 over 6, which is 1, so the error bound is just the z-score itself, 1.645. The interval runs from 34.355 to 37.645, a width of 3.29 — twice the error bound, as always.
Worked example
Example 8.3, where the sample mean has to be computed first.
\[ n = 30, \; \sigma = 0.337, \; \text{CL} = 0.98 \]
Compute the point estimate
Why: The mean of the 30 values.
\[ 1.024 \]
Find the z-score
Why: For 98 percent.
\[ 2.326 \]
Compute the error bound
Why: 2.326 times 0.337 over root 30.
\[ 0.1431 \]
Build the interval
Why: 1.024 minus and plus it.
\[ (0.8809, 1.1671) \]
Figure (svg): The solution to Worked example a 98 percent interval for SAR levels shown as a ladder of expressions, one row per legal move
\[ 1.024 \pm 0.1431 = (0.8809, 1.1671) \]
Verify: confirm where the book's last digits come from
Why: The exact sample mean of these 30 values is 1.023733, and the book rounds it to 1.024 before subtracting — which is what produces its 0.8809 rather than the 0.8806 the unrounded mean gives. The difference is 0.0003 and changes nothing, but it is worth recognising as intermediate rounding rather than an error in either calculation. Carrying full precision until the final step is the better habit.
OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 411-412
Trap
\[ \text{EBM} = (1.645)(3) = 4.935 \]
Multiply the z-score by the population standard deviation
Why: Sigma is the number the problem supplies.
\[ (63.07, 72.94) \text{ instead of } (67.18, 68.82) \]
That interval describes where individual exam scores fall, not where the population mean lies.
\[ \text{EBM} = (1.645)\left(\frac{3}{\sqrt{36}}\right) = 0.8225 \]
Divide sigma by the square root of n first
Why: The parameter being estimated is a mean, so the mean's spread applies.
The book raises this in its own words: the standard deviation used must be appropriate for the parameter we are estimating. It is the same distinction chapter 7 spent a whole section on, and it produces the same size of error — a factor of the square root of n, which here is six.
Discrimination
The parameter being estimated decides.
Sort into buckets
Sort each quantity by whether it belongs in the error bound for a mean.
Two truths and a lie
All three concern the construction.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The error bound uses the STANDARD ERROR, sigma over the square root of n. Using sigma directly gives an interval about the square root of n times too wide, which describes where individual values fall rather than where the mean lies.
Estimation
A 95 percent interval is built from sigma = 10 and n = 100.
Predict first
Roughly how wide is the interval?
Correct: About 3.9.
Why: The standard error is 10 over 10, which is 1, so the error bound is about 1.96 and the width is twice that, about 3.9. The option 39 uses sigma directly instead of the standard error — a factor of the square root of 100, which is exactly the error this idea's trap describes.
Section
Section 4
Concept
Increasing the confidence level increases the error bound and makes the interval wider; decreasing it makes the interval narrower. Increasing the sample size causes the error bound to decrease and makes the interval narrower; decreasing the sample size makes it wider.
the width trade-off — Confidence and precision pull against each other. More confidence costs width, and the only way to buy back precision without losing confidence is a larger sample.
\[ \text{EBM} = z_{\alpha/2}\,\frac{\sigma}{\sqrt{n}}: \quad z \uparrow \Rightarrow \text{wider}, \quad n \uparrow \Rightarrow \text{narrower} \]
Both effects are visible in the formula, and they are the only two things a researcher controls. Sigma belongs to the population and cannot be chosen; the level and the sample size can. Because n enters under a square root, buying precision is expensive — halving the width requires quadrupling the sample, which is exactly the diminishing return chapter 7 established.
Figure (svg): Two horizontal intervals of different widths centred on the same sample mean, the ninety-five percent one visibly wider than the ninety percent one
OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 413-415 — Examples 8.4 and 8.5, and the two summaries
Picture it
Examples 8.2 and 8.4: a 90 percent interval and a 95 percent one.
Figure (svg): Two horizontal intervals of different widths centred on the same sample mean, the ninety-five percent one visibly wider than the ninety percent one
Nothing about the sample changed between these two — same mean, same sigma, same n. Only the level was raised, and the interval widened from 1.645 to 1.96 error bounds. That is the honest cost of confidence, and it explains why a 100 percent interval, which would have to be infinitely wide, would carry no information at all.
Worked example
Example 8.4, changing only the level.
\[ \sigma = 3, \; n = 36, \; \bar{x} = 68, \; \text{CL} = 0.95 \]
Find alpha over two
Why: 0.05 halved.
\[ 0.025 \]
Find the z-score
Why: invNorm at 0.975.
\[ 1.96 \]
Compute the error bound
Why: 1.96 times 0.5.
\[ 0.98 \]
Build the interval
Why: 68 minus and plus it.
\[ (67.02, 68.98) \]
Figure (svg): The solution to Worked example raising the level to 95 percent shown as a ladder of expressions, one row per legal move
\[ 68 \pm (1.96)(0.5) = (67.02, 68.98) \]
Verify: confirm the wider interval is the more confident one
Why: The 95 percent interval is wider by about 0.32 in total, and the book explains why that makes sense: because the area 0.95 is larger than the area 0.90, and to be more confident that the interval contains the true value it necessarily needs to be wider. An interval that got NARROWER as the level rose would signal that alpha had been halved wrongly.
OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 413-414
Sorting
Start from a 90 percent interval built on n = 36.
Sort into buckets
Sort each change by its effect on the interval's width.
The two levers pull in opposite directions and can be traded against each other: raising the level from 90 to 95 percent widens by a factor of 1.96 over 1.645, and that can be paid for with a sample about 42 percent larger. Study design is largely this trade.
Worked example
Example 8.5, holding the level at 90 percent.
\[ \sigma = 3, \; \text{CL} = 0.90; \; n = 100 \text{ then } n = 25 \]
At n = 100
Why: 1.645 times 3 over 10.
\[ 0.4935 \]
Compare with n = 36
Why: The original.
\[ 0.8225 \]
At n = 25
Why: 1.645 times 3 over 5.
\[ 0.987 \]
State the pattern
Why: Larger n, smaller bound.
Figure (svg): The solution to Worked example changing the sample size shown as a ladder of expressions, one row per legal move
\[ n = 100 \Rightarrow 0.4935, \qquad n = 25 \Rightarrow 0.987 \]
Verify: confirm the error bounds scale with one over root n
Why: Going from n equal to 25 to n equal to 100 quadruples the sample and exactly halves the error bound, from 0.987 to 0.4935 — the square root law from chapter 7, appearing here as the cost of precision. Any pair of error bounds from the same sigma and level must stand in the inverse ratio of the square roots of their sample sizes.
OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 414-415
Trap
\[ \text{raise the level to } 99 \text{ percent and keep the interval} \]
Treat the confidence level as a label on a fixed interval
Why: The data has not changed, so the interval should not either.
\[ (67.18, 68.82) \text{ described as } 99 \text{ percent confident} \]
That interval is only 90 percent confident; calling it 99 percent overstates what the data supports.
\[ 99 \text{ percent needs } z = 2.576, \text{ giving } (66.71, 69.29) \]
Recompute the error bound whenever the level changes
Why: The level enters through the z-score.
Confidence and width are locked together by the formula, and the only escape is a larger sample. This is why published studies quote the level alongside every interval: an interval without its level is uninterpretable, since the same numbers could be a modest claim at 99 percent or a strong one at 80 percent.
Faded example
An error bound of 0.987 comes from n = 25 at a 90 percent level.
Fill in the blanks
\text4 n \text100 ___, \text___ n = ___
Why: Halving the error bound needs the square root of n doubled, so n itself is quadrupled: 25 becomes 100, and the bound falls from 0.987 to 0.4935 — exactly Example 8.5's two values.
Two truths and a lie
All three concern width.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, or at least badly incomplete. A wider interval is more likely to contain mu but says less about where mu is — in the limit, the interval from minus infinity to infinity is certain and useless. Better means the right trade between confidence and precision, which depends on what the estimate is for.
Prediction
Commit before reasoning.
Predict first
A researcher wants a narrower interval without losing confidence. What must they do?
Correct: Collect a larger sample.
Why: The error bound has exactly three inputs: the z-score set by the level, sigma set by the population, and n set by the researcher. Holding the level fixed rules out the first, sigma is not a choice, so only the sample size remains — and because it enters under a square root, the gain is expensive.
Section
Section 5
Concept
A published study may give only the interval. The error bound is then the upper value minus the sample mean, or half the difference between the endpoints; the sample mean is the upper value minus the error bound, or the average of the two endpoints. Solving the error bound formula for n gives the sample size a study needs.
the sample size formula — n equals z sigma over EBM, all squared. Because a sample size must be a whole number, the book instructs that the answer is always rounded UP to the next higher integer to ensure the sample is large enough.
\[ n = \left(\frac{z_{\alpha/2}\,\sigma}{\text{EBM}}\right)^2 \]
The rounding instruction is worth taking literally. Rounding 216.09 down to 216 would give an error bound slightly larger than the two years asked for, which fails the requirement; rounding up to 217 satisfies it with a little to spare. This is one of the few places in the book where ordinary rounding rules are deliberately overridden, and the reason is that the requirement is one-sided.
Figure (svg): A four-column table listing confidence levels of ninety to ninety-nine percent with their alpha values and z-scores
OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 415-416 — working backwards and Example 8.7
Picture it
Each one feeds both the interval and the sample-size formula.
Figure (svg): A four-column table listing confidence levels of ninety to ninety-nine percent with their alpha values and z-scores
Because the z-score is squared in the sample-size formula, moving from 95 to 99 percent multiplies the required sample by 2.576 over 1.96 squared, which is about 1.73. Raising confidence is therefore substantially more expensive at the planning stage than the interval widths alone suggest.
Worked example
Example 8.6, both ways.
\[ \text{the interval } (67.18, 68.82) \]
Half the width
Why: 68.82 minus 67.18, halved.
\[ EBM = 0.82 \]
Or subtract a known mean
Why: 68.82 minus 68.
\[ 0.82\text{ again} \]
Average the endpoints
Why: 67.18 plus 68.82, halved.
\[ \text{mean } = 68 \]
Or subtract the bound
Why: 68.82 minus 0.82.
\[ 68\text{ again} \]
Figure (svg): The solution to Worked example recovering the parts of an interval shown as a ladder of expressions, one row per legal move
\[ \text{EBM} = \frac{68.82 - 67.18}{2} = 0.82, \qquad \bar{x} = \frac{67.18 + 68.82}{2} = 68 \]
Verify: confirm the two routes must agree
Why: Both work because the interval is symmetrical: the mean sits exactly at the midpoint and the bound is exactly half the width. The book offers each calculation two ways precisely so that whichever information you happen to have is enough. If the two routes disagreed, the interval would not be symmetrical — which for this chapter's methods would mean an arithmetic error.
OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 415-416
Faded example
A published confidence interval is (42.12, 47.88).
Fill in the blanks
\text2.88 = \frac45___ = ___, \qquad \bar___ = \frac______ = ___
Why: Half the width of 5.76 is 2.88, and the midpoint of the endpoints is 45. Both follow from the interval being symmetrical about its point estimate.
Worked example
Example 8.7, planning a study.
\[ \sigma = 15, \; \text{EBM} = 2, \; \text{CL} = 0.95 \]
Find the z-score
Why: For 95 percent.
\[ 1.96 \]
Form the ratio
Why: 1.96 times 15, over 2.
\[ 14.7 \]
Square it
Why: The sample size formula.
\[ 216.09 \]
Round UP
Why: To the next whole number.
\[ 217 \]
Figure (svg): The solution to Worked example how many students to survey shown as a ladder of expressions, one row per legal move
\[ n = \left(\frac{(1.96)(15)}{2}\right)^2 = 216.09 \;\to\; 217 \]
Verify: confirm that rounding down would fail the requirement
Why: With n equal to 216 the error bound is 1.96 times 15 over the square root of 216, about 2.0004 years — just over the two years required. With 217 it is about 1.9958, just under. The margin is tiny, but the requirement was to be within two years, so rounding up is what satisfies it. This is why the book insists on always rounding up rather than to the nearest integer.
OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, p. 416
Trap
\[ n = 216.09 \;\to\; 216 \]
Round to the nearest integer, as usual
Why: 0.09 is far below a half.
\[ \text{EBM} = 2.0004 > 2 \]
The requirement was to be within two years, and 216 students misses it — narrowly, but it misses.
\[ n = 216.09 \;\to\; 217 \]
Always round a sample size UP
Why: The book states this as a rule, and the reason is that the requirement is one-sided.
The general point is that not every rounding decision is symmetric. A sample size has to be at least large enough, so any fractional part demands another whole observation. The same reasoning appears wherever a resource must cover a requirement — and here the cost of the extra student is trivial against the cost of missing the target.
Faded example
Heights have sigma = 3 inches, and we want 95 percent confidence within one inch.
Fill in the blanks
n = \left(\frac34.5735\right)^2 = ___ \;\to\; ___
Why: 5.88 squared is about 34.57, which rounds up to 35 students. Rounding down to 34 would give an error bound slightly over the one inch required.
Two truths and a lie
All three concern working backwards.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The error bound is squared in the sample-size formula, so halving it QUADRUPLES the required n — the same square root law as everywhere else in these two chapters. Doubling would only reduce the bound by a factor of about 1.41.
Estimation
A study needs n = 216 for 95 percent confidence within two years.
Predict first
Roughly what n would 99 percent confidence within the same two years need?
Correct: About 374.
Why: The z-score rises from 1.96 to 2.576, and it is squared in the formula, so n grows by a factor of about 1.73 — from 216 to about 374. Raising confidence is considerably more expensive at the planning stage than the modest widening of an interval suggests.
Comparison
Fill the blanks. Only one of the three is a researcher's free choice at fixed confidence.
Comparison matrix
| Input | Where it comes from | Effect of increasing it |
|---|---|---|
| z sub alpha over two | the confidence level chosen | wider interval |
| sigma | the population; not a choice | wider interval |
| n | the researcher's sample | narrower interval |
| Cost of halving the width | quadruple n | because n is under a square root |
The last row is the practical summary of the whole chapter so far. Precision is bought with sample size, at a square-root rate, and there is no other way to buy it without giving up confidence.
Pattern
Six steps, and the last two are the ones that get skipped.
Check the interval by confirming its midpoint is the sample mean and its half-width is the error bound. Check the z-score against the ordering 1.645, 1.96, 2.326, 2.576.
OpenStax Introductory Business Statistics 2e, §8.1 A Confidence Interval When the Population Standard Deviation Is Known or Large Sample Size §8.1 A Confidence Interval When the Population Standard Deviation Is Known or Large Sample Size
Check
The z-score.
Check your understanding
What z-score is used for a 95 percent confidence interval?
Answer: A
Why: Alpha is 0.05, alpha over two is 0.025, and the z-score with 0.975 to its left is 1.96.
Check
The error bound.
Check your understanding
For sigma = 3, n = 36 and a 90 percent level, what is the error bound?
Answer: A
Why: The standard error is 3 over 6, which is 0.5, and 1.645 times 0.5 is 0.8225.
Check
Interpretation.
Check your understanding
A 95 percent confidence interval for a mean is (4.5, 9.5). Which statement is correct?
Answer: A
Why: The confidence level describes the procedure across repeated samples. The interval is the random thing; the parameter is fixed.
Real world
A polling firm reports that support for a proposal is 46 percent with a margin of error of 3 percentage points at 95 percent confidence. A commentator writes that since the interval runs from 43 to 49 and is entirely below 50, the proposal is certain to fail. A second commentator says the poll is worthless because the true figure could be anything.
Discussion prompt
Assess both claims, and say what the poll does and does not establish.
Hint: One commentator overstates the certainty and the other understates the information.
Answer:
The first commentator overstates it. The interval from 43 to 49 does lie entirely below 50, which is genuine evidence that support is short of a majority — but a 95 percent interval misses the truth about one time in twenty, and that failure rate is not a remote technicality. Certain is the wrong word for a claim resting on a procedure that is wrong 5 percent of the time.
The second commentator understates it badly. The poll narrows the plausible range to six percentage points out of a hundred, which is a great deal of information. Dismissing an interval because it is not a point is the mirror image of the first error.
\[ 46 \pm 3 \;\Rightarrow\; (43, 49), \qquad \text{width} = 2\,\text{EBM} \]
What the poll establishes is a range and a procedure, not a verdict. The honest reading is that support is probably in the low-to-high forties, that a majority looks unlikely on this evidence, and that the estimate would be sharper with a larger sample — narrowing the margin to 1.5 points would require quadrupling the sample, by the square root law.
Two cautions belong in a careful answer, and both fall outside the arithmetic entirely. The margin of error quantifies only SAMPLING variability: it says nothing about people who were not reachable, who declined to answer, or who answered differently from how they will vote, and those non-sampling errors are frequently larger than the quoted margin. And a single poll is one interval from one sample — the reason polling aggregators combine many is precisely that the confidence level is a statement about a collection of intervals rather than about any one of them.
Commit first
Answer, then rate your confidence honestly.
Predict first
What does a 95 percent confidence level actually describe?
Correct: The percent of intervals containing mu across repeated samples.
\[ \alpha + \text{CL} = 1, \qquad \text{EBM} = z_{\alpha/2}\,\frac{\sigma}{\sqrt{n}} \]
Why: This is the book's own careful wording, and it is careful for a reason: the interval is a random variable while the population parameter is fixed. Once a sample has been taken, the interval either contains mu or it does not, and no probability attaches to that particular case. The level describes how the procedure behaves, not how this instance turned out.
Explain it
They built a 95 percent interval for a mean using invNorm(0.95), getting 1.645 as the z-score.
Discussion prompt
In two sentences or fewer, locate the error.
Hint: Ask how much area their z-score leaves in each tail.
Answer:
Their 1.645 leaves 0.05 in the upper tail, so the interval between plus and minus it holds 90 percent rather than 95.
Alpha of 0.05 has to be halved before it is used, so the area to the left is 0.975 and the z-score is 1.96.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, the interval is random and the parameter is fixed. For the second, halve alpha and use one minus alpha over two as the left area. For the third, divide sigma by the square root of n before multiplying by z. For the fourth, square the ratio and always round up. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Twenty minutes.
Draw it
At the top, write the interval's standard form as point estimate minus and plus the error bound, and beneath it the error bound formula with its three inputs labelled by where each comes from. Below that, draw a standard normal curve, shade the central 90 percent, mark 1.645 on both sides, and label the two tails 0.05 each — then write the four-row table of confidence level, alpha, alpha over two and z for 90, 95, 98 and 99 percent. In the middle of the page, work Example 8.2 completely: standard error, z-score, error bound and interval, ending with the full interpreting sentence. Beside it, redo the error bound at 95 percent and draw the two intervals one above the other to scale, so the wider one is visibly wider. Then redo it at n = 100 and n = 25, and write one sentence on which lever each change pulled. Near the bottom, draw twenty short horizontal intervals scattered around one vertical line for mu, marking one or two that miss it, and write beside it the sentence that the interval is random and the parameter is fixed. Finish with Example 8.7's sample size calculation, showing the squaring and the rounding up, and one sentence on why rounding down would fail.
Check the middle section by confirming the midpoint of each interval you draw is exactly 68, and that the 95 percent one is wider than the 90 percent one by the ratio 1.96 over 1.645. Check the sample size by computing the error bound at your answer and confirming it is just UNDER the two years required, not just over.
Recap
Five things, and the first is the one examiners actually test.
| If you see | Then |
|---|---|
| A stated confidence level | alpha is one minus it, halved between the tails |
| A 95 percent interval | The z-score is 1.96, from invNorm(0.975) |
| sigma and n given | The error bound uses sigma over root n |
| A higher confidence level | A larger z, so a wider interval |
| A published interval only | The mean is its midpoint, the bound its half-width |
| A target error bound | Square z sigma over EBM, and round UP |
| A claim that mu is probably in this interval | Reword: the procedure captures mu 95 percent of the time |
Section 8.2 removes the assumption that makes this section work. In practice the population standard deviation is almost never known, and replacing it with the sample standard deviation turns out to need a different distribution — Gosset's t, developed at the Guinness brewery precisely because small samples made the normal approximation unreliable.
OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 406-416 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.