8.1 A Single Population Mean using the Normal Distribution

The first act of statistical inference in the course. Chapter 7 asked how far a sample mean is likely to fall from a known population mean; this section reverses the question and asks which values of an unknown population mean are consistent with an observed sample mean. The answer is not a single number but an interval, formed as the point estimate plus and minus a margin of error called the error bound, which equals a z-score chosen by the confidence level multiplied by the standard error of the mean. The confidence level and its complement alpha determine that z-score, since alpha is split equally between the two tails. Raising the confidence level or shrinking the sample widens the interval, and the error bound formula can be solved backwards to find the sample size a study needs. The interpretation is the section's real difficulty: the interval is the random thing and the parameter is fixed, so the confidence level describes the procedure across repeated samples rather than any one interval.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 8.1 A Single Population Mean using the Normal Distribution

Title

Statistics · Chapter 8 — Confidence Intervals

A Single Population Mean using the Normal Distribution

2. By the end of this lesson you can

Objectives

Five outcomes, and the last two sentences of the first one are where most marks are lost.

OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 406-416 — the section these objectives are drawn from

3. What you already have

Warm-up

Chapter 7 established that a sample mean of n values has standard error sigma over the square root of n.

Discussion prompt

Songs downloaded per month have a known population standard deviation of 1, and a sample of 100 gives a mean of 2. What can be said about the unknown population mean?

Hint: How far is a sample mean typically from mu, and can that statement be turned around?

Answer:

The standard error is 1 over the square root of 100, which is 0.1. By the Empirical Rule, about 95 percent of samples give a mean within two standard errors — 0.2 units — of mu.

Now turn it round. If the sample mean is within 0.2 of mu, then mu is within 0.2 of the sample mean, since being within a distance is symmetric. So mu lies between 1.8 and 2.2 for 95 percent of samples.

That interval is the book's example, and the reversal in the second step is the whole idea of this chapter. Nothing new has been computed — the standard error came from chapter 7 and the two came from chapter 6 — but the question being answered has changed direction entirely.

4. An interval of reasonable values, not a single number

Concept

A confidence interval is an estimate that is an interval of numbers rather than just one. It provides a range of reasonable values in which we expect the population parameter to fall. There is no guarantee that a given confidence interval does capture the parameter, but there is a predictable probability of success.

point estimate and interval estimate — The sample mean is the point estimate of the population mean, and the sample standard deviation is the point estimate of the population standard deviation. A confidence interval surrounds a point estimate with a margin of error to give a range instead.

\[ (\text{point estimate} - \text{EBM}, \;\text{point estimate} + \text{EBM}) \]

The book's warning about interpretation belongs here rather than later, because everything else in the chapter depends on it: the confidence interval is a random variable, and it is the population parameter that is fixed. Each new sample produces a different interval; mu never moves. That is why a confidence level is a statement about how the procedure behaves over many samples and not about the one interval in front of you.

Figure (svg): A number line showing a point estimate at sixty-eight with an error bound either side, giving an interval from sixty-seven point one eight to sixty-eight point eight two

The sample mean sits in the middle and the error bound is the same distance either side, which is why the book calls these symmetrical confidence intervals.

OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, p. 407

5. Point estimate and margin of error

Section

Section 1

6. One number, then a range around it

Concept

The sample mean is the point estimate for the population mean. A confidence interval takes that point estimate and extends it by a margin of error, called here the error bound for a population mean, in both directions.

error bound for a population mean — The margin of error, abbreviated EBM. The interval is the point estimate minus EBM to the point estimate plus EBM, and because the same amount is added and subtracted these are called symmetrical confidence intervals.

\[ \left(\bar{x} - \text{EBM},\; \bar{x} + \text{EBM}\right) \]

The book notes that some reports use the phrase margin of error while others give an interval, and that these are two ways of expressing the same concept. It also adds a caveat worth remembering: although the text only covers symmetrical confidence intervals, there are non-symmetrical ones — a confidence interval for a standard deviation is its example, and section 11.6 will produce one.

Figure (svg): A number line showing a point estimate at sixty-eight with an error bound either side, giving an interval from sixty-seven point one eight to sixty-eight point eight two

The sample mean sits in the middle and the error bound is the same distance either side, which is why the book calls these symmetrical confidence intervals.

OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 407-408 — the form of a confidence interval and Example 8.1

7. The shape of every interval in the chapter

Picture it

Example 8.2: a sample mean of 68 with an error bound of 0.8225.

Figure (svg): A number line showing a point estimate at sixty-eight with an error bound either side, giving an interval from sixty-seven point one eight to sixty-eight point eight two

The sample mean sits in the middle and the error bound is the same distance either side, which is why the book calls these symmetrical confidence intervals.

Sections 8.2 and 8.3 change what goes into the error bound but not this shape. A point estimate in the middle with an equal margin either side is what a confidence interval looks like throughout the rest of the book, so recognising the shape is worth more than memorising any one formula.

8. Worked example: building an interval from its parts

Worked example

Example 8.1, which supplies both pieces directly.

\[ \bar{x} = 7, \; \text{EBM} = 2.5, \; \text{CL} = 95\text{ percent} \]

Subtract the error bound

Why: Seven minus 2.5.

\[ 4.5 \]

Add the error bound

Why: Seven plus 2.5.

\[ 9.5 \]

Write the interval

Why: Lower then upper.

\[ (4.5, 9.5) \]

Write the sentence

Why: In the book's template.

\[ 95 \%\text{ confidence} \]

Figure (svg): The solution to Worked example building an interval from its parts shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ (7 - 2.5,\; 7 + 2.5) = (4.5, 9.5) \]

Verify: confirm the point estimate sits at the centre

Why: The midpoint of 4.5 and 9.5 is 7, which is the sample mean — as it must be for a symmetrical interval. That check works in both directions and is used in the last idea to recover a sample mean from a published interval. An interval whose midpoint is not the point estimate has been assembled wrongly.

OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, p. 407

9. Assemble an interval

Faded example

A sample mean is 15 and the error bound is 3.2.

Fill in the blanks

\left(15 - 3.2,\; 15 + 3.2\right) = (11.8, 18.2)

Why: The interval runs from 11.8 to 18.2, and its midpoint is 15 — the point estimate, as it must be for a symmetrical interval.

10. Worked example: the Apple Music interval

Worked example

The chapter's opening illustration, built from the Empirical Rule.

\[ \sigma = 1, \; n = 100, \; \bar{x} = 2 \]

Find the standard error

Why: One over root 100.

\[ 0.1 \]

Take two of them

Why: The Empirical Rule's 95 percent.

\[ 0.2 \]

Subtract and add

Why: Two minus and plus 0.2.

\[ 1.8\text{ and } 2.2 \]

State the interval

Why: With its level.

\[ (1.8, 2.2) \]

Figure (svg): The solution to Worked example the Apple Music interval shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 2 \pm (2)(0.1) \;\Rightarrow\; (1.8, 2.2) \]

Verify: confirm the two possibilities the book names

Why: The book spells out what the interval implies: either the interval contains the true mean, or the sample produced a mean that is not within 0.2 units of it. The first happens for 95 percent of well-chosen samples, and the second happens for 5 percent even though correct procedures are followed. Recognising that the second case is not an error but an expected outcome is central to reading any confidence interval.

OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 406-407

11. Trap: reading the level as a probability about this interval

Trap

The trap

\[ (1.8, 2.2) \;\Rightarrow\; P(1.8 < \mu < 2.2) = 0.95 \]

Treat the interval as fixed and mu as random

Why: The interval is the thing written down, so mu seems to be the unknown that moves.

\[ \text{but } \mu \text{ is a fixed number} \]

Either mu is in that interval or it is not; there is no randomness left once the sample has been taken.

The fix

\[ \text{95 percent of intervals built this way contain } \mu \]

Locate the randomness in the procedure, not in the parameter

Why: The book: the confidence interval is a random variable, and it is the population parameter that is fixed.

The book's own phrasing is worth quoting because it is careful: the confidence level is the percent of confidence intervals that contain the true population parameter when repeated samples are taken. That sentence is longer than the wrong one and says something different, and the next idea's figure is what makes the difference visible.

12. One of these is false

Two truths and a lie

All three concern what an interval is.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The sample mean is the point estimate for mu
  • C. There is no guarantee a given interval contains the parameter
  • B. A confidence interval always contains the population mean

Survives elimination: B

Why: The survivor is false, and the book contradicts it directly: there is no guarantee that a given confidence interval does capture the parameter, but there is a predictable probability of success. An interval that missed was not produced by a mistake — missing is what the other 5 percent of correct procedure looks like.

13. Where is the randomness?

Prediction

Commit before reasoning.

Predict first

After a sample has been taken and an interval computed, what is random?

  • Nothing: the interval is fixed and mu was always fixed
  • The population mean
  • Both the interval and mu
  • The confidence level

Correct: Nothing.

Why: Once the sample is drawn, the interval is a pair of specific numbers and mu was never random. The randomness lives in the procedure — in which sample you happened to get — which is why the confidence level describes what happens across repeated samples rather than making a probability statement about the interval in front of you.

14. The standard form

Fill the middle

How every confidence interval in this chapter is written.

Fill in the blanks

\left(\textEBM - \text___,\; \text___ + ___\right)

Why: The same error bound is added and subtracted, which is what makes these symmetrical confidence intervals and what puts the point estimate exactly at the midpoint.

15. The confidence level and alpha

Section

Section 2

16. A level in the middle, alpha split between the tails

Concept

The confidence level is the area in the middle of the standard normal distribution. Alpha is the probability that the interval does not contain the parameter, and alpha plus the confidence level equals one. Alpha is split equally between the two tails, so each holds alpha over two.

z sub alpha over two — The z-score with an area of alpha over two to its right. For a 95 percent level, alpha is 0.05 and alpha over two is 0.025, so the z-score wanted is the one with 0.975 to its left, namely 1.96.

\[ \alpha + \text{CL} = 1, \qquad \text{each tail holds } \frac{\alpha}{2} \]

This is exactly section 6.2's middle-region arithmetic: subtract the level from one, halve the remainder, and read off the boundaries. The book's own note is a practical one — remember to use the area to the LEFT when reversing a cumulative, so a 95 percent interval needs 0.975 rather than 0.025 typed into a calculator.

Figure (svg): A bell curve with the central ninety percent shaded green and two red tails of five percent each, marked at plus and minus one point six four five

This is section 6.2's middle-region arithmetic, reused: subtract from one, halve, and read off the two boundaries.

OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 408-409 — the relationship between CL and alpha, and finding the z-score

17. Ninety percent in the middle

Picture it

The area the confidence level names, and the two tails alpha leaves behind.

Figure (svg): A bell curve with the central ninety percent shaded green and two red tails of five percent each, marked at plus and minus one point six four five

This is section 6.2's middle-region arithmetic, reused: subtract from one, halve, and read off the two boundaries.

The 1.645 marked on this picture is the same number section 7.2's percentile check used, and the same one that will appear in every 90 percent interval and every one-tailed 5 percent test from chapter 9 onward. It is worth recognising rather than recomputing.

18. Worked example: the z-score for 90 percent

Worked example

Example 8.2's second step.

\[ \text{CL} = 0.90; \text{ find } z_{\alpha/2} \]

Find alpha

Why: One minus 0.90.

\[ 0.10 \]

Halve it

Why: Each tail's area.

\[ 0.05 \]

Find the area to the left

Why: One minus 0.05.

\[ 0.95 \]

Reverse the cumulative

Why: On the standard normal.

\[ 1.645 \]

Figure (svg): The solution to Worked example the z-score for 90 percent shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ z_{0.05} = 1.645 \]

Verify: confirm the z-score is consistent with the level's size

Why: A 90 percent level should need a z between 1 and 2, since the Empirical Rule puts 68 percent within one standard deviation and 95 percent within two — and 1.645 sits between them, closer to the upper end as 90 is closer to 95 than to 68. Any z-score outside that bracket for a level between 68 and 95 percent signals that alpha was halved wrongly or not at all.

OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 409-410

19. From level to z

Faded example

A 99 percent confidence level.

Fill in the blanks

\alpha = 0.01, \quad \frac2.576___ = 0.005, \quad z = \text___(0.995) \approx ___

Why: Alpha is 0.01, each tail holds 0.005, and the area to the left of the boundary is 0.995 — giving about 2.576, the largest of the four z-scores in common use.

20. Worked example: the z-score for 98 percent

Worked example

Example 8.3's step, with a smaller alpha.

\[ \text{CL} = 0.98; \text{ find } z_{\alpha/2} \]

Find alpha

Why: One minus 0.98.

\[ 0.02 \]

Halve it

Why: Each tail.

\[ 0.01 \]

Find the area to the left

Why: One minus 0.01.

\[ 0.99 \]

Reverse the cumulative

Why: On the standard normal.

\[ 2.326 \]

Figure (svg): The solution to Worked example the z-score for 98 percent shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ z_{0.01} = 2.326 \]

Verify: confirm it exceeds the 95 percent value, and why it must

Why: A 98 percent level leaves less in the tails than a 95 percent one, so its boundary must sit further out: 2.326 against 1.96. Confidence levels and z-scores move together, always, which makes the ordering of these four numbers — 1.645, 1.96, 2.326, 2.576 — a check on any of them. A higher level with a smaller z is impossible.

OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, p. 412

21. Error analysis: four attempts at the z-score for a 95 percent level

Error analysis

The correct value is 1.96.

Annotate

On: \( \begin{aligned} &(1)\; \text{invNorm}(0.95) = 1.645 \\ &(2)\; \text{invNorm}(0.05) = -1.645 \\ &(3)\; \text{invNorm}(0.025) = -1.96 \\ &(4)\; \text{invNorm}(0.975) = 1.96 \end{aligned} \)

  • (1) forgets to halve alpha, using the level itself as the left area. It gives the 90 percent z-score instead.
  • (2) uses alpha as a left area without halving, and returns a negative boundary.
  • (3) halves alpha correctly but takes the lower tail, giving the negative twin of the right answer.
  • (4) is correct: one minus alpha over two as the area to the left.

Error (1) is the dangerous one, because 1.645 is a perfectly legitimate z-score for a different level, so nothing about it looks wrong. The defence is the ordering check: a 95 percent interval must be wider than a 90 percent one, so its z cannot be the smaller of the two.

22. Level to z-score

Matching

Match each confidence level to its z.

Match the pairs

  • l1. 90 percent
  • l2. 95 percent
  • l3. 98 percent
  • l4. 99 percent
  • r1. 1.645
  • r2. 1.96
  • r3. 2.326
  • r4. 2.576

Why: These four recur constantly from here to the end of the book. Note that they increase with the level, which is the ordering check: a more confident interval always needs a larger multiplier and so comes out wider.

23. One of these is false

Two truths and a lie

All three concern the level and alpha.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Alpha plus the confidence level equals one
  • C. Alpha is split equally between the two tails
  • B. The z-score for a 95 percent level is invNorm(0.95)

Survives elimination: B

Why: The survivor is false and is the section's commonest slip. The left area wanted is one minus alpha over two, which is 0.975 rather than 0.95 — and invNorm(0.95) returns 1.645, the z-score for a 90 percent interval.

24. What does a bigger level do to z?

Prediction

Commit before reasoning.

Predict first

Moving from a 90 percent to a 99 percent confidence level, what happens to the z-score?

  • It rises, from 1.645 to 2.576
  • It falls, since alpha is smaller
  • It is unchanged
  • It depends on the sample size

Correct: It rises, from 1.645 to 2.576.

Why: A higher level leaves less area in the tails, so the boundary must move further out. The second option confuses a smaller alpha with a smaller boundary — they move in opposite directions. The sample size does not enter the z-score at all; it enters the standard error.

25. Constructing the interval

Section

Section 3

26. z times the standard error, added and subtracted

Concept

The error bound is the z-score for the confidence level multiplied by the standard error of the mean. The book stresses that the standard deviation used must be appropriate for the parameter being estimated, so it is sigma over the square root of n rather than sigma itself.

EBM — The error bound for a population mean, equal to z sub alpha over two times sigma over the square root of n. The interval is then the sample mean plus and minus that quantity.

\[ \text{EBM} = z_{\alpha/2}\left(\frac{\sigma}{\sqrt{n}}\right) \]

The whole construction depends on chapter 7. It is because the sample mean is normally distributed with that standard error that a z-score can be used at all — and the book says so, summarising that as a result of the central limit theorem, X-bar is normally distributed, and when sigma is known we use a normal distribution to calculate the error bound. Section 8.2 changes exactly that second clause.

Figure (svg): The five steps for constructing and interpreting a confidence interval

The same five steps carry through sections 8.2 and 8.3 with only the second one changing.

OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 408-411 — the steps and Example 8.2

27. Five steps, and the fifth is not optional

Picture it

The book's own procedure.

Figure (svg): The five steps for constructing and interpreting a confidence interval

The same five steps carry through sections 8.2 and 8.3 with only the second one changing.

The last step carries a template worth using verbatim at first: we estimate with a given percent confidence that the true population mean, described in the words of the problem, is between two values with their units. It forces the level, the parameter and both endpoints into one sentence, which is exactly what a reader needs.

28. Worked example: a 90 percent interval for exam scores

Worked example

Example 8.2, the chapter's central worked case.

\[ \sigma = 3, \; n = 36, \; \bar{x} = 68, \; \text{CL} = 0.90 \]

Find the standard error

Why: Three over root 36.

\[ 0.5 \]

Find the z-score

Why: For a 90 percent level.

\[ 1.645 \]

Compute the error bound

Why: 1.645 times 0.5.

\[ 0.8225 \]

Build the interval

Why: 68 minus and plus it.

\[ (67.18, 68.82) \]

Figure (svg): The solution to Worked example a 90 percent interval for exam scores shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 68 \pm (1.645)\left(\frac{3}{\sqrt{36}}\right) = (67.18, 68.82) \]

Verify: confirm the standard error rather than sigma was used

Why: Using sigma of 3 in place of the standard error of 0.5 would give an error bound of 4.935 and an interval from 63.07 to 72.94 — six times too wide. The book flags this explicitly, saying the standard deviation used must be appropriate for the parameter being estimated. Since the parameter is a mean, the relevant spread is the mean's, which is the standard error.

OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 409-411

29. Build an interval

Faded example

Pizza delivery times have sigma = 6 minutes; a sample of 36 gives a mean of 36 minutes. Use 90 percent.

Fill in the blanks

\text1.645 = (1.645)\left(\frac3.29___}\right) = ___, \quad \text___ = (36 - ___,\; 36 + ___) \text___ ___

Why: The standard error is 6 over 6, which is 1, so the error bound is just the z-score itself, 1.645. The interval runs from 34.355 to 37.645, a width of 3.29 — twice the error bound, as always.

30. Worked example: a 98 percent interval for SAR levels

Worked example

Example 8.3, where the sample mean has to be computed first.

\[ n = 30, \; \sigma = 0.337, \; \text{CL} = 0.98 \]

Compute the point estimate

Why: The mean of the 30 values.

\[ 1.024 \]

Find the z-score

Why: For 98 percent.

\[ 2.326 \]

Compute the error bound

Why: 2.326 times 0.337 over root 30.

\[ 0.1431 \]

Build the interval

Why: 1.024 minus and plus it.

\[ (0.8809, 1.1671) \]

Figure (svg): The solution to Worked example a 98 percent interval for SAR levels shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 1.024 \pm 0.1431 = (0.8809, 1.1671) \]

Verify: confirm where the book's last digits come from

Why: The exact sample mean of these 30 values is 1.023733, and the book rounds it to 1.024 before subtracting — which is what produces its 0.8809 rather than the 0.8806 the unrounded mean gives. The difference is 0.0003 and changes nothing, but it is worth recognising as intermediate rounding rather than an error in either calculation. Carrying full precision until the final step is the better habit.

OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 411-412

31. Trap: using sigma instead of the standard error

Trap

The trap

\[ \text{EBM} = (1.645)(3) = 4.935 \]

Multiply the z-score by the population standard deviation

Why: Sigma is the number the problem supplies.

\[ (63.07, 72.94) \text{ instead of } (67.18, 68.82) \]

That interval describes where individual exam scores fall, not where the population mean lies.

The fix

\[ \text{EBM} = (1.645)\left(\frac{3}{\sqrt{36}}\right) = 0.8225 \]

Divide sigma by the square root of n first

Why: The parameter being estimated is a mean, so the mean's spread applies.

The book raises this in its own words: the standard deviation used must be appropriate for the parameter we are estimating. It is the same distinction chapter 7 spent a whole section on, and it produces the same size of error — a factor of the square root of n, which here is six.

32. Which spread belongs in the formula?

Discrimination

The parameter being estimated decides.

Sort into buckets

Sort each quantity by whether it belongs in the error bound for a mean.

Belongs in the formula
sigma over the square root of n; the standard error of the mean
Does not
sigma alone; sigma divided by n; sigma times the square root of n
yes
This is the standard deviation of the sample mean, which is the parameter being estimated.
no
These describe individual values, or use the wrong power of n altogether.

33. One of these is false

Two truths and a lie

All three concern the construction.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The interval's width is twice the error bound
  • C. The construction relies on the central limit theorem
  • B. The error bound uses the population standard deviation directly

Survives elimination: B

Why: The survivor is false. The error bound uses the STANDARD ERROR, sigma over the square root of n. Using sigma directly gives an interval about the square root of n times too wide, which describes where individual values fall rather than where the mean lies.

34. How wide, roughly?

Estimation

A 95 percent interval is built from sigma = 10 and n = 100.

Predict first

Roughly how wide is the interval?

  • About 3.9
  • About 39
  • About 2
  • About 20

Correct: About 3.9.

Why: The standard error is 10 over 10, which is 1, so the error bound is about 1.96 and the width is twice that, about 3.9. The option 39 uses sigma directly instead of the standard error — a factor of the square root of 100, which is exactly the error this idea's trap describes.

35. What changes the width

Section

Section 4

36. The level and the sample size, in opposite directions

Concept

Increasing the confidence level increases the error bound and makes the interval wider; decreasing it makes the interval narrower. Increasing the sample size causes the error bound to decrease and makes the interval narrower; decreasing the sample size makes it wider.

the width trade-off — Confidence and precision pull against each other. More confidence costs width, and the only way to buy back precision without losing confidence is a larger sample.

\[ \text{EBM} = z_{\alpha/2}\,\frac{\sigma}{\sqrt{n}}: \quad z \uparrow \Rightarrow \text{wider}, \quad n \uparrow \Rightarrow \text{narrower} \]

Both effects are visible in the formula, and they are the only two things a researcher controls. Sigma belongs to the population and cannot be chosen; the level and the sample size can. Because n enters under a square root, buying precision is expensive — halving the width requires quadrupling the sample, which is exactly the diminishing return chapter 7 established.

Figure (svg): Two horizontal intervals of different widths centred on the same sample mean, the ninety-five percent one visibly wider than the ninety percent one

Examples 8.2 and 8.4: identical data, different levels, and the wider interval is the more confident one.

OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 413-415 — Examples 8.4 and 8.5, and the two summaries

37. The same data at two levels

Picture it

Examples 8.2 and 8.4: a 90 percent interval and a 95 percent one.

Figure (svg): Two horizontal intervals of different widths centred on the same sample mean, the ninety-five percent one visibly wider than the ninety percent one

Examples 8.2 and 8.4: identical data, different levels, and the wider interval is the more confident one.

Nothing about the sample changed between these two — same mean, same sigma, same n. Only the level was raised, and the interval widened from 1.645 to 1.96 error bounds. That is the honest cost of confidence, and it explains why a 100 percent interval, which would have to be infinitely wide, would carry no information at all.

38. Worked example: raising the level to 95 percent

Worked example

Example 8.4, changing only the level.

\[ \sigma = 3, \; n = 36, \; \bar{x} = 68, \; \text{CL} = 0.95 \]

Find alpha over two

Why: 0.05 halved.

\[ 0.025 \]

Find the z-score

Why: invNorm at 0.975.

\[ 1.96 \]

Compute the error bound

Why: 1.96 times 0.5.

\[ 0.98 \]

Build the interval

Why: 68 minus and plus it.

\[ (67.02, 68.98) \]

Figure (svg): The solution to Worked example raising the level to 95 percent shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 68 \pm (1.96)(0.5) = (67.02, 68.98) \]

Verify: confirm the wider interval is the more confident one

Why: The 95 percent interval is wider by about 0.32 in total, and the book explains why that makes sense: because the area 0.95 is larger than the area 0.90, and to be more confident that the interval contains the true value it necessarily needs to be wider. An interval that got NARROWER as the level rose would signal that alpha had been halved wrongly.

OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 413-414

39. Wider or narrower?

Sorting

Start from a 90 percent interval built on n = 36.

Sort into buckets

Sort each change by its effect on the interval's width.

Wider
raise the level to 99 percent; decrease the sample to 25
Narrower
increase the sample to 100; lower the level to 80 percent; quadruple the sample size
wide
Either the z-score grows or the standard error does.
narrow
Either the z-score shrinks or a larger sample shrinks the standard error.

The two levers pull in opposite directions and can be traded against each other: raising the level from 90 to 95 percent widens by a factor of 1.96 over 1.645, and that can be paid for with a sample about 42 percent larger. Study design is largely this trade.

40. Worked example: changing the sample size

Worked example

Example 8.5, holding the level at 90 percent.

\[ \sigma = 3, \; \text{CL} = 0.90; \; n = 100 \text{ then } n = 25 \]

At n = 100

Why: 1.645 times 3 over 10.

\[ 0.4935 \]

Compare with n = 36

Why: The original.

\[ 0.8225 \]

At n = 25

Why: 1.645 times 3 over 5.

\[ 0.987 \]

State the pattern

Why: Larger n, smaller bound.

Figure (svg): The solution to Worked example changing the sample size shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ n = 100 \Rightarrow 0.4935, \qquad n = 25 \Rightarrow 0.987 \]

Verify: confirm the error bounds scale with one over root n

Why: Going from n equal to 25 to n equal to 100 quadruples the sample and exactly halves the error bound, from 0.987 to 0.4935 — the square root law from chapter 7, appearing here as the cost of precision. Any pair of error bounds from the same sigma and level must stand in the inverse ratio of the square roots of their sample sizes.

OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 414-415

41. Trap: expecting more confidence to come free

Trap

The trap

\[ \text{raise the level to } 99 \text{ percent and keep the interval} \]

Treat the confidence level as a label on a fixed interval

Why: The data has not changed, so the interval should not either.

\[ (67.18, 68.82) \text{ described as } 99 \text{ percent confident} \]

That interval is only 90 percent confident; calling it 99 percent overstates what the data supports.

The fix

\[ 99 \text{ percent needs } z = 2.576, \text{ giving } (66.71, 69.29) \]

Recompute the error bound whenever the level changes

Why: The level enters through the z-score.

Confidence and width are locked together by the formula, and the only escape is a larger sample. This is why published studies quote the level alongside every interval: an interval without its level is uninterpretable, since the same numbers could be a modest claim at 99 percent or a strong one at 80 percent.

42. The square root cost

Faded example

An error bound of 0.987 comes from n = 25 at a 90 percent level.

Fill in the blanks

\text4 n \text100 ___, \text___ n = ___

Why: Halving the error bound needs the square root of n doubled, so n itself is quadrupled: 25 becomes 100, and the bound falls from 0.987 to 0.4935 — exactly Example 8.5's two values.

43. One of these is false

Two truths and a lie

All three concern width.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. A higher confidence level gives a wider interval
  • C. A larger sample gives a narrower interval
  • B. A wider interval is a better estimate

Survives elimination: B

Why: The survivor is false, or at least badly incomplete. A wider interval is more likely to contain mu but says less about where mu is — in the limit, the interval from minus infinity to infinity is certain and useless. Better means the right trade between confidence and precision, which depends on what the estimate is for.

44. Which lever is cheaper?

Prediction

Commit before reasoning.

Predict first

A researcher wants a narrower interval without losing confidence. What must they do?

  • Collect a larger sample
  • Lower the confidence level
  • Use a larger sigma
  • Nothing can be done

Correct: Collect a larger sample.

Why: The error bound has exactly three inputs: the z-score set by the level, sigma set by the population, and n set by the researcher. Holding the level fixed rules out the first, sigma is not a choice, so only the sample size remains — and because it enters under a square root, the gain is expensive.

45. Working backwards, and sample size

Section

Section 5

46. From an interval to its parts, and from a target to a sample

Concept

A published study may give only the interval. The error bound is then the upper value minus the sample mean, or half the difference between the endpoints; the sample mean is the upper value minus the error bound, or the average of the two endpoints. Solving the error bound formula for n gives the sample size a study needs.

the sample size formula — n equals z sigma over EBM, all squared. Because a sample size must be a whole number, the book instructs that the answer is always rounded UP to the next higher integer to ensure the sample is large enough.

\[ n = \left(\frac{z_{\alpha/2}\,\sigma}{\text{EBM}}\right)^2 \]

The rounding instruction is worth taking literally. Rounding 216.09 down to 216 would give an error bound slightly larger than the two years asked for, which fails the requirement; rounding up to 217 satisfies it with a little to spare. This is one of the few places in the book where ordinary rounding rules are deliberately overridden, and the reason is that the requirement is one-sided.

Figure (svg): A four-column table listing confidence levels of ninety to ninety-nine percent with their alpha values and z-scores

These four z-scores are worth recognising on sight; 1.645 and 1.96 in particular recur in every chapter from here on.

OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 415-416 — working backwards and Example 8.7

47. The four z-scores you will keep needing

Picture it

Each one feeds both the interval and the sample-size formula.

Figure (svg): A four-column table listing confidence levels of ninety to ninety-nine percent with their alpha values and z-scores

These four z-scores are worth recognising on sight; 1.645 and 1.96 in particular recur in every chapter from here on.

Because the z-score is squared in the sample-size formula, moving from 95 to 99 percent multiplies the required sample by 2.576 over 1.96 squared, which is about 1.73. Raising confidence is therefore substantially more expensive at the planning stage than the interval widths alone suggest.

48. Worked example: recovering the parts of an interval

Worked example

Example 8.6, both ways.

\[ \text{the interval } (67.18, 68.82) \]

Half the width

Why: 68.82 minus 67.18, halved.

\[ EBM = 0.82 \]

Or subtract a known mean

Why: 68.82 minus 68.

\[ 0.82\text{ again} \]

Average the endpoints

Why: 67.18 plus 68.82, halved.

\[ \text{mean } = 68 \]

Or subtract the bound

Why: 68.82 minus 0.82.

\[ 68\text{ again} \]

Figure (svg): The solution to Worked example recovering the parts of an interval shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{EBM} = \frac{68.82 - 67.18}{2} = 0.82, \qquad \bar{x} = \frac{67.18 + 68.82}{2} = 68 \]

Verify: confirm the two routes must agree

Why: Both work because the interval is symmetrical: the mean sits exactly at the midpoint and the bound is exactly half the width. The book offers each calculation two ways precisely so that whichever information you happen to have is enough. If the two routes disagreed, the interval would not be symmetrical — which for this chapter's methods would mean an arithmetic error.

OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 415-416

49. Recover the parts

Faded example

A published confidence interval is (42.12, 47.88).

Fill in the blanks

\text2.88 = \frac45___ = ___, \qquad \bar___ = \frac______ = ___

Why: Half the width of 5.76 is 2.88, and the midpoint of the endpoints is 45. Both follow from the interval being symmetrical about its point estimate.

50. Worked example: how many students to survey

Worked example

Example 8.7, planning a study.

\[ \sigma = 15, \; \text{EBM} = 2, \; \text{CL} = 0.95 \]

Find the z-score

Why: For 95 percent.

\[ 1.96 \]

Form the ratio

Why: 1.96 times 15, over 2.

\[ 14.7 \]

Square it

Why: The sample size formula.

\[ 216.09 \]

Round UP

Why: To the next whole number.

\[ 217 \]

Figure (svg): The solution to Worked example how many students to survey shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ n = \left(\frac{(1.96)(15)}{2}\right)^2 = 216.09 \;\to\; 217 \]

Verify: confirm that rounding down would fail the requirement

Why: With n equal to 216 the error bound is 1.96 times 15 over the square root of 216, about 2.0004 years — just over the two years required. With 217 it is about 1.9958, just under. The margin is tiny, but the requirement was to be within two years, so rounding up is what satisfies it. This is why the book insists on always rounding up rather than to the nearest integer.

OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, p. 416

51. Trap: rounding a sample size to the nearest whole number

Trap

The trap

\[ n = 216.09 \;\to\; 216 \]

Round to the nearest integer, as usual

Why: 0.09 is far below a half.

\[ \text{EBM} = 2.0004 > 2 \]

The requirement was to be within two years, and 216 students misses it — narrowly, but it misses.

The fix

\[ n = 216.09 \;\to\; 217 \]

Always round a sample size UP

Why: The book states this as a rule, and the reason is that the requirement is one-sided.

The general point is that not every rounding decision is symmetric. A sample size has to be at least large enough, so any fractional part demands another whole observation. The same reasoning appears wherever a resource must cover a requirement — and here the cost of the extra student is trivial against the cost of missing the target.

52. Plan a sample size

Faded example

Heights have sigma = 3 inches, and we want 95 percent confidence within one inch.

Fill in the blanks

n = \left(\frac34.5735\right)^2 = ___ \;\to\; ___

Why: 5.88 squared is about 34.57, which rounds up to 35 students. Rounding down to 34 would give an error bound slightly over the one inch required.

53. One of these is false

Two truths and a lie

All three concern working backwards.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The sample mean is the midpoint of the interval
  • C. A sample size is always rounded up
  • B. Halving the target error bound doubles the required sample size

Survives elimination: B

Why: The survivor is false. The error bound is squared in the sample-size formula, so halving it QUADRUPLES the required n — the same square root law as everywhere else in these two chapters. Doubling would only reduce the bound by a factor of about 1.41.

54. The cost of more confidence

Estimation

A study needs n = 216 for 95 percent confidence within two years.

Predict first

Roughly what n would 99 percent confidence within the same two years need?

  • About 374
  • About 216
  • About 280
  • About 108

Correct: About 374.

Why: The z-score rises from 1.96 to 2.576, and it is squared in the formula, so n grows by a factor of about 1.73 — from 216 to about 374. Raising confidence is considerably more expensive at the planning stage than the modest widening of an interval suggests.

55. The three inputs to an error bound

Comparison

Fill the blanks. Only one of the three is a researcher's free choice at fixed confidence.

Comparison matrix

InputWhere it comes fromEffect of increasing it
z sub alpha over twothe confidence level chosenwider interval
sigmathe population; not a choicewider interval
nthe researcher's samplenarrower interval
Cost of halving the widthquadruple nbecause n is under a square root

The last row is the practical summary of the whole chapter so far. Precision is bought with sample size, at a square-root rate, and there is no other way to buy it without giving up confidence.

56. Building a confidence interval for a mean, sigma known

Pattern

Six steps, and the last two are the ones that get skipped.

  1. Confirm the population standard deviation is known, so a z-score rather than a t is appropriate.
  2. Compute the sample mean from the data if it is not given; this is the point estimate.
  3. Convert the confidence level to alpha, halve it, and find the z-score with one minus alpha over two to its left.
  4. Compute the error bound as that z-score times sigma over the square root of n.
  5. Build the interval as the point estimate minus and plus the error bound.
  6. Write the interpretation naming the level, the parameter in context, both endpoints and the units.

Check the interval by confirming its midpoint is the sample mean and its half-width is the error bound. Check the z-score against the ordering 1.645, 1.96, 2.326, 2.576.

OpenStax Introductory Business Statistics 2e, §8.1 A Confidence Interval When the Population Standard Deviation Is Known or Large Sample Size §8.1 A Confidence Interval When the Population Standard Deviation Is Known or Large Sample Size

57. Check yourself 1 of 3

Check

The z-score.

Check your understanding

What z-score is used for a 95 percent confidence interval?

  • A. 1.96 (correct)
  • B. 1.645
  • C. 0.95
  • D. 2.576

Answer: A

Why: Alpha is 0.05, alpha over two is 0.025, and the z-score with 0.975 to its left is 1.96.

Why B tempts people
That is the 90 percent value, obtained by using the level itself as the left area instead of one minus alpha over two.
Why C tempts people
That is the confidence level, not a z-score.
Why D tempts people
That is the 99 percent value, which would give too wide an interval.

58. Check yourself 2 of 3

Check

The error bound.

Check your understanding

For sigma = 3, n = 36 and a 90 percent level, what is the error bound?

  • A. 0.8225 (correct)
  • B. 4.935
  • C. 0.1371
  • D. 1.645

Answer: A

Why: The standard error is 3 over 6, which is 0.5, and 1.645 times 0.5 is 0.8225.

Why B tempts people
That multiplies the z-score by sigma directly, ignoring the division by root n.
Why C tempts people
That divides sigma by n rather than by its square root.
Why D tempts people
That is the z-score alone, before multiplying by the standard error.

59. Check yourself 3 of 3

Check

Interpretation.

Check your understanding

A 95 percent confidence interval for a mean is (4.5, 9.5). Which statement is correct?

  • A. 95 percent of intervals built this way contain the true mean (correct)
  • B. There is a 95 percent probability that mu is between 4.5 and 9.5
  • C. 95 percent of the data lies between 4.5 and 9.5
  • D. The true mean is definitely between 4.5 and 9.5

Answer: A

Why: The confidence level describes the procedure across repeated samples. The interval is the random thing; the parameter is fixed.

Why B tempts people
This treats mu as random. Once the sample is taken, mu is either in the interval or it is not.
Why C tempts people
The interval estimates the MEAN, not the spread of individual values, and is far narrower than the data.
Why D tempts people
There is no guarantee; 5 percent of correctly built 95 percent intervals miss.

60. Where this shows up outside the textbook

Real world

A polling firm reports that support for a proposal is 46 percent with a margin of error of 3 percentage points at 95 percent confidence. A commentator writes that since the interval runs from 43 to 49 and is entirely below 50, the proposal is certain to fail. A second commentator says the poll is worthless because the true figure could be anything.

Discussion prompt

Assess both claims, and say what the poll does and does not establish.

Hint: One commentator overstates the certainty and the other understates the information.

Answer:

The first commentator overstates it. The interval from 43 to 49 does lie entirely below 50, which is genuine evidence that support is short of a majority — but a 95 percent interval misses the truth about one time in twenty, and that failure rate is not a remote technicality. Certain is the wrong word for a claim resting on a procedure that is wrong 5 percent of the time.

The second commentator understates it badly. The poll narrows the plausible range to six percentage points out of a hundred, which is a great deal of information. Dismissing an interval because it is not a point is the mirror image of the first error.

\[ 46 \pm 3 \;\Rightarrow\; (43, 49), \qquad \text{width} = 2\,\text{EBM} \]

What the poll establishes is a range and a procedure, not a verdict. The honest reading is that support is probably in the low-to-high forties, that a majority looks unlikely on this evidence, and that the estimate would be sharper with a larger sample — narrowing the margin to 1.5 points would require quadrupling the sample, by the square root law.

Two cautions belong in a careful answer, and both fall outside the arithmetic entirely. The margin of error quantifies only SAMPLING variability: it says nothing about people who were not reachable, who declined to answer, or who answered differently from how they will vote, and those non-sampling errors are frequently larger than the quoted margin. And a single poll is one interval from one sample — the reason polling aggregators combine many is precisely that the confidence level is a statement about a collection of intervals rather than about any one of them.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

What does a 95 percent confidence level actually describe?

  • The probability that mu lies in this particular interval
  • The percent of confidence intervals that contain mu when repeated samples are taken
  • The proportion of the data inside the interval
  • The probability that the sample mean is correct

Correct: The percent of intervals containing mu across repeated samples.

\[ \alpha + \text{CL} = 1, \qquad \text{EBM} = z_{\alpha/2}\,\frac{\sigma}{\sqrt{n}} \]

Why: This is the book's own careful wording, and it is careful for a reason: the interval is a random variable while the population parameter is fixed. Once a sample has been taken, the interval either contains mu or it does not, and no probability attaches to that particular case. The level describes how the procedure behaves, not how this instance turned out.

62. Explain it to someone a year behind you

Explain it

They built a 95 percent interval for a mean using invNorm(0.95), getting 1.645 as the z-score.

Discussion prompt

In two sentences or fewer, locate the error.

Hint: Ask how much area their z-score leaves in each tail.

Answer:

Their 1.645 leaves 0.05 in the upper tail, so the interval between plus and minus it holds 90 percent rather than 95.

Alpha of 0.05 has to be halved before it is used, so the area to the left is 0.975 and the z-score is 1.96.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Stating what a confidence level means without making mu random
  • Getting the z-score right from a stated confidence level
  • Using the standard error rather than sigma in the error bound
  • Working backwards to a sample size, with the rounding rule

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the first, the interval is random and the parameter is fixed. For the second, halve alpha and use one minus alpha over two as the left area. For the third, divide sigma by the square root of n before multiplying by z. For the fourth, square the ratio and always round up. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Twenty minutes.

Draw it

At the top, write the interval's standard form as point estimate minus and plus the error bound, and beneath it the error bound formula with its three inputs labelled by where each comes from. Below that, draw a standard normal curve, shade the central 90 percent, mark 1.645 on both sides, and label the two tails 0.05 each — then write the four-row table of confidence level, alpha, alpha over two and z for 90, 95, 98 and 99 percent. In the middle of the page, work Example 8.2 completely: standard error, z-score, error bound and interval, ending with the full interpreting sentence. Beside it, redo the error bound at 95 percent and draw the two intervals one above the other to scale, so the wider one is visibly wider. Then redo it at n = 100 and n = 25, and write one sentence on which lever each change pulled. Near the bottom, draw twenty short horizontal intervals scattered around one vertical line for mu, marking one or two that miss it, and write beside it the sentence that the interval is random and the parameter is fixed. Finish with Example 8.7's sample size calculation, showing the squaring and the rounding up, and one sentence on why rounding down would fail.

Check the middle section by confirming the midpoint of each interval you draw is exactly 68, and that the 95 percent one is wider than the 90 percent one by the ratio 1.96 over 1.645. Check the sample size by computing the error bound at your answer and confirming it is just UNDER the two years required, not just over.

65. What you can do now

Recap

Five things, and the first is the one examiners actually test.

If you seeThen
A stated confidence levelalpha is one minus it, halved between the tails
A 95 percent intervalThe z-score is 1.96, from invNorm(0.975)
sigma and n givenThe error bound uses sigma over root n
A higher confidence levelA larger z, so a wider interval
A published interval onlyThe mean is its midpoint, the bound its half-width
A target error boundSquare z sigma over EBM, and round UP
A claim that mu is probably in this intervalReword: the procedure captures mu 95 percent of the time

Section 8.2 removes the assumption that makes this section work. In practice the population standard deviation is almost never known, and replacing it with the sample standard deviation turns out to need a different distribution — Gosset's t, developed at the Guinness brewery precisely because small samples made the normal approximation unreliable.

OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution §8.1, pp. 406-416 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §8.1 A Single Population Mean using the Normal Distribution — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 406-416
  2. OpenStax Introductory Business Statistics 2e, §8.1 A Confidence Interval When the Population Standard Deviation Is Known or Large Sample Size — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108