8.3 A Population Proportion

Sections 8.1 and 8.2 estimated a population mean. This one changes the parameter rather than the multiplier: when the data are categorical, the quantity of interest is the proportion of the population in one category, and the underlying distribution is the binomial of section 4.3. The point estimate is the sample proportion, written p prime, the number of successes divided by the number of trials, and its error bound is a z-score times the square root of p prime times q prime over n. No population standard deviation appears anywhere, because a binomial's variability is determined once its probability is. The section also introduces the plus-four adjustment, which adds two successes and two failures to the observed counts before proceeding, and gives more accurate intervals when the confidence level is at least ninety percent and the sample at least ten. Solving the error bound for n gives a planning formula, and using one half for the unknown proportion makes it as large as it can be.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 8.3 A Population Proportion

Title

Statistics · Chapter 8 — Confidence Intervals

A Population Proportion

2. By the end of this lesson you can

Objectives

Five outcomes, and the first is the recognition test the book supplies.

OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, pp. 420-426 — the section these objectives are drawn from

3. What you already have

Warm-up

Sections 8.1 and 8.2 built intervals for a mean from a point estimate and a standard error.

Discussion prompt

Of 500 adults surveyed, 421 own a smartphone. What is the point estimate for the population proportion, and where would its standard error come from — there is no sigma anywhere in the problem.

Hint: Section 4.3 gave the binomial its own standard deviation.

Answer:

The point estimate is 421 over 500, which is 0.842 — the sample proportion, and the natural analogue of the sample mean.

For the standard error there is no separate spread to estimate, and none is needed. A binomial's variability is fixed once its probability is, so the spread of a sample proportion is the square root of p times q over n — and substituting the estimate p prime for the unknown p gives something computable from the one number already found.

That is the whole structural difference from the previous two sections. Estimating a mean needed two quantities, a centre and a spread; estimating a proportion needs only one, because the spread comes free with the centre.

4. One estimate does the work of two

Concept

To form a proportion, take X, the random variable for the number of successes, and divide it by n, the number of trials. That random variable, read p prime, is the sample proportion and the point estimate of the population proportion. When n is large and p is not close to zero or one, the normal distribution approximates the binomial, and the interval takes the familiar form.

p prime — The sample proportion, x over n. Sometimes written p hat. It is the point estimate for the population proportion p, and it also determines the standard error, since q prime is simply one minus it.

\[ p' = \frac{x}{n}, \qquad \text{EBP} = z_{\alpha/2}\sqrt{\frac{p'q'}{n}} \]

The normal approximation is the same one section 4.6 raised for the Poisson and section 6.2 will be relied on throughout: when n is large and p is not close to zero or one, the binomial's shape is close enough to normal for a z-score to apply. That condition matters — a proportion very near zero or one from a small sample produces an interval that can run outside the range from zero to one, which is one of the reasons the plus-four adjustment exists.

Figure (svg): A card defining the sample proportion as the number of successes over the number of trials, with the error bound formula

Example 8.10: 421 of 500 gives p prime of 0.842, and everything else follows from that one number.

OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, pp. 420-421

5. Recognising a proportion problem

Section

Section 1

6. Categorical data, and no mention of a mean

Concept

The book gives a two-part test. First, the underlying distribution is a binomial distribution — the data record successes among trials. Second, there is no mention of a mean or average anywhere in the problem.

the recognition test — Binomial structure plus the absence of any average. If the question asks what fraction, what percent, or how many out of, the parameter is a proportion rather than a mean.

\[ X \sim B(n, p) \;\Longrightarrow\; p' = \frac{X}{n} \]

The examples the book opens with are all polling and market research: the proportion of voters favouring a candidate, the proportion of stocks that rise in a week, the proportion of households owning a personal computer. All are counts of yes-or-no outcomes, and none has a natural average — asking for the average vote would not mean anything.

Figure (svg): Two columns contrasting a problem about a mean with a problem about a proportion

The book's own test: the underlying distribution is a binomial distribution, and there is no mention of a mean or average.

OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, p. 420 — the recognition test and the opening examples

7. Two kinds of question

Picture it

What distinguishes a proportion problem from a mean problem.

Figure (svg): Two columns contrasting a problem about a mean with a problem about a proportion

The book's own test: the underlying distribution is a binomial distribution, and there is no mention of a mean or average.

The fourth row is the structural difference and the one worth remembering. For a mean the spread is a separate quantity that has to be given or estimated; for a proportion it is determined by the estimate itself, which is why these problems supply no sigma and need none.

8. Worked example: identifying the parameter

Worked example

Example 8.10's setup.

\[ \text{of } 500 \text{ adults surveyed, } 421 \text{ own smartphones} \]

Check the data type

Why: Owns one, or does not.

Check for a mean

Why: None mentioned.

Name the model

Why: Successes among trials.

Form the point estimate

Why: 421 over 500.

\[ p' = 0.842 \]

Figure (svg): The solution to Worked example identifying the parameter shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ X \sim B(500, p), \qquad p' = \frac{421}{500} = 0.842 \]

Verify: confirm the answer will be a fraction rather than a quantity

Why: A proportion must lie between zero and one, so an answer outside that range would be impossible on its face — unlike a mean, which can be any number the data's units allow. That bound is a genuine check available at the end of every problem in this section, and it also warns when the normal approximation is straining.

OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, pp. 420-422

9. Mean or proportion?

Sorting

Ask what kind of data the question is about.

Sort into buckets

Sort each question.

A proportion
what fraction of adults own a smartphone; the percent of students registered to vote; how many out of 50 would consider an electric vehicle
A mean
the average sensory rate of 15 subjects; the mean number of chemicals in cord blood
prop
The data are yes-or-no categories and the question asks for a fraction or percent.
mean
The data are measurements or counts per unit, and the question names an average.

Item (d) is the one worth pausing on: it counts chemicals, so it looks categorical, but each infant contributes a NUMBER rather than a yes or no — and the question asks for a mean of those numbers. Counting is not the same as classifying.

10. Worked example: a percentage question

Worked example

Example 8.11, where the question asks for a percent.

\[ \text{of } 500 \text{ students surveyed, } 300 \text{ are registered voters} \]

Read what is asked

Why: The percent registered.

Form p prime

Why: Three hundred over 500.

\[ 0.600 \]

Note q prime

Why: One minus it.

\[ 0.400 \]

Note what is absent

Why: No sigma, no s.

Figure (svg): The solution to Worked example a percentage question shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ p' = \frac{300}{500} = 0.600, \qquad q' = 0.400 \]

Verify: confirm percent and proportion are the same parameter

Why: A question asking for a percent and one asking for a proportion are identical apart from a factor of a hundred, and the book reports its answer both ways — between 56.4 percent and 63.6 percent, or the interval (0.564, 0.636). Working in proportions throughout and converting only in the final sentence avoids the error of putting a percentage into the formula, where it would inflate the error bound tenfold.

OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, p. 423

11. Trap: putting a percentage into the formula

Trap

The trap

\[ p' = 60, \quad q' = 40, \quad \sqrt{\frac{(60)(40)}{500}} \approx 2.19 \]

Use the percentages as they are stated

Why: The question is phrased in percents.

\[ \text{an error bound of over } 3, \text{ for a quantity bounded by } 1 \]

The formula is written in proportions, so 60 percent has to enter as 0.60.

The fix

\[ p' = 0.60, \quad q' = 0.40, \quad \sqrt{\frac{(0.60)(0.40)}{500}} \approx 0.0219 \]

Convert every percentage to a proportion before substituting

Why: Divide by a hundred first.

The error is caught instantly by the range check: an error bound of 2.19 for a proportion is impossible, since the whole parameter lies between zero and one. Converting to percentages only in the final interpreting sentence keeps the formulas clean and makes this mistake hard to make.

12. The point estimate

Fill the middle

How the sample proportion is formed.

Fill in the blanks

p' = \fracsuccesses___, \text___ x \text___ ___

Why: Successes — the count in the category of interest, divided by the number of trials. The label success is neutral: in Example 8.12 a success is a student who smoked.

13. One of these is false

Two truths and a lie

All three concern recognising the problem type.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The underlying distribution is binomial
  • C. No mention of a mean or average appears
  • B. A population standard deviation must be given

Survives elimination: B

Why: The survivor is false, and its falsity is the section's structural point. A binomial's spread is determined by its probability, so the standard error is built from p prime itself — no sigma is given because none is needed.

14. What must the answer look like?

Prediction

Commit before reasoning.

Predict first

What range must a confidence interval for a proportion lie in?

  • Between 0 and 1
  • Between minus 1 and 1
  • Any range, depending on the data
  • Between 0 and n

Correct: Between 0 and 1.

Why: A proportion is a fraction of a whole, so it cannot be negative or exceed one. That gives a check unavailable for means: an endpoint outside the range signals either a percentage entered as a whole number or a normal approximation straining on a small sample near zero or one.

15. Building the interval

Section

Section 2

16. The spread comes from the estimate itself

Concept

The error bound for a proportion is the z-score for the confidence level multiplied by the square root of p prime times q prime over n. The z-score is found exactly as in section 8.1; only the standard error is different.

EBP — The error bound for a population proportion. Unlike the error bound for a mean, it needs no separate estimate of spread, because a binomial's variability follows from its probability.

\[ \left(p' - \text{EBP},\; p' + \text{EBP}\right), \qquad \text{EBP} = z_{\alpha/2}\sqrt{\frac{p'q'}{n}} \]

It is worth noticing what has NOT changed. The z-score comes from the confidence level exactly as before, the interval is symmetric about its point estimate exactly as before, and the interpretation template is the same. The book says the procedure is similar to that for the population mean but the formulas are different, and the difference is confined to a single square root.

Figure (svg): The procedure for a confidence interval for a proportion

Only steps two and four differ from section 8.1; the z-score and the interpretation template are unchanged.

OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, pp. 421-423 — the formula and Examples 8.10 and 8.11

17. Five steps, two of them new

Picture it

The section 8.1 procedure with a proportion's point estimate and standard error.

Figure (svg): The procedure for a confidence interval for a proportion

Only steps two and four differ from section 8.1; the z-score and the interpretation template are unchanged.

The note at the bottom matters for the next idea: when the plus-four adjustment is used, the counts are changed before the procedure starts, and nothing about the procedure itself is altered. That is the whole appeal of the method — it costs no new formula.

18. Worked example: smartphone ownership

Worked example

Example 8.10, at a 95 percent level.

\[ x = 421, \; n = 500, \; \text{CL} = 0.95 \]

Form the point estimate

Why: 421 over 500.

\[ 0.842 \]

Form q prime

Why: One minus it.

\[ 0.158 \]

Find the z-score

Why: For 95 percent.

\[ 1.96 \]

Compute the error bound

Why: 1.96 times root of 0.842 times 0.158 over 500.

\[ 0.032 \]

Figure (svg): The solution to Worked example smartphone ownership shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 0.842 \pm 1.96\sqrt{\frac{(0.842)(0.158)}{500}} = (0.810, 0.874) \]

Verify: confirm the interval lies inside the range a proportion can take

Why: Both endpoints are between zero and one, as they must be, and the interval is roughly six percentage points wide on a sample of 500 — which is the familiar order of magnitude for a poll of that size. An error bound of about 0.03 at n equal to 500 is worth carrying as a benchmark, since it is close to the margin of error most published polls report.

OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, p. 422

19. Build the interval

Faded example

Of 250 people surveyed, 98 own a tablet. Use 95 percent.

Fill in the blanks

p' = \frac0.3920.061 = ___, \quad \text___ = 1.96\sqrt______} \approx ___

Why: The point estimate is 0.392, and the error bound about 0.061 — so the interval runs from roughly 0.331 to 0.453. The bound is larger than Example 8.10's because the sample is half the size.

20. Worked example: registered voters at 90 percent

Worked example

Example 8.11, with a different level and a proportion near one half.

\[ x = 300, \; n = 500, \; \text{CL} = 0.90 \]

Form the point estimate

Why: Three hundred over 500.

\[ 0.600 \]

Find the z-score

Why: For 90 percent.

\[ 1.645 \]

Compute the error bound

Why: 1.645 times root of 0.24 over 500.

\[ 0.036 \]

Convert to percent

Why: For the interpretation.

\[ 56.4\text{ to } 63.6 \]

Figure (svg): The solution to Worked example registered voters at 90 percent shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 0.600 \pm 1.645\sqrt{\frac{(0.6)(0.4)}{500}} = (0.564, 0.636) \]

Verify: compare the error bound with Example 8.10's, and explain the difference

Why: This bound is 0.036 against 0.032 for the smartphone example, even though the level here is LOWER and the sample the same size. The reason is the product p prime q prime: 0.6 times 0.4 is 0.24, against 0.842 times 0.158 which is 0.133 — so a proportion near one half is intrinsically harder to pin down, and that is exactly what the next idea's planning formula exploits.

OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, p. 423

21. Error analysis: four attempts at the error bound for 300 of 500 at 90 percent

Error analysis

The correct value is about 0.036.

Annotate

On: \( \begin{aligned} &(1)\; 1.645\sqrt{\tfrac{(60)(40)}{500}} \approx 3.60 \\ &(2)\; 1.645\,\tfrac{(0.6)(0.4)}{500} \approx 0.0008 \\ &(3)\; 1.96\sqrt{\tfrac{(0.6)(0.4)}{500}} \approx 0.043 \\ &(4)\; 1.645\sqrt{\tfrac{(0.6)(0.4)}{500}} \approx 0.036 \end{aligned} \)

  • (1) enters percentages rather than proportions, giving a bound of 3.6 for a quantity bounded by 1.
  • (2) omits the square root, giving a bound about forty times too small.
  • (3) uses the 95 percent z on a 90 percent problem, giving an interval that is too wide.
  • (4) is correct: 1.645 times the square root of p prime q prime over n.

Errors (1) and (2) are both caught by the range check, since one exceeds 1 and the other implies an implausibly precise interval. Error (3) produces a perfectly plausible number and is caught only by checking the level against the z-score.

22. One of these is false

Two truths and a lie

All three concern the formula.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The z-score is found exactly as in section 8.1
  • C. The standard error uses p prime and q prime
  • B. A t distribution is used when n is small

Survives elimination: B

Why: The survivor is false. The t distribution exists to handle an estimated standard deviation for a MEAN; a proportion's spread comes from p prime itself, so no t is involved. Small samples are handled instead by the plus-four adjustment, which is the next idea.

23. A benchmark margin

Estimation

A poll of about 1,000 people reports a proportion near one half.

Predict first

Roughly what is its 95 percent margin of error?

  • About 3 percentage points
  • About 1 point
  • About 10 points
  • About 0.3 points

Correct: About 3 percentage points.

Why: 1.96 times the square root of 0.25 over 1,000 is about 0.031. That is why so many published polls quote a margin of about three points: a sample near a thousand at a proportion near one half is the industry's standard configuration, and this calculation is where the number comes from.

24. Which proportion is hardest to pin down?

Prediction

Commit before reasoning.

Predict first

For a fixed sample size, which sample proportion gives the widest interval?

  • 0.5
  • 0.1
  • 0.9
  • They are all the same

Correct: 0.5.

Why: The error bound depends on the product p prime q prime, which is largest at one half: 0.25 there against 0.09 at either 0.1 or 0.9. A proportion near the extremes is intrinsically easier to estimate, and the next idea turns that fact into a conservative planning rule.

25. The plus-four adjustment

Section

Section 3

26. Pretend four more observations

Concept

Because we do not know the true proportion, we are forced to use point estimates to calculate the standard deviation of the sampling distribution, and studies have shown that the resulting estimate can be flawed. The adjustment is to pretend we have four additional observations: two successes and two failures. The new sample size is n plus four and the new count of successes is x plus two.

the plus-four method — Adding two successes and two failures to the observed counts before computing the interval. It should be used when the confidence level desired is at least 90 percent and the sample size is at least ten.

\[ x \to x + 2, \qquad n \to n + 4 \]

What the adjustment does is pull the point estimate toward one half and widen the interval slightly, and both effects are corrections in the right direction. The ordinary interval is known to be too narrow on average for small samples, and to behave particularly badly when the observed proportion is near zero or one — where it can even produce endpoints outside the range a proportion can occupy.

Figure (svg): Two horizontal intervals for the same data, the plus-four one shifted slightly toward the centre and a little wider

Computer studies have demonstrated the effectiveness of this method, and the calculation is otherwise unchanged.

OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, pp. 424-425 — the plus-four method and Examples 8.12 and 8.13

27. The same data, adjusted and not

Picture it

Example 8.12: six of twenty-five statistics students reported smoking.

Figure (svg): Two horizontal intervals for the same data, the plus-four one shifted slightly toward the centre and a little wider

Computer studies have demonstrated the effectiveness of this method, and the calculation is otherwise unchanged.

The adjusted interval is a little wider and sits a little higher, because 8 over 29 is 0.276 against 6 over 25 at 0.240. On a sample of twenty-five that shift is meaningful; on a sample of a thousand the same four observations would move nothing, which is why the method is aimed at small samples.

28. Worked example: smoking among statistics students

Worked example

Example 8.12, using the plus-four method.

\[ x = 6, \; n = 25, \; \text{CL} = 0.95 \]

Adjust the counts

Why: Add two and four.

\[ x = 8, n = 29 \]

Form the point estimate

Why: Eight over 29.

\[ 0.276 \]

Find the z-score

Why: For 95 percent.

\[ 1.96 \]

Compute the error bound

Why: From the adjusted numbers.

\[ 0.163 \]

Figure (svg): The solution to Worked example smoking among statistics students shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 0.276 \pm 0.163 = (0.113, 0.439) \]

Verify: confirm the adjustment moved the estimate in the expected direction

Why: The observed proportion was 6 over 25, which is 0.240, and the adjusted one is 0.276 — pulled toward one half, as adding one success and one failure in equal measure must do. The interval is also slightly wider than the unadjusted one. Both changes are corrections toward honesty about how little twenty-five observations pin down.

OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, pp. 424-425

29. Apply the adjustment

Faded example

Of 65 first-year students, 31 have declared a major. Use the plus-four method.

Fill in the blanks

x = 31 + 2 = 33, \quad n = 65 + 4 = 69, \quad p' = \frac______ \approx 0.478

Why: Thirty-three successes out of sixty-nine trials gives about 0.478, barely moved from the observed 0.477 because that proportion was already close to one half — the adjustment has least effect exactly there.

30. Worked example: electric vehicle interest

Worked example

Example 8.13, at a 90 percent level.

\[ x = 13, \; n = 50, \; \text{CL} = 0.90 \]

Adjust the counts

Why: Fifteen out of 54.

\[ x = 15, n = 54 \]

Form the point estimate

Why: Fifteen over 54.

\[ 0.278 \]

Find the z-score

Why: For 90 percent.

\[ 1.645 \]

Compute the error bound

Why: From the adjusted numbers.

\[ 0.100 \]

Figure (svg): The solution to Worked example electric vehicle interest shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 0.278 \pm 0.100 = (0.178, 0.378) \]

Verify: confirm the conditions for using the method are met

Why: The confidence level is 90 percent, which meets the at-least-90 requirement, and the sample of 50 comfortably exceeds ten. Both conditions hold, so the adjustment is appropriate. The book's own follow-up asks the same question of a sample of 588, where the four extra observations barely move anything — which is the honest way to see that the method is a small-sample correction.

OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, p. 425

31. Trap: adjusting the counts but not the sample size

Trap

The trap

\[ x = 6 + 2 = 8, \quad n = 25 \]

Add the two successes and leave n alone

Why: Only the successes were mentioned as changing.

\[ p' = \frac{8}{25} = 0.32, \text{ instead of } 0.276 \]

Four observations were pretended, not two, so the denominator must grow by four.

The fix

\[ x = 8, \quad n = 29, \quad p' = \frac{8}{29} \approx 0.276 \]

Add two successes AND two failures, so n grows by four

Why: The book's reminder: the method assumes an additional four trials.

The check is that the adjustment should always pull the estimate TOWARD one half, never past it or away from it. Here 0.240 moves to 0.276, which is toward 0.5; the erroneous 0.32 overshoots because the two pretend failures were never counted.

32. One of these is false

Two truths and a lie

All three concern the plus-four method.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Two successes and two failures are added
  • C. It is used when the level is at least 90 percent and n at least ten
  • B. The calculation changes once the counts are adjusted

Survives elimination: B

Why: The survivor is false, and the book's reminder says so: you do not need to change the process for calculating the confidence interval; simply update the values of x and n. The whole appeal of the method is that it introduces no new formula.

33. Which way does it move the estimate?

Prediction

Commit before reasoning.

Predict first

An observed proportion is 0.10. What does the plus-four adjustment do to it?

  • Raises it, pulling it toward one half
  • Lowers it further
  • Leaves it unchanged
  • It depends on the confidence level

Correct: Raises it, toward one half.

Why: Adding one pretend success and one pretend failure in equal numbers always drags a proportion toward the middle, and the effect is largest when the observed proportion is extreme. That is precisely where the ordinary interval behaves worst, which is why the correction is aimed there.

34. When does it stop mattering?

Estimation

The same four extra observations, on two sample sizes.

Predict first

For which sample does the plus-four adjustment change the estimate more?

  • n = 25
  • n = 588
  • Equally for both
  • Neither is affected

Correct: n = 25.

Why: Four extra observations are a sixth of a sample of 25 and under a percent of a sample of 588, so the shift is large in the first case and negligible in the second. The book makes exactly this comparison in its follow-up to Example 8.13, and it is the clearest way to see that the method is a small-sample correction.

35. Planning the sample size

Section

Section 4

36. Use one half when the proportion is unknown

Concept

Solving the error bound formula for n gives the sample size a study needs. But n depends on p prime, which is not yet known — so the book substitutes one half for it, because that makes the product p prime q prime as large as it can be and therefore gives the largest, safest sample size.

the conservative substitution — Setting p prime and q prime both to 0.5, giving a product of 0.25. Since 0.25 is the largest possible value of that product, the resulting sample size is large enough whatever the true proportion turns out to be.

\[ n = \frac{z^2\,p'q'}{\text{EBP}^2}, \qquad p'q' \le 0.25 \]

The book checks the claim by trial rather than by algebra: 0.6 times 0.4 is 0.24, 0.3 times 0.7 is 0.21, 0.2 times 0.8 is 0.16, and so on — all below the 0.25 that one half gives. The consequence is that a study planned this way is never under-powered, though it may be somewhat larger than strictly necessary if the true proportion turns out to be extreme.

Figure (svg): A downward-opening curve peaking at nought point two five when the proportion is one half

The book checks it by trial: 0.6 times 0.4 is 0.24, 0.3 times 0.7 is 0.21, 0.2 times 0.8 is 0.16 — all below 0.25.

OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, p. 426 — the sample size formula and Example 8.14

37. Why one half is the safe choice

Picture it

The product of a proportion and its complement, across the whole range.

Figure (svg): A downward-opening curve peaking at nought point two five when the proportion is one half

The book checks it by trial: 0.6 times 0.4 is 0.24, 0.3 times 0.7 is 0.21, 0.2 times 0.8 is 0.16 — all below 0.25.

The curve is flat near its peak, which has a practical consequence worth noticing: any proportion between about 0.3 and 0.7 gives a product within about fifteen percent of the maximum, so the conservative choice costs little unless the true proportion is genuinely extreme.

38. Worked example: how many customers to survey

Worked example

Example 8.14, planning a study.

\[ \text{EBP} = 0.03, \; \text{CL} = 0.90 \]

Find the z-score

Why: For 90 percent.

\[ 1.645 \]

Choose p prime

Why: Unknown, so use one half.

\[ p' q' = 0.25 \]

Apply the formula

Why: z squared times 0.25 over EBP squared.

\[ 751.67 \]

Round UP

Why: As for any sample size.

\[ 752 \]

Figure (svg): The solution to Worked example how many customers to survey shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ n = \frac{(1.645)^2(0.25)}{(0.03)^2} \approx 751.7 \;\to\; 752 \]

Verify: confirm the choice of one half really is the safe one

Why: Had the true proportion turned out to be 0.2, the product would be 0.16 and the required sample only about 481 — so planning for 752 over-samples by more than half. But had a researcher optimistically assumed 0.2 and the truth been 0.5, the study would have missed its target margin. Over-sampling costs money; under-sampling costs the answer, which is why the conservative choice is standard.

OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, p. 426

39. Plan a sample

Faded example

A 95 percent interval within two percentage points, with the proportion unknown.

Fill in the blanks

n = \frac24012401 \approx ___ \;\to\; ___

Why: About 2,401 respondents. Tightening the margin from three points to two, at a higher confidence level, has more than tripled the sample from Example 8.14's 752.

40. Worked example: the cost of a tighter margin

Worked example

The same study at two target margins.

\[ \text{EBP} = 0.03 \text{ against } 0.015, \; \text{CL} = 0.90 \]

At a 3 point margin

Why: The original.

\[ 752 \]

Halve the margin

Why: To 1.5 points.

\[ EBP = 0.015 \]

Note the squaring

Why: EBP is squared below.

Compute

Why: Four times 752.

\[ \text{about } 3, 007 \]

Figure (svg): The solution to Worked example the cost of a tighter margin shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ n \propto \frac{1}{\text{EBP}^2} \]

Verify: confirm this matches section 8.1's rule for means

Why: Section 8.1's sample-size formula also squared the ratio, so halving the error bound quadrupled n there too — and both trace back to chapter 7's standard error falling like one over the square root of n. The rule is the same whatever the parameter: precision costs sample size at a square rate, which is why very tight margins are rarely attempted.

OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, p. 426

41. Trap: guessing a small proportion to save sample size

Trap

The trap

\[ \text{expect about } 20 \text{ percent, so use } p'q' = 0.16 \]

Use an optimistic guess to reduce the required sample

Why: A smaller product means a smaller n.

\[ n \approx 481, \text{ but if the truth is } 0.5 \text{ the margin is missed} \]

The margin achieved would be about 0.0375 rather than the 0.03 promised.

The fix

\[ \text{use } p'q' = 0.25, \text{ giving } n = 752 \]

Take the largest possible product, so the target holds whatever the truth is

Why: One half maximises it.

A prior estimate is legitimate when it is genuinely well founded — from an earlier survey of the same population, say — and then a smaller sample is defensible. What is not defensible is guessing optimistically to reduce cost, because the guess is precisely what the study was commissioned to establish.

42. One of these is false

Two truths and a lie

All three concern planning.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. One half maximises the product p prime q prime
  • C. A sample size is always rounded up
  • B. Halving the target margin doubles the sample size

Survives elimination: B

Why: The survivor is false: the error bound is squared in the denominator, so halving it QUADRUPLES the sample. From 752 to about 3,007 for Example 8.14's study, which is the same square-rate cost that governs every sample-size calculation in the book.

43. Where does the poll's margin come from?

Estimation

A national poll of 1,000 quotes a margin of about 3 points at 95 percent.

Predict first

What proportion was assumed in that calculation?

  • One half, the conservative choice
  • The observed proportion
  • Zero
  • It cannot be determined

Correct: One half.

Why: 1.96 times the square root of 0.25 over 1,000 gives 0.031 — about three points. Published polls almost always quote the conservative margin, which is why the same figure appears whatever the poll actually found. A poll reporting 20 percent support has a genuinely smaller margin than the headline number suggests.

44. Justify the conservative choice

Explain it

A classmate asks why anyone would plan for the widest possible interval rather than a realistic one.

Discussion prompt

In two sentences or fewer, justify it.

Hint: Ask what the study is being run to find out.

Answer:

The proportion is exactly what the study is commissioned to discover, so any assumption about it made beforehand could be wrong in the direction that matters.

Planning at one half guarantees the promised margin whatever the truth turns out to be, and over-sampling costs money while under-sampling costs the answer.

45. Reading the result

Section

Section 5

46. A percentage of the population, stated in context

Concept

The interpretation follows the same template as for a mean, converted to percentages where the question was asked that way. The book gives an alternate wording for proportions: we estimate with a stated confidence that between one percentage and another of all members of the population are in the category.

interpreting a proportion interval — Naming the level, the population, the category and both endpoints. The book's own sentences say all students or all adult residents of this city, which fixes exactly what population is being described.

\[ (0.564, 0.636) \;\longrightarrow\; \text{“between 56.4 and 63.6 percent of ALL students”} \]

The words all students matter more here than they might seem to. A confidence interval extrapolates from the surveyed group to the whole population, and that extrapolation is only licensed by random sampling — which is a stronger requirement for polls than for measurements, because non-response tends to correlate with the very opinions being measured.

Figure (svg): A number line showing a sample proportion of nought point eight four two with its error bound either side

We estimate with 95 percent confidence that between 81 percent and 87.4 percent of all adult residents of this city have smartphones.

OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, pp. 422-423 — the interpretations and the alternate wording

47. The smartphone interval

Picture it

Example 8.10's result on a number line.

Figure (svg): A number line showing a sample proportion of nought point eight four two with its error bound either side

We estimate with 95 percent confidence that between 81 percent and 87.4 percent of all adult residents of this city have smartphones.

Stated in full: we estimate with 95 percent confidence that between 81 percent and 87.4 percent of all adult residents of this city have smartphones. The phrase all adult residents of this city is doing real work, since the survey covered 500 of them and the claim is about every one.

48. Worked example: two correct wordings

Worked example

Example 8.11 gives both.

\[ (0.564, 0.636) \text{ at } 90 \text{ percent} \]

The parameter wording

Why: About the percent itself.

The alternate wording

Why: About the students.

\[ \text{between } 56.4\text{ and } 63.6 \%\text{ of } ALL\text{ students} \]

Name the level

Why: Ninety percent.

Name the population

Why: All students at the university.

Figure (svg): The solution to Worked example two correct wordings shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ (0.564, 0.636) \equiv (56.4\%, 63.6\%) \]

Verify: confirm the two wordings say the same thing

Why: They differ only in whether the percentage attaches to the parameter or to the population, and both name the level, the group and both endpoints. The book offers the second because it reads more naturally in a report, and the important thing is that neither omits the population — an interval quoted without saying who it describes is not interpretable.

OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, p. 423

49. One of these is false

Two truths and a lie

All three concern interpretation.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The interval describes the population, not the sample
  • C. The level means the same thing as it did for a mean
  • B. The interval says 95 percent of people fall in that range

Survives elimination: B

Why: The survivor is false and confuses an interval for a parameter with a range for individuals. Each person either owns a smartphone or does not; there is no sense in which 95 percent of people fall between 0.810 and 0.874. The interval locates the population's overall proportion.

50. Worked example: what the level describes

Worked example

The book's explanation, in the same terms as section 8.1.

\[ \text{a } 95 \text{ percent interval for a proportion} \]

Identify what is random

Why: The interval.

Identify what is fixed

Why: The population proportion.

State the level's meaning

Why: Across repeated samples.

\[ 95 \%\text{ contain } p \]

Note what it excludes

Why: A claim about this interval.

Figure (svg): The solution to Worked example what the level describes shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{the interval varies; } p \text{ does not} \]

Verify: confirm this matches section 8.1's account exactly

Why: The book's explanation for proportions is word-for-word parallel to its explanation for means, which is worth noticing: the meaning of a confidence level does not depend on what parameter is being estimated. Learning it once in section 8.1 was enough, and every later chapter reuses the same idea without restating it.

OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, pp. 422-423

51. Trap: describing the sample instead of the population

Trap

The trap

\[ \text{between } 81 \text{ and } 87.4 \text{ percent of the } 500 \text{ surveyed own smartphones} \]

Attach the interval to the people actually asked

Why: They are the ones the data came from.

\[ \text{but } 421 \text{ of } 500 \text{ is exactly } 84.2 \text{ percent, with no uncertainty} \]

The sample's own proportion is known precisely; the interval exists to describe the population it was drawn from.

The fix

\[ \text{between } 81 \text{ and } 87.4 \text{ percent of ALL adult residents of this city} \]

Attach the interval to the population, and say which population

Why: That extrapolation is the whole purpose of the interval.

This error empties the interval of content, since there is nothing uncertain about the sample. It is worth catching because the corrected sentence also forces the question of WHICH population — all adults in the city, or all adults reachable by this survey method, and those may not be the same thing.

52. Convert to percentages

Faded example

An interval for a proportion is (0.178, 0.378).

Fill in the blanks

\text17.8 37.8 \text___ ___ \text___

Why: Between 17.8 and 37.8 percent — Example 8.13's result for adults aged 18 to 29 who would consider an electric vehicle. Converting at the end keeps percentages out of the formula, where they would inflate the error bound.

53. Which sentence is right?

Discrimination

For the interval (0.810, 0.874) from a survey of 500 city residents.

Sort into buckets

Sort each statement.

A correct interpretation
between 81 and 87.4 percent of all adult residents own smartphones; the true proportion for the city is between 0.810 and 0.874; we are 95 percent confident the city's proportion lies in this range
Wrong
between 81 and 87.4 percent of the 500 surveyed own smartphones; 95 percent of residents own between 0.810 and 0.874 smartphones
ok
The statement is about the population's proportion, with the level attached.
no
Either it describes the sample, whose proportion is known exactly, or it treats a proportion as a count per person.

54. What licenses the extrapolation?

Prediction

Commit before reasoning.

Predict first

What allows an interval from 500 people to describe all adult residents of a city?

  • Random sampling from that population
  • The large sample size
  • The normal approximation
  • The confidence level being high

Correct: Random sampling.

Why: Only random selection makes the sample representative of the population, and nothing else substitutes for it. A large non-random sample, a good normal approximation and a 99 percent level would all still describe whoever happened to be reachable — which is section 8.2's point about the two separate assumptions, and it applies just as sharply here.

55. The three intervals of chapter 8

Comparison

Fill the blanks. One shape, three sets of parts.

Comparison matrix

Point estimateError bound
Mean, sigma knownx-barz times sigma over root n
Mean, sigma unknownx-bart times s over root n
Proportionp primez times root of p'q' over n
What is estimated separatelythe spread, for both meansnothing: p' gives both

The last row is the structural difference worth carrying out of the chapter. Estimating a mean requires a centre and a spread; estimating a proportion requires only the one number, because a binomial's variability follows from its probability.

56. Building an interval for a proportion, in order

Pattern

Six steps, and the first is the recognition test.

  1. Confirm the data are categorical and no mean is mentioned, so the parameter is a proportion.
  2. If the plus-four method applies, add two to the successes and four to the sample size before anything else.
  3. Compute p prime as successes over trials, and q prime as one minus it, as PROPORTIONS rather than percentages.
  4. Find the z-score for the confidence level, exactly as in section 8.1.
  5. Compute the error bound as z times the square root of p prime q prime over n, and build the interval.
  6. Check both endpoints lie between zero and one, then interpret as a percentage of the named population.

For planning, solve for n with p prime q prime set to 0.25, and round up. Halving the target margin quadruples the sample.

OpenStax Introductory Business Statistics 2e, §8.3 A Confidence Interval for A Population Proportion §8.3 A Confidence Interval for A Population Proportion

57. Check yourself 1 of 3

Check

The point estimate.

Check your understanding

Of 250 people surveyed, 98 own a tablet. What is p prime?

  • A. 0.392 (correct)
  • B. 98
  • C. 2.551
  • D. 0.608

Answer: A

Why: The sample proportion is the successes over the trials: 98 divided by 250 is 0.392.

Why B tempts people
That is the count of successes, before dividing by the sample size.
Why C tempts people
That divides n by x rather than x by n, and exceeds one so cannot be a proportion.
Why D tempts people
That is q prime, the proportion who do NOT own a tablet.

58. Check yourself 2 of 3

Check

The plus-four method.

Check your understanding

Six of 25 students smoke. What counts does the plus-four method use?

  • A. x = 8 and n = 29 (correct)
  • B. x = 8 and n = 25
  • C. x = 6 and n = 29
  • D. x = 10 and n = 29

Answer: A

Why: Two successes and two failures are added, so x rises by two and n rises by four.

Why B tempts people
That adds the successes without adding the four trials they belong to.
Why C tempts people
That adds the four trials without counting the two extra successes among them.
Why D tempts people
That adds four successes rather than two.

59. Check yourself 3 of 3

Check

Planning a sample.

Check your understanding

Why is p prime taken to be 0.5 when planning a sample size?

  • A. It maximises p prime q prime, giving the largest and therefore safest n (correct)
  • B. Most proportions are near one half
  • C. It makes the arithmetic easier
  • D. It minimises the required sample size

Answer: A

Why: The product is largest at one half, so the sample is large enough whatever the true proportion turns out to be.

Why B tempts people
Many are not, but the point is safety rather than typicality.
Why C tempts people
It does simplify the arithmetic, but that is incidental to the reason.
Why D tempts people
It maximises rather than minimises n; an optimistic guess would minimise it and risk missing the target.

60. Where this shows up outside the textbook

Real world

A hospital reviews 40 randomly selected surgical cases and finds 2 with a post-operative complication. A quality report states the complication rate as 5 percent and, using the ordinary formula, gives a 95 percent interval of minus 1.8 to 11.8 percent. A manager asks how a complication rate can be negative.

Discussion prompt

Explain what went wrong, rebuild the interval properly, and say what the result does and does not support.

Hint: The proportion is close to zero and the sample is small — exactly the case the section warns about.

Answer:

Nothing went wrong arithmetically; the method itself broke down. With p prime of 0.05 and n of 40, the error bound is 1.96 times the square root of 0.05 times 0.95 over 40, about 0.0676 — so the ordinary interval genuinely does run from minus 1.8 to 11.8 percent. A negative endpoint is impossible for a proportion, and its appearance is the signal that the normal approximation has failed.

It failed for the reason the section names: the approximation needs n large and p not close to zero or one, and here the proportion is very near zero on a small sample. Only two successes were observed, which is far too few for a bell-shaped approximation to hold.

\[ \text{plus-four: } x = 4, \; n = 44, \; p' = 0.0909, \quad \text{EBP} \approx 0.0850 \]

The plus-four method is exactly the repair. Adding two successes and two failures gives 4 out of 44, a point estimate of 0.0909 and an error bound of about 0.085, so the interval runs from about 0.6 percent to 17.6 percent — entirely inside the legal range and, conditions met, more accurate.

What the result supports is very little, and saying so is the honest conclusion. Forty cases with two complications is consistent with a true rate anywhere from well under one percent to nearly eighteen — a range too wide to act on. The manager's real question should be how many cases would be needed: at a target margin of two percentage points and a plausible rate near 5 percent, the planning formula gives roughly 456 cases, and using the conservative 0.25 would give about 2,401. Reporting the point estimate of 5 percent without that context would suggest a precision the data does not have.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

Why does a confidence interval for a proportion need no separate estimate of spread?

  • Because proportions have no variability
  • Because a binomial's variability is determined once its probability is, so p prime supplies both
  • Because the sample size is always large
  • Because the spread is assumed to be one

Correct: Because p prime supplies both the centre and the spread.

\[ \text{EBP} = z_{\alpha/2}\sqrt{\frac{p'q'}{n}}, \qquad q' = 1 - p' \]

Why: A binomial distribution is fixed by n and p, so its standard deviation follows from p rather than being a free parameter. Substituting the estimate p prime therefore gives a computable standard error with no second quantity to estimate — which is why these problems supply no sigma and why no t distribution is involved.

62. Explain it to someone a year behind you

Explain it

They computed an error bound using p prime of 60 and q prime of 40, and got 3.60.

Discussion prompt

In two sentences or fewer, locate the error.

Hint: Ask what range a proportion can occupy.

Answer:

An error bound of 3.60 is impossible for a proportion, since the whole parameter lies between zero and one.

They put percentages into a formula written in proportions: 60 percent has to enter as 0.60, which gives a bound of about 0.036 instead.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Recognising a proportion problem and forming p prime
  • Computing the error bound with the right square root
  • Applying the plus-four adjustment to both counts
  • Planning a sample size with the conservative substitution

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the first, look for categorical data with no mean mentioned. For the second, use proportions rather than percentages and keep the square root. For the third, x rises by two and n by four. For the fourth, use 0.25 for the product and round up. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the chapter on one page

Connect it up

Paper. Twenty-five minutes. This one closes the chapter, so make it a comparison.

Draw it

Down the left, write the three intervals of chapter 8 as three rows: a mean with sigma known, a mean with sigma unknown, and a proportion. Beside each write its point estimate and its error bound, and circle what differs between them. In the middle of the page, work Example 8.10 completely: form p prime from 421 over 500, find the z-score, compute the error bound, build the interval, and write the interpreting sentence naming all adult residents of this city. Below it, work Example 8.11 at 90 percent and note that its error bound is LARGER despite the lower level, then write one sentence explaining why in terms of the product p prime q prime. To the right, sketch the curve of p prime times q prime from 0 to 1, mark its peak of 0.25 at one half, and mark the values at 0.2, 0.3 and 0.6 to confirm the peak. Beneath that, work Example 8.12 twice — once as observed with 6 of 25 and once plus-four with 8 of 29 — draw both intervals to scale one above the other, and write which way the adjustment moved the estimate and why. At the bottom, do Example 8.14's sample size calculation showing the conservative substitution and the rounding up, then redo it with the margin halved and note the factor.

Check the middle section by confirming Example 8.11's error bound really does exceed Example 8.10's — if it does not, a z-score has been swapped. Check the plus-four pair by confirming the adjusted estimate lies between the observed one and 0.5, never outside that span.

65. What you can do now

Recap

Five things, and the first is the recognition test.

If you seeThen
Categorical data, no mean mentionedA proportion: the model is binomial
A count out of a totalp prime is that count over the total
Percentages in the problemConvert to proportions before substituting
An endpoint outside 0 to 1The approximation has failed; use plus-four
A small sample at 90 percent or betterThe plus-four method applies
A sample size to planUse p'q' = 0.25 and round up
A margin to halveQuadruple the sample

That closes chapter 8, and with it estimation. Chapter 9 turns to the other half of inference: instead of asking what values a parameter might take, it starts from a specific claim about the parameter and asks whether the data are consistent with it. The machinery is the same standard error and the same distributions, put to a different question.

OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, pp. 420-426 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §8.3 A Population Proportion — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 420-426
  2. OpenStax Introductory Business Statistics 2e, §8.3 A Confidence Interval for A Population Proportion — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108