Sections 8.1 and 8.2 estimated a population mean. This one changes the parameter rather than the multiplier: when the data are categorical, the quantity of interest is the proportion of the population in one category, and the underlying distribution is the binomial of section 4.3. The point estimate is the sample proportion, written p prime, the number of successes divided by the number of trials, and its error bound is a z-score times the square root of p prime times q prime over n. No population standard deviation appears anywhere, because a binomial's variability is determined once its probability is. The section also introduces the plus-four adjustment, which adds two successes and two failures to the observed counts before proceeding, and gives more accurate intervals when the confidence level is at least ninety percent and the sample at least ten. Solving the error bound for n gives a planning formula, and using one half for the unknown proportion makes it as large as it can be.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 8 — Confidence Intervals
A Population Proportion
Objectives
Five outcomes, and the first is the recognition test the book supplies.
OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, pp. 420-426 — the section these objectives are drawn from
Warm-up
Sections 8.1 and 8.2 built intervals for a mean from a point estimate and a standard error.
Discussion prompt
Of 500 adults surveyed, 421 own a smartphone. What is the point estimate for the population proportion, and where would its standard error come from — there is no sigma anywhere in the problem.
Hint: Section 4.3 gave the binomial its own standard deviation.
Answer:
The point estimate is 421 over 500, which is 0.842 — the sample proportion, and the natural analogue of the sample mean.
For the standard error there is no separate spread to estimate, and none is needed. A binomial's variability is fixed once its probability is, so the spread of a sample proportion is the square root of p times q over n — and substituting the estimate p prime for the unknown p gives something computable from the one number already found.
That is the whole structural difference from the previous two sections. Estimating a mean needed two quantities, a centre and a spread; estimating a proportion needs only one, because the spread comes free with the centre.
Concept
To form a proportion, take X, the random variable for the number of successes, and divide it by n, the number of trials. That random variable, read p prime, is the sample proportion and the point estimate of the population proportion. When n is large and p is not close to zero or one, the normal distribution approximates the binomial, and the interval takes the familiar form.
p prime — The sample proportion, x over n. Sometimes written p hat. It is the point estimate for the population proportion p, and it also determines the standard error, since q prime is simply one minus it.
\[ p' = \frac{x}{n}, \qquad \text{EBP} = z_{\alpha/2}\sqrt{\frac{p'q'}{n}} \]
The normal approximation is the same one section 4.6 raised for the Poisson and section 6.2 will be relied on throughout: when n is large and p is not close to zero or one, the binomial's shape is close enough to normal for a z-score to apply. That condition matters — a proportion very near zero or one from a small sample produces an interval that can run outside the range from zero to one, which is one of the reasons the plus-four adjustment exists.
Figure (svg): A card defining the sample proportion as the number of successes over the number of trials, with the error bound formula
OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, pp. 420-421
Section
Section 1
Concept
The book gives a two-part test. First, the underlying distribution is a binomial distribution — the data record successes among trials. Second, there is no mention of a mean or average anywhere in the problem.
the recognition test — Binomial structure plus the absence of any average. If the question asks what fraction, what percent, or how many out of, the parameter is a proportion rather than a mean.
\[ X \sim B(n, p) \;\Longrightarrow\; p' = \frac{X}{n} \]
The examples the book opens with are all polling and market research: the proportion of voters favouring a candidate, the proportion of stocks that rise in a week, the proportion of households owning a personal computer. All are counts of yes-or-no outcomes, and none has a natural average — asking for the average vote would not mean anything.
Figure (svg): Two columns contrasting a problem about a mean with a problem about a proportion
OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, p. 420 — the recognition test and the opening examples
Picture it
What distinguishes a proportion problem from a mean problem.
Figure (svg): Two columns contrasting a problem about a mean with a problem about a proportion
The fourth row is the structural difference and the one worth remembering. For a mean the spread is a separate quantity that has to be given or estimated; for a proportion it is determined by the estimate itself, which is why these problems supply no sigma and need none.
Worked example
Example 8.10's setup.
\[ \text{of } 500 \text{ adults surveyed, } 421 \text{ own smartphones} \]
Check the data type
Why: Owns one, or does not.
Check for a mean
Why: None mentioned.
Name the model
Why: Successes among trials.
Form the point estimate
Why: 421 over 500.
\[ p' = 0.842 \]
Figure (svg): The solution to Worked example identifying the parameter shown as a ladder of expressions, one row per legal move
\[ X \sim B(500, p), \qquad p' = \frac{421}{500} = 0.842 \]
Verify: confirm the answer will be a fraction rather than a quantity
Why: A proportion must lie between zero and one, so an answer outside that range would be impossible on its face — unlike a mean, which can be any number the data's units allow. That bound is a genuine check available at the end of every problem in this section, and it also warns when the normal approximation is straining.
OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, pp. 420-422
Sorting
Ask what kind of data the question is about.
Sort into buckets
Sort each question.
Item (d) is the one worth pausing on: it counts chemicals, so it looks categorical, but each infant contributes a NUMBER rather than a yes or no — and the question asks for a mean of those numbers. Counting is not the same as classifying.
Worked example
Example 8.11, where the question asks for a percent.
\[ \text{of } 500 \text{ students surveyed, } 300 \text{ are registered voters} \]
Read what is asked
Why: The percent registered.
Form p prime
Why: Three hundred over 500.
\[ 0.600 \]
Note q prime
Why: One minus it.
\[ 0.400 \]
Note what is absent
Why: No sigma, no s.
Figure (svg): The solution to Worked example a percentage question shown as a ladder of expressions, one row per legal move
\[ p' = \frac{300}{500} = 0.600, \qquad q' = 0.400 \]
Verify: confirm percent and proportion are the same parameter
Why: A question asking for a percent and one asking for a proportion are identical apart from a factor of a hundred, and the book reports its answer both ways — between 56.4 percent and 63.6 percent, or the interval (0.564, 0.636). Working in proportions throughout and converting only in the final sentence avoids the error of putting a percentage into the formula, where it would inflate the error bound tenfold.
OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, p. 423
Trap
\[ p' = 60, \quad q' = 40, \quad \sqrt{\frac{(60)(40)}{500}} \approx 2.19 \]
Use the percentages as they are stated
Why: The question is phrased in percents.
\[ \text{an error bound of over } 3, \text{ for a quantity bounded by } 1 \]
The formula is written in proportions, so 60 percent has to enter as 0.60.
\[ p' = 0.60, \quad q' = 0.40, \quad \sqrt{\frac{(0.60)(0.40)}{500}} \approx 0.0219 \]
Convert every percentage to a proportion before substituting
Why: Divide by a hundred first.
The error is caught instantly by the range check: an error bound of 2.19 for a proportion is impossible, since the whole parameter lies between zero and one. Converting to percentages only in the final interpreting sentence keeps the formulas clean and makes this mistake hard to make.
Fill the middle
How the sample proportion is formed.
Fill in the blanks
p' = \fracsuccesses___, \text___ x \text___ ___
Why: Successes — the count in the category of interest, divided by the number of trials. The label success is neutral: in Example 8.12 a success is a student who smoked.
Two truths and a lie
All three concern recognising the problem type.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and its falsity is the section's structural point. A binomial's spread is determined by its probability, so the standard error is built from p prime itself — no sigma is given because none is needed.
Prediction
Commit before reasoning.
Predict first
What range must a confidence interval for a proportion lie in?
Correct: Between 0 and 1.
Why: A proportion is a fraction of a whole, so it cannot be negative or exceed one. That gives a check unavailable for means: an endpoint outside the range signals either a percentage entered as a whole number or a normal approximation straining on a small sample near zero or one.
Section
Section 2
Concept
The error bound for a proportion is the z-score for the confidence level multiplied by the square root of p prime times q prime over n. The z-score is found exactly as in section 8.1; only the standard error is different.
EBP — The error bound for a population proportion. Unlike the error bound for a mean, it needs no separate estimate of spread, because a binomial's variability follows from its probability.
\[ \left(p' - \text{EBP},\; p' + \text{EBP}\right), \qquad \text{EBP} = z_{\alpha/2}\sqrt{\frac{p'q'}{n}} \]
It is worth noticing what has NOT changed. The z-score comes from the confidence level exactly as before, the interval is symmetric about its point estimate exactly as before, and the interpretation template is the same. The book says the procedure is similar to that for the population mean but the formulas are different, and the difference is confined to a single square root.
Figure (svg): The procedure for a confidence interval for a proportion
OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, pp. 421-423 — the formula and Examples 8.10 and 8.11
Picture it
The section 8.1 procedure with a proportion's point estimate and standard error.
Figure (svg): The procedure for a confidence interval for a proportion
The note at the bottom matters for the next idea: when the plus-four adjustment is used, the counts are changed before the procedure starts, and nothing about the procedure itself is altered. That is the whole appeal of the method — it costs no new formula.
Worked example
Example 8.10, at a 95 percent level.
\[ x = 421, \; n = 500, \; \text{CL} = 0.95 \]
Form the point estimate
Why: 421 over 500.
\[ 0.842 \]
Form q prime
Why: One minus it.
\[ 0.158 \]
Find the z-score
Why: For 95 percent.
\[ 1.96 \]
Compute the error bound
Why: 1.96 times root of 0.842 times 0.158 over 500.
\[ 0.032 \]
Figure (svg): The solution to Worked example smartphone ownership shown as a ladder of expressions, one row per legal move
\[ 0.842 \pm 1.96\sqrt{\frac{(0.842)(0.158)}{500}} = (0.810, 0.874) \]
Verify: confirm the interval lies inside the range a proportion can take
Why: Both endpoints are between zero and one, as they must be, and the interval is roughly six percentage points wide on a sample of 500 — which is the familiar order of magnitude for a poll of that size. An error bound of about 0.03 at n equal to 500 is worth carrying as a benchmark, since it is close to the margin of error most published polls report.
OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, p. 422
Faded example
Of 250 people surveyed, 98 own a tablet. Use 95 percent.
Fill in the blanks
p' = \frac0.3920.061 = ___, \quad \text___ = 1.96\sqrt______} \approx ___
Why: The point estimate is 0.392, and the error bound about 0.061 — so the interval runs from roughly 0.331 to 0.453. The bound is larger than Example 8.10's because the sample is half the size.
Worked example
Example 8.11, with a different level and a proportion near one half.
\[ x = 300, \; n = 500, \; \text{CL} = 0.90 \]
Form the point estimate
Why: Three hundred over 500.
\[ 0.600 \]
Find the z-score
Why: For 90 percent.
\[ 1.645 \]
Compute the error bound
Why: 1.645 times root of 0.24 over 500.
\[ 0.036 \]
Convert to percent
Why: For the interpretation.
\[ 56.4\text{ to } 63.6 \]
Figure (svg): The solution to Worked example registered voters at 90 percent shown as a ladder of expressions, one row per legal move
\[ 0.600 \pm 1.645\sqrt{\frac{(0.6)(0.4)}{500}} = (0.564, 0.636) \]
Verify: compare the error bound with Example 8.10's, and explain the difference
Why: This bound is 0.036 against 0.032 for the smartphone example, even though the level here is LOWER and the sample the same size. The reason is the product p prime q prime: 0.6 times 0.4 is 0.24, against 0.842 times 0.158 which is 0.133 — so a proportion near one half is intrinsically harder to pin down, and that is exactly what the next idea's planning formula exploits.
OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, p. 423
Error analysis
The correct value is about 0.036.
Annotate
On: \( \begin{aligned} &(1)\; 1.645\sqrt{\tfrac{(60)(40)}{500}} \approx 3.60 \\ &(2)\; 1.645\,\tfrac{(0.6)(0.4)}{500} \approx 0.0008 \\ &(3)\; 1.96\sqrt{\tfrac{(0.6)(0.4)}{500}} \approx 0.043 \\ &(4)\; 1.645\sqrt{\tfrac{(0.6)(0.4)}{500}} \approx 0.036 \end{aligned} \)
Errors (1) and (2) are both caught by the range check, since one exceeds 1 and the other implies an implausibly precise interval. Error (3) produces a perfectly plausible number and is caught only by checking the level against the z-score.
Two truths and a lie
All three concern the formula.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The t distribution exists to handle an estimated standard deviation for a MEAN; a proportion's spread comes from p prime itself, so no t is involved. Small samples are handled instead by the plus-four adjustment, which is the next idea.
Estimation
A poll of about 1,000 people reports a proportion near one half.
Predict first
Roughly what is its 95 percent margin of error?
Correct: About 3 percentage points.
Why: 1.96 times the square root of 0.25 over 1,000 is about 0.031. That is why so many published polls quote a margin of about three points: a sample near a thousand at a proportion near one half is the industry's standard configuration, and this calculation is where the number comes from.
Prediction
Commit before reasoning.
Predict first
For a fixed sample size, which sample proportion gives the widest interval?
Correct: 0.5.
Why: The error bound depends on the product p prime q prime, which is largest at one half: 0.25 there against 0.09 at either 0.1 or 0.9. A proportion near the extremes is intrinsically easier to estimate, and the next idea turns that fact into a conservative planning rule.
Section
Section 3
Concept
Because we do not know the true proportion, we are forced to use point estimates to calculate the standard deviation of the sampling distribution, and studies have shown that the resulting estimate can be flawed. The adjustment is to pretend we have four additional observations: two successes and two failures. The new sample size is n plus four and the new count of successes is x plus two.
the plus-four method — Adding two successes and two failures to the observed counts before computing the interval. It should be used when the confidence level desired is at least 90 percent and the sample size is at least ten.
\[ x \to x + 2, \qquad n \to n + 4 \]
What the adjustment does is pull the point estimate toward one half and widen the interval slightly, and both effects are corrections in the right direction. The ordinary interval is known to be too narrow on average for small samples, and to behave particularly badly when the observed proportion is near zero or one — where it can even produce endpoints outside the range a proportion can occupy.
Figure (svg): Two horizontal intervals for the same data, the plus-four one shifted slightly toward the centre and a little wider
OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, pp. 424-425 — the plus-four method and Examples 8.12 and 8.13
Picture it
Example 8.12: six of twenty-five statistics students reported smoking.
Figure (svg): Two horizontal intervals for the same data, the plus-four one shifted slightly toward the centre and a little wider
The adjusted interval is a little wider and sits a little higher, because 8 over 29 is 0.276 against 6 over 25 at 0.240. On a sample of twenty-five that shift is meaningful; on a sample of a thousand the same four observations would move nothing, which is why the method is aimed at small samples.
Worked example
Example 8.12, using the plus-four method.
\[ x = 6, \; n = 25, \; \text{CL} = 0.95 \]
Adjust the counts
Why: Add two and four.
\[ x = 8, n = 29 \]
Form the point estimate
Why: Eight over 29.
\[ 0.276 \]
Find the z-score
Why: For 95 percent.
\[ 1.96 \]
Compute the error bound
Why: From the adjusted numbers.
\[ 0.163 \]
Figure (svg): The solution to Worked example smoking among statistics students shown as a ladder of expressions, one row per legal move
\[ 0.276 \pm 0.163 = (0.113, 0.439) \]
Verify: confirm the adjustment moved the estimate in the expected direction
Why: The observed proportion was 6 over 25, which is 0.240, and the adjusted one is 0.276 — pulled toward one half, as adding one success and one failure in equal measure must do. The interval is also slightly wider than the unadjusted one. Both changes are corrections toward honesty about how little twenty-five observations pin down.
OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, pp. 424-425
Faded example
Of 65 first-year students, 31 have declared a major. Use the plus-four method.
Fill in the blanks
x = 31 + 2 = 33, \quad n = 65 + 4 = 69, \quad p' = \frac______ \approx 0.478
Why: Thirty-three successes out of sixty-nine trials gives about 0.478, barely moved from the observed 0.477 because that proportion was already close to one half — the adjustment has least effect exactly there.
Worked example
Example 8.13, at a 90 percent level.
\[ x = 13, \; n = 50, \; \text{CL} = 0.90 \]
Adjust the counts
Why: Fifteen out of 54.
\[ x = 15, n = 54 \]
Form the point estimate
Why: Fifteen over 54.
\[ 0.278 \]
Find the z-score
Why: For 90 percent.
\[ 1.645 \]
Compute the error bound
Why: From the adjusted numbers.
\[ 0.100 \]
Figure (svg): The solution to Worked example electric vehicle interest shown as a ladder of expressions, one row per legal move
\[ 0.278 \pm 0.100 = (0.178, 0.378) \]
Verify: confirm the conditions for using the method are met
Why: The confidence level is 90 percent, which meets the at-least-90 requirement, and the sample of 50 comfortably exceeds ten. Both conditions hold, so the adjustment is appropriate. The book's own follow-up asks the same question of a sample of 588, where the four extra observations barely move anything — which is the honest way to see that the method is a small-sample correction.
OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, p. 425
Trap
\[ x = 6 + 2 = 8, \quad n = 25 \]
Add the two successes and leave n alone
Why: Only the successes were mentioned as changing.
\[ p' = \frac{8}{25} = 0.32, \text{ instead of } 0.276 \]
Four observations were pretended, not two, so the denominator must grow by four.
\[ x = 8, \quad n = 29, \quad p' = \frac{8}{29} \approx 0.276 \]
Add two successes AND two failures, so n grows by four
Why: The book's reminder: the method assumes an additional four trials.
The check is that the adjustment should always pull the estimate TOWARD one half, never past it or away from it. Here 0.240 moves to 0.276, which is toward 0.5; the erroneous 0.32 overshoots because the two pretend failures were never counted.
Two truths and a lie
All three concern the plus-four method.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and the book's reminder says so: you do not need to change the process for calculating the confidence interval; simply update the values of x and n. The whole appeal of the method is that it introduces no new formula.
Prediction
Commit before reasoning.
Predict first
An observed proportion is 0.10. What does the plus-four adjustment do to it?
Correct: Raises it, toward one half.
Why: Adding one pretend success and one pretend failure in equal numbers always drags a proportion toward the middle, and the effect is largest when the observed proportion is extreme. That is precisely where the ordinary interval behaves worst, which is why the correction is aimed there.
Estimation
The same four extra observations, on two sample sizes.
Predict first
For which sample does the plus-four adjustment change the estimate more?
Correct: n = 25.
Why: Four extra observations are a sixth of a sample of 25 and under a percent of a sample of 588, so the shift is large in the first case and negligible in the second. The book makes exactly this comparison in its follow-up to Example 8.13, and it is the clearest way to see that the method is a small-sample correction.
Section
Section 4
Concept
Solving the error bound formula for n gives the sample size a study needs. But n depends on p prime, which is not yet known — so the book substitutes one half for it, because that makes the product p prime q prime as large as it can be and therefore gives the largest, safest sample size.
the conservative substitution — Setting p prime and q prime both to 0.5, giving a product of 0.25. Since 0.25 is the largest possible value of that product, the resulting sample size is large enough whatever the true proportion turns out to be.
\[ n = \frac{z^2\,p'q'}{\text{EBP}^2}, \qquad p'q' \le 0.25 \]
The book checks the claim by trial rather than by algebra: 0.6 times 0.4 is 0.24, 0.3 times 0.7 is 0.21, 0.2 times 0.8 is 0.16, and so on — all below the 0.25 that one half gives. The consequence is that a study planned this way is never under-powered, though it may be somewhat larger than strictly necessary if the true proportion turns out to be extreme.
Figure (svg): A downward-opening curve peaking at nought point two five when the proportion is one half
OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, p. 426 — the sample size formula and Example 8.14
Picture it
The product of a proportion and its complement, across the whole range.
Figure (svg): A downward-opening curve peaking at nought point two five when the proportion is one half
The curve is flat near its peak, which has a practical consequence worth noticing: any proportion between about 0.3 and 0.7 gives a product within about fifteen percent of the maximum, so the conservative choice costs little unless the true proportion is genuinely extreme.
Worked example
Example 8.14, planning a study.
\[ \text{EBP} = 0.03, \; \text{CL} = 0.90 \]
Find the z-score
Why: For 90 percent.
\[ 1.645 \]
Choose p prime
Why: Unknown, so use one half.
\[ p' q' = 0.25 \]
Apply the formula
Why: z squared times 0.25 over EBP squared.
\[ 751.67 \]
Round UP
Why: As for any sample size.
\[ 752 \]
Figure (svg): The solution to Worked example how many customers to survey shown as a ladder of expressions, one row per legal move
\[ n = \frac{(1.645)^2(0.25)}{(0.03)^2} \approx 751.7 \;\to\; 752 \]
Verify: confirm the choice of one half really is the safe one
Why: Had the true proportion turned out to be 0.2, the product would be 0.16 and the required sample only about 481 — so planning for 752 over-samples by more than half. But had a researcher optimistically assumed 0.2 and the truth been 0.5, the study would have missed its target margin. Over-sampling costs money; under-sampling costs the answer, which is why the conservative choice is standard.
OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, p. 426
Faded example
A 95 percent interval within two percentage points, with the proportion unknown.
Fill in the blanks
n = \frac24012401 \approx ___ \;\to\; ___
Why: About 2,401 respondents. Tightening the margin from three points to two, at a higher confidence level, has more than tripled the sample from Example 8.14's 752.
Worked example
The same study at two target margins.
\[ \text{EBP} = 0.03 \text{ against } 0.015, \; \text{CL} = 0.90 \]
At a 3 point margin
Why: The original.
\[ 752 \]
Halve the margin
Why: To 1.5 points.
\[ EBP = 0.015 \]
Note the squaring
Why: EBP is squared below.
Compute
Why: Four times 752.
\[ \text{about } 3, 007 \]
Figure (svg): The solution to Worked example the cost of a tighter margin shown as a ladder of expressions, one row per legal move
\[ n \propto \frac{1}{\text{EBP}^2} \]
Verify: confirm this matches section 8.1's rule for means
Why: Section 8.1's sample-size formula also squared the ratio, so halving the error bound quadrupled n there too — and both trace back to chapter 7's standard error falling like one over the square root of n. The rule is the same whatever the parameter: precision costs sample size at a square rate, which is why very tight margins are rarely attempted.
OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, p. 426
Trap
\[ \text{expect about } 20 \text{ percent, so use } p'q' = 0.16 \]
Use an optimistic guess to reduce the required sample
Why: A smaller product means a smaller n.
\[ n \approx 481, \text{ but if the truth is } 0.5 \text{ the margin is missed} \]
The margin achieved would be about 0.0375 rather than the 0.03 promised.
\[ \text{use } p'q' = 0.25, \text{ giving } n = 752 \]
Take the largest possible product, so the target holds whatever the truth is
Why: One half maximises it.
A prior estimate is legitimate when it is genuinely well founded — from an earlier survey of the same population, say — and then a smaller sample is defensible. What is not defensible is guessing optimistically to reduce cost, because the guess is precisely what the study was commissioned to establish.
Two truths and a lie
All three concern planning.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false: the error bound is squared in the denominator, so halving it QUADRUPLES the sample. From 752 to about 3,007 for Example 8.14's study, which is the same square-rate cost that governs every sample-size calculation in the book.
Estimation
A national poll of 1,000 quotes a margin of about 3 points at 95 percent.
Predict first
What proportion was assumed in that calculation?
Correct: One half.
Why: 1.96 times the square root of 0.25 over 1,000 gives 0.031 — about three points. Published polls almost always quote the conservative margin, which is why the same figure appears whatever the poll actually found. A poll reporting 20 percent support has a genuinely smaller margin than the headline number suggests.
Explain it
A classmate asks why anyone would plan for the widest possible interval rather than a realistic one.
Discussion prompt
In two sentences or fewer, justify it.
Hint: Ask what the study is being run to find out.
Answer:
The proportion is exactly what the study is commissioned to discover, so any assumption about it made beforehand could be wrong in the direction that matters.
Planning at one half guarantees the promised margin whatever the truth turns out to be, and over-sampling costs money while under-sampling costs the answer.
Section
Section 5
Concept
The interpretation follows the same template as for a mean, converted to percentages where the question was asked that way. The book gives an alternate wording for proportions: we estimate with a stated confidence that between one percentage and another of all members of the population are in the category.
interpreting a proportion interval — Naming the level, the population, the category and both endpoints. The book's own sentences say all students or all adult residents of this city, which fixes exactly what population is being described.
\[ (0.564, 0.636) \;\longrightarrow\; \text{“between 56.4 and 63.6 percent of ALL students”} \]
The words all students matter more here than they might seem to. A confidence interval extrapolates from the surveyed group to the whole population, and that extrapolation is only licensed by random sampling — which is a stronger requirement for polls than for measurements, because non-response tends to correlate with the very opinions being measured.
Figure (svg): A number line showing a sample proportion of nought point eight four two with its error bound either side
OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, pp. 422-423 — the interpretations and the alternate wording
Picture it
Example 8.10's result on a number line.
Figure (svg): A number line showing a sample proportion of nought point eight four two with its error bound either side
Stated in full: we estimate with 95 percent confidence that between 81 percent and 87.4 percent of all adult residents of this city have smartphones. The phrase all adult residents of this city is doing real work, since the survey covered 500 of them and the claim is about every one.
Worked example
Example 8.11 gives both.
\[ (0.564, 0.636) \text{ at } 90 \text{ percent} \]
The parameter wording
Why: About the percent itself.
The alternate wording
Why: About the students.
\[ \text{between } 56.4\text{ and } 63.6 \%\text{ of } ALL\text{ students} \]
Name the level
Why: Ninety percent.
Name the population
Why: All students at the university.
Figure (svg): The solution to Worked example two correct wordings shown as a ladder of expressions, one row per legal move
\[ (0.564, 0.636) \equiv (56.4\%, 63.6\%) \]
Verify: confirm the two wordings say the same thing
Why: They differ only in whether the percentage attaches to the parameter or to the population, and both name the level, the group and both endpoints. The book offers the second because it reads more naturally in a report, and the important thing is that neither omits the population — an interval quoted without saying who it describes is not interpretable.
OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, p. 423
Two truths and a lie
All three concern interpretation.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and confuses an interval for a parameter with a range for individuals. Each person either owns a smartphone or does not; there is no sense in which 95 percent of people fall between 0.810 and 0.874. The interval locates the population's overall proportion.
Worked example
The book's explanation, in the same terms as section 8.1.
\[ \text{a } 95 \text{ percent interval for a proportion} \]
Identify what is random
Why: The interval.
Identify what is fixed
Why: The population proportion.
State the level's meaning
Why: Across repeated samples.
\[ 95 \%\text{ contain } p \]
Note what it excludes
Why: A claim about this interval.
Figure (svg): The solution to Worked example what the level describes shown as a ladder of expressions, one row per legal move
\[ \text{the interval varies; } p \text{ does not} \]
Verify: confirm this matches section 8.1's account exactly
Why: The book's explanation for proportions is word-for-word parallel to its explanation for means, which is worth noticing: the meaning of a confidence level does not depend on what parameter is being estimated. Learning it once in section 8.1 was enough, and every later chapter reuses the same idea without restating it.
OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, pp. 422-423
Trap
\[ \text{between } 81 \text{ and } 87.4 \text{ percent of the } 500 \text{ surveyed own smartphones} \]
Attach the interval to the people actually asked
Why: They are the ones the data came from.
\[ \text{but } 421 \text{ of } 500 \text{ is exactly } 84.2 \text{ percent, with no uncertainty} \]
The sample's own proportion is known precisely; the interval exists to describe the population it was drawn from.
\[ \text{between } 81 \text{ and } 87.4 \text{ percent of ALL adult residents of this city} \]
Attach the interval to the population, and say which population
Why: That extrapolation is the whole purpose of the interval.
This error empties the interval of content, since there is nothing uncertain about the sample. It is worth catching because the corrected sentence also forces the question of WHICH population — all adults in the city, or all adults reachable by this survey method, and those may not be the same thing.
Faded example
An interval for a proportion is (0.178, 0.378).
Fill in the blanks
\text17.8 37.8 \text___ ___ \text___
Why: Between 17.8 and 37.8 percent — Example 8.13's result for adults aged 18 to 29 who would consider an electric vehicle. Converting at the end keeps percentages out of the formula, where they would inflate the error bound.
Discrimination
For the interval (0.810, 0.874) from a survey of 500 city residents.
Sort into buckets
Sort each statement.
Prediction
Commit before reasoning.
Predict first
What allows an interval from 500 people to describe all adult residents of a city?
Correct: Random sampling.
Why: Only random selection makes the sample representative of the population, and nothing else substitutes for it. A large non-random sample, a good normal approximation and a 99 percent level would all still describe whoever happened to be reachable — which is section 8.2's point about the two separate assumptions, and it applies just as sharply here.
Comparison
Fill the blanks. One shape, three sets of parts.
Comparison matrix
| Point estimate | Error bound | |
|---|---|---|
| Mean, sigma known | x-bar | z times sigma over root n |
| Mean, sigma unknown | x-bar | t times s over root n |
| Proportion | p prime | z times root of p'q' over n |
| What is estimated separately | the spread, for both means | nothing: p' gives both |
The last row is the structural difference worth carrying out of the chapter. Estimating a mean requires a centre and a spread; estimating a proportion requires only the one number, because a binomial's variability follows from its probability.
Pattern
Six steps, and the first is the recognition test.
For planning, solve for n with p prime q prime set to 0.25, and round up. Halving the target margin quadruples the sample.
OpenStax Introductory Business Statistics 2e, §8.3 A Confidence Interval for A Population Proportion §8.3 A Confidence Interval for A Population Proportion
Check
The point estimate.
Check your understanding
Of 250 people surveyed, 98 own a tablet. What is p prime?
Answer: A
Why: The sample proportion is the successes over the trials: 98 divided by 250 is 0.392.
Check
The plus-four method.
Check your understanding
Six of 25 students smoke. What counts does the plus-four method use?
Answer: A
Why: Two successes and two failures are added, so x rises by two and n rises by four.
Check
Planning a sample.
Check your understanding
Why is p prime taken to be 0.5 when planning a sample size?
Answer: A
Why: The product is largest at one half, so the sample is large enough whatever the true proportion turns out to be.
Real world
A hospital reviews 40 randomly selected surgical cases and finds 2 with a post-operative complication. A quality report states the complication rate as 5 percent and, using the ordinary formula, gives a 95 percent interval of minus 1.8 to 11.8 percent. A manager asks how a complication rate can be negative.
Discussion prompt
Explain what went wrong, rebuild the interval properly, and say what the result does and does not support.
Hint: The proportion is close to zero and the sample is small — exactly the case the section warns about.
Answer:
Nothing went wrong arithmetically; the method itself broke down. With p prime of 0.05 and n of 40, the error bound is 1.96 times the square root of 0.05 times 0.95 over 40, about 0.0676 — so the ordinary interval genuinely does run from minus 1.8 to 11.8 percent. A negative endpoint is impossible for a proportion, and its appearance is the signal that the normal approximation has failed.
It failed for the reason the section names: the approximation needs n large and p not close to zero or one, and here the proportion is very near zero on a small sample. Only two successes were observed, which is far too few for a bell-shaped approximation to hold.
\[ \text{plus-four: } x = 4, \; n = 44, \; p' = 0.0909, \quad \text{EBP} \approx 0.0850 \]
The plus-four method is exactly the repair. Adding two successes and two failures gives 4 out of 44, a point estimate of 0.0909 and an error bound of about 0.085, so the interval runs from about 0.6 percent to 17.6 percent — entirely inside the legal range and, conditions met, more accurate.
What the result supports is very little, and saying so is the honest conclusion. Forty cases with two complications is consistent with a true rate anywhere from well under one percent to nearly eighteen — a range too wide to act on. The manager's real question should be how many cases would be needed: at a target margin of two percentage points and a plausible rate near 5 percent, the planning formula gives roughly 456 cases, and using the conservative 0.25 would give about 2,401. Reporting the point estimate of 5 percent without that context would suggest a precision the data does not have.
Commit first
Answer, then rate your confidence honestly.
Predict first
Why does a confidence interval for a proportion need no separate estimate of spread?
Correct: Because p prime supplies both the centre and the spread.
\[ \text{EBP} = z_{\alpha/2}\sqrt{\frac{p'q'}{n}}, \qquad q' = 1 - p' \]
Why: A binomial distribution is fixed by n and p, so its standard deviation follows from p rather than being a free parameter. Substituting the estimate p prime therefore gives a computable standard error with no second quantity to estimate — which is why these problems supply no sigma and why no t distribution is involved.
Explain it
They computed an error bound using p prime of 60 and q prime of 40, and got 3.60.
Discussion prompt
In two sentences or fewer, locate the error.
Hint: Ask what range a proportion can occupy.
Answer:
An error bound of 3.60 is impossible for a proportion, since the whole parameter lies between zero and one.
They put percentages into a formula written in proportions: 60 percent has to enter as 0.60, which gives a bound of about 0.036 instead.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, look for categorical data with no mean mentioned. For the second, use proportions rather than percentages and keep the square root. For the third, x rises by two and n by four. For the fourth, use 0.25 for the product and round up. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Twenty-five minutes. This one closes the chapter, so make it a comparison.
Draw it
Down the left, write the three intervals of chapter 8 as three rows: a mean with sigma known, a mean with sigma unknown, and a proportion. Beside each write its point estimate and its error bound, and circle what differs between them. In the middle of the page, work Example 8.10 completely: form p prime from 421 over 500, find the z-score, compute the error bound, build the interval, and write the interpreting sentence naming all adult residents of this city. Below it, work Example 8.11 at 90 percent and note that its error bound is LARGER despite the lower level, then write one sentence explaining why in terms of the product p prime q prime. To the right, sketch the curve of p prime times q prime from 0 to 1, mark its peak of 0.25 at one half, and mark the values at 0.2, 0.3 and 0.6 to confirm the peak. Beneath that, work Example 8.12 twice — once as observed with 6 of 25 and once plus-four with 8 of 29 — draw both intervals to scale one above the other, and write which way the adjustment moved the estimate and why. At the bottom, do Example 8.14's sample size calculation showing the conservative substitution and the rounding up, then redo it with the margin halved and note the factor.
Check the middle section by confirming Example 8.11's error bound really does exceed Example 8.10's — if it does not, a z-score has been swapped. Check the plus-four pair by confirming the adjusted estimate lies between the observed one and 0.5, never outside that span.
Recap
Five things, and the first is the recognition test.
| If you see | Then |
|---|---|
| Categorical data, no mean mentioned | A proportion: the model is binomial |
| A count out of a total | p prime is that count over the total |
| Percentages in the problem | Convert to proportions before substituting |
| An endpoint outside 0 to 1 | The approximation has failed; use plus-four |
| A small sample at 90 percent or better | The plus-four method applies |
| A sample size to plan | Use p'q' = 0.25 and round up |
| A margin to halve | Quadruple the sample |
That closes chapter 8, and with it estimation. Chapter 9 turns to the other half of inference: instead of asking what values a parameter might take, it starts from a specific claim about the parameter and asks whether the data are consistent with it. The machinery is the same standard error and the same distributions, put to a different question.
OpenStax Introductory Statistics 2e, §8.3 A Population Proportion §8.3, pp. 420-426 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.