The same comparison as the previous two sections, with categorical data. The parameter is the difference of two population proportions, the point estimate is the difference of the two sample proportions, and because the difference of two proportions follows an approximate normal distribution, the test statistic is a z-score with no degrees of freedom. One feature is genuinely new and is the section's central idea: because the null hypothesis generally states that the two proportions are the same, the two samples are estimating one common value, so their successes and their trials are combined into a single pooled proportion before the standard error is formed. That is the opposite of the instruction in section 10.1, and the reason is that a binomial's spread follows from its probability — so a null of equal proportions is also a null of equal spreads, while a null of equal means says nothing about the two standard deviations.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 10 — Hypothesis Testing with Two Samples
Comparing Two Independent Population Proportions
Objectives
Five outcomes, and the second explains why this section pools where section 10.1 forbade it.
OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 523-527 — the section these objectives are drawn from
Warm-up
Section 10.1 was emphatic that two sample variances must not be pooled.
Discussion prompt
A null hypothesis says two population proportions are equal. If that is true, are the two populations' standard deviations also equal?
Hint: Section 8.3 established that a binomial's spread comes from its probability.
Answer:
Yes, necessarily. A proportion's variability is p times q over n, and q is one minus p — so fixing p fixes the spread completely. Two populations with the same proportion cannot differ in variability.
That is not true for means. Two populations can have identical means and wildly different standard deviations, which is exactly why section 10.1 refused to pool: the null said nothing about the spreads, so nothing licensed combining them.
Here the null does say something about the spreads, because it says everything about the proportions. So the two samples really are estimating one common value, and pooling them is the correct thing to do rather than a shortcut.
Concept
Comparing two proportions is common, and if two estimated proportions differ it may be due to a difference in the populations or it may be due to chance. Generally the null hypothesis states that the two proportions are the same, and to conduct the test we use a pooled proportion.
the pooled proportion — Written pc, and computed as the total successes over the total trials across both samples. It is the single best estimate of the common proportion the null asserts.
\[ p_c = \frac{x_1 + x_2}{n_1 + n_2}, \qquad \text{SE} = \sqrt{p_c q_c\left(\frac{1}{n_1} + \frac{1}{n_2}\right)} \]
The logic is worth stating carefully because it is the reverse of section 10.1's. There the null constrained only the means, so the two spreads had to be estimated separately. Here the null constrains the proportions, and since a proportion determines its own spread, the null constrains the spreads too — which is exactly what makes one combined estimate the right one.
Figure (svg): A card showing the successes of two samples combined into a single pooled proportion
OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, p. 523
Section
Section 1
Concept
Let A and B be the subscripts for the two groups. The random variable is the difference of the two sample proportions, and the null generally states that the population proportions are equal — equivalently that their difference is zero.
the two notations again — H-nought as p-A equals p-B, or as p-A minus p-B equals zero. As in section 10.1, the second makes explicit that the null names a single value for the parameter.
\[ H_0: p_A = p_B \;\Longleftrightarrow\; H_0: p_A - p_B = 0 \]
The recognition test from section 8.3 applies unchanged: the data are categorical, the underlying distributions are binomial, and no mean is mentioned. What is new is only that there are two of them, so the question becomes whether one binomial's probability differs from another's rather than whether one differs from a fixed number.
Figure (svg): Two columns contrasting why means are not pooled with why proportions are
OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, p. 523 — Example 10.8's hypotheses
Picture it
The instruction reverses, and the reason is in the null.
Figure (svg): Two columns contrasting why means are not pooled with why proportions are
The second row of each column is the whole explanation. A null about means leaves the spreads free; a null about proportions does not, because a binomial has no spread parameter independent of its probability. Two apparently contradictory rules follow from one principle applied to two different situations.
Worked example
Two medications for hives, tested for a difference.
\[ \text{is there a DIFFERENCE in the proportions still reacting?} \]
Identify the parameter
Why: Proportions still reacting.
Read the claim
Why: Is a difference.
Write the alternative
Why: Two-sided.
\[ H a: p A \ne p B \]
Write the null
Why: Its contradiction.
\[ H 0: p A = p B \]
Figure (svg): The solution to Worked example setting up Example 10.8 shown as a ladder of expressions, one row per legal move
\[ H_0: p_A = p_B, \qquad H_a: p_A \ne p_B \]
Verify: confirm the parameter is a proportion rather than a mean
Why: The data record whether each adult still had hives after thirty minutes — a yes-or-no outcome counted across a sample, with no average of any measured quantity. Section 8.3's test applies: the underlying distribution is binomial and no mean is mentioned. Had the study measured how long the hives lasted, the parameter would have been a mean and section 10.1's test would apply instead.
OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, p. 523
Sorting
Ask what each individual contributes.
Sort into buckets
Sort each comparison.
The distinction is exactly section 8.3's, applied twice over. Getting it wrong sends the whole problem to the wrong standard error and the wrong pooling rule, so it is worth settling before anything else.
Worked example
Example 10.10, where the claim points one way.
\[ \text{do MORE younger adults own electric vehicles than older adults?} \]
Identify the groups
Why: Younger and older.
Read the claim
Why: More younger.
Write the alternative
Why: The claim.
\[ H a: p Y > p O \]
Read the tail
Why: Greater than.
Figure (svg): The solution to Worked example a directional comparison shown as a ladder of expressions, one row per legal move
\[ H_0: p_Y \le p_O, \qquad H_a: p_Y > p_O \]
Verify: confirm which group is subtracted from which
Why: Writing the alternative as p-Y greater than p-O means the difference p-Y minus p-O is positive, so the right tail is wanted and the sample proportions must be subtracted in that order. Reversing them would flip the statistic's sign and send the test to the wrong tail — the same discipline section 10.1 required, and equally consequential here.
OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 526-527
Trap
\[ 20 \text{ of } 200 \text{ and } 12 \text{ of } 200 \;\Rightarrow\; \text{compare mean reactions} \]
Treat the counts as measurements to average
Why: They are numbers, after all.
\[ \text{but each subject contributes a yes or a no} \]
There is no measured quantity to average; each person is in a category, and the parameter is the fraction in it.
\[ p'_A = \tfrac{20}{200} = 0.10, \quad p'_B = \tfrac{12}{200} = 0.06 \]
Form proportions and compare those
Why: Section 8.3's recognition test, with two groups.
The tell is what each individual contributes. A measured quantity per subject — a time, a weight, a score — gives a mean; a category per subject gives a proportion. Counts of subjects in a category are the raw material of a proportion, not observations to be averaged.
Fill the middle
The equivalent form.
Fill in the blanks
H_0: p_A = p_B \text0 p_A - p_B = ___
Why: Zero. As in section 10.1, writing the null this way shows that a single value is being named, which is what allows a sampling distribution to be centred on it.
Two truths and a lie
All three concern the setup.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. Example 10.10 compares 1,343 with 232 and works perfectly well; the standard error formula takes each sample size separately. Equal sizes are convenient but never required by any test in this chapter.
Prediction
Commit before reasoning.
Predict first
For H-a that p-Y exceeds p-O, which difference should be computed?
Correct: p-Y minus p-O.
Why: The alternative claims that difference is positive, so computing it that way puts the evidence in the right tail. Reversing it would flip the sign and the test would look in the wrong tail, reporting one minus the correct p-value — an error that turns a rejection into a comfortable non-rejection.
Section
Section 2
Concept
The pooled proportion combines the successes from both samples over the trials from both samples. It is the single best estimate of the common proportion the null claims the two populations share.
pc and qc — The pooled proportion and its complement. Since the null says both populations share a proportion, one estimate built from all the data is better than either sample's alone.
\[ p_c = \frac{x_A + x_B}{n_A + n_B}, \qquad q_c = 1 - p_c \]
It is worth noticing that the pooled proportion is used only for the STANDARD ERROR, never for the point estimate. The difference at the top of the statistic is still the difference of the two separate sample proportions — because that is what was observed. Pooling belongs to the denominator, where the null's assumption is being applied.
Figure (svg): A card showing the successes of two samples combined into a single pooled proportion
OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, p. 523 — the pooled proportion and the distribution
Picture it
The successes and the trials both combined.
Figure (svg): A card showing the successes of two samples combined into a single pooled proportion
The last line of the figure is the one to carry. Pooling here is not a convenience or an approximation — it follows from what the null actually says, and refusing to pool would mean ignoring information the null hypothesis provides.
Worked example
Twenty of two hundred against twelve of two hundred.
\[ x_A = 20, \; n_A = 200; \quad x_B = 12, \; n_B = 200 \]
Add the successes
Why: Twenty plus twelve.
\[ 32 \]
Add the trials
Why: Two hundred each.
\[ 400 \]
Divide
Why: The pooled proportion.
\[ 0.08 \]
Form the standard error
Why: Root of pc qc times the reciprocals.
\[ 0.02713 \]
Figure (svg): The solution to Worked example pooling for Example 10.8 shown as a ladder of expressions, one row per legal move
\[ p_c = \frac{32}{400} = 0.08, \quad \text{SE} = \sqrt{(0.08)(0.92)\left(\tfrac{1}{200}+\tfrac{1}{200}\right)} \approx 0.0271 \]
Verify: confirm the pooled value lies between the two sample proportions
Why: The two samples gave 0.10 and 0.06, and the pooled 0.08 sits between them — as a weighted average of the two must. A pooled proportion outside that range would signal an arithmetic error, and with equal sample sizes it lands exactly midway, as it does here.
OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 523-524
Faded example
Fifteen of 100 valve A cracked; six of 100 valve B cracked.
Fill in the blanks
p_c = \frac210.105 = \frac___}___ = ___
Why: Twenty-one successes in two hundred trials gives a pooled proportion of 0.105, which sits between the two sample proportions of 0.15 and 0.06.
Worked example
Example 10.10, where one group is nearly six times the other.
\[ 135 \text{ of } 1343; \quad 12 \text{ of } 232 \]
Add the successes
Why: 135 plus 12.
\[ 147 \]
Add the trials
Why: 1343 plus 232.
\[ 1575 \]
Divide
Why: The pooled proportion.
\[ 0.0933 \]
Compare with the samples
Why: 0.1005 and 0.0517.
Figure (svg): The solution to Worked example pooling with unequal samples shown as a ladder of expressions, one row per legal move
\[ p_c = \frac{147}{1575} \approx 0.0933 \]
Verify: confirm why the larger sample dominates
Why: The pooled proportion is a weighted average with the sample sizes as weights, so a group of 1,343 pulls it far more than one of 232. That is correct: the larger sample carries more information about the common proportion the null asserts, and the pooled estimate should reflect that. With equal samples the weighting disappears and the pooled value is the simple midpoint.
OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 526-527
Error analysis
The correct value is about 0.0271.
Annotate
On: \( \begin{aligned} &(1)\; \sqrt{\tfrac{(0.1)(0.9)}{200} + \tfrac{(0.06)(0.94)}{200}} \approx 0.0263 \\ &(2)\; \sqrt{(0.08)(0.92)\left(\tfrac{1}{200}\right)} \approx 0.0192 \\ &(3)\; (0.08)(0.92)\left(\tfrac{1}{200}+\tfrac{1}{200}\right) \approx 0.00074 \\ &(4)\; \sqrt{(0.08)(0.92)\left(\tfrac{1}{200}+\tfrac{1}{200}\right)} \approx 0.0271 \end{aligned} \)
Error (1) is the interesting one because it is section 10.1's method applied here. It is not absurd — it is a legitimate standard error for a confidence interval on the difference — but for a TEST, where the null supplies a common proportion, pooling is the correct choice.
Two truths and a lie
All three concern pooling.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The statistic's numerator is the observed difference of the two separate sample proportions — that is the evidence. Pooling applies only to the denominator, where the null's claim of a common proportion is being used.
Prediction
Commit before reasoning.
Predict first
Samples of 1,343 and 232 give proportions of 0.10 and 0.05. Where does the pooled value sit?
Correct: Much closer to 0.10.
Why: The pooled proportion weights each sample by its size, so a group nearly six times larger dominates — the value comes out at about 0.0933. Only with equal sample sizes does the simple midpoint apply, which is why Example 10.8's pooled 0.08 is exactly between 0.10 and 0.06.
Explain it
A classmate asks why section 10.1 forbade pooling and this section requires it.
Discussion prompt
In two sentences or fewer, explain.
Hint: Ask what each null hypothesis actually constrains.
Answer:
A null about two means says nothing about their spreads, so the two standard deviations have to be estimated separately.
A null about two proportions constrains the spreads too, because a binomial's variability is determined by its probability — so under the null both samples really are estimating one common value.
Section
Section 3
Concept
The standard error is the square root of the pooled proportion times its complement, multiplied by the sum of the reciprocals of the two sample sizes. The test statistic is the observed difference of sample proportions divided by it.
the two-proportion standard error — Built from a single pooled proportion and both sample sizes. The reciprocals are summed because each sample contributes uncertainty in inverse proportion to its size.
\[ z = \frac{p'_A - p'_B}{\sqrt{p_c q_c\left(\frac{1}{n_A} + \frac{1}{n_B}\right)}} \]
The structure is worth comparing with section 8.3's single-proportion standard error, which was the root of p prime q prime over n. Here one proportion serves both groups, and the single n is replaced by the sum of two reciprocals — which behaves as expected: two samples of 200 give the same standard error as a single sample of 100, since combining information is not the same as having twice as much.
Figure (svg): A bell curve with both tails shaded beyond plus and minus one point four seven
OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 523-524 — the distribution and the test statistic
Picture it
Example 10.8: is there a difference between the two medications?
Figure (svg): A bell curve with both tails shaded beyond plus and minus one point four seven
The book draws the picture on the scale of the difference itself, marking 0.04 and its mirror at minus 0.04 — the same construction section 9.5 used for a single proportion. Both tails are counted because the alternative claims a difference without saying which way.
Worked example
Example 10.8, computed.
\[ p'_A = 0.10, \; p'_B = 0.06, \; \text{SE} = 0.0271 \]
Compute the difference
Why: 0.10 minus 0.06.
\[ 0.04 \]
Divide by the standard error
Why: 0.04 over 0.0271.
\[ z = 1.47 \]
Take both tails
Why: A two-tailed test.
Find the p-value
Why: The book's value.
\[ 0.1404 \]
Figure (svg): The solution to Worked example the hives test statistic shown as a ladder of expressions, one row per legal move
\[ z = \frac{0.04}{0.0271} \approx 1.47, \quad p = 0.1404 \]
Verify: confirm the doubling was applied
Why: One tail beyond 1.47 is about 0.0702, and 0.1404 is exactly twice it — the check section 9.5 introduced for two-tailed tests. Halving a two-tailed p-value always makes the evidence look stronger, and here it would have given 0.07, still above the 1 percent level but a materially different-looking figure.
OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, p. 524
Faded example
A pooled proportion of 0.105 with samples of 100 and 100.
Fill in the blanks
\text0.00188 = \sqrt0.0434___+\frac______\right)} = \sqrt___} \approx ___
Why: 0.09398 times 0.02 gives about 0.00188, whose square root is about 0.0434 — the standard error for comparing the two valve types.
Worked example
How the standard error's structure differs from section 8.3's.
\[ \text{one sample of } 200 \text{ against two of } 200 \]
One sample of 200
Why: Root of pq over 200.
\[ 0.0192 \]
Two samples of 200
Why: Root of pq times 2 over 200.
\[ 0.0271 \]
Compare
Why: Larger by root 2.
\[ \text{about } 41 \% \]
Say why
Why: Two uncertain quantities.
Figure (svg): The solution to Worked example comparing with a single-proportion test shown as a ladder of expressions, one row per legal move
\[ \sqrt{2} \times 0.0192 \approx 0.0271 \]
Verify: confirm this matches the pattern from means
Why: Section 10.1 found the same thing: a difference of two quantities is more variable than either alone, so comparing two groups needs more data than describing one. Here the factor is exactly the square root of two when the samples are equal — so detecting a difference between two groups of 200 is roughly as hard as pinning down one group of 100.
OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 523-524
Trap
\[ \text{SE} = \sqrt{\tfrac{(0.1)(0.9)}{200} + \tfrac{(0.06)(0.94)}{200}} \approx 0.0263 \]
Build the standard error from each sample's own proportion
Why: It mirrors section 10.1's formula for means.
\[ \text{but the null says both estimate ONE proportion} \]
Ignoring that discards information the null hypothesis supplies, and gives a standard error computed under no particular hypothesis.
\[ \text{SE} = \sqrt{p_c q_c\left(\tfrac{1}{n_A} + \tfrac{1}{n_B}\right)} \approx 0.0271 \]
Pool first, because the test computes everything under the null
Why: The whole point of a p-value is what would happen IF the null held.
The unpooled version is not nonsense — it is the right standard error for a CONFIDENCE INTERVAL on the difference, where no null is being assumed. The distinction is that a test conditions on the null and an interval does not, which is why the two procedures use different formulas for what looks like the same quantity.
Two truths and a lie
All three concern the standard error.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The unpooled standard error is correct for a confidence interval on the difference, where no null is assumed. For a hypothesis test the pooled version is right, because the test computes everything under the assumption that the null holds.
Estimation
A difference of 0.04 with a standard error of 0.0271.
Predict first
Roughly how many standard errors apart are the two proportions?
Correct: About 1.5.
Why: 0.04 divided by 0.0271 is about 1.47 — under two standard errors, which for a two-tailed test gives a p-value around 0.14. Making this division before computing predicts the outcome immediately, and a statistic under about 1.96 will not reject at the 5 percent level two-tailed.
Prediction
Commit before reasoning.
Predict first
Two samples of 200 are compared. Roughly how does the standard error compare with a single sample of 200?
Correct: Larger by about root two.
Why: The single-sample formula has one over n and the two-sample formula has two over n when the sizes are equal, so the standard error grows by the square root of two — about 41 percent. Comparing two groups is genuinely harder than describing one, which is the same conclusion section 10.1 reached for means.
Section
Section 4
Concept
The two independent samples must be simple random samples that are independent. The number of successes must be at least five and the number of failures at least five for each of the samples. And growing literature states that the population must be at least ten or twenty times the size of the sample.
four counts to check — Successes and failures in each of the two samples. A group with plenty of successes may still fail on failures, and either failure makes the normal approximation unsound.
\[ x_A \ge 5, \; n_A - x_A \ge 5, \; x_B \ge 5, \; n_B - x_B \ge 5 \]
The third condition is new to this section and is worth understanding. Sampling a large fraction of a small population makes the observations less independent than the formula assumes — the same finite-population effect section 4.5 met as the hypergeometric distribution. Keeping the sample under about a tenth of the population keeps that effect negligible.
Figure (svg): The three conditions a two-proportion test requires
OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, p. 523 — the three characteristics
Picture it
What the test requires of its data.
Figure (svg): The three conditions a two-proportion test requires
The fourth line makes the connection to section 8.3 explicit: this is the np and nq check applied twice. The book states it in terms of counts rather than products, which is the same requirement in a form that can be read straight off a data table.
Worked example
Twenty and twelve successes out of two hundred each.
\[ x_A = 20, \; n_A = 200; \quad x_B = 12, \; n_B = 200 \]
A's successes
Why: Twenty.
A's failures
Why: 180.
B's successes
Why: Twelve.
B's failures
Why: 188.
Figure (svg): The solution to Worked example checking Example 10.8 shown as a ladder of expressions, one row per legal move
\[ 20, 180, 12, 188 \;\ge\; 5 \]
Verify: confirm which count would fail first if the samples shrank
Why: The smallest of the four is B's twelve successes, so halving both samples to a hundred each would bring it to six — still just acceptable — and quartering them would fail it. The scarcest category in the smaller sample is always the binding constraint, and checking that one first is the quickest route to an answer.
OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 523-524
Discrimination
Check all four counts in each case.
Sort into buckets
Sort each comparison.
Worked example
A rare outcome in a modest sample.
\[ 3 \text{ of } 80 \text{ against } 9 \text{ of } 120 \]
First sample's successes
Why: Three.
First sample's failures
Why: Seventy-seven.
Second sample
Why: Nine and 111.
Conclude
Why: One count fails.
Figure (svg): The solution to Worked example a case that fails shown as a ladder of expressions, one row per legal move
\[ x_A = 3 \;<\; 5 \]
Verify: confirm what the failure means in practice
Why: With three successes the first sample's binomial is bunched hard against zero and strongly skewed, exactly as section 9.3's figure showed — and a normal curve misstates its tail, which is where the p-value lives. The remedy is an exact method rather than a different curve, and collecting more data is the usual practical answer.
OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, p. 523
Trap
\[ 190 \text{ of } 200 \text{ and } 185 \text{ of } 200: \text{ both above five} \]
Check that each sample has enough successes
Why: The condition names successes first.
\[ \text{but the failures are } 10 \text{ and } 15 \]
Those happen to pass here, but a proportion near one fails on failures exactly as one near zero fails on successes.
\[ \text{check all FOUR counts: successes and failures, both samples} \]
The condition is symmetric in successes and failures
Why: A proportion near either boundary is skewed.
This is section 8.3's np and nq check with two groups, and the same trap: checking one product passes nearly everything. Four counts is only two more multiplications than two, and it is the difference between a check that works and one that only appears to.
Two truths and a lie
All three concern the conditions.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. Example 10.10's total is over fifteen hundred, and its smaller group still has only twelve successes — comfortably above five, but not because the total was large. A very rare outcome can fail the count even in an enormous sample.
Prediction
Commit before reasoning.
Predict first
Why does the book require the population to be at least ten or twenty times the sample?
Correct: It makes observations less independent.
Why: Drawing a large fraction without replacement is the hypergeometric situation of section 4.5, where each selection changes what remains. The binomial model assumes independence, and keeping the sample under about a tenth of the population keeps the departure negligible — which is why the book says over-sampling causes incorrect results.
Faded example
Comparing 9 of 120 with 40 of 300.
Fill in the blanks
\text9 above, \text___ ___ \text___
Why: The four counts are 9, 111, 40 and 260, so the smallest is 9 — above five, so the condition holds. Checking the scarcest category first settles most cases immediately.
Section
Section 5
Concept
With the conditions checked and the pooled standard error computed, the test proceeds as every test has since chapter 9: read the tail, find the p-value on the normal, compare with alpha, decide, and write the conclusion in context.
the unchanged procedure — Only the standard error's construction is new. The hypotheses, the tail, the decision rule and the conclusion wording are chapter 9's throughout.
\[ z \;\to\; p \;\to\; \text{compare with } \alpha \;\to\; \text{conclude} \]
Example 10.10 is worth working because it shows the method on very unequal samples and because it illustrates a rounding artefact worth recognising. The book's text reports a p-value of 0.0077 while its own calculator note reports 0.0092 with a z of 2.33 — and the difference comes from whether the sample proportions were rounded to 0.10 and 0.05 before the statistic was formed.
Figure (svg): A normal curve with the right tail beyond two point three six shaded
OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 524-527 — Examples 10.8 and 10.10 worked
Picture it
Example 10.10: 1,343 younger adults against 232 older ones.
Figure (svg): A normal curve with the right tail beyond two point three six shaded
The statistic of about 2.36 leaves roughly 0.009 in the right tail, so the null is rejected at the 5 percent level and a larger proportion of younger adults own electric vehicles. The unequal sample sizes affect the standard error but nothing about the procedure.
Worked example
Example 10.8, at the 1 percent level.
\[ 20 \text{ of } 200 \text{ against } 12 \text{ of } 200; \; \alpha = 0.01 \]
Set the hypotheses
Why: Is a difference.
Pool and form the SE
Why: 32 over 400.
\[ 0.0271 \]
Compute the statistic
Why: 0.04 over 0.0271.
\[ z = 1.47 \]
Find the p-value and decide
Why: Both tails; alpha below it.
\[ 0.1404;\text{ do not reject} \]
Figure (svg): The solution to Worked example the hives test in full shown as a ladder of expressions, one row per legal move
\[ p = 0.1404 > 0.01 \]
Verify: confirm the conclusion does not claim the medications are equivalent
Why: The book's wording is not sufficient evidence to conclude that there is a difference, which leaves open that one exists. With 200 per group the study could reliably detect only fairly large differences, so a real advantage of a few percentage points could easily have gone undetected — the same point section 9.1 made about accepting a null.
OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, p. 524
Faded example
A two-tailed test with a z of 1.4744.
Fill in the blanks
\text2 \approx 0.0702, \text0.1404 ___ \times 0.0702 = ___
Why: Doubling the one-sided area gives 0.1404, the book's p-value — well above any conventional alpha, so the null survives.
Worked example
The book prints 0.0077 in its text and 0.0092 in its calculator note.
\[ 135 \text{ of } 1343 \text{ against } 12 \text{ of } 232 \]
Use the exact counts
Why: 0.10052 and 0.05172.
\[ z = 2.359, p = 0.00915 \]
Use rounded proportions
Why: 0.10 and 0.05.
\[ z = 2.417, p = 0.00781 \]
Compare with the book
Why: 0.0092 and 0.0077.
Say which is accurate
Why: The exact counts.
\[ 0.0092 \]
Figure (svg): The solution to Worked example resolving Example 10.10's two p-values shown as a ladder of expressions, one row per legal move
\[ p_{\text{exact}} \approx 0.0092, \quad p_{\text{rounded}} \approx 0.0078 \]
Verify: confirm nothing turns on it here, and when it would
Why: Both figures reject comfortably at the 5 percent level, so the conclusion is unaffected. The pattern is worth recognising because it recurs — Examples 8.3, 8.8 and 9.15 all show it — and because a p-value near a threshold could be pushed across by exactly this kind of intermediate rounding. Carrying full precision until the final step removes the question.
OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 526-527
Trap
\[ p'_Y = 0.10, \; p'_O = 0.05 \;\Rightarrow\; z = 2.42, \; p = 0.0078 \]
Round the sample proportions to two decimals first
Why: They are close to those values.
\[ \text{the exact counts give } z = 2.36, \; p = 0.0092 \]
Rounding 0.05172 down to 0.05 exaggerates the gap by about 3 percent, and the p-value moves by nearly a fifth.
\[ \text{use } \tfrac{135}{1343} \text{ and } \tfrac{12}{232} \text{ throughout} \]
Carry full precision until the final answer
Why: Round once, at the end.
A shift from 0.0092 to 0.0078 changes nothing at the 5 percent level, but the same proportional shift applied to a p-value near 0.05 could reverse a decision. Rounding at the end costs nothing and removes an entire class of question about whether an answer is right.
Two truths and a lie
All three concern running the test.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and it is section 9.1's rule about accepting a null. With 200 per group the hives study could only detect fairly large differences, so a real gap of a few points could easily survive undetected.
Estimation
Example 10.8's standard error is 0.0271, at a two-tailed 5 percent level.
Predict first
Roughly what difference in proportions would the study reliably detect?
Correct: About 0.053.
Why: The critical value of 1.96 times the standard error of 0.0271 gives about 0.053. The observed gap of 0.04 falls short of that, which is precisely why the test did not reject — and reporting this figure alongside a non-result tells a reader what the study was capable of.
Explain it
A classmate is confused that the book gives both 0.0077 and 0.0092 for Example 10.10.
Discussion prompt
In two sentences or fewer, explain.
Hint: Ask what proportions each figure was computed from.
Answer:
The 0.0077 comes from rounding the sample proportions to 0.10 and 0.05 before computing, while 0.0092 comes from the exact counts of 135 out of 1,343 and 12 out of 232.
The calculator's 0.0092 is the accurate one, and both reject at 5 percent — but carrying full precision until the end avoids the discrepancy entirely.
Comparison
Fill the blanks. The pooling rule reverses between the first two and the third.
Comparison matrix
| Test | Standard error | Distribution |
|---|---|---|
| Two means, sigma unknown | each s-squared over its own n, added | Student t, Aspin-Welch df |
| Two means, sigma known | each sigma-squared over its own n, added | normal |
| Two proportions | one pooled p, times the summed reciprocals | normal |
| Why pooling differs | a null on means leaves spreads free | a null on proportions fixes them |
The last row is the idea worth carrying out of the chapter. Two rules that look contradictory follow from one principle: pool exactly what the null hypothesis says is common, and nothing else.
Pattern
Six steps, and the third is the only new one.
Use the exact counts throughout rather than rounded proportions, and round only the final answer.
OpenStax Introductory Business Statistics 2e, §10.4 Comparing Two Independent Population Proportions §10.4 Comparing Two Independent Population Proportions
Check
The pooled proportion.
Check your understanding
For 20 of 200 and 12 of 200, what is the pooled proportion?
Answer: A
Why: Total successes over total trials: 32 divided by 400 is 0.08, midway between the two sample proportions since the sample sizes are equal.
Check
Why pooling applies here.
Check your understanding
Why are the samples pooled here when section 10.1 forbade pooling?
Answer: A
Why: A binomial's variability follows from its probability, so equal proportions means equal spreads — and both samples estimate one common value.
Check
The conditions.
Check your understanding
A comparison has 196 successes out of 200 in one group. Does the condition hold?
Answer: A
Why: Both successes and failures must be at least five in each sample, and 200 minus 196 is only four.
Real world
A hospital compares infection rates after two sterilisation protocols. Protocol A gives 8 infections in 400 procedures; protocol B gives 3 in 150. An administrator computes the rates as 2.0 percent and 2.0 percent, concludes they are identical, and proposes adopting whichever is cheaper.
Discussion prompt
Assess the reasoning, and say what the comparison can and cannot establish.
Hint: Check the conditions before checking the arithmetic.
Answer:
The rates really are close, and the arithmetic is right. Eight of 400 is exactly 2.0 percent and three of 150 is exactly 2.0 percent, so the observed difference is zero and no test would reject — the statistic is zero and the p-value is one.
But the second sample fails the conditions. Protocol B has only three infections, below the required five, so the normal approximation is unsound. The test that produced the reassuring answer should not have been run in this form at all.
\[ x_B = 3 \;<\; 5, \qquad \text{so the normal approximation does not apply} \]
More importantly, identical observed rates do not establish identical true rates. With these sample sizes the standard error of the difference is about 1.4 percentage points, so the study could only reliably detect a difference of around 2.8 points — larger than either rate. Protocol B's true rate could plausibly be double protocol A's and this comparison would not have seen it.
Two practical points follow. Rare-outcome comparisons need very large samples: to detect a doubling from 2 percent to 4 percent with reasonable power would take well over a thousand procedures per protocol, which is a design fact worth knowing before the study rather than after. And an exact method — Fisher's test or an exact binomial approach — would be the right analysis here, since it needs no count condition; the administrator's conclusion may even turn out to be defensible, but not on this evidence.
Commit first
Answer, then rate your confidence honestly.
Predict first
Why does a two-proportion test pool the samples when a two-mean test does not?
Correct: Because the null fixes the spreads too.
\[ p_c = \frac{x_1+x_2}{n_1+n_2}, \qquad \text{SE} = \sqrt{p_c q_c\left(\tfrac{1}{n_1}+\tfrac{1}{n_2}\right)} \]
Why: A binomial's variability is p times q, determined entirely by its probability — so if the null says the two proportions are equal, it says the two spreads are equal as well, and one pooled estimate is correct. A null about means constrains only the centres and leaves the two standard deviations free, which is why section 10.1 forbids pooling. One principle, two opposite instructions.
Explain it
They built the standard error from each sample's own proportion, as in section 10.1.
Discussion prompt
In two sentences or fewer, correct them.
Hint: Ask what their null hypothesis says about the two spreads.
Answer:
Their null says the two proportions are equal, and a binomial's spread comes from its proportion — so the null says the two spreads are equal as well.
Since a test computes everything under the null, both samples should be pooled into one estimate of that common proportion before the standard error is formed.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, pool exactly what the null says is common. For the second, total successes over total trials, then root of pc qc times the summed reciprocals. For the third, successes and failures in BOTH samples. For the fourth, double the one-sided area whenever the alternative says different. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top, draw two boxes for the two samples feeding into one pooled box, and write the pooled proportion formula beneath. Beside it, write two columns headed two means and two proportions, and in each write what the null constrains and whether pooling follows — this is the one idea the lesson turns on. In the middle of the page, work Example 10.8 completely: the two sample proportions of 0.10 and 0.06, the pooled 0.08, the standard error of 0.0271, the statistic of 1.47, a normal curve with BOTH tails shaded, the doubled p-value of 0.1404, and a conclusion naming the 1 percent level. Beneath it, write the four counts and confirm each is at least five. To the right, work Example 10.10 with the exact counts to get 0.0092, then redo it with the proportions rounded to 0.10 and 0.05 to get 0.0078 — and write one sentence on which is accurate and why the book prints both. At the bottom, list the three conditions and mark which one is new to this section.
Check your pooled proportion by confirming it lies between the two sample proportions, and exactly midway when the sample sizes are equal. Check the p-value by halving it and confirming the result matches the single-tail area, which is the guard against forgetting to double.
Recap
Five things, and the second is the idea that distinguishes this section.
| If you see | Then |
|---|---|
| Categorical data in two groups | A difference of proportions |
| A null of equal proportions | Pool the successes and trials |
| A null of equal means | Do NOT pool: see section 10.1 |
| Fewer than five in any of four counts | The normal approximation is unsound |
| A sample near a tenth of its population | The independence assumption is straining |
| A not-equal alternative | Double the one-sided area |
| Rounded proportions in a calculation | Recompute from the exact counts |
Section 10.4 closes the chapter with a design rather than a new parameter. When the two measurements come from the same subjects, the samples are not independent and none of this chapter's formulas apply — but taking the differences collapses the problem to a one-sample t test on those differences, which chapter 9 already covered.
OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 523-527 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.