10.3 Comparing Two Independent Population Proportions

The same comparison as the previous two sections, with categorical data. The parameter is the difference of two population proportions, the point estimate is the difference of the two sample proportions, and because the difference of two proportions follows an approximate normal distribution, the test statistic is a z-score with no degrees of freedom. One feature is genuinely new and is the section's central idea: because the null hypothesis generally states that the two proportions are the same, the two samples are estimating one common value, so their successes and their trials are combined into a single pooled proportion before the standard error is formed. That is the opposite of the instruction in section 10.1, and the reason is that a binomial's spread follows from its probability — so a null of equal proportions is also a null of equal spreads, while a null of equal means says nothing about the two standard deviations.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 10.3 Comparing Two Independent Population Proportions

Title

Statistics · Chapter 10 — Hypothesis Testing with Two Samples

Comparing Two Independent Population Proportions

2. By the end of this lesson you can

Objectives

Five outcomes, and the second explains why this section pools where section 10.1 forbade it.

OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 523-527 — the section these objectives are drawn from

3. What you already have

Warm-up

Section 10.1 was emphatic that two sample variances must not be pooled.

Discussion prompt

A null hypothesis says two population proportions are equal. If that is true, are the two populations' standard deviations also equal?

Hint: Section 8.3 established that a binomial's spread comes from its probability.

Answer:

Yes, necessarily. A proportion's variability is p times q over n, and q is one minus p — so fixing p fixes the spread completely. Two populations with the same proportion cannot differ in variability.

That is not true for means. Two populations can have identical means and wildly different standard deviations, which is exactly why section 10.1 refused to pool: the null said nothing about the spreads, so nothing licensed combining them.

Here the null does say something about the spreads, because it says everything about the proportions. So the two samples really are estimating one common value, and pooling them is the correct thing to do rather than a shortcut.

4. Under the null, both samples estimate the same proportion

Concept

Comparing two proportions is common, and if two estimated proportions differ it may be due to a difference in the populations or it may be due to chance. Generally the null hypothesis states that the two proportions are the same, and to conduct the test we use a pooled proportion.

the pooled proportion — Written pc, and computed as the total successes over the total trials across both samples. It is the single best estimate of the common proportion the null asserts.

\[ p_c = \frac{x_1 + x_2}{n_1 + n_2}, \qquad \text{SE} = \sqrt{p_c q_c\left(\frac{1}{n_1} + \frac{1}{n_2}\right)} \]

The logic is worth stating carefully because it is the reverse of section 10.1's. There the null constrained only the means, so the two spreads had to be estimated separately. Here the null constrains the proportions, and since a proportion determines its own spread, the null constrains the spreads too — which is exactly what makes one combined estimate the right one.

Figure (svg): A card showing the successes of two samples combined into a single pooled proportion

The successes are added and the trials are added, giving one estimate of the common proportion the null asserts.

OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, p. 523

5. Setting up the comparison

Section

Section 1

6. A difference of proportions, tested against zero

Concept

Let A and B be the subscripts for the two groups. The random variable is the difference of the two sample proportions, and the null generally states that the population proportions are equal — equivalently that their difference is zero.

the two notations again — H-nought as p-A equals p-B, or as p-A minus p-B equals zero. As in section 10.1, the second makes explicit that the null names a single value for the parameter.

\[ H_0: p_A = p_B \;\Longleftrightarrow\; H_0: p_A - p_B = 0 \]

The recognition test from section 8.3 applies unchanged: the data are categorical, the underlying distributions are binomial, and no mean is mentioned. What is new is only that there are two of them, so the question becomes whether one binomial's probability differs from another's rather than whether one differs from a fixed number.

Figure (svg): Two columns contrasting why means are not pooled with why proportions are

The two sections give opposite instructions for a reason. A binomial's spread follows from its probability, so a null of equal proportions is also a null of equal spreads.

OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, p. 523 — Example 10.8's hypotheses

7. Why the two sections disagree about pooling

Picture it

The instruction reverses, and the reason is in the null.

Figure (svg): Two columns contrasting why means are not pooled with why proportions are

The two sections give opposite instructions for a reason. A binomial's spread follows from its probability, so a null of equal proportions is also a null of equal spreads.

The second row of each column is the whole explanation. A null about means leaves the spreads free; a null about proportions does not, because a binomial has no spread parameter independent of its probability. Two apparently contradictory rules follow from one principle applied to two different situations.

8. Worked example: setting up Example 10.8

Worked example

Two medications for hives, tested for a difference.

\[ \text{is there a DIFFERENCE in the proportions still reacting?} \]

Identify the parameter

Why: Proportions still reacting.

Read the claim

Why: Is a difference.

Write the alternative

Why: Two-sided.

\[ H a: p A \ne p B \]

Write the null

Why: Its contradiction.

\[ H 0: p A = p B \]

Figure (svg): The solution to Worked example setting up Example 10.8 shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ H_0: p_A = p_B, \qquad H_a: p_A \ne p_B \]

Verify: confirm the parameter is a proportion rather than a mean

Why: The data record whether each adult still had hives after thirty minutes — a yes-or-no outcome counted across a sample, with no average of any measured quantity. Section 8.3's test applies: the underlying distribution is binomial and no mean is mentioned. Had the study measured how long the hives lasted, the parameter would have been a mean and section 10.1's test would apply instead.

OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, p. 523

9. Two means or two proportions?

Sorting

Ask what each individual contributes.

Sort into buckets

Sort each comparison.

Two means
hours of sport per day, girls against boys; months two floor waxes last
Two proportions
the fraction still reacting to two medications; the percent of two age groups owning electric vehicles; the proportion of two valve types cracking under pressure
mean
Each individual contributes a measured quantity, and the comparison is of averages.
prop
Each individual falls into a category, and the comparison is of fractions.

The distinction is exactly section 8.3's, applied twice over. Getting it wrong sends the whole problem to the wrong standard error and the wrong pooling rule, so it is worth settling before anything else.

10. Worked example: a directional comparison

Worked example

Example 10.10, where the claim points one way.

\[ \text{do MORE younger adults own electric vehicles than older adults?} \]

Identify the groups

Why: Younger and older.

Read the claim

Why: More younger.

Write the alternative

Why: The claim.

\[ H a: p Y > p O \]

Read the tail

Why: Greater than.

Figure (svg): The solution to Worked example a directional comparison shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ H_0: p_Y \le p_O, \qquad H_a: p_Y > p_O \]

Verify: confirm which group is subtracted from which

Why: Writing the alternative as p-Y greater than p-O means the difference p-Y minus p-O is positive, so the right tail is wanted and the sample proportions must be subtracted in that order. Reversing them would flip the statistic's sign and send the test to the wrong tail — the same discipline section 10.1 required, and equally consequential here.

OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 526-527

11. Trap: treating a count as a mean

Trap

The trap

\[ 20 \text{ of } 200 \text{ and } 12 \text{ of } 200 \;\Rightarrow\; \text{compare mean reactions} \]

Treat the counts as measurements to average

Why: They are numbers, after all.

\[ \text{but each subject contributes a yes or a no} \]

There is no measured quantity to average; each person is in a category, and the parameter is the fraction in it.

The fix

\[ p'_A = \tfrac{20}{200} = 0.10, \quad p'_B = \tfrac{12}{200} = 0.06 \]

Form proportions and compare those

Why: Section 8.3's recognition test, with two groups.

The tell is what each individual contributes. A measured quantity per subject — a time, a weight, a score — gives a mean; a category per subject gives a proportion. Counts of subjects in a category are the raw material of a proportion, not observations to be averaged.

12. The null as a difference

Fill the middle

The equivalent form.

Fill in the blanks

H_0: p_A = p_B \text0 p_A - p_B = ___

Why: Zero. As in section 10.1, writing the null this way shows that a single value is being named, which is what allows a sampling distribution to be centred on it.

13. One of these is false

Two truths and a lie

All three concern the setup.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The underlying distributions are binomial
  • C. The difference of two proportions is approximately normal
  • B. The comparison requires equal sample sizes

Survives elimination: B

Why: The survivor is false. Example 10.10 compares 1,343 with 232 and works perfectly well; the standard error formula takes each sample size separately. Equal sizes are convenient but never required by any test in this chapter.

14. Which subtraction order?

Prediction

Commit before reasoning.

Predict first

For H-a that p-Y exceeds p-O, which difference should be computed?

  • p-Y minus p-O, so a positive result supports the claim
  • p-O minus p-Y
  • Either, since the test is symmetric
  • The absolute difference

Correct: p-Y minus p-O.

Why: The alternative claims that difference is positive, so computing it that way puts the evidence in the right tail. Reversing it would flip the sign and the test would look in the wrong tail, reporting one minus the correct p-value — an error that turns a rejection into a comfortable non-rejection.

15. The pooled proportion

Section

Section 2

16. One estimate of the value the null asserts

Concept

The pooled proportion combines the successes from both samples over the trials from both samples. It is the single best estimate of the common proportion the null claims the two populations share.

pc and qc — The pooled proportion and its complement. Since the null says both populations share a proportion, one estimate built from all the data is better than either sample's alone.

\[ p_c = \frac{x_A + x_B}{n_A + n_B}, \qquad q_c = 1 - p_c \]

It is worth noticing that the pooled proportion is used only for the STANDARD ERROR, never for the point estimate. The difference at the top of the statistic is still the difference of the two separate sample proportions — because that is what was observed. Pooling belongs to the denominator, where the null's assumption is being applied.

Figure (svg): A card showing the successes of two samples combined into a single pooled proportion

The successes are added and the trials are added, giving one estimate of the common proportion the null asserts.

OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, p. 523 — the pooled proportion and the distribution

17. Two samples, one estimate

Picture it

The successes and the trials both combined.

Figure (svg): A card showing the successes of two samples combined into a single pooled proportion

The successes are added and the trials are added, giving one estimate of the common proportion the null asserts.

The last line of the figure is the one to carry. Pooling here is not a convenience or an approximation — it follows from what the null actually says, and refusing to pool would mean ignoring information the null hypothesis provides.

18. Worked example: pooling for Example 10.8

Worked example

Twenty of two hundred against twelve of two hundred.

\[ x_A = 20, \; n_A = 200; \quad x_B = 12, \; n_B = 200 \]

Add the successes

Why: Twenty plus twelve.

\[ 32 \]

Add the trials

Why: Two hundred each.

\[ 400 \]

Divide

Why: The pooled proportion.

\[ 0.08 \]

Form the standard error

Why: Root of pc qc times the reciprocals.

\[ 0.02713 \]

Figure (svg): The solution to Worked example pooling for Example 10.8 shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ p_c = \frac{32}{400} = 0.08, \quad \text{SE} = \sqrt{(0.08)(0.92)\left(\tfrac{1}{200}+\tfrac{1}{200}\right)} \approx 0.0271 \]

Verify: confirm the pooled value lies between the two sample proportions

Why: The two samples gave 0.10 and 0.06, and the pooled 0.08 sits between them — as a weighted average of the two must. A pooled proportion outside that range would signal an arithmetic error, and with equal sample sizes it lands exactly midway, as it does here.

OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 523-524

19. Compute the pooled proportion

Faded example

Fifteen of 100 valve A cracked; six of 100 valve B cracked.

Fill in the blanks

p_c = \frac210.105 = \frac___}___ = ___

Why: Twenty-one successes in two hundred trials gives a pooled proportion of 0.105, which sits between the two sample proportions of 0.15 and 0.06.

20. Worked example: pooling with unequal samples

Worked example

Example 10.10, where one group is nearly six times the other.

\[ 135 \text{ of } 1343; \quad 12 \text{ of } 232 \]

Add the successes

Why: 135 plus 12.

\[ 147 \]

Add the trials

Why: 1343 plus 232.

\[ 1575 \]

Divide

Why: The pooled proportion.

\[ 0.0933 \]

Compare with the samples

Why: 0.1005 and 0.0517.

Figure (svg): The solution to Worked example pooling with unequal samples shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ p_c = \frac{147}{1575} \approx 0.0933 \]

Verify: confirm why the larger sample dominates

Why: The pooled proportion is a weighted average with the sample sizes as weights, so a group of 1,343 pulls it far more than one of 232. That is correct: the larger sample carries more information about the common proportion the null asserts, and the pooled estimate should reflect that. With equal samples the weighting disappears and the pooled value is the simple midpoint.

OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 526-527

21. Error analysis: four standard errors for Example 10.8

Error analysis

The correct value is about 0.0271.

Annotate

On: \( \begin{aligned} &(1)\; \sqrt{\tfrac{(0.1)(0.9)}{200} + \tfrac{(0.06)(0.94)}{200}} \approx 0.0263 \\ &(2)\; \sqrt{(0.08)(0.92)\left(\tfrac{1}{200}\right)} \approx 0.0192 \\ &(3)\; (0.08)(0.92)\left(\tfrac{1}{200}+\tfrac{1}{200}\right) \approx 0.00074 \\ &(4)\; \sqrt{(0.08)(0.92)\left(\tfrac{1}{200}+\tfrac{1}{200}\right)} \approx 0.0271 \end{aligned} \)

  • (1) keeps the two proportions separate rather than pooling. It gives a plausible number and ignores what the null tells us.
  • (2) uses only one sample size, halving the term inside the root.
  • (3) omits the square root, giving the variance rather than the standard error.
  • (4) is correct: the pooled proportion, times its complement, times the sum of the reciprocal sample sizes.

Error (1) is the interesting one because it is section 10.1's method applied here. It is not absurd — it is a legitimate standard error for a confidence interval on the difference — but for a TEST, where the null supplies a common proportion, pooling is the correct choice.

22. One of these is false

Two truths and a lie

All three concern pooling.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The pooled proportion lies between the two sample proportions
  • C. Pooling is used for the standard error, not the point estimate
  • B. The pooled proportion replaces both sample proportions everywhere

Survives elimination: B

Why: The survivor is false. The statistic's numerator is the observed difference of the two separate sample proportions — that is the evidence. Pooling applies only to the denominator, where the null's claim of a common proportion is being used.

23. Which sample dominates?

Prediction

Commit before reasoning.

Predict first

Samples of 1,343 and 232 give proportions of 0.10 and 0.05. Where does the pooled value sit?

  • Much closer to 0.10, since the larger sample carries more weight
  • Exactly midway at 0.075
  • Much closer to 0.05
  • Outside the range of the two

Correct: Much closer to 0.10.

Why: The pooled proportion weights each sample by its size, so a group nearly six times larger dominates — the value comes out at about 0.0933. Only with equal sample sizes does the simple midpoint apply, which is why Example 10.8's pooled 0.08 is exactly between 0.10 and 0.06.

24. Explain the reversal

Explain it

A classmate asks why section 10.1 forbade pooling and this section requires it.

Discussion prompt

In two sentences or fewer, explain.

Hint: Ask what each null hypothesis actually constrains.

Answer:

A null about two means says nothing about their spreads, so the two standard deviations have to be estimated separately.

A null about two proportions constrains the spreads too, because a binomial's variability is determined by its probability — so under the null both samples really are estimating one common value.

25. The standard error and the statistic

Section

Section 3

26. One proportion, two sample sizes

Concept

The standard error is the square root of the pooled proportion times its complement, multiplied by the sum of the reciprocals of the two sample sizes. The test statistic is the observed difference of sample proportions divided by it.

the two-proportion standard error — Built from a single pooled proportion and both sample sizes. The reciprocals are summed because each sample contributes uncertainty in inverse proportion to its size.

\[ z = \frac{p'_A - p'_B}{\sqrt{p_c q_c\left(\frac{1}{n_A} + \frac{1}{n_B}\right)}} \]

The structure is worth comparing with section 8.3's single-proportion standard error, which was the root of p prime q prime over n. Here one proportion serves both groups, and the single n is replaced by the sum of two reciprocals — which behaves as expected: two samples of 200 give the same standard error as a single sample of 100, since combining information is not the same as having twice as much.

Figure (svg): A bell curve with both tails shaded beyond plus and minus one point four seven

The book: half the p-value is below minus 0.04, and half is above 0.04, on the scale of the difference in proportions.

OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 523-524 — the distribution and the test statistic

27. A two-tailed test on the difference

Picture it

Example 10.8: is there a difference between the two medications?

Figure (svg): A bell curve with both tails shaded beyond plus and minus one point four seven

The book: half the p-value is below minus 0.04, and half is above 0.04, on the scale of the difference in proportions.

The book draws the picture on the scale of the difference itself, marking 0.04 and its mirror at minus 0.04 — the same construction section 9.5 used for a single proportion. Both tails are counted because the alternative claims a difference without saying which way.

28. Worked example: the hives test statistic

Worked example

Example 10.8, computed.

\[ p'_A = 0.10, \; p'_B = 0.06, \; \text{SE} = 0.0271 \]

Compute the difference

Why: 0.10 minus 0.06.

\[ 0.04 \]

Divide by the standard error

Why: 0.04 over 0.0271.

\[ z = 1.47 \]

Take both tails

Why: A two-tailed test.

Find the p-value

Why: The book's value.

\[ 0.1404 \]

Figure (svg): The solution to Worked example the hives test statistic shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ z = \frac{0.04}{0.0271} \approx 1.47, \quad p = 0.1404 \]

Verify: confirm the doubling was applied

Why: One tail beyond 1.47 is about 0.0702, and 0.1404 is exactly twice it — the check section 9.5 introduced for two-tailed tests. Halving a two-tailed p-value always makes the evidence look stronger, and here it would have given 0.07, still above the 1 percent level but a materially different-looking figure.

OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, p. 524

29. Compute the standard error

Faded example

A pooled proportion of 0.105 with samples of 100 and 100.

Fill in the blanks

\text0.00188 = \sqrt0.0434___+\frac______\right)} = \sqrt___} \approx ___

Why: 0.09398 times 0.02 gives about 0.00188, whose square root is about 0.0434 — the standard error for comparing the two valve types.

30. Worked example: comparing with a single-proportion test

Worked example

How the standard error's structure differs from section 8.3's.

\[ \text{one sample of } 200 \text{ against two of } 200 \]

One sample of 200

Why: Root of pq over 200.

\[ 0.0192 \]

Two samples of 200

Why: Root of pq times 2 over 200.

\[ 0.0271 \]

Compare

Why: Larger by root 2.

\[ \text{about } 41 \% \]

Say why

Why: Two uncertain quantities.

Figure (svg): The solution to Worked example comparing with a single-proportion test shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \sqrt{2} \times 0.0192 \approx 0.0271 \]

Verify: confirm this matches the pattern from means

Why: Section 10.1 found the same thing: a difference of two quantities is more variable than either alone, so comparing two groups needs more data than describing one. Here the factor is exactly the square root of two when the samples are equal — so detecting a difference between two groups of 200 is roughly as hard as pinning down one group of 100.

OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 523-524

31. Trap: using the separate proportions in the standard error

Trap

The trap

\[ \text{SE} = \sqrt{\tfrac{(0.1)(0.9)}{200} + \tfrac{(0.06)(0.94)}{200}} \approx 0.0263 \]

Build the standard error from each sample's own proportion

Why: It mirrors section 10.1's formula for means.

\[ \text{but the null says both estimate ONE proportion} \]

Ignoring that discards information the null hypothesis supplies, and gives a standard error computed under no particular hypothesis.

The fix

\[ \text{SE} = \sqrt{p_c q_c\left(\tfrac{1}{n_A} + \tfrac{1}{n_B}\right)} \approx 0.0271 \]

Pool first, because the test computes everything under the null

Why: The whole point of a p-value is what would happen IF the null held.

The unpooled version is not nonsense — it is the right standard error for a CONFIDENCE INTERVAL on the difference, where no null is being assumed. The distinction is that a test conditions on the null and an interval does not, which is why the two procedures use different formulas for what looks like the same quantity.

32. One of these is false

Two truths and a lie

All three concern the standard error.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The pooled proportion appears only in the denominator
  • C. The reciprocals of both sample sizes are summed
  • B. The unpooled formula is simply wrong in every context

Survives elimination: B

Why: The survivor is false. The unpooled standard error is correct for a confidence interval on the difference, where no null is assumed. For a hypothesis test the pooled version is right, because the test computes everything under the assumption that the null holds.

33. Predict the outcome

Estimation

A difference of 0.04 with a standard error of 0.0271.

Predict first

Roughly how many standard errors apart are the two proportions?

  • About 1.5
  • About 4
  • About 0.7
  • About 15

Correct: About 1.5.

Why: 0.04 divided by 0.0271 is about 1.47 — under two standard errors, which for a two-tailed test gives a p-value around 0.14. Making this division before computing predicts the outcome immediately, and a statistic under about 1.96 will not reject at the 5 percent level two-tailed.

34. How much harder is a comparison?

Prediction

Commit before reasoning.

Predict first

Two samples of 200 are compared. Roughly how does the standard error compare with a single sample of 200?

  • Larger by a factor of about the square root of two
  • Half as large
  • The same
  • Twice as large

Correct: Larger by about root two.

Why: The single-sample formula has one over n and the two-sample formula has two over n when the sizes are equal, so the standard error grows by the square root of two — about 41 percent. Comparing two groups is genuinely harder than describing one, which is the same conclusion section 10.1 reached for means.

35. The three conditions

Section

Section 4

36. Independence, five each way, and a population large enough

Concept

The two independent samples must be simple random samples that are independent. The number of successes must be at least five and the number of failures at least five for each of the samples. And growing literature states that the population must be at least ten or twenty times the size of the sample.

four counts to check — Successes and failures in each of the two samples. A group with plenty of successes may still fail on failures, and either failure makes the normal approximation unsound.

\[ x_A \ge 5, \; n_A - x_A \ge 5, \; x_B \ge 5, \; n_B - x_B \ge 5 \]

The third condition is new to this section and is worth understanding. Sampling a large fraction of a small population makes the observations less independent than the formula assumes — the same finite-population effect section 4.5 met as the hypergeometric distribution. Keeping the sample under about a tenth of the population keeps that effect negligible.

Figure (svg): The three conditions a two-proportion test requires

The third condition is new to this section, and it is why national surveys of small populations need care.

OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, p. 523 — the three characteristics

37. Five lines, three conditions

Picture it

What the test requires of its data.

Figure (svg): The three conditions a two-proportion test requires

The third condition is new to this section, and it is why national surveys of small populations need care.

The fourth line makes the connection to section 8.3 explicit: this is the np and nq check applied twice. The book states it in terms of counts rather than products, which is the same requirement in a form that can be read straight off a data table.

38. Worked example: checking Example 10.8

Worked example

Twenty and twelve successes out of two hundred each.

\[ x_A = 20, \; n_A = 200; \quad x_B = 12, \; n_B = 200 \]

A's successes

Why: Twenty.

A's failures

Why: 180.

B's successes

Why: Twelve.

B's failures

Why: 188.

Figure (svg): The solution to Worked example checking Example 10.8 shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 20, 180, 12, 188 \;\ge\; 5 \]

Verify: confirm which count would fail first if the samples shrank

Why: The smallest of the four is B's twelve successes, so halving both samples to a hundred each would bring it to six — still just acceptable — and quartering them would fail it. The scarcest category in the smaller sample is always the binding constraint, and checking that one first is the quickest route to an answer.

OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 523-524

39. Do the conditions hold?

Discrimination

Check all four counts in each case.

Sort into buckets

Sort each comparison.

All four counts at least five
20 of 200 against 12 of 200; 15 of 100 against 6 of 100; 135 of 1343 against 12 of 232
At least one count fails
3 of 80 against 9 of 120; 196 of 200 against 190 of 200
ok
Both samples have at least five successes and at least five failures.
no
One category is too scarce — either too few successes or too few failures.

40. Worked example: a case that fails

Worked example

A rare outcome in a modest sample.

\[ 3 \text{ of } 80 \text{ against } 9 \text{ of } 120 \]

First sample's successes

Why: Three.

First sample's failures

Why: Seventy-seven.

Second sample

Why: Nine and 111.

Conclude

Why: One count fails.

Figure (svg): The solution to Worked example a case that fails shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ x_A = 3 \;<\; 5 \]

Verify: confirm what the failure means in practice

Why: With three successes the first sample's binomial is bunched hard against zero and strongly skewed, exactly as section 9.3's figure showed — and a normal curve misstates its tail, which is where the p-value lives. The remedy is an exact method rather than a different curve, and collecting more data is the usual practical answer.

OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, p. 523

41. Trap: checking only the successes

Trap

The trap

\[ 190 \text{ of } 200 \text{ and } 185 \text{ of } 200: \text{ both above five} \]

Check that each sample has enough successes

Why: The condition names successes first.

\[ \text{but the failures are } 10 \text{ and } 15 \]

Those happen to pass here, but a proportion near one fails on failures exactly as one near zero fails on successes.

The fix

\[ \text{check all FOUR counts: successes and failures, both samples} \]

The condition is symmetric in successes and failures

Why: A proportion near either boundary is skewed.

This is section 8.3's np and nq check with two groups, and the same trap: checking one product passes nearly everything. Four counts is only two more multiplications than two, and it is the difference between a check that works and one that only appears to.

42. One of these is false

Two truths and a lie

All three concern the conditions.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Four counts must be checked, not two
  • C. The population should be at least ten times the sample
  • B. A large total sample size guarantees the conditions hold

Survives elimination: B

Why: The survivor is false. Example 10.10's total is over fifteen hundred, and its smaller group still has only twelve successes — comfortably above five, but not because the total was large. A very rare outcome can fail the count even in an enormous sample.

43. Why must the population be much larger?

Prediction

Commit before reasoning.

Predict first

Why does the book require the population to be at least ten or twenty times the sample?

  • Sampling a large fraction of a population makes observations less independent
  • Small populations are harder to survey
  • The formula divides by the population size
  • It ensures the sample is random

Correct: It makes observations less independent.

Why: Drawing a large fraction without replacement is the hypergeometric situation of section 4.5, where each selection changes what remains. The binomial model assumes independence, and keeping the sample under about a tenth of the population keeps the departure negligible — which is why the book says over-sampling causes incorrect results.

44. Find the binding count

Faded example

Comparing 9 of 120 with 40 of 300.

Fill in the blanks

\text9 above, \text___ ___ \text___

Why: The four counts are 9, 111, 40 and 260, so the smallest is 9 — above five, so the condition holds. Checking the scarcest category first settles most cases immediately.

45. Running the test

Section

Section 5

46. Six steps, with a pooled denominator

Concept

With the conditions checked and the pooled standard error computed, the test proceeds as every test has since chapter 9: read the tail, find the p-value on the normal, compare with alpha, decide, and write the conclusion in context.

the unchanged procedure — Only the standard error's construction is new. The hypotheses, the tail, the decision rule and the conclusion wording are chapter 9's throughout.

\[ z \;\to\; p \;\to\; \text{compare with } \alpha \;\to\; \text{conclude} \]

Example 10.10 is worth working because it shows the method on very unequal samples and because it illustrates a rounding artefact worth recognising. The book's text reports a p-value of 0.0077 while its own calculator note reports 0.0092 with a z of 2.33 — and the difference comes from whether the sample proportions were rounded to 0.10 and 0.05 before the statistic was formed.

Figure (svg): A normal curve with the right tail beyond two point three six shaded

The p-value is about 0.009, so at the 5 percent level the null is rejected — a larger proportion of younger adults own electric vehicles.

OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 524-527 — Examples 10.8 and 10.10 worked

47. A right-tailed test on very unequal samples

Picture it

Example 10.10: 1,343 younger adults against 232 older ones.

Figure (svg): A normal curve with the right tail beyond two point three six shaded

The p-value is about 0.009, so at the 5 percent level the null is rejected — a larger proportion of younger adults own electric vehicles.

The statistic of about 2.36 leaves roughly 0.009 in the right tail, so the null is rejected at the 5 percent level and a larger proportion of younger adults own electric vehicles. The unequal sample sizes affect the standard error but nothing about the procedure.

48. Worked example: the hives test in full

Worked example

Example 10.8, at the 1 percent level.

\[ 20 \text{ of } 200 \text{ against } 12 \text{ of } 200; \; \alpha = 0.01 \]

Set the hypotheses

Why: Is a difference.

Pool and form the SE

Why: 32 over 400.

\[ 0.0271 \]

Compute the statistic

Why: 0.04 over 0.0271.

\[ z = 1.47 \]

Find the p-value and decide

Why: Both tails; alpha below it.

\[ 0.1404;\text{ do not reject} \]

Figure (svg): The solution to Worked example the hives test in full shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ p = 0.1404 > 0.01 \]

Verify: confirm the conclusion does not claim the medications are equivalent

Why: The book's wording is not sufficient evidence to conclude that there is a difference, which leaves open that one exists. With 200 per group the study could reliably detect only fairly large differences, so a real advantage of a few percentage points could easily have gone undetected — the same point section 9.1 made about accepting a null.

OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, p. 524

49. Complete the test

Faded example

A two-tailed test with a z of 1.4744.

Fill in the blanks

\text2 \approx 0.0702, \text0.1404 ___ \times 0.0702 = ___

Why: Doubling the one-sided area gives 0.1404, the book's p-value — well above any conventional alpha, so the null survives.

50. Worked example: resolving Example 10.10's two p-values

Worked example

The book prints 0.0077 in its text and 0.0092 in its calculator note.

\[ 135 \text{ of } 1343 \text{ against } 12 \text{ of } 232 \]

Use the exact counts

Why: 0.10052 and 0.05172.

\[ z = 2.359, p = 0.00915 \]

Use rounded proportions

Why: 0.10 and 0.05.

\[ z = 2.417, p = 0.00781 \]

Compare with the book

Why: 0.0092 and 0.0077.

Say which is accurate

Why: The exact counts.

\[ 0.0092 \]

Figure (svg): The solution to Worked example resolving Example 10.10's two p-values shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ p_{\text{exact}} \approx 0.0092, \quad p_{\text{rounded}} \approx 0.0078 \]

Verify: confirm nothing turns on it here, and when it would

Why: Both figures reject comfortably at the 5 percent level, so the conclusion is unaffected. The pattern is worth recognising because it recurs — Examples 8.3, 8.8 and 9.15 all show it — and because a p-value near a threshold could be pushed across by exactly this kind of intermediate rounding. Carrying full precision until the final step removes the question.

OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 526-527

51. Trap: rounding the proportions before computing

Trap

The trap

\[ p'_Y = 0.10, \; p'_O = 0.05 \;\Rightarrow\; z = 2.42, \; p = 0.0078 \]

Round the sample proportions to two decimals first

Why: They are close to those values.

\[ \text{the exact counts give } z = 2.36, \; p = 0.0092 \]

Rounding 0.05172 down to 0.05 exaggerates the gap by about 3 percent, and the p-value moves by nearly a fifth.

The fix

\[ \text{use } \tfrac{135}{1343} \text{ and } \tfrac{12}{232} \text{ throughout} \]

Carry full precision until the final answer

Why: Round once, at the end.

A shift from 0.0092 to 0.0078 changes nothing at the 5 percent level, but the same proportional shift applied to a p-value near 0.05 could reverse a decision. Rounding at the end costs nothing and removes an entire class of question about whether an answer is right.

52. One of these is false

Two truths and a lie

All three concern running the test.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Unequal sample sizes change the standard error but not the procedure
  • C. A two-tailed p-value is twice the one-sided area
  • B. A non-significant result shows the two proportions are equal

Survives elimination: B

Why: The survivor is false, and it is section 9.1's rule about accepting a null. With 200 per group the hives study could only detect fairly large differences, so a real gap of a few points could easily survive undetected.

53. How large a difference could it detect?

Estimation

Example 10.8's standard error is 0.0271, at a two-tailed 5 percent level.

Predict first

Roughly what difference in proportions would the study reliably detect?

  • About 0.053, or five and a third percentage points
  • About 0.027
  • About 0.01
  • Any difference at all

Correct: About 0.053.

Why: The critical value of 1.96 times the standard error of 0.0271 gives about 0.053. The observed gap of 0.04 falls short of that, which is precisely why the test did not reject — and reporting this figure alongside a non-result tells a reader what the study was capable of.

54. Explain the two p-values

Explain it

A classmate is confused that the book gives both 0.0077 and 0.0092 for Example 10.10.

Discussion prompt

In two sentences or fewer, explain.

Hint: Ask what proportions each figure was computed from.

Answer:

The 0.0077 comes from rounding the sample proportions to 0.10 and 0.05 before computing, while 0.0092 comes from the exact counts of 135 out of 1,343 and 12 out of 232.

The calculator's 0.0092 is the accurate one, and both reject at 5 percent — but carrying full precision until the end avoids the discrepancy entirely.

55. The three two-sample tests so far

Comparison

Fill the blanks. The pooling rule reverses between the first two and the third.

Comparison matrix

TestStandard errorDistribution
Two means, sigma unknowneach s-squared over its own n, addedStudent t, Aspin-Welch df
Two means, sigma knowneach sigma-squared over its own n, addednormal
Two proportionsone pooled p, times the summed reciprocalsnormal
Why pooling differsa null on means leaves spreads freea null on proportions fixes them

The last row is the idea worth carrying out of the chapter. Two rules that look contradictory follow from one principle: pool exactly what the null hypothesis says is common, and nothing else.

56. A two-proportion test, in order

Pattern

Six steps, and the third is the only new one.

  1. Confirm the data are categorical and the two samples are independent simple random samples.
  2. Check all four counts: successes and failures in each sample must be at least five.
  3. Compute the pooled proportion as total successes over total trials, and form the standard error from it.
  4. Write the hypotheses, fixing which group is A, and read the tail from the alternative.
  5. Compute the statistic as the observed difference of sample proportions over the standard error, and find the p-value — doubling it for a two-tailed test.
  6. Compare with alpha, decide, and write the conclusion naming the level, the direction and the context.

Use the exact counts throughout rather than rounded proportions, and round only the final answer.

OpenStax Introductory Business Statistics 2e, §10.4 Comparing Two Independent Population Proportions §10.4 Comparing Two Independent Population Proportions

57. Check yourself 1 of 3

Check

The pooled proportion.

Check your understanding

For 20 of 200 and 12 of 200, what is the pooled proportion?

  • A. 0.08 (correct)
  • B. 0.10
  • C. 0.04
  • D. 0.16

Answer: A

Why: Total successes over total trials: 32 divided by 400 is 0.08, midway between the two sample proportions since the sample sizes are equal.

Why B tempts people
That is sample A's own proportion, not the pooled estimate.
Why C tempts people
That is the difference between the two sample proportions.
Why D tempts people
That adds the two proportions rather than pooling the counts.

58. Check yourself 2 of 3

Check

Why pooling applies here.

Check your understanding

Why are the samples pooled here when section 10.1 forbade pooling?

  • A. Because a null of equal proportions also fixes the two spreads (correct)
  • B. Because proportions are easier to pool
  • C. Because the sample sizes are equal
  • D. Because the distribution is normal

Answer: A

Why: A binomial's variability follows from its probability, so equal proportions means equal spreads — and both samples estimate one common value.

Why B tempts people
Ease has nothing to do with it; the justification is what the null constrains.
Why C tempts people
Example 10.10 pools samples of 1,343 and 232 with no difficulty.
Why D tempts people
The normal distribution follows from the central limit theorem, not from pooling.

59. Check yourself 3 of 3

Check

The conditions.

Check your understanding

A comparison has 196 successes out of 200 in one group. Does the condition hold?

  • A. No: only four failures, which is below five (correct)
  • B. Yes: 196 successes is plenty
  • C. Yes: the sample size exceeds 30
  • D. It cannot be judged without the other group

Answer: A

Why: Both successes and failures must be at least five in each sample, and 200 minus 196 is only four.

Why B tempts people
Checking successes alone passes almost any proportion; the failures are the binding count here.
Why C tempts people
The condition constrains the four counts, not the sample size.
Why D tempts people
This group fails on its own, whatever the other group looks like.

60. Where this shows up outside the textbook

Real world

A hospital compares infection rates after two sterilisation protocols. Protocol A gives 8 infections in 400 procedures; protocol B gives 3 in 150. An administrator computes the rates as 2.0 percent and 2.0 percent, concludes they are identical, and proposes adopting whichever is cheaper.

Discussion prompt

Assess the reasoning, and say what the comparison can and cannot establish.

Hint: Check the conditions before checking the arithmetic.

Answer:

The rates really are close, and the arithmetic is right. Eight of 400 is exactly 2.0 percent and three of 150 is exactly 2.0 percent, so the observed difference is zero and no test would reject — the statistic is zero and the p-value is one.

But the second sample fails the conditions. Protocol B has only three infections, below the required five, so the normal approximation is unsound. The test that produced the reassuring answer should not have been run in this form at all.

\[ x_B = 3 \;<\; 5, \qquad \text{so the normal approximation does not apply} \]

More importantly, identical observed rates do not establish identical true rates. With these sample sizes the standard error of the difference is about 1.4 percentage points, so the study could only reliably detect a difference of around 2.8 points — larger than either rate. Protocol B's true rate could plausibly be double protocol A's and this comparison would not have seen it.

Two practical points follow. Rare-outcome comparisons need very large samples: to detect a doubling from 2 percent to 4 percent with reasonable power would take well over a thousand procedures per protocol, which is a design fact worth knowing before the study rather than after. And an exact method — Fisher's test or an exact binomial approach — would be the right analysis here, since it needs no count condition; the administrator's conclusion may even turn out to be defensible, but not on this evidence.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

Why does a two-proportion test pool the samples when a two-mean test does not?

  • Because proportions are always similar
  • Because a null of equal proportions also fixes the two spreads, while a null of equal means does not
  • Because the sample sizes are usually equal
  • Because pooling is simpler

Correct: Because the null fixes the spreads too.

\[ p_c = \frac{x_1+x_2}{n_1+n_2}, \qquad \text{SE} = \sqrt{p_c q_c\left(\tfrac{1}{n_1}+\tfrac{1}{n_2}\right)} \]

Why: A binomial's variability is p times q, determined entirely by its probability — so if the null says the two proportions are equal, it says the two spreads are equal as well, and one pooled estimate is correct. A null about means constrains only the centres and leaves the two standard deviations free, which is why section 10.1 forbids pooling. One principle, two opposite instructions.

62. Explain it to someone a year behind you

Explain it

They built the standard error from each sample's own proportion, as in section 10.1.

Discussion prompt

In two sentences or fewer, correct them.

Hint: Ask what their null hypothesis says about the two spreads.

Answer:

Their null says the two proportions are equal, and a binomial's spread comes from its proportion — so the null says the two spreads are equal as well.

Since a test computes everything under the null, both samples should be pooled into one estimate of that common proportion before the standard error is formed.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Explaining why this section pools and section 10.1 does not
  • Computing the pooled proportion and standard error
  • Checking all four counts
  • Running the test and doubling for a two-tailed p-value

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the first, pool exactly what the null says is common. For the second, total successes over total trials, then root of pc qc times the summed reciprocals. For the third, successes and failures in BOTH samples. For the fourth, double the one-sided area whenever the alternative says different. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Fifteen minutes.

Draw it

At the top, draw two boxes for the two samples feeding into one pooled box, and write the pooled proportion formula beneath. Beside it, write two columns headed two means and two proportions, and in each write what the null constrains and whether pooling follows — this is the one idea the lesson turns on. In the middle of the page, work Example 10.8 completely: the two sample proportions of 0.10 and 0.06, the pooled 0.08, the standard error of 0.0271, the statistic of 1.47, a normal curve with BOTH tails shaded, the doubled p-value of 0.1404, and a conclusion naming the 1 percent level. Beneath it, write the four counts and confirm each is at least five. To the right, work Example 10.10 with the exact counts to get 0.0092, then redo it with the proportions rounded to 0.10 and 0.05 to get 0.0078 — and write one sentence on which is accurate and why the book prints both. At the bottom, list the three conditions and mark which one is new to this section.

Check your pooled proportion by confirming it lies between the two sample proportions, and exactly midway when the sample sizes are equal. Check the p-value by halving it and confirming the result matches the single-tail area, which is the guard against forgetting to double.

65. What you can do now

Recap

Five things, and the second is the idea that distinguishes this section.

If you seeThen
Categorical data in two groupsA difference of proportions
A null of equal proportionsPool the successes and trials
A null of equal meansDo NOT pool: see section 10.1
Fewer than five in any of four countsThe normal approximation is unsound
A sample near a tenth of its populationThe independence assumption is straining
A not-equal alternativeDouble the one-sided area
Rounded proportions in a calculationRecompute from the exact counts

Section 10.4 closes the chapter with a design rather than a new parameter. When the two measurements come from the same subjects, the samples are not independent and none of this chapter's formulas apply — but taking the differences collapses the problem to a one-sample t test on those differences, which chapter 9 already covered.

OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions §10.3, pp. 523-527 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §10.3 Comparing Two Independent Population Proportions — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 523-527
  2. OpenStax Introductory Business Statistics 2e, §10.4 Comparing Two Independent Population Proportions — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108