The binomial's independence characteristic relaxed. A hypergeometric experiment samples without replacement from two groups, so each pick changes what remains, the picks are not independent, and there is no single success probability to raise to a power. The count of items drawn from the group of interest has a hypergeometric distribution, written X follows H of r, b and n, and its probabilities are computed by counting combinations: the ways of choosing x from the group of interest, times the ways of choosing the rest from the second group, over the ways of choosing the sample at all. The possible values of X are limited by the group sizes as well as by the sample size, and the mean is the sample size times the proportion in the group of interest.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 4 — Discrete Random Variables
Hypergeometric Distribution
Objectives
Five outcomes. The first is a recognition problem and the rest follow from it.
OpenStax Introductory Statistics 2e, §4.5 Hypergeometric Distribution §4.5, pp. 247-250 — the section these objectives are drawn from
Warm-up
Section 4.3's third characteristic required independent trials with a constant p, and warned that sampling without replacement breaks it.
Discussion prompt
A committee of four is chosen from six men and five women. What is the probability the first two chosen are men, and does that probability stay the same for the third pick?
Hint: Work out the chance for each pick in turn, updating the group after each one.
Answer:
The first pick is a man with probability six elevenths. If it is, the second is a man with probability five tenths, because one man and one person have gone. The third would be four ninths.
So the probability changes at every draw, and it changes differently depending on what has already been picked — exactly the failure section 4.3 flagged. There is no single p to raise to a power, so the binomial formula has nothing to work with.
What is still true is that every possible committee of four is equally likely, so the probability of any event is a count of favourable committees over a count of all of them. That counting is what this section's formula does, and it needs no p at all.
Concept
In a hypergeometric experiment you take samples from two groups, you are concerned with one of them called the group of interest, and you sample without replacement from the combined groups. Each pick is not independent, and you are not dealing with Bernoulli trials. The random variable X is the number of items from the group of interest.
hypergeometric experiment — Sampling without replacement from two combined groups, with X counting the items drawn from the group of interest. Written X follows H(r, b, n), where r is the size of the group of interest, b the size of the second group, and n the size of the sample.
\[ X \sim H(r, b, n) \]
The fifth characteristic is worth pausing on: you are not dealing with Bernoulli trials. A Bernoulli trial has a fixed success probability, and here there is none — the chance of drawing from the group of interest depends on what has already been drawn. That is why the formula counts combinations instead of multiplying probabilities, and why no p appears in it anywhere.
Figure (svg): The five characteristics of a hypergeometric experiment listed as a numbered procedure
OpenStax Introductory Statistics 2e, §4.5 Hypergeometric Distribution §4.5, p. 247
Section
Section 1
Concept
You take samples from two groups; you are concerned with a group of interest, called the first group; you sample without replacement from the combined groups; each pick is not independent, since sampling is without replacement; and you are not dealing with Bernoulli trials.
the group of interest — The first of the two groups, and the one X counts. Which group is of interest is decided by the question: if it asks about defective laptops, the defectives are the group of interest, whatever their size.
\[ r = \text{group of interest}, \quad b = \text{second group}, \quad n = \text{sample size} \]
The book's example for the third characteristic is a softball team of ten chosen from eleven men and thirteen women, and it spells the dependence out: the probability of picking a woman first is thirteen twenty-fourths, while the probability of picking a man second is eleven twenty-thirds if a woman was picked first and ten twenty-thirds if a man was. The probability of the second pick depends on what happened in the first, which is the definition of dependence.
Figure (svg): The five characteristics of a hypergeometric experiment listed as a numbered procedure
OpenStax Introductory Statistics 2e, §4.5 Hypergeometric Distribution §4.5, p. 247 — the five characteristics and the softball example
Picture it
Characteristics three and four say the same thing twice.
Figure (svg): The five characteristics of a hypergeometric experiment listed as a numbered procedure
Sampling without replacement is the cause and non-independence is the consequence, so recognising a hypergeometric problem really means noticing one thing: that the items are not put back. Everything else in the family follows from it, including the absence of a p and the need for combination counting.
Worked example
Example 4.22. The question decides which group is of interest.
\[ 100 \text{ jelly beans and } 80 \text{ gumdrops; } 50 \text{ candies picked at random} \]
Find the two groups
Why: Jelly beans and gumdrops.
Read what the question asks about
Why: The probability of picking gumdrops.
Set r and b
Why: Eighty gumdrops, one hundred jelly beans.
\[ r = 80, b = 100 \]
Set n and define X
Why: Fifty candies picked.
\[ n = 50, X =\text{ gumdrops among them} \]
Figure (svg): The solution to Worked example naming the parts shown as a ladder of expressions, one row per legal move
\[ X \sim H(80, 100, 50), \quad \text{find } P(x = 35) \]
Verify: confirm that r need not be the larger group
Why: The group of interest here has eighty members and the other has a hundred, so r is smaller than b — and that is fine. Which group is called first is decided entirely by the question, not by size. Had the question asked about jelly beans, r would be 100 and b would be 80, and the answer would be about a different count. Getting this backwards is the commonest setup error in the section.
OpenStax Introductory Statistics 2e, §4.5 Hypergeometric Distribution §4.5, p. 248
Sorting
Ask whether items are replaced between selections.
Sort into buckets
Sort each experiment.
The tell for the left-hand column is that the same item cannot be chosen twice — a person cannot hold two committee seats and a card cannot be dealt twice. The tell for the right is a mechanism that resets, like a coin or a die.
Worked example
Example 4.24, checked against all five characteristics.
\[ \text{A committee of seven from } 18 \text{ women and } 15 \text{ men; interested in the men} \]
Two groups?
Why: Women and men.
A group of interest?
Why: The question asks about men.
\[ r = 15 m e n \]
Without replacement?
Why: A person cannot serve twice.
Independent picks?
Why: The composition changes after each choice.
Figure (svg): The solution to Worked example five reasons it is hypergeometric shown as a ladder of expressions, one row per legal move
\[ X \sim H(15, 18, 7), \quad x = 0, 1, \ldots, 7 \]
Verify: confirm why the committee cannot be chosen with replacement
Why: One person cannot occupy two committee seats, so the sampling is necessarily without replacement — it is forced by the situation rather than chosen. That is typical of committee, team and hand-of-cards problems, and it is why they are the standard examples of this family. Any problem where the same item cannot be selected twice is hypergeometric rather than binomial.
OpenStax Introductory Statistics 2e, §4.5 Hypergeometric Distribution §4.5, pp. 248-249
Trap
\[ 100 \text{ jelly beans and } 80 \text{ gumdrops} \]
Take r to be 100, the larger group
Why: The first group ought to be the bigger one.
\[ \text{but the question asks about GUMDROPS} \]
With r set to 100 the variable counts jelly beans, and P(x = 35) then answers a different question entirely.
\[ r = 80 \text{ gumdrops, because the question asks about gumdrops} \]
Read the question first, then assign r to whatever it counts
Why: Group of interest means interesting to the question, not large.
The book is explicit: since the probability question asks for the probability of picking gumdrops, the group of interest is gumdrops. The two labellings give genuinely different distributions — counting gumdrops out of fifty and counting jelly beans out of fifty are complementary, so P(35 gumdrops) equals P(15 jelly beans) rather than P(35 jelly beans). Writing the definition of X as a sentence before assigning r prevents it.
Fill the middle
The notation for a hypergeometric distribution.
Fill in the blanks
X \sim H(r, b, n): \quad r \textgroup of interest ___
Why: The group of interest, which is whichever group the question asks about rather than whichever is larger. The other group's size is b, and n is the size of the sample drawn from the two combined.
Two truths and a lie
All three concern the characteristics.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. In Example 4.22 the group of interest is the 80 gumdrops against 100 jelly beans, and in Example 4.23 it is the 10 defective laptops against 90 good ones. The question decides which group is of interest, and it is frequently the smaller one.
Prediction
Commit before reasoning.
Predict first
Why does the hypergeometric formula contain no success probability p?
Correct: Because there is no single success probability.
Why: The chance of drawing from the group of interest depends on what has already been drawn, so no one number describes every trial and there is nothing to raise to a power. The formula sidesteps the issue entirely by counting equally likely samples instead, which needs no per-trial probability at all.
Section
Section 2
Concept
X counts the items from the group of interest in a sample of n, so it can be at most n. But it can also be at most r, since the sample cannot contain more of the group than exists — and it must be at least n minus b, since any shortfall from the second group must be made up from the first.
the range of X — X runs from the larger of zero and n minus b, up to the smaller of n and r. A binomial always runs from zero to n; a hypergeometric can be restricted at either end by the group sizes.
\[ \max(0, n-b) \le x \le \min(n, r) \]
The book flags the upper limit explicitly in Example 4.23: a sample of twelve laptops is drawn from a shipment containing ten defective ones, and it says X may not take on the values 11 or 12, since the sample size is 12 but there are only 10 defective laptops. That is a genuine structural difference from the binomial, whose count of successes always runs the full range from zero to n.
Figure (svg): A diagram showing that the possible values of a hypergeometric variable can be limited by the group sizes
OpenStax Introductory Statistics 2e, §4.5 Hypergeometric Distribution §4.5, p. 248 — Example 4.23 and the restricted range
Picture it
One example where every value is possible, and one where two are not.
Figure (svg): A diagram showing that the possible values of a hypergeometric variable can be limited by the group sizes
The lower limit bites less often but is just as real: drawing a sample of twelve from a population with only five in the second group forces at least seven from the group of interest, so X cannot be below seven. Both limits come from the same source — a sample cannot contain more of a group than the group contains.
Worked example
Example 4.23. The book states the restriction directly.
\[ 100 \text{ laptops of which } 10 \text{ are defective; } 12 \text{ are inspected} \]
Identify the group of interest
Why: The question asks about defectives.
\[ r = 10 \]
Note the second group and the sample
Why: Ninety good laptops, twelve inspected.
\[ b = 90, n = 12 \]
Find the upper limit
Why: The smaller of n and r.
\[ \min(12, 10) = 10 \]
Find the lower limit
Why: n minus b is negative, so zero.
\[ \max(0, -78) = 0 \]
Figure (svg): The solution to Worked example a restricted upper limit shown as a ladder of expressions, one row per legal move
\[ x = 0, 1, 2, \ldots, 10 \]
Verify: confirm the restriction is real rather than a technicality
Why: A sample of twelve laptops cannot contain eleven defective ones when the whole shipment holds only ten, so P(11) and P(12) are exactly zero. The formula agrees: C(10, 11) is zero, because there is no way to choose eleven items from ten. So the formula enforces the restriction automatically, and the reason for stating it explicitly is that a student listing the values as 0 through 12 would be wrong about the distribution's shape.
OpenStax Introductory Statistics 2e, §4.5 Hypergeometric Distribution §4.5, p. 248
Faded example
A gross of 144 eggs contains 12 cracked ones; 15 are inspected.
Fill in the blanks
x \text12 0 \text12 \min(15, ___) = ___
Why: Only twelve cracked eggs exist, so a sample of fifteen can contain at most twelve of them. The upper limit is the smaller of the sample size and the group of interest, which is 12 here.
Worked example
Example 4.25, where both limits are slack.
\[ \text{a committee of } 4 \text{ from } 6 \text{ men and } 5 \text{ women} \]
Set the parameters
Why: Men are the group of interest.
\[ r = 6, b = 5, n = 4 \]
Find the upper limit
Why: The smaller of 4 and 6.
\[ 4 \]
Find the lower limit
Why: Four minus five is negative.
\[ 0 \]
List the values
Why: Every count from none to all four.
\[ 0, 1, 2, 3, 4 \]
Figure (svg): The solution to Worked example an unrestricted range shown as a ladder of expressions, one row per legal move
\[ \max(0, 4-5) = 0 \le x \le \min(4, 6) = 4 \]
Verify: confirm by checking the probabilities sum to one
Why: The five probabilities are about 0.0455, 0.2727, 0.4545, 0.2273 and 0.0455, and they total exactly 1.0000 — section 4.1's second condition, and a genuine check that no value has been omitted. Had the range been listed as 0 through 6, the extra two probabilities would both be zero and the total would be unchanged, so the sum check alone does not catch a range that is too wide; listing the limits deliberately is what does.
OpenStax Introductory Statistics 2e, §4.5 Hypergeometric Distribution §4.5, p. 249
Trap
\[ n = 12 \text{, so } x = 0, 1, 2, \ldots, 12 \]
Take the range from the sample size, as for a binomial
Why: Every binomial count ran from zero to n.
\[ \text{but only } 10 \text{ defective laptops exist} \]
Eleven and twelve defectives in a sample of twelve are impossible, so the distribution has eleven values rather than thirteen.
\[ x = 0, 1, \ldots, \min(n, r) = 10 \]
Compare the sample size with the group of interest and take the smaller
Why: The sample cannot contain more of a group than the group holds.
The formula gives zero for the impossible values anyway, so the arithmetic survives the error — which is exactly why it is worth stating the range explicitly. A student who lists thirteen values and computes thirteen probabilities gets the right answers with two wasted lines, but one who reasons about the SHAPE of the distribution from a wrong range will describe it wrongly.
Discrimination
Compare n with r, and n with b.
Sort into buckets
Sort each setup by whether X can take every value from 0 to n.
Two truths and a lie
All three concern the range.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and the book contradicts it explicitly for Example 4.23: X may not take the values 11 or 12 when only ten defective laptops exist. Both ends of the range can be cut by the group sizes, which is a structural difference from the binomial.
Prediction
Commit before reasoning.
Predict first
A sample of 10 is drawn from 15 men and 3 women, counting men. What is the smallest possible value of X?
Correct: 7.
Why: At most three of the ten can be women, since only three exist, so at least seven must be men. The rule is that x is at least n minus b, which is 10 minus 3. This lower limit is easy to miss because it never arises in binomial problems, where zero successes is always possible however large the sample.
Section
Section 3
Concept
Every sample of size n from the combined groups is equally likely, so the probability of any particular composition is the number of samples with that composition divided by the number of samples altogether.
the hypergeometric formula — P(x) equals the number of ways to choose x from the group of interest, times the number of ways to choose the remaining n minus x from the second group, all divided by the number of ways to choose n from the combined groups.
\[ P(x) = \frac{\binom{r}{x} \binom{b}{n-x}}{\binom{r+b}{n}} \]
The structure is a counting argument rather than a probability one. The denominator counts every possible sample, and the numerator counts the favourable ones by choosing the two parts separately and multiplying — the fundamental counting principle. Because every sample is equally likely, section 3.1's counting rule applies directly, which is why no per-trial probability is needed.
Figure (svg): The hypergeometric probability formula shown as three combination counts
OpenStax Introductory Statistics 2e, §4.5 Hypergeometric Distribution §4.5, pp. 247-249 — the hypergeometric distribution and Example 4.25
Picture it
Three combinations, and what each one counts.
Figure (svg): The hypergeometric probability formula shown as three combination counts
Notice that the two numerator terms must use up the whole sample: x from the first group and n minus x from the second, totalling n. That is the same accounting check section 4.3 used on the binomial's exponents, and it catches the same class of error — a numerator whose two parts do not total n has miscounted the sample.
Worked example
Example 4.25. The book gives the answer as 0.4545.
\[ \text{a committee of } 4 \text{ from } 6 \text{ men and } 5 \text{ women; find } P(x = 2) \]
Choose the men
Why: Two from six.
\[ C(6, 2) = 15 \]
Choose the rest from the women
Why: The other two seats from five women.
\[ C(5, 2) = 10 \]
Count all possible committees
Why: Four from eleven people.
\[ C(11, 4) = 330 \]
Divide
Why: Favourable over total.
\[ \frac{150}{330} \]
Figure (svg): The solution to Worked example the site committee shown as a ladder of expressions, one row per legal move
\[ P(2) = \frac{\binom{6}{2}\binom{5}{2}}{\binom{11}{4}} = \frac{15 \times 10}{330} \approx 0.4545 \]
Verify: confirm the two numerator terms account for the whole committee
Why: Two men plus two women is four seats, matching n. Had the second term been C(5, 3) the committee would have five members, and the answer would be wrong with nothing obviously amiss in the arithmetic. Checking that x plus n minus x totals n is the fastest test of a hypergeometric substitution, exactly as the exponent check was for the binomial.
OpenStax Introductory Statistics 2e, §4.5 Hypergeometric Distribution §4.5, p. 249
Faded example
Three men on a committee of four, from six men and five women.
Fill in the blanks
P(3) = \frac14\binom______}}______}, \text___ 3 + ___ = ___
Why: Three men and one woman fill the four seats, so the second combination chooses one from five. The two numerator terms must always total n, and checking that is the quickest test of a substitution.
Worked example
The same setup, treated wrongly as binomial, to see the size of the error.
\[ \text{binomial with } n = 4 \text{ and } p = \frac{6}{11} \]
Compute the binomial value
Why: Six choose two, times p squared, times q squared.
\[ \text{about } 0.3688 \]
Recall the correct value
Why: The hypergeometric answer.
\[ 0.4545 \]
Compare
Why: The binomial understates it.
\[ \text{off by about } 0.086 \]
Say why
Why: Without replacement, a mixed committee is more likely.
Figure (svg): The solution to Worked example comparing with the binomial shown as a ladder of expressions, one row per legal move
\[ 0.3688 \text{ against } 0.4545 \]
Verify: confirm the direction of the discrepancy
Why: Choosing a man makes the remaining pool proportionally more female, so a committee of two men and two women is MORE likely without replacement than with — the draws pull against each other. The binomial, assuming independence, misses that and understates the balanced outcome. On a population of eleven the error is a quarter of the answer; the next idea shows how fast it shrinks as the population grows.
OpenStax Introductory Statistics 2e, §4.5 Hypergeometric Distribution §4.5, p. 249
Error analysis
A committee of four from six men and five women; the probability of exactly two men.
Annotate
On: \( \begin{aligned} &(1)\; \tfrac{\binom{6}{2}}{\binom{11}{4}} \\ &(2)\; \tfrac{\binom{6}{2}\binom{5}{3}}{\binom{11}{4}} \\ &(3)\; \binom{4}{2}\left(\tfrac{6}{11}\right)^2\left(\tfrac{5}{11}\right)^2 \\ &(4)\; \tfrac{\binom{6}{2}\binom{5}{2}}{\binom{11}{4}} \end{aligned} \)
Errors (1) and (2) are both caught by the seat count, and error (3) is caught by asking whether anyone can serve twice. The third is the dangerous one, because its answer is in the right region and nothing about it looks wrong.
Matching
The three combinations in the formula.
Match the pairs
Why: The fourth row is the fundamental counting principle at work: choosing the two parts independently and multiplying gives every favourable sample exactly once. The whole formula is then section 3.1's counting rule — favourable over total — applied to equally likely samples.
Two truths and a lie
All three concern the formula.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false as stated. The two agree when the POPULATION is large relative to the sample, not when the sample is small in absolute terms — a sample of four from eleven people gives an error of a quarter, while the same sample of four from eleven thousand is essentially binomial. It is the ratio that matters.
Estimation
A committee of four from six men and five women.
Predict first
Which is more likely: exactly two men, or exactly four men?
Correct: Exactly two men, by a wide margin.
Why: Two men has probability 0.4545 and four men has 0.0455, a factor of ten apart. The intuition is that balanced compositions can be reached many more ways: fifteen ways to choose two men times ten ways to choose two women is 150 committees, against just fifteen all-male ones. Distributions of this kind peak near the proportion in the population, which here is a little over half.
Section
Section 4
Concept
The mean of a hypergeometric distribution is the sample size multiplied by the proportion of the combined groups that belongs to the group of interest.
the hypergeometric mean — Mu equals n times r divided by r plus b: the sample size times the fraction of the population in the group of interest. It has the same form as the binomial's np, with the population proportion playing the part of p.
\[ \mu = \frac{nr}{r+b} \]
The formula is the binomial's np in disguise: r over r plus b is the proportion of the population in the group of interest, which is what p would have been had the sampling been with replacement. So the two families have the SAME mean and differ only in their spread — sampling without replacement is more predictable, because the pool is being used up.
Figure (svg): The hypergeometric distribution for a committee of four chosen from six men and five women, peaking at two men
OpenStax Introductory Statistics 2e, §4.5 Hypergeometric Distribution §4.5, pp. 249-250 — the formula for the mean, and Example 4.25
Picture it
Example 4.25's distribution, with the mean at 2.18.
Figure (svg): The hypergeometric distribution for a committee of four chosen from six men and five women, peaking at two men
The mean of 2.18 sits just above the tallest stem at two, which is the usual arrangement for a mildly right-leaning distribution. Six of the eleven people are men, so a committee of four should contain about four times six elevenths, and that is exactly what the formula computes.
Worked example
Example 4.25's second question.
\[ \text{a committee of } 4 \text{ from } 6 \text{ men and } 5 \text{ women} \]
Find the proportion of men
Why: Six of eleven people.
\[ \frac{6}{11} \]
Multiply by the sample size
Why: Four committee seats.
\[ 4 x 6 / 11 \]
Evaluate
Why: Twenty-four elevenths.
\[ \text{about } 2.18 \]
Interpret
Why: The long-run average across many committees.
Figure (svg): The solution to Worked example expected men on the committee shown as a ladder of expressions, one row per legal move
\[ \mu = \frac{nr}{r+b} = \frac{4 \times 6}{11} \approx 2.18 \]
Verify: confirm the answer against the proportion
Why: Men are a little over half the pool, so a committee of four should contain a little over two — and 2.18 is exactly that. The sanity check is that mu divided by n must equal the population proportion: 2.18 over 4 is 0.545, which is six elevenths. Any hypergeometric mean failing that check has used the wrong group as r.
OpenStax Introductory Statistics 2e, §4.5 Hypergeometric Distribution §4.5, p. 250
Faded example
A shipment of 100 laptops contains 10 defective; 12 are inspected.
Fill in the blanks
\mu = \frac1001.2} = ___
Why: One tenth of the shipment is defective, so twelve inspected laptops should contain about 1.2 defective ones. The denominator is the whole population of 100, not the 90 good laptops.
Worked example
Comparing the two families' means and spreads.
\[ \text{H}(6, 5, 4) \text{ against } B\left(4, \tfrac{6}{11}\right) \]
Compute the hypergeometric mean
Why: n times the proportion.
\[ 2.18 \]
Compute the binomial mean
Why: np with p = 6/11.
\[ \text{also } 2.18 \]
Compare the spreads
Why: The hypergeometric is slightly narrower.
Say why
Why: Using up the pool constrains the outcome.
Figure (svg): The solution to Worked example the same mean as a binomial shown as a ladder of expressions, one row per legal move
\[ \mu_H = \mu_B = 2.18, \quad \sigma_H < \sigma_B \]
Verify: confirm the intuition behind the smaller spread
Why: Drawing without replacement pulls the sample toward the population's composition: taking several men early leaves proportionally more women for the later picks, which damps extreme compositions. With replacement nothing damps them, so all-male committees are relatively more likely. The effect is the finite population correction, and it is largest when the sample is a big fraction of the population — here four of eleven, which is a very large fraction.
OpenStax Introductory Statistics 2e, §4.5 Hypergeometric Distribution §4.5, pp. 249-250
Trap
\[ \mu = \frac{nr}{b} = \frac{4 \times 6}{5} = 4.8 \]
Divide by the second group's size
Why: The formula has r on top and a group size below, so b looks plausible.
\[ 4.8 \text{ men on a committee of } 4 \quad \text{(impossible)} \]
The denominator must be the whole population, since the proportion is of everybody rather than of the other group.
\[ \mu = \frac{nr}{r+b} = \frac{4 \times 6}{11} \approx 2.18 \]
Use the population proportion: the group of interest over EVERYBODY
Why: The mean is the sample size times a proportion, and a proportion divides by the total.
The error announces itself: a mean of 4.8 exceeds the sample size of 4, which is impossible since X counts a subset of the sample. That range check — the mean must lie between the smallest and largest possible values of X — is the same one section 4.2 used and it catches this instantly.
Two truths and a lie
All three concern the mean.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. X counts a subset of the sample, so it is at most n and its mean must be too — a mean above n is proof of an arithmetic error, usually dividing by b rather than by r plus b. Section 4.2's range check applies here exactly.
Estimation
A palette of 200 milk cartons contains 10 leaking; 18 are inspected.
Predict first
About how many leaking cartons should the inspection find?
Correct: About 0.9.
Why: Five percent of the palette leaks, so eighteen cartons should contain about 0.9 leaking ones — eighteen times ten over two hundred. The other options either report a group size or use the wrong proportion; the sanity check is that the answer must be well below the sample size of eighteen, since leaking cartons are a small minority.
Prediction
Commit before reasoning.
Predict first
A hypergeometric and a matching binomial have the same mean. Why does the hypergeometric have the smaller standard deviation?
Correct: Because drawing without replacement damps extreme compositions.
Why: Taking several from the group of interest early leaves proportionally fewer for the later draws, so runs are self-limiting and the sample is pulled toward the population's composition. With replacement nothing pushes back and extreme samples are relatively more likely. The finiteness of the population is the underlying reason but not the mechanism, and the effect is largest when the sample is a large fraction of the population.
Section
Section 5
Concept
Section 1.2 observed that sampling without replacement is approximately the same as sampling with replacement when the population is large and the sample is small in comparison, because the chance of the composition changing appreciably is very low. That observation is what licenses using a binomial in place of a hypergeometric.
the binomial approximation — When the sample is a small fraction of the population, the hypergeometric probabilities are very close to the binomial ones with p equal to the population proportion. The approximation improves as the population grows relative to the sample.
\[ \frac{n}{r+b} \text{ small} \;\Longrightarrow\; H(r,b,n) \approx B\left(n, \frac{r}{r+b}\right) \]
This closes a loop opened in section 1.2 and reopened in sections 3.2 and 4.3. Every survey samples people without replacement, so strictly every survey is hypergeometric — and every survey is analysed with the binomial, because a sample of a thousand from a population of millions changes the composition by a millionth. The approximation is not a shortcut so much as the reason the binomial is usable at all in practice.
Figure (svg): A table showing the hypergeometric probability converging to the binomial value as the population grows from eleven to one hundred and ten thousand
OpenStax Introductory Statistics 2e, §4.5 Hypergeometric Distribution §4.5, pp. 247-249 — sampling without replacement from two groups
Picture it
The same proportions in populations of eleven, a hundred and ten, and more.
Figure (svg): A table showing the hypergeometric probability converging to the binomial value as the population grows from eleven to one hundred and ten thousand
At a population of eleven the hypergeometric and binomial answers differ by about 0.086, which is about a fifth of the value; at 110 the difference is about 0.007; at 110,000 it is invisible. The rule of thumb usually quoted is that the binomial is acceptable when the sample is under about five percent of the population, and the table shows why.
Worked example
The same composition at four population sizes.
\[ P(x = 2) \text{ for a sample of } 4 \text{ with } \tfrac{6}{11} \text{ in the group of interest} \]
Population of 11
Why: The book's own committee.
\[ 0.4545 \]
Population of 110
Why: Same proportions, ten times larger.
\[ \text{about } 0.3756 \]
Population of 1,100
Why: Larger again.
\[ \text{about } 0.3695 \]
The binomial value
Why: Four trials at six elevenths.
\[ \text{about } 0.3688 \]
Figure (svg): The solution to Worked example how large is large enough shown as a ladder of expressions, one row per legal move
\[ 0.4545 \to 0.3756 \to 0.3695 \to 0.3688 \]
Verify: confirm the direction of the convergence
Why: Each row is closer to the binomial value than the one above, and the differences shrink by roughly a factor of ten each time the population does — which matches the intuition that the composition changes by about n over the population size. The first row is the outlier: a sample of four from eleven is over a third of the population, far outside any rule of thumb, and the 19 percent error reflects that.
OpenStax Introductory Statistics 2e, §4.5 Hypergeometric Distribution §4.5, p. 247
Sorting
Compute the sample as a fraction of the population.
Sort into buckets
Sort each situation.
Item (e) is the classic case: five cards from 52 is nearly ten percent, which is why card problems are always hypergeometric and never binomial. Item (c) at twelve percent is well outside the rule of thumb too, which is why the book poses it as a hypergeometric problem.
Worked example
Applying the criterion to a practical case.
\[ 1\,000 \text{ voters sampled from } 50 \text{ million} \]
Compute the sampling fraction
Why: A thousand out of fifty million.
\[ 0.00002 \]
Compare with the rule of thumb
Why: Far below five percent.
Say what the composition does
Why: Removing 1,000 changes it by a fiftieth of a percent.
Choose the model
Why: The binomial is essentially exact.
\[ \text{use } B(1000, p) \]
Figure (svg): The solution to Worked example which model for a real survey shown as a ladder of expressions, one row per legal move
\[ \frac{1000}{50\,000\,000} = 0.00002 \;\ll\; 0.05 \]
Verify: confirm the practical stakes of the choice
Why: The hypergeometric would be correct and is essentially uncomputable at this scale, requiring combinations of fifty million things. The binomial is exact to more decimal places than any survey's sampling error, so nothing is lost. Every chapter from 7 onward assumes the binomial for survey data, and this calculation is the justification — which is worth knowing, because the assumption is invisible in those chapters.
OpenStax Introductory Statistics 2e, §4.5 Hypergeometric Distribution §4.5, pp. 247-249
Trap
\[ \text{a committee of } 4 \text{ from } 11 \text{ people, treated as binomial} \]
Use p equal to six elevenths and the binomial formula
Why: The proportion is known, so a p is available.
\[ 0.3688 \text{ against the correct } 0.4545 \quad \text{(off by about a fifth)} \]
Four people out of eleven is over a third of the population, so the composition changes dramatically as the committee is filled.
\[ \text{compute the sampling fraction first: } \tfrac{4}{11} \approx 0.36 \]
Check n over the population before approximating
Why: Under about five percent is the usual threshold.
The rule of thumb is worth applying explicitly rather than by feel, because the error grows quickly and is invisible in the answer. A sampling fraction of 36 percent gives an error of 19 percent here; a fraction of five percent would give well under one percent. When the fraction is large the hypergeometric must be used, and it is no harder — three combinations rather than a formula with powers.
Faded example
A sample of 50 is drawn from a population of 100,000.
Fill in the blanks
\frac0.00055 = ___, \text___ ___ \text___
Why: A sampling fraction of 0.05 percent, a hundred times below the five percent threshold, so the binomial is essentially exact. The rule of thumb is a guideline rather than a theorem, and the further below it a fraction sits the safer the approximation.
Two truths and a lie
All three concern the approximation.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false because absolute size is the wrong criterion. A sample of four is small in absolute terms and is a third of a population of eleven, where the approximation fails badly. What matters is n divided by the population, and the usual threshold is about five percent.
Explain it
A classmate asks why survey data are analysed with the binomial when nobody is surveyed twice.
Discussion prompt
In three sentences or fewer, justify the practice.
Hint: Ask what removing one respondent does to a population of millions.
Answer:
Point out that strictly they are right: surveys sample without replacement, so the exact model is hypergeometric.
But removing a thousand people from fifty million changes the population's composition by about two thousandths of a percent, so the probability facing the second respondent is essentially identical to the first's.
The binomial answer then agrees with the hypergeometric to far more decimal places than the survey's own sampling error, so nothing is lost and a great deal of computation is saved.
Comparison
Fill the blanks. Each relaxes a different binomial characteristic.
Comparison matrix
| Family | What it relaxes | Parameters |
|---|---|---|
| Binomial | nothing: all three characteristics hold | n and p |
| Geometric | the fixed number of trials | p alone |
| Hypergeometric | independence: sampling is without replacement | r, b and n |
| The tell | an n given; the word until; items not replaced | read the sampling scheme first |
Section 4.3's three characteristics are the map: each family in this chapter is the binomial with one of them removed, and identifying which one fails identifies the family. That is why the characteristics were worth checking explicitly rather than by feel.
Pattern
Six steps, and the second is the one that decides everything after it.
Before all of this, compute n over the population. If it is under about five percent, the binomial with p equal to r over r plus b will do and is far less work.
OpenStax Introductory Business Statistics 2e, §4.1 Hypergeometric Distribution §4.1 Hypergeometric Distribution
Check
Identify the family.
Check your understanding
Which of these is hypergeometric?
Answer: A
Why: The committee is drawn without replacement from two groups, since one person cannot fill two seats, so the picks are not independent.
Check
The range of X.
Check your understanding
A sample of 12 is drawn from a shipment of 100 containing 10 defective items. What values can X, the number of defectives, take?
Answer: A
Why: The sample cannot contain more defectives than the ten that exist, so the upper limit is 10 rather than the sample size of 12.
Check
The mean.
Check your understanding
A committee of four is chosen from six men and five women. What is the expected number of men?
Answer: A
Why: The mean is n times r over r plus b, which is four times six over eleven, about 2.18.
Real world
A warehouse receives a pallet of 200 components, of which an unknown number are defective. The acceptance rule is to inspect 20 at random and reject the pallet if two or more are defective. A supplier claims the pallet contains 10 defective components and complains that the rule rejects too often.
Discussion prompt
Model the inspection, compute the rejection probability under the supplier's claim, and say what the rule is really doing.
Hint: Check the sampling fraction before choosing a model.
Answer:
The model is hypergeometric. Twenty components are inspected from 200 without replacement, so the sampling fraction is ten percent — twice the usual five percent threshold, and the binomial would be a noticeably poor approximation here.
Under the supplier's claim the rejection probability is about 0.26. With r equal to 10, b equal to 190 and n equal to 20, the probability of fewer than two defectives is about 0.74, so the pallet is rejected roughly a quarter of the time. The expected number found is n times r over the population, which is 20 times 10 over 200, exactly 1.
The rule is doing what it was designed to do, and the supplier's complaint has a point. A pallet with 5 percent defectives is rejected about a quarter of the time, which is a substantial rate for a batch the supplier may consider acceptable. Whether that is too often is a business question about what defect rate the warehouse is willing to accept, not a statistical error.
\[ \mu = \frac{20 \times 10}{200} = 1, \qquad P(\text{reject}) = P(x \ge 2) \approx 0.26 \]
Two further points belong in the answer. The rule cannot distinguish a pallet with 10 defectives from one with 15 on a single inspection, because both produce overlapping distributions of the count — which is why acceptance sampling standards specify both an acceptable quality level and the probability of rejecting it. And the hypergeometric is the right model precisely because the pallet is finite and the sample is a tenth of it; treating it as binomial would understate the variability and misstate the rejection rate.
Commit first
Answer, then rate your confidence honestly.
Predict first
Why does the hypergeometric formula use combinations rather than powers of p?
Correct: Because there is no single p.
\[ P(x) = \frac{\binom{r}{x}\binom{b}{n-x}}{\binom{r+b}{n}}: \quad \text{no } p \text{ anywhere} \]
Why: Sampling without replacement makes the success probability depend on what has already been drawn, so no one number describes every trial. The formula avoids the problem by counting equally likely samples — favourable over total — which needs no per-trial probability. The finiteness of the population is the underlying cause of the changing probability rather than the reason for the counting approach.
Explain it
They used the binomial for a committee of four chosen from eleven people and cannot see the problem.
Discussion prompt
In three sentences or fewer, show them where the assumption breaks.
Hint: Ask them for the probability that the second pick is a man.
Answer:
Ask them what the probability of picking a man is on the second choice: it is five tenths if the first was a man and six tenths if the first was a woman, so there is no single p.
The binomial formula raises one p to a power, and there is nothing here to raise — the number they used, six elevenths, is only correct for the very first pick.
The hypergeometric counts committees instead: fifteen ways to pick two men times ten ways to pick two women, over 330 committees altogether, which gives 0.4545 against their 0.3688.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For recognition, ask whether the same item could be selected twice. For the group of interest, read what the question counts rather than which group is larger. For the range, take the smaller of n and r at the top and n minus b at the bottom. For the approximation, compute n over the population and compare with five percent. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top, write the five characteristics of a hypergeometric experiment and mark which two say the same thing. Beside them, write the three binomial characteristics from section 4.3 and draw an arrow from the one this family relaxes. In the middle, work Example 4.25 completely: a committee of four from six men and five women, with the men as the group of interest. Write r, b and n, list the values X can take with the min and max rule, and compute all five probabilities as three combinations each — then check that they total one. Draw the five stems to scale and mark the mean at 2.18. Below that, compute the same P(x = 2) with the binomial using p equal to six elevenths, write both answers side by side, and write one sentence on why they differ and in which direction. At the bottom, make a three-row table of binomial, geometric and hypergeometric with columns for which characteristic is relaxed, what the parameters are, and one phrase in a problem that signals it. Finish by computing the sampling fraction for four from eleven and for a thousand from fifty million, and saying which one licenses a binomial approximation.
Check your five probabilities by adding them: they must total exactly one. And check the mean by dividing 2.18 by the sample size of four — it should give six elevenths, the population proportion. If it does not, r and b were assigned the wrong way round.
Recap
Five things, and the first is a recognition skill rather than a calculation.
| If you see | Then |
|---|---|
| Items drawn and not replaced | Hypergeometric: no single p exists |
| A committee, a hand of cards, an inspected batch | Hypergeometric by construction |
| A question naming one of two groups | That group is r, whatever its size |
| A sample larger than the group of interest | The range of X stops at r, not n |
| A sample under about 5 percent of the population | The binomial approximation is fine |
| A mean above the sample size | An error: probably divided by b instead of r + b |
| A numerator whose parts do not total n | A substitution error in the formula |
Section 4.6 closes the chapter with a family that counts occurrences rather than successes among trials. When events happen at a known average rate in a fixed interval of time or space, independently of when the last one occurred, their count has a Poisson distribution — and it doubles as an approximation to the binomial when n is large and p small.
OpenStax Introductory Statistics 2e, §4.5 Hypergeometric Distribution §4.5, pp. 247-250 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.