The last discrete family in the chapter and the only one that does not count successes among trials. A Poisson variable counts occurrences in a fixed interval of time or space, when those events happen with a known average rate and independently of the time since the last one. It is written X follows P of mu, where mu is the mean number of occurrences per interval, and mu scales in proportion when the interval is lengthened or shortened. Its single parameter is both its mean and its variance, so the standard deviation is the square root of mu, and the family doubles as an approximation to the binomial when the number of trials is large and the success probability small.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 4 — Discrete Random Variables
Poisson Distribution
Objectives
Five outcomes. The second is where most of the arithmetic errors live.
OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 250-254 — the section these objectives are drawn from
Warm-up
Every family so far counted successes among a fixed or repeated set of trials.
Discussion prompt
A call centre receives about six calls in two hours. What is the probability of exactly two calls in the next fifteen minutes — and what would you use as n and p?
Hint: Try to identify the trials. How many chances to receive a call are there in fifteen minutes?
Answer:
There is no n. A caller may ring at any instant, so there is no list of trials to count successes among — and without an n the binomial has nothing to work with, whatever value of p you might invent.
What the situation does give is a rate: six calls per two hours, which is three per hour and 0.75 per quarter hour. That single number is enough, because the count of calls in an interval is determined by the average rate and the length of the window.
This section's family takes exactly that as its one parameter. It is written X follows P(mu), where mu is the mean number of occurrences in the interval of interest, and it answers the question with no n and no p at all.
Concept
A Poisson experiment has two main characteristics. The first is that it gives the probability of a number of events occurring in a fixed interval of time or space, if these events happen with a known average rate and independently of the time since the last event. The second is that it may be used to approximate the binomial when the probability of success is small and the number of trials is large. The random variable X is the number of occurrences in the interval of interest.
Poisson distribution — The distribution of the number of occurrences in a fixed interval of time or space, when occurrences happen at a known average rate and independently of the time since the last one. Written X follows P(mu), where mu is the mean for the interval of interest.
\[ X \sim P(\mu) \]
The second characteristic is worth noticing, because it is unusual: it does not describe the experiment at all, but says what the family is good for. The book states it as part of the definition, which is a fair reflection of how the Poisson is actually used — it is both a model in its own right and the standard stand-in for an unwieldy binomial. This lesson's last idea returns to it with the book's own worked comparison.
Figure (svg): The two characteristics of a Poisson experiment listed as a numbered procedure
OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, p. 250
Section
Section 1
Concept
The first characteristic carries three requirements at once: a fixed interval of time or space, a known average rate at which events happen, and independence from the time since the last event. That last phrase is what makes the family memoryless.
independently of the time since the last event — Knowing that an hour has passed with no call tells you nothing about the next fifteen minutes. The process does not become due for an event, and it does not tire — each interval faces the same distribution as any other of the same length.
\[ P(\text{event in } I_1) = P(\text{event in } I_2) \text{ whenever } |I_1| = |I_2| \]
The book's illustration is a book editor counting misspelled words, averaging five per hundred pages, and it names the interval explicitly: the interval is the 100 pages. Naming the interval is not a formality — every later step depends on it, and the book asks for it first in each of its examples. Its other setups follow the same shape: a bank expecting six bad checks per day, a news reporter saying uh twice per broadcast, an emergency room seeing five patients per hour.
Figure (svg): The two characteristics of a Poisson experiment listed as a numbered procedure
OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 250-251 — the two characteristics, and Examples 4.26 through 4.28
Picture it
A rate over a fixed interval, and a use the family can be put to.
Figure (svg): The two characteristics of a Poisson experiment listed as a numbered procedure
Notice what is absent from the first characteristic: any mention of trials, of n, or of a success probability. That absence is the family's defining feature and the reason it needed a section of its own, since none of the earlier formulas could be adapted to a situation with nothing to count trials over.
Worked example
Example 4.29, checked against both characteristics.
\[ \text{Leah answers about six calls in a two-hour period} \]
Look for trials
Why: There is no fixed number of chances to receive a call.
Find the rate and the window
Why: Six calls per two hours.
Check the first characteristic
Why: Any two intervals of equal length are alike.
Check the second
Why: A call does not make the next more or less likely.
Figure (svg): The solution to Worked example identifying a Poisson setup shown as a ladder of expressions, one row per legal move
\[ X \sim P(\mu), \quad \mu \text{ from the rate and the interval} \]
Verify: confirm that no binomial reading is available
Why: To use a binomial you would need to name n, and there is no natural candidate: a call may come at any instant of the two hours, so the number of opportunities is not finite. Some texts motivate the Poisson as the limit of a binomial with n going to infinity and p to zero, which is a fair description of what it does — but here it is simply the right model from the start, with a rate as its only input.
OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 251-252
Sorting
Ask whether you can point at an identifiable trial.
Sort into buckets
Sort each experiment.
The upper limit is a quick secondary test. Twenty flips cannot give twenty-one heads, but a page could in principle carry any number of errors — an unbounded count almost always signals a Poisson.
Worked example
Examples 4.26 through 4.28. The book asks for the interval first every time.
\[ \text{five misspelled words per } 100 \text{ pages; six bad checks per day; two uhs per broadcast} \]
Name it for the editor
Why: A hundred pages of text.
Name it for the bank
Why: One day of checks.
Name it for the reporter
Why: One broadcast.
Say what all three share
Why: A known rate and a fixed window, with no trials.
Figure (svg): The solution to Worked example naming the interval shown as a ladder of expressions, one row per legal move
\[ \text{interval} \in \{\text{time}, \text{space}\} \]
Verify: confirm the rate must be stated per unit of that same interval
Why: A rate of five misspellings per hundred pages and a question about two hundred pages are compatible; the same rate and a question about five chapters are not, until a chapter's length is known. Every Poisson problem needs the rate and the question expressed in the same units, and mismatched units are the commonest source of a wrong mu — which the next idea takes up directly.
OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 250-251
Trap
\[ \text{six calls in two hours} \;\to\; n = 120 \text{ minutes}, \; p = \tfrac{6}{120} \]
Treat each minute as a trial that either brings a call or does not
Why: That produces an n and a p, so the binomial formula applies.
\[ \text{but two calls can arrive in the same minute} \]
The trials are not really binary, so the model undercounts: it caps the answer at one call per minute when the real process has no such limit.
\[ X \sim P(\mu) \text{ with } \mu \text{ from the rate} \]
Use the rate directly, with no trials at all
Why: The Poisson was built for exactly this situation.
The invented binomial is not absurd — with fine enough divisions it converges on the Poisson, which is one way of deriving the family. But it is unnecessary work with a built-in error, and the error grows as the intervals get coarser. If you cannot point at a trial, the answer is not to manufacture one.
Fill the middle
How a Poisson distribution is written.
Fill in the blanks
X \sim P(\mu), \textinterval ___
Why: Per interval — and specifically the interval of interest, which is the one the question asks about rather than the one the rate was quoted over. Getting those two to agree is the next idea.
Two truths and a lie
All three concern the characteristics.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and its falsity is the point of the whole section. A Poisson experiment has no trials at all, which is why neither n nor p appears in the family's own formula and why X has no upper limit. The one place n and p do appear is the second characteristic, where a binomial is being approximated rather than described.
Prediction
Commit before reasoning.
Predict first
Which situation violates the requirement of a known average rate holding across intervals?
Correct: Calls arriving at a rate that varies by time of day.
Why: Two intervals of the same length then face different rates, which is exactly what the first characteristic forbids. The usual fix is to model the busy and quiet periods separately, each with its own mu. The other options are all consistent with a Poisson: an above-average page is ordinary variation, simultaneous calls are expected, and an unbounded count is a feature of the family.
Section
Section 2
Concept
The rate is usually quoted over one interval and the question asked about another. Because occurrences happen at a constant average rate, the mean scales in direct proportion: halve the interval and you halve mu.
the interval of interest — The window the question asks about. The stated rate must be rescaled onto it before any probability is computed, by multiplying by the ratio of the two interval lengths.
\[ \mu_{\text{new}} = \mu_{\text{stated}} \times \frac{\text{new interval}}{\text{stated interval}} \]
This is the step where most Poisson errors are made, and they are invisible afterwards: a probability computed from the wrong mu looks entirely reasonable. The habit worth building is to write the rate as a fraction with units, then multiply, so that the units cancel and leave a pure count. Six calls per 120 minutes, times 15 minutes, gives 0.75 calls.
Figure (svg): Four bars of decreasing length showing the mean number of calls falling in proportion as the interval shortens from two hours to fifteen minutes
OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 251-252 — Example 4.29 and the rescaling
Picture it
Six calls in two hours, expressed over shorter intervals.
Figure (svg): Four bars of decreasing length showing the mean number of calls falling in proportion as the interval shortens from two hours to fifteen minutes
Nothing about the process changes down the four rows — only the window being asked about. That is worth being clear on, because the phrase changing the mean sounds like changing the situation, and it is not: the rate of three calls an hour is constant throughout.
Worked example
Example 4.29's setup. The rate is given over two hours.
\[ \text{six calls per two hours; the interval of interest is } 15 \text{ minutes} \]
Write the rate with units
Why: Six calls per 120 minutes.
\[ \frac{6}{120}\text{ calls per minute} \]
Multiply by the new interval
Why: Fifteen minutes.
\[ 6 x 15 / 120 \]
Evaluate
Why: Six times one eighth.
\[ 0.75 \]
State the distribution
Why: The mean over a quarter hour.
Figure (svg): The solution to Worked example Leah's fifteen-minute break shown as a ladder of expressions, one row per legal move
\[ \mu = 6 \times \frac{15}{120} = 0.75 \]
Verify: confirm the direction of the change
Why: Fifteen minutes is an eighth of two hours, so the mean must be an eighth of six — smaller, as a shorter window should be. A common slip is to divide when multiplying is called for, giving 48, which is absurdly larger than the two-hour mean and is caught instantly by asking whether the answer should have gone up or down. Always check the direction before computing anything from mu.
OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, p. 252
Faded example
A machine produces 12 defective items per 8-hour shift. Find mu for a 2-hour period.
Fill in the blanks
\mu = 12 \times \frac83} = ___
Why: Two hours is a quarter of the shift, so the mean is a quarter of twelve, which is 3. The denominator is the interval the rate was quoted over, and the numerator the one asked about.
Worked example
Example 4.31, where the rescaling and the probability are asked together.
\[ 41.5 \text{ texts per day; find } P(x = 2) \text{ and } P(x > 2) \text{ per hour} \]
Rescale onto one hour
Why: Forty-one and a half over twenty-four.
\[ \mu = 1.7292 \]
Compute the exact value
Why: Exactly two texts in the hour.
\[ 0.2653 \]
Compute the cumulative value
Why: At most two.
\[ 0.7495 \]
Subtract for the tail
Why: One minus that.
\[ 0.2505 \]
Figure (svg): The solution to Worked example text messages per hour shown as a ladder of expressions, one row per legal move
\[ \mu = \frac{41.5}{24} \approx 1.7292, \quad P(x=2) \approx 0.2653, \quad P(x>2) \approx 0.2505 \]
Verify: confirm that the units cancel to leave a pure count
Why: Texts per day, times days, gives texts — a count with no units attached, which is what mu must always be. If the units do not cancel, the rate and the interval were not expressed compatibly, and that is the signal to convert one before multiplying. The direction check also passes: an hour is a twenty-fourth of a day, so the mean falls from 41.5 to about 1.73.
OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, p. 254
Trap
\[ \text{six calls in two hours} \;\to\; X \sim P(6) \]
Take mu straight from the rate
Why: Six is the number the problem gives.
\[ \text{but the question asks about } 15 \text{ minutes} \]
A mean of six calls in a quarter hour describes a very different call centre, and every probability computed from it is wrong.
\[ \mu = 6 \times \tfrac{15}{120} = 0.75, \quad X \sim P(0.75) \]
Rescale onto the interval the question asks about
Why: The rate is per two hours; the question is per quarter hour.
The damage is quiet. With mu equal to six, P(x > 1) comes out around 0.98 instead of 0.17 — a completely different conclusion about Leah's break, reached with correct arithmetic on the wrong parameter. Reading the interval out of the question BEFORE looking at the rate is the habit that prevents it.
Estimation
A rate of 20 customers per hour; the question asks about 10 minutes.
Predict first
Should mu be larger or smaller than 20, and roughly what?
Correct: Smaller, about 3.3.
Why: Ten minutes is a sixth of an hour, so the mean is twenty divided by six, about 3.33. The direction check comes first and rules out two of the options immediately: a shorter window must expect fewer customers, so anything at or above 20 is wrong before any arithmetic is done.
Two truths and a lie
All three concern rescaling.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The rate of three calls an hour is a property of the call centre and does not change when you ask about a different window — only the expected count over that window changes. Confusing the two makes the scaling feel arbitrary when it is simply multiplication.
Explain it
A classmate got mu = 48 for Leah's fifteen minutes by dividing instead of multiplying.
Discussion prompt
In two sentences or fewer, show them the error without doing the arithmetic for them.
Hint: Ask whether a shorter break should expect more calls or fewer.
Answer:
Ask whether fifteen minutes should bring more calls or fewer than two hours: fewer, obviously, so any answer above six is wrong before the arithmetic is checked.
Then point at the fraction: fifteen over a hundred and twenty is less than one, so multiplying by it must shrink the six — and 0.75 is what that gives.
Section
Section 3
Concept
Once mu is on the right interval, the probability of exactly x occurrences follows from mu alone. Questions asking for more than, at least or at most are answered by adding the relevant exact values, usually through a complement.
cumulative probability — P(x is at most k), the sum of the exact probabilities from zero up to k. Because a Poisson variable has no upper limit, questions about more than k are always answered as one minus a cumulative value rather than by adding a tail.
\[ P(x) = \frac{\mu^x e^{-\mu}}{x!} \]
The absence of an upper limit changes how the complement is used. With a binomial you could in principle add the tail directly, since it is finite; with a Poisson the tail is infinite and you have no choice but to subtract from one. That makes the complement rule from section 3.2 not a convenience here but a necessity.
Figure (svg): The Poisson distribution for a mean of nought point seven five calls, falling steeply from its peak at zero
OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 251-253 — the notation and Example 4.29
Picture it
Six stems, with the four that make up P(x > 1) shaded.
Figure (svg): The Poisson distribution for a mean of nought point seven five calls, falling steeply from its peak at zero
The shaded region looks small, and it is: 0.1734, a little over one chance in six. That is the answer to whether Leah is likely to be interrupted more than once, and its smallness is worth noticing — with a mean of 0.75 calls, being interrupted twice is genuinely unusual.
Worked example
Example 4.29's question. The book gives 0.1734.
\[ X \sim P(0.75); \text{ find } P(x > 1) \]
Recognise the infinite tail
Why: More than one means 2, 3, 4, and on forever.
Write the complement
Why: One minus at most one.
\[ 1 - P(x \le 1) \]
Compute the two exact values
Why: Zero calls and one call.
\[ 0.4724 + 0.3543 \]
Subtract
Why: One minus their sum.
\[ 1 - 0.8266 \]
Figure (svg): The solution to Worked example more than one call shown as a ladder of expressions, one row per legal move
\[ P(x > 1) = 1 - P(x \le 1) \approx 1 - 0.8266 = 0.1734 \]
Verify: confirm the boundary is excluded correctly
Why: More than one excludes one itself, so the complement is at most one and includes both zero and one. Had the question said at least one, the complement would have been zero alone and the answer 0.5276 — three times larger. The strictness of the inequality is the whole difference, and section 4.3's habit of writing out which values are in and which are out applies here unchanged.
OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, p. 252
Faded example
X follows P(3). Find P(x > 2).
Fill in the blanks
P(x > 2) = 1 - P(x \le 2) = 1 - [P(0) + P(1) + P(2)]
Why: More than two excludes two, so the complement is at most two and contains three exact values: zero, one and two. The tail above is infinite, so subtracting is the only route.
Worked example
Example 4.30: an email user receives 147 emails a day, on average.
\[ X \sim P(147); \text{ find } P(x = 160) \text{ and } P(x \le 160) \]
Compute the exact value
Why: The probability of precisely 160.
\[ 0.0180 \]
Compute the cumulative value
Why: At most 160.
\[ 0.8666 \]
Compute the standard deviation
Why: The square root of the mean.
\[ \text{about } 12.12 \]
Locate 160 in the distribution
Why: Thirteen above a mean of 147.
Figure (svg): The solution to Worked example a large mean shown as a ladder of expressions, one row per legal move
\[ P(160) \approx 0.0180, \quad P(x \le 160) \approx 0.8666, \quad \sigma \approx 12.12 \]
Verify: confirm the two answers are consistent with each other
Why: The cumulative value of 0.8666 says 160 sits at about the 87th percentile, and 160 is about 1.07 standard deviations above the mean — which for a roughly symmetric distribution should put it near the 86th percentile. The two agree, which is a real check on both. The exact value of 0.0180 is also sensible: with a spread of about twelve, the probability is shared among roughly fifty plausible values, so no single one can be large.
OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 253-254
Error analysis
A mean of 0.75 calls in a fifteen-minute interval.
Annotate
On: \( \begin{aligned} &(1)\; 1 - P(x = 1) \\ &(2)\; 1 - P(x \le 1) \text{ with } \mu = 6 \\ &(3)\; P(x = 2) + P(x = 3) + P(x = 4) \\ &(4)\; 1 - P(x \le 1) \text{ with } \mu = 0.75 \end{aligned} \)
Error (2) is the one to fear. Errors (1) and (3) produce answers that look off once you think about the shape, but (2) is internally consistent and wrong only in its input — which is why the interval belongs in the very first line of the working.
Discrimination
Match each phrase to the cumulative value you would subtract from one.
Sort into buckets
Sort each phrase by its complement.
Two truths and a lie
All three concern computing Poisson probabilities.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. A Poisson variable takes every whole number from zero upward with positive probability, however far above the mean — the values merely become very small. That unboundedness is what forces the complement, and it distinguishes the family from every other one in this chapter.
Prediction
Commit before reasoning.
Predict first
For X following P(0.75), which single value of x is most likely?
Correct: 0.
Why: With a mean below one, no occurrences is the likeliest outcome, at about 0.4724 against 0.3543 for exactly one. The mean of 0.75 is not itself a possible value, since X counts whole occurrences — a reminder from section 4.2 that a mean need not be attainable. Whenever mu is below one the distribution peaks at zero and falls from there.
Section
Section 4
Concept
The mean of a Poisson distribution is mu, its variance is also mu, and so its standard deviation is the square root of mu. No second parameter is needed or available.
equidispersion — The property that the variance equals the mean. It is unique among the families in this chapter, and it gives the Poisson a testable signature: real counts whose variance far exceeds their mean are not Poisson.
\[ \mu_X = \mu, \quad \sigma^2_X = \mu, \quad \sigma_X = \sqrt{\mu} \]
The binomial needed n and p to give np and the square root of npq; the geometric needed p. The Poisson needs only mu, which is why a single number fully specifies the distribution. The practical consequence is that a Poisson count with a large mean is proportionally tighter: with mu equal to 147 the standard deviation is about 12.12, which is only eight percent of the mean, while with mu equal to 4 the standard deviation of 2 is fully half of it.
Figure (svg): A diagram showing that the Poisson mean and variance are both mu, so the standard deviation is the square root of mu
OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 253-254 — Example 4.30 and the standard deviation
Picture it
Compared with the two parameters the other families required.
Figure (svg): A diagram showing that the Poisson mean and variance are both mu, so the standard deviation is the square root of mu
That the variance equals the mean is a strong claim about real data, and it can be checked. Counts that clump — accidents at a junction, say, where one crash causes others — show a variance well above the mean, and that excess is the standard evidence that a Poisson model does not fit.
Worked example
Example 4.30's third part.
\[ X \sim P(147) \]
Recall the variance
Why: It equals the mean.
\[ 147 \]
Take the square root
Why: The standard deviation.
\[ \text{about } 12.12 \]
Express it as a fraction of the mean
Why: Twelve over 147.
\[ \text{about } 8 \% \]
Interpret
Why: Counts near 147 give or take a dozen.
Figure (svg): The solution to Worked example the standard deviation at mu 147 shown as a ladder of expressions, one row per legal move
\[ \sigma = \sqrt{147} \approx 12.12 \]
Verify: confirm against the cumulative value already computed
Why: The earlier part found P(x at most 160) to be 0.8666. Since 160 is 13 above the mean, it sits about 1.07 standard deviations up, and for a roughly symmetric distribution that should correspond to a percentile in the middle eighties — which 0.8666 is. The two results confirm each other, and a standard deviation that failed this check would signal an error in one of them.
OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, p. 254
Faded example
A hospital ward admits an average of 25 patients per day.
Fill in the blanks
\sigma = \sqrt25} = 5
Why: The variance equals the mean of 25, so the standard deviation is its square root, 5. Daily admissions should typically run 25 give or take about 5.
Worked example
The same family at a small mean and a large one.
\[ \mu = 4 \text{ against } \mu = 400 \]
Standard deviation at 4
Why: The square root of four.
\[ 2 \]
As a fraction of the mean
Why: Two over four.
\[ 50 \% \]
Standard deviation at 400
Why: The square root of four hundred.
\[ 20 \]
As a fraction of the mean
Why: Twenty over four hundred.
\[ 5 \% \]
Figure (svg): The solution to Worked example relative spread at two means shown as a ladder of expressions, one row per legal move
\[ \frac{\sqrt{\mu}}{\mu} = \frac{1}{\sqrt{\mu}} \]
Verify: confirm the general rule behind the two cases
Why: The relative spread is the square root of mu over mu, which simplifies to one over the square root of mu — so it falls as the mean grows, and falls slowly. Multiplying the mean by a hundred divides the relative spread by ten. That is why counting over a longer interval gives a proportionally more reliable estimate of a rate, and it is the same square-root behaviour that will govern sample means in chapter 7.
OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 253-254
Trap
\[ \mu = 147 \;\Rightarrow\; \sigma = 147 \]
Read the variance equals the mean as the sd equals the mean
Why: Both are called spread, so they blur together.
\[ \text{a mean of } 147 \text{ with a spread of } 147 \]
That would put a count of zero within one standard deviation of the mean, which does not describe the distribution at all.
\[ \sigma^2 = \mu = 147 \;\Rightarrow\; \sigma = \sqrt{147} \approx 12.12 \]
It is the VARIANCE that equals the mean, so take a square root
Why: Section 2.7's distinction between variance and standard deviation, applied here.
The sanity check is the cumulative value: 160 should be a bit over one standard deviation above 147, and it comes out at 1.07 with sigma equal to 12.12. With sigma equal to 147 it would be less than a tenth of a standard deviation up, which cannot square with a percentile of 87.
Two truths and a lie
All three concern the mean and spread.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and it confuses the variance with the standard deviation. The variance equals mu; the standard deviation is its square root. At mu equal to 147 that is the difference between 147 and 12.12.
Estimation
A junction averages 16 crossings per minute.
Predict first
Would a count of 28 in one minute be unusual?
Correct: Yes: three standard deviations above.
Why: The standard deviation is the square root of sixteen, which is 4, so 28 sits twelve above the mean and therefore three standard deviations up. Section 2.7's rule of thumb calls that unusual, and it is the kind of judgement the Poisson makes possible from a single number.
Prediction
Commit before reasoning.
Predict first
A count of accidents has mean 5 and variance 22. What does that suggest?
Correct: The events clump, so independence fails.
Why: A Poisson requires the variance to equal the mean, so a variance four times the mean is direct evidence against the model. The usual cause is dependence between occurrences — one accident causing others, or a rate that varies between intervals, breaking the first characteristic's requirement of a known rate holding independently of the past. This test is the practical reason equidispersion is worth remembering.
Section
Section 5
Concept
This is the second of the two characteristics, returned to as a technique. When a binomial has a large number of trials and a small success probability, its probabilities are very close to those of a Poisson with mu equal to np. The book gives the conditions as n large, greater than 20, and p small, less than 0.05.
the Poisson approximation to the binomial — Replacing B(n, p) by P(np) when n is large and p small. The two agree closely, and the Poisson is far easier to compute because it needs no combinations.
\[ n > 20, \; p < 0.05 \;\Longrightarrow\; B(n, p) \approx P(np) \]
The characteristic itself is looser than the rule of thumb: it says the probability of success should be small, such as 0.01, and the number of trials large, such as 1,000. The sharper thresholds come with the book's worked comparison at the end of the section, where it justifies the agreement by checking n above 20 and p below 0.05. The approximation also runs the opposite way from section 4.5's, where a hypergeometric was replaced by a binomial when the population was large.
Figure (svg): A three-column table comparing binomial probabilities at n equals two hundred and p equals nought point nought one with Poisson probabilities at mu equals two, showing close agreement
OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 250-254 — the second characteristic, and Example 4.32
Picture it
Two hundred trials at one percent, against a Poisson with mean two.
Figure (svg): A three-column table comparing binomial probabilities at n equals two hundred and p equals nought point nought one with Poisson probabilities at mu equals two, showing close agreement
The agreement is not coincidence. As n grows and p shrinks with np held fixed, the binomial formula converges term by term on the Poisson one — which is the standard derivation of the family, and the reason the Poisson describes rare events among many opportunities so well.
Worked example
Example 4.32, worked both ways as the book does.
\[ n = 200 \text{ days, } p = 0.0102; \text{ find } P(x = 10) \text{ both ways} \]
Compute the binomial
Why: Two hundred days at just over one percent.
\[ \text{about } 0.000039 \]
Compute the approximating mean
Why: n times p.
\[ \mu = 2.04 \]
Compute the Poisson
Why: Ten occurrences at that mean.
\[ \text{about } 0.000045 \]
Check the conditions
Why: n above 20 and p below 0.05.
Figure (svg): The solution to Worked example the seismic-activity comparison shown as a ladder of expressions, one row per legal move
\[ \text{binomial } 0.000039 \quad \text{against} \quad \text{Poisson } 0.000045 \]
Verify: confirm the conditions the book gives for expecting agreement
Why: The book's own justification is that the approximation should be good because n is large, greater than 20, and p is small, less than 0.05 — and here n is 200 and p is 0.0102, so both hold comfortably. Note also what close means at this scale: the two differ by about fifteen percent of each other, but both round to zero at any practical number of decimal places, so the difference cannot affect a conclusion.
OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, p. 254
Sorting
Check n at least 20 and p at most 0.05.
Sort into buckets
Sort each binomial.
Item (d) fails only on n, and mildly: with ten trials at one percent the exact binomial is trivial to compute anyway, so there is nothing to gain. The approximation earns its place when the exact computation is awkward, which needs n to be genuinely large.
Worked example
The same n with a much larger p.
\[ n = 200, \; p = 0.4 \]
Check n
Why: Two hundred trials.
Check p
Why: Four tenths.
\[ \text{far above } 0.05 \]
Say what goes wrong
Why: The binomial's variance is npq, not np.
Compare
Why: npq is 48 against a Poisson variance of 80.
Figure (svg): The solution to Worked example when the conditions fail shown as a ladder of expressions, one row per legal move
\[ npq = 48 \quad \text{against} \quad \mu = 80 \]
Verify: confirm why small p is the essential condition
Why: The binomial variance is npq and the Poisson variance is np, so the two agree only when q is close to one — that is, when p is close to zero. At p equal to 0.004 the value of q is 0.996 and the variances differ by under half a percent; at p equal to 0.4 they differ by forty percent. The condition on p is doing the real work, and the condition on n merely ensures the mean is not too small for the shape to settle.
OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 253-254
Trap
\[ n = 200, \; p = 0.4 \;\to\; \text{use } P(80) \]
Apply the approximation because n is large
Why: Two hundred trials certainly counts as many.
\[ npq = 48 \quad \text{but} \quad \sigma^2_{\text{Poisson}} = 80 \]
The Poisson has no q to shrink its variance, so it overstates the spread by two thirds whenever p is not small.
\[ \text{check BOTH: } n \ge 20 \text{ and } p \le 0.05 \]
Verify the small-p condition, which is the binding one
Why: The variances agree only when q is near one.
Both conditions must hold, and it is p that usually decides. A binomial with two hundred trials at p equal to 0.4 is a well-behaved distribution that should simply be computed as a binomial, or handled by the normal approximation of chapter 6 — which is the tool for exactly the case where p is not small.
Faded example
A binomial with 400 trials and a success probability of 0.01.
Fill in the blanks
\mu = np = 400 \times 0.01 = 4
Why: The approximating Poisson uses mu equal to np, which is 4. Both conditions hold comfortably here, so the two families agree to several decimal places.
Two truths and a lie
All three concern the approximation.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. With n equal to 200 and p equal to 0.4 the trials are plentiful but the variances differ by forty percent, since npq is 48 against the Poisson's 80. Both conditions are needed and the one on p does the real work.
Prediction
Commit before reasoning.
Predict first
When a Poisson approximates a binomial, whose variance is larger?
Correct: The Poisson's.
Why: The binomial variance npq is smaller than np whenever q is below one, which is always. So the Poisson is slightly wider, and the gap is negligible when p is small and substantial when it is not. Knowing the direction is useful: a Poisson approximation errs by spreading the probability a little too far into the tails.
Comparison
Fill the blanks. The chapter's whole map on one table.
Comparison matrix
| Family | What X counts | Parameters |
|---|---|---|
| Binomial | successes in n independent trials | n and p |
| Geometric | trials until the first success | p alone |
| Hypergeometric | items from the group of interest, drawn without replacement | r, b and n |
| Poisson | occurrences in an interval | mu alone |
Three of the four count successes among trials, and the Poisson counts occurrences with no trials at all. That is why it is the one family whose recognition question is different: instead of asking which binomial characteristic fails, ask whether there are any trials to speak of.
Pattern
Six steps, and the second is where the errors hide.
If the problem does give a fixed n and a small p, both this family and the binomial apply. Check n at least 20 and p at most 0.05, and use whichever is less work.
OpenStax Introductory Business Statistics 2e, §4.4 Poisson Distribution §4.4 Poisson Distribution
Check
Identify the family.
Check your understanding
Which of these is a Poisson experiment?
Answer: A
Why: Emails arrive at a rate over an interval with no identifiable trials, and the count has no upper limit.
Check
Rescaling the mean.
Check your understanding
A shop serves an average of 24 customers in a 3-hour period. What is mu for a 30-minute interval?
Answer: A
Why: Thirty minutes is a sixth of three hours, so the mean is 24 divided by 6, which is 4.
Check
The standard deviation.
Check your understanding
X follows P(64). What is the standard deviation?
Answer: A
Why: The variance equals the mean of 64, so the standard deviation is its square root, 8.
Real world
A hospital emergency department admits an average of 9 patients between midnight and 6 a.m. The night staffing plan can handle up to 4 admissions in any two-hour block without calling in support, and the manager wants to know how often support will be needed.
Discussion prompt
Model the admissions, compute the probability that a given two-hour block exceeds the plan, and say what the model assumes that a real emergency department might violate.
Hint: Rescale first, then use a complement.
Answer:
Rescale onto the two-hour block. Nine admissions over six hours is a rate of 1.5 per hour, so a two-hour block has a mean of 3. The distribution is X following P(3), and the standard deviation is the square root of three, about 1.73.
The probability of exceeding the plan is about 0.185. More than four admissions means one minus the cumulative probability at four, which is one minus about 0.8153. So roughly one two-hour block in five needs support — across three blocks a night, support is needed on most nights at least once.
\[ \mu = 9 \times \tfrac{2}{6} = 3, \qquad P(x > 4) = 1 - P(x \le 4) \approx 0.185 \]
The assumptions are the interesting part, and the first characteristic fails twice over here. It requires a known average rate, and emergency admissions are not uniform through the night — the hours right after midnight are typically busier than those before dawn. Modelling the night as three blocks with a single mean of 3 therefore understates the risk early and overstates it late.
It also requires independence from the time since the last event, and a multi-casualty incident brings several admissions at once, which is precisely a violation. Both failures push in the same direction: real admission counts would show a variance above their mean, so the true probability of exceeding the plan is higher than 0.185. The honest conclusion is that the Poisson gives a floor on how often support is needed, not an estimate — and the way to check it is to compare the observed variance of past blocks against their mean.
Commit first
Answer, then rate your confidence honestly.
Predict first
Why must a more-than question about a Poisson variable be answered with a complement?
Correct: Because the upper tail is infinite.
\[ P(x > k) = 1 - P(x \le k) = 1 - \sum_{i=0}^{k} \frac{\mu^i e^{-\mu}}{i!} \]
Why: A Poisson variable takes every whole number from zero upward, so there is no last value at which a tail sum could stop. Subtracting a finite cumulative probability from one is the only exact route. That is a real difference from the binomial, where the tail is finite and could be added directly if one chose to.
Explain it
They got 0.983 for Leah's break by using a mean of six, and cannot see what went wrong.
Discussion prompt
In three sentences or fewer, locate the error.
Hint: Ask them what interval the six refers to.
Answer:
Ask what period the six calls covers: two hours, while the question asks about a fifteen-minute break.
Point out that their answer says Leah is interrupted more than once in almost every break, which cannot be right for someone taking six calls across a whole two hours.
Rescaling gives a mean of 0.75 for the quarter hour, and the answer drops to 0.1734 — about one break in six, which matches the intuition.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For recognition, ask whether you can point at a trial. For rescaling, write the rate as a fraction with units and check the direction before computing. For complements, write out which values are in and which are out. For the approximation, check n at least 20 and p at most 0.05, remembering that p is the binding condition. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Twenty minutes. This one closes the chapter, so make it a map.
Draw it
Down the left, list the four families: binomial, geometric, hypergeometric and Poisson. Beside each, write what X counts, what parameters it needs, and one phrase in a problem that signals it. Draw an arrow from the binomial to each of the other three, and label each arrow with what it relaxes. In the middle of the page, work Example 4.29 completely: six calls in two hours, rescaled to a fifteen-minute interval, with the mean, the six exact probabilities from zero to five, and P(x > 1) as a complement. Draw the stems to scale and shade the four that make up the answer. Below that, write the mean and standard deviation formulas for all four families side by side and circle the one that needs only a single number. Finish with two approximation rules: hypergeometric to binomial when the sample is under five percent of the population, and binomial to Poisson when n is at least twenty and p at most 0.05 — and beside each, write which quantity is being matched and which is only approximately preserved.
Check your six probabilities by adding them: they should total about 0.9999 rather than exactly one, because the tail above five is small but not empty. That shortfall is itself worth a sentence — it is the clearest possible reminder that a Poisson variable has no largest value.
Recap
Five things, and the chapter's whole map besides.
| If you see | Then |
|---|---|
| A rate over time, distance, area or volume | Poisson: X counts occurrences |
| No n anywhere in the problem | Poisson rather than a trial-counting family |
| A rate quoted over a different interval | Rescale by the ratio of lengths first |
| More than, or greater than | One minus a cumulative value: the tail is infinite |
| A request for the standard deviation | The square root of mu, not mu |
| A variance far above the mean in real data | The Poisson does not fit: look for clumping |
| n at least 20 and p at most 0.05 | A Poisson with mu equal to np will serve |
That closes the discrete families. Chapter 5 changes the question entirely: when a variable can take any value in an interval rather than isolated whole numbers, single values have probability zero and probability becomes area under a curve. The uniform and exponential distributions are where that idea is built, and the normal distribution of chapter 6 is where it pays off.
OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 250-254 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.