4.6 Poisson Distribution

The last discrete family in the chapter and the only one that does not count successes among trials. A Poisson variable counts occurrences in a fixed interval of time or space, when those events happen with a known average rate and independently of the time since the last one. It is written X follows P of mu, where mu is the mean number of occurrences per interval, and mu scales in proportion when the interval is lengthened or shortened. Its single parameter is both its mean and its variance, so the standard deviation is the square root of mu, and the family doubles as an approximation to the binomial when the number of trials is large and the success probability small.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 4.6 Poisson Distribution

Title

Statistics · Chapter 4 — Discrete Random Variables

Poisson Distribution

2. By the end of this lesson you can

Objectives

Five outcomes. The second is where most of the arithmetic errors live.

OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 250-254 — the section these objectives are drawn from

3. What you already have

Warm-up

Every family so far counted successes among a fixed or repeated set of trials.

Discussion prompt

A call centre receives about six calls in two hours. What is the probability of exactly two calls in the next fifteen minutes — and what would you use as n and p?

Hint: Try to identify the trials. How many chances to receive a call are there in fifteen minutes?

Answer:

There is no n. A caller may ring at any instant, so there is no list of trials to count successes among — and without an n the binomial has nothing to work with, whatever value of p you might invent.

What the situation does give is a rate: six calls per two hours, which is three per hour and 0.75 per quarter hour. That single number is enough, because the count of calls in an interval is determined by the average rate and the length of the window.

This section's family takes exactly that as its one parameter. It is written X follows P(mu), where mu is the mean number of occurrences in the interval of interest, and it answers the question with no n and no p at all.

4. Count occurrences in an interval, not successes among trials

Concept

A Poisson experiment has two main characteristics. The first is that it gives the probability of a number of events occurring in a fixed interval of time or space, if these events happen with a known average rate and independently of the time since the last event. The second is that it may be used to approximate the binomial when the probability of success is small and the number of trials is large. The random variable X is the number of occurrences in the interval of interest.

Poisson distribution — The distribution of the number of occurrences in a fixed interval of time or space, when occurrences happen at a known average rate and independently of the time since the last one. Written X follows P(mu), where mu is the mean for the interval of interest.

\[ X \sim P(\mu) \]

The second characteristic is worth noticing, because it is unusual: it does not describe the experiment at all, but says what the family is good for. The book states it as part of the definition, which is a fair reflection of how the Poisson is actually used — it is both a model in its own right and the standard stand-in for an unwieldy binomial. This lesson's last idea returns to it with the book's own worked comparison.

Figure (svg): The two characteristics of a Poisson experiment listed as a numbered procedure

X is the number of occurrences in the interval of interest, and mu is the mean for that interval.

OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, p. 250

5. The two characteristics

Section

Section 1

6. A known average rate, and no memory

Concept

The first characteristic carries three requirements at once: a fixed interval of time or space, a known average rate at which events happen, and independence from the time since the last event. That last phrase is what makes the family memoryless.

independently of the time since the last event — Knowing that an hour has passed with no call tells you nothing about the next fifteen minutes. The process does not become due for an event, and it does not tire — each interval faces the same distribution as any other of the same length.

\[ P(\text{event in } I_1) = P(\text{event in } I_2) \text{ whenever } |I_1| = |I_2| \]

The book's illustration is a book editor counting misspelled words, averaging five per hundred pages, and it names the interval explicitly: the interval is the 100 pages. Naming the interval is not a formality — every later step depends on it, and the book asks for it first in each of its examples. Its other setups follow the same shape: a bank expecting six bad checks per day, a news reporter saying uh twice per broadcast, an emergency room seeing five patients per hour.

Figure (svg): The two characteristics of a Poisson experiment listed as a numbered procedure

X is the number of occurrences in the interval of interest, and mu is the mean for that interval.

OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 250-251 — the two characteristics, and Examples 4.26 through 4.28

7. Two characteristics, only one of which describes the experiment

Picture it

A rate over a fixed interval, and a use the family can be put to.

Figure (svg): The two characteristics of a Poisson experiment listed as a numbered procedure

X is the number of occurrences in the interval of interest, and mu is the mean for that interval.

Notice what is absent from the first characteristic: any mention of trials, of n, or of a success probability. That absence is the family's defining feature and the reason it needed a section of its own, since none of the earlier formulas could be adapted to a situation with nothing to count trials over.

8. Worked example: identifying a Poisson setup

Worked example

Example 4.29, checked against both characteristics.

\[ \text{Leah answers about six calls in a two-hour period} \]

Look for trials

Why: There is no fixed number of chances to receive a call.

Find the rate and the window

Why: Six calls per two hours.

Check the first characteristic

Why: Any two intervals of equal length are alike.

Check the second

Why: A call does not make the next more or less likely.

Figure (svg): The solution to Worked example identifying a Poisson setup shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ X \sim P(\mu), \quad \mu \text{ from the rate and the interval} \]

Verify: confirm that no binomial reading is available

Why: To use a binomial you would need to name n, and there is no natural candidate: a call may come at any instant of the two hours, so the number of opportunities is not finite. Some texts motivate the Poisson as the limit of a binomial with n going to infinity and p to zero, which is a fair description of what it does — but here it is simply the right model from the start, with a rate as its only input.

OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 251-252

9. Trials, or occurrences?

Sorting

Ask whether you can point at an identifiable trial.

Sort into buckets

Sort each experiment.

Poisson: occurrences in an interval
typographical errors on a page; cars passing a junction in a minute; earthquakes in a region in a year
Trial-counting family
heads in 20 coin flips; defective items in a sample of 30
pois
There is a rate and a window but no list of trials, and no upper limit on the count.
trial
A fixed set of attempts exists, each of which succeeds or fails.

The upper limit is a quick secondary test. Twenty flips cannot give twenty-one heads, but a page could in principle carry any number of errors — an unbounded count almost always signals a Poisson.

10. Worked example: naming the interval

Worked example

Examples 4.26 through 4.28. The book asks for the interval first every time.

\[ \text{five misspelled words per } 100 \text{ pages; six bad checks per day; two uhs per broadcast} \]

Name it for the editor

Why: A hundred pages of text.

Name it for the bank

Why: One day of checks.

Name it for the reporter

Why: One broadcast.

Say what all three share

Why: A known rate and a fixed window, with no trials.

Figure (svg): The solution to Worked example naming the interval shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{interval} \in \{\text{time}, \text{space}\} \]

Verify: confirm the rate must be stated per unit of that same interval

Why: A rate of five misspellings per hundred pages and a question about two hundred pages are compatible; the same rate and a question about five chapters are not, until a chapter's length is known. Every Poisson problem needs the rate and the question expressed in the same units, and mismatched units are the commonest source of a wrong mu — which the next idea takes up directly.

OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 250-251

11. Trap: inventing an n to force a binomial

Trap

The trap

\[ \text{six calls in two hours} \;\to\; n = 120 \text{ minutes}, \; p = \tfrac{6}{120} \]

Treat each minute as a trial that either brings a call or does not

Why: That produces an n and a p, so the binomial formula applies.

\[ \text{but two calls can arrive in the same minute} \]

The trials are not really binary, so the model undercounts: it caps the answer at one call per minute when the real process has no such limit.

The fix

\[ X \sim P(\mu) \text{ with } \mu \text{ from the rate} \]

Use the rate directly, with no trials at all

Why: The Poisson was built for exactly this situation.

The invented binomial is not absurd — with fine enough divisions it converges on the Poisson, which is one way of deriving the family. But it is unnecessary work with a built-in error, and the error grows as the intervals get coarser. If you cannot point at a trial, the answer is not to manufacture one.

12. The notation

Fill the middle

How a Poisson distribution is written.

Fill in the blanks

X \sim P(\mu), \textinterval ___

Why: Per interval — and specifically the interval of interest, which is the one the question asks about rather than the one the rate was quoted over. Getting those two to agree is the next idea.

13. One of these is false

Two truths and a lie

All three concern the characteristics.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The interval may be one of time or of space
  • C. The events occur independently of the time since the last event
  • B. A Poisson experiment has a fixed number of trials

Survives elimination: B

Why: The survivor is false, and its falsity is the point of the whole section. A Poisson experiment has no trials at all, which is why neither n nor p appears in the family's own formula and why X has no upper limit. The one place n and p do appear is the second characteristic, where a binomial is being approximated rather than described.

14. What breaks the first characteristic?

Prediction

Commit before reasoning.

Predict first

Which situation violates the requirement of a known average rate holding across intervals?

  • Calls arriving at a rate that is higher at 9 a.m. than at 3 p.m.
  • A page with more errors than average
  • Two calls arriving in the same minute
  • A count with no upper limit

Correct: Calls arriving at a rate that varies by time of day.

Why: Two intervals of the same length then face different rates, which is exactly what the first characteristic forbids. The usual fix is to model the busy and quiet periods separately, each with its own mu. The other options are all consistent with a Poisson: an above-average page is ordinary variation, simultaneous calls are expected, and an unbounded count is a feature of the family.

15. Getting mu onto the right interval

Section

Section 2

16. Mu scales with the length of the window

Concept

The rate is usually quoted over one interval and the question asked about another. Because occurrences happen at a constant average rate, the mean scales in direct proportion: halve the interval and you halve mu.

the interval of interest — The window the question asks about. The stated rate must be rescaled onto it before any probability is computed, by multiplying by the ratio of the two interval lengths.

\[ \mu_{\text{new}} = \mu_{\text{stated}} \times \frac{\text{new interval}}{\text{stated interval}} \]

This is the step where most Poisson errors are made, and they are invisible afterwards: a probability computed from the wrong mu looks entirely reasonable. The habit worth building is to write the rate as a fraction with units, then multiply, so that the units cancel and leave a pure count. Six calls per 120 minutes, times 15 minutes, gives 0.75 calls.

Figure (svg): Four bars of decreasing length showing the mean number of calls falling in proportion as the interval shortens from two hours to fifteen minutes

Getting mu onto the interval the question asks about is the first step of every Poisson problem, and the commonest place to go wrong.

OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 251-252 — Example 4.29 and the rescaling

17. The same rate, four windows

Picture it

Six calls in two hours, expressed over shorter intervals.

Figure (svg): Four bars of decreasing length showing the mean number of calls falling in proportion as the interval shortens from two hours to fifteen minutes

Getting mu onto the interval the question asks about is the first step of every Poisson problem, and the commonest place to go wrong.

Nothing about the process changes down the four rows — only the window being asked about. That is worth being clear on, because the phrase changing the mean sounds like changing the situation, and it is not: the rate of three calls an hour is constant throughout.

18. Worked example: Leah's fifteen-minute break

Worked example

Example 4.29's setup. The rate is given over two hours.

\[ \text{six calls per two hours; the interval of interest is } 15 \text{ minutes} \]

Write the rate with units

Why: Six calls per 120 minutes.

\[ \frac{6}{120}\text{ calls per minute} \]

Multiply by the new interval

Why: Fifteen minutes.

\[ 6 x 15 / 120 \]

Evaluate

Why: Six times one eighth.

\[ 0.75 \]

State the distribution

Why: The mean over a quarter hour.

Figure (svg): The solution to Worked example Leah's fifteen-minute break shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \mu = 6 \times \frac{15}{120} = 0.75 \]

Verify: confirm the direction of the change

Why: Fifteen minutes is an eighth of two hours, so the mean must be an eighth of six — smaller, as a shorter window should be. A common slip is to divide when multiplying is called for, giving 48, which is absurdly larger than the two-hour mean and is caught instantly by asking whether the answer should have gone up or down. Always check the direction before computing anything from mu.

OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, p. 252

19. Rescale the mean

Faded example

A machine produces 12 defective items per 8-hour shift. Find mu for a 2-hour period.

Fill in the blanks

\mu = 12 \times \frac83} = ___

Why: Two hours is a quarter of the shift, so the mean is a quarter of twelve, which is 3. The denominator is the interval the rate was quoted over, and the numerator the one asked about.

20. Worked example: text messages per hour

Worked example

Example 4.31, where the rescaling and the probability are asked together.

\[ 41.5 \text{ texts per day; find } P(x = 2) \text{ and } P(x > 2) \text{ per hour} \]

Rescale onto one hour

Why: Forty-one and a half over twenty-four.

\[ \mu = 1.7292 \]

Compute the exact value

Why: Exactly two texts in the hour.

\[ 0.2653 \]

Compute the cumulative value

Why: At most two.

\[ 0.7495 \]

Subtract for the tail

Why: One minus that.

\[ 0.2505 \]

Figure (svg): The solution to Worked example text messages per hour shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \mu = \frac{41.5}{24} \approx 1.7292, \quad P(x=2) \approx 0.2653, \quad P(x>2) \approx 0.2505 \]

Verify: confirm that the units cancel to leave a pure count

Why: Texts per day, times days, gives texts — a count with no units attached, which is what mu must always be. If the units do not cancel, the rate and the interval were not expressed compatibly, and that is the signal to convert one before multiplying. The direction check also passes: an hour is a twenty-fourth of a day, so the mean falls from 41.5 to about 1.73.

OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, p. 254

21. Trap: using the rate as stated

Trap

The trap

\[ \text{six calls in two hours} \;\to\; X \sim P(6) \]

Take mu straight from the rate

Why: Six is the number the problem gives.

\[ \text{but the question asks about } 15 \text{ minutes} \]

A mean of six calls in a quarter hour describes a very different call centre, and every probability computed from it is wrong.

The fix

\[ \mu = 6 \times \tfrac{15}{120} = 0.75, \quad X \sim P(0.75) \]

Rescale onto the interval the question asks about

Why: The rate is per two hours; the question is per quarter hour.

The damage is quiet. With mu equal to six, P(x > 1) comes out around 0.98 instead of 0.17 — a completely different conclusion about Leah's break, reached with correct arithmetic on the wrong parameter. Reading the interval out of the question BEFORE looking at the rate is the habit that prevents it.

22. Which direction?

Estimation

A rate of 20 customers per hour; the question asks about 10 minutes.

Predict first

Should mu be larger or smaller than 20, and roughly what?

  • Smaller, about 3.3
  • Larger, about 120
  • Smaller, about 2
  • The same, 20

Correct: Smaller, about 3.3.

Why: Ten minutes is a sixth of an hour, so the mean is twenty divided by six, about 3.33. The direction check comes first and rules out two of the options immediately: a shorter window must expect fewer customers, so anything at or above 20 is wrong before any arithmetic is done.

23. One of these is false

Two truths and a lie

All three concern rescaling.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Mu scales in direct proportion to the interval length
  • C. Mu is a pure count with no units attached
  • B. Rescaling the interval changes the underlying rate

Survives elimination: B

Why: The survivor is false. The rate of three calls an hour is a property of the call centre and does not change when you ask about a different window — only the expected count over that window changes. Confusing the two makes the scaling feel arbitrary when it is simply multiplication.

24. Explain the scaling

Explain it

A classmate got mu = 48 for Leah's fifteen minutes by dividing instead of multiplying.

Discussion prompt

In two sentences or fewer, show them the error without doing the arithmetic for them.

Hint: Ask whether a shorter break should expect more calls or fewer.

Answer:

Ask whether fifteen minutes should bring more calls or fewer than two hours: fewer, obviously, so any answer above six is wrong before the arithmetic is checked.

Then point at the fraction: fifteen over a hundred and twenty is less than one, so multiplying by it must shrink the six — and 0.75 is what that gives.

25. Computing probabilities

Section

Section 3

26. Exact values, and cumulative ones

Concept

Once mu is on the right interval, the probability of exactly x occurrences follows from mu alone. Questions asking for more than, at least or at most are answered by adding the relevant exact values, usually through a complement.

cumulative probability — P(x is at most k), the sum of the exact probabilities from zero up to k. Because a Poisson variable has no upper limit, questions about more than k are always answered as one minus a cumulative value rather than by adding a tail.

\[ P(x) = \frac{\mu^x e^{-\mu}}{x!} \]

The absence of an upper limit changes how the complement is used. With a binomial you could in principle add the tail directly, since it is finite; with a Poisson the tail is infinite and you have no choice but to subtract from one. That makes the complement rule from section 3.2 not a convenience here but a necessity.

Figure (svg): The Poisson distribution for a mean of nought point seven five calls, falling steeply from its peak at zero

With a mean below one, zero occurrences is the single likeliest outcome, and the tail thins fast.

OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 251-253 — the notation and Example 4.29

27. Leah's distribution, and the tail in question

Picture it

Six stems, with the four that make up P(x > 1) shaded.

Figure (svg): The Poisson distribution for a mean of nought point seven five calls, falling steeply from its peak at zero

With a mean below one, zero occurrences is the single likeliest outcome, and the tail thins fast.

The shaded region looks small, and it is: 0.1734, a little over one chance in six. That is the answer to whether Leah is likely to be interrupted more than once, and its smallness is worth noticing — with a mean of 0.75 calls, being interrupted twice is genuinely unusual.

28. Worked example: more than one call

Worked example

Example 4.29's question. The book gives 0.1734.

\[ X \sim P(0.75); \text{ find } P(x > 1) \]

Recognise the infinite tail

Why: More than one means 2, 3, 4, and on forever.

Write the complement

Why: One minus at most one.

\[ 1 - P(x \le 1) \]

Compute the two exact values

Why: Zero calls and one call.

\[ 0.4724 + 0.3543 \]

Subtract

Why: One minus their sum.

\[ 1 - 0.8266 \]

Figure (svg): The solution to Worked example more than one call shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(x > 1) = 1 - P(x \le 1) \approx 1 - 0.8266 = 0.1734 \]

Verify: confirm the boundary is excluded correctly

Why: More than one excludes one itself, so the complement is at most one and includes both zero and one. Had the question said at least one, the complement would have been zero alone and the answer 0.5276 — three times larger. The strictness of the inequality is the whole difference, and section 4.3's habit of writing out which values are in and which are out applies here unchanged.

OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, p. 252

29. Set up a complement

Faded example

X follows P(3). Find P(x > 2).

Fill in the blanks

P(x > 2) = 1 - P(x \le 2) = 1 - [P(0) + P(1) + P(2)]

Why: More than two excludes two, so the complement is at most two and contains three exact values: zero, one and two. The tail above is infinite, so subtracting is the only route.

30. Worked example: a large mean

Worked example

Example 4.30: an email user receives 147 emails a day, on average.

\[ X \sim P(147); \text{ find } P(x = 160) \text{ and } P(x \le 160) \]

Compute the exact value

Why: The probability of precisely 160.

\[ 0.0180 \]

Compute the cumulative value

Why: At most 160.

\[ 0.8666 \]

Compute the standard deviation

Why: The square root of the mean.

\[ \text{about } 12.12 \]

Locate 160 in the distribution

Why: Thirteen above a mean of 147.

Figure (svg): The solution to Worked example a large mean shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(160) \approx 0.0180, \quad P(x \le 160) \approx 0.8666, \quad \sigma \approx 12.12 \]

Verify: confirm the two answers are consistent with each other

Why: The cumulative value of 0.8666 says 160 sits at about the 87th percentile, and 160 is about 1.07 standard deviations above the mean — which for a roughly symmetric distribution should put it near the 86th percentile. The two agree, which is a real check on both. The exact value of 0.0180 is also sensible: with a spread of about twelve, the probability is shared among roughly fifty plausible values, so no single one can be large.

OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 253-254

31. Error analysis: four attempts at P(x > 1) for Leah's break

Error analysis

A mean of 0.75 calls in a fifteen-minute interval.

Annotate

On: \( \begin{aligned} &(1)\; 1 - P(x = 1) \\ &(2)\; 1 - P(x \le 1) \text{ with } \mu = 6 \\ &(3)\; P(x = 2) + P(x = 3) + P(x = 4) \\ &(4)\; 1 - P(x \le 1) \text{ with } \mu = 0.75 \end{aligned} \)

  • (1) subtracts only P(x = 1), forgetting that zero calls is also not more than one. It gives about 0.646.
  • (2) uses the two-hour mean without rescaling. It gives about 0.983 — a completely different conclusion, reached with correct arithmetic.
  • (3) adds a tail that has no end, and stopping at four silently discards every larger value. It undershoots.
  • (4) is correct: one minus the cumulative probability at one, with mu on the right interval, giving 0.1734.

Error (2) is the one to fear. Errors (1) and (3) produce answers that look off once you think about the shape, but (2) is internally consistent and wrong only in its input — which is why the interval belongs in the very first line of the working.

32. Which complement?

Discrimination

Match each phrase to the cumulative value you would subtract from one.

Sort into buckets

Sort each phrase by its complement.

Subtract P(x at most 3)
more than 3; greater than 3
Subtract P(x at most 2), or read it directly
at least 3; greater than or equal to 3; fewer than 3
three
The phrase excludes three itself, so three belongs in the complement.
two
The phrase includes three, so the complement stops at two.

33. One of these is false

Two truths and a lie

All three concern computing Poisson probabilities.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. A more-than question must be answered by a complement
  • C. P(x = 0) is e to the minus mu
  • B. The probabilities stop at x equal to mu rounded up

Survives elimination: B

Why: The survivor is false. A Poisson variable takes every whole number from zero upward with positive probability, however far above the mean — the values merely become very small. That unboundedness is what forces the complement, and it distinguishes the family from every other one in this chapter.

34. Where is the peak?

Prediction

Commit before reasoning.

Predict first

For X following P(0.75), which single value of x is most likely?

  • 0
  • 1
  • 0.75
  • 2

Correct: 0.

Why: With a mean below one, no occurrences is the likeliest outcome, at about 0.4724 against 0.3543 for exactly one. The mean of 0.75 is not itself a possible value, since X counts whole occurrences — a reminder from section 4.2 that a mean need not be attainable. Whenever mu is below one the distribution peaks at zero and falls from there.

35. The mean and the standard deviation

Section

Section 4

36. One parameter for both

Concept

The mean of a Poisson distribution is mu, its variance is also mu, and so its standard deviation is the square root of mu. No second parameter is needed or available.

equidispersion — The property that the variance equals the mean. It is unique among the families in this chapter, and it gives the Poisson a testable signature: real counts whose variance far exceeds their mean are not Poisson.

\[ \mu_X = \mu, \quad \sigma^2_X = \mu, \quad \sigma_X = \sqrt{\mu} \]

The binomial needed n and p to give np and the square root of npq; the geometric needed p. The Poisson needs only mu, which is why a single number fully specifies the distribution. The practical consequence is that a Poisson count with a large mean is proportionally tighter: with mu equal to 147 the standard deviation is about 12.12, which is only eight percent of the mean, while with mu equal to 4 the standard deviation of 2 is fully half of it.

Figure (svg): A diagram showing that the Poisson mean and variance are both mu, so the standard deviation is the square root of mu

Because the variance equals the mean, a Poisson count with a large mean is proportionally tighter than one with a small mean.

OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 253-254 — Example 4.30 and the standard deviation

37. Two numbers from one

Picture it

Compared with the two parameters the other families required.

Figure (svg): A diagram showing that the Poisson mean and variance are both mu, so the standard deviation is the square root of mu

Because the variance equals the mean, a Poisson count with a large mean is proportionally tighter than one with a small mean.

That the variance equals the mean is a strong claim about real data, and it can be checked. Counts that clump — accidents at a junction, say, where one crash causes others — show a variance well above the mean, and that excess is the standard evidence that a Poisson model does not fit.

38. Worked example: the standard deviation at mu = 147

Worked example

Example 4.30's third part.

\[ X \sim P(147) \]

Recall the variance

Why: It equals the mean.

\[ 147 \]

Take the square root

Why: The standard deviation.

\[ \text{about } 12.12 \]

Express it as a fraction of the mean

Why: Twelve over 147.

\[ \text{about } 8 \% \]

Interpret

Why: Counts near 147 give or take a dozen.

Figure (svg): The solution to Worked example the standard deviation at mu 147 shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \sigma = \sqrt{147} \approx 12.12 \]

Verify: confirm against the cumulative value already computed

Why: The earlier part found P(x at most 160) to be 0.8666. Since 160 is 13 above the mean, it sits about 1.07 standard deviations up, and for a roughly symmetric distribution that should correspond to a percentile in the middle eighties — which 0.8666 is. The two results confirm each other, and a standard deviation that failed this check would signal an error in one of them.

OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, p. 254

39. Compute the spread

Faded example

A hospital ward admits an average of 25 patients per day.

Fill in the blanks

\sigma = \sqrt25} = 5

Why: The variance equals the mean of 25, so the standard deviation is its square root, 5. Daily admissions should typically run 25 give or take about 5.

40. Worked example: relative spread at two means

Worked example

The same family at a small mean and a large one.

\[ \mu = 4 \text{ against } \mu = 400 \]

Standard deviation at 4

Why: The square root of four.

\[ 2 \]

As a fraction of the mean

Why: Two over four.

\[ 50 \% \]

Standard deviation at 400

Why: The square root of four hundred.

\[ 20 \]

As a fraction of the mean

Why: Twenty over four hundred.

\[ 5 \% \]

Figure (svg): The solution to Worked example relative spread at two means shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \frac{\sqrt{\mu}}{\mu} = \frac{1}{\sqrt{\mu}} \]

Verify: confirm the general rule behind the two cases

Why: The relative spread is the square root of mu over mu, which simplifies to one over the square root of mu — so it falls as the mean grows, and falls slowly. Multiplying the mean by a hundred divides the relative spread by ten. That is why counting over a longer interval gives a proportionally more reliable estimate of a rate, and it is the same square-root behaviour that will govern sample means in chapter 7.

OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 253-254

41. Trap: taking the standard deviation to be mu

Trap

The trap

\[ \mu = 147 \;\Rightarrow\; \sigma = 147 \]

Read the variance equals the mean as the sd equals the mean

Why: Both are called spread, so they blur together.

\[ \text{a mean of } 147 \text{ with a spread of } 147 \]

That would put a count of zero within one standard deviation of the mean, which does not describe the distribution at all.

The fix

\[ \sigma^2 = \mu = 147 \;\Rightarrow\; \sigma = \sqrt{147} \approx 12.12 \]

It is the VARIANCE that equals the mean, so take a square root

Why: Section 2.7's distinction between variance and standard deviation, applied here.

The sanity check is the cumulative value: 160 should be a bit over one standard deviation above 147, and it comes out at 1.07 with sigma equal to 12.12. With sigma equal to 147 it would be less than a tenth of a standard deviation up, which cannot square with a percentile of 87.

42. One of these is false

Two truths and a lie

All three concern the mean and spread.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The variance equals the mean
  • C. The relative spread falls as the mean grows
  • B. The standard deviation equals the mean

Survives elimination: B

Why: The survivor is false, and it confuses the variance with the standard deviation. The variance equals mu; the standard deviation is its square root. At mu equal to 147 that is the difference between 147 and 12.12.

43. Is the count unusual?

Estimation

A junction averages 16 crossings per minute.

Predict first

Would a count of 28 in one minute be unusual?

  • Yes: it is three standard deviations above the mean
  • No: it is within one standard deviation
  • Yes: any count above the mean is unusual
  • It cannot be judged without more information

Correct: Yes: three standard deviations above.

Why: The standard deviation is the square root of sixteen, which is 4, so 28 sits twelve above the mean and therefore three standard deviations up. Section 2.7's rule of thumb calls that unusual, and it is the kind of judgement the Poisson makes possible from a single number.

44. What does an excess variance mean?

Prediction

Commit before reasoning.

Predict first

A count of accidents has mean 5 and variance 22. What does that suggest?

  • The mean was computed wrongly
  • The events clump, so the Poisson's independence assumption fails
  • The interval was too short
  • Nothing: variance and mean are unrelated

Correct: The events clump, so independence fails.

Why: A Poisson requires the variance to equal the mean, so a variance four times the mean is direct evidence against the model. The usual cause is dependence between occurrences — one accident causing others, or a rate that varies between intervals, breaking the first characteristic's requirement of a known rate holding independently of the past. This test is the practical reason equidispersion is worth remembering.

45. Approximating a binomial

Section

Section 5

46. Large n, small p

Concept

This is the second of the two characteristics, returned to as a technique. When a binomial has a large number of trials and a small success probability, its probabilities are very close to those of a Poisson with mu equal to np. The book gives the conditions as n large, greater than 20, and p small, less than 0.05.

the Poisson approximation to the binomial — Replacing B(n, p) by P(np) when n is large and p small. The two agree closely, and the Poisson is far easier to compute because it needs no combinations.

\[ n > 20, \; p < 0.05 \;\Longrightarrow\; B(n, p) \approx P(np) \]

The characteristic itself is looser than the rule of thumb: it says the probability of success should be small, such as 0.01, and the number of trials large, such as 1,000. The sharper thresholds come with the book's worked comparison at the end of the section, where it justifies the agreement by checking n above 20 and p below 0.05. The approximation also runs the opposite way from section 4.5's, where a hypergeometric was replaced by a binomial when the population was large.

Figure (svg): A three-column table comparing binomial probabilities at n equals two hundred and p equals nought point nought one with Poisson probabilities at mu equals two, showing close agreement

When n is large and p small, the binomial is hard to compute and the Poisson is easy, and they give the same answers.

OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 250-254 — the second characteristic, and Example 4.32

47. Two columns of numbers that agree

Picture it

Two hundred trials at one percent, against a Poisson with mean two.

Figure (svg): A three-column table comparing binomial probabilities at n equals two hundred and p equals nought point nought one with Poisson probabilities at mu equals two, showing close agreement

When n is large and p small, the binomial is hard to compute and the Poisson is easy, and they give the same answers.

The agreement is not coincidence. As n grows and p shrinks with np held fixed, the binomial formula converges term by term on the Poisson one — which is the standard derivation of the family, and the reason the Poisson describes rare events among many opportunities so well.

48. Worked example: the seismic-activity comparison

Worked example

Example 4.32, worked both ways as the book does.

\[ n = 200 \text{ days, } p = 0.0102; \text{ find } P(x = 10) \text{ both ways} \]

Compute the binomial

Why: Two hundred days at just over one percent.

\[ \text{about } 0.000039 \]

Compute the approximating mean

Why: n times p.

\[ \mu = 2.04 \]

Compute the Poisson

Why: Ten occurrences at that mean.

\[ \text{about } 0.000045 \]

Check the conditions

Why: n above 20 and p below 0.05.

Figure (svg): The solution to Worked example the seismic-activity comparison shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{binomial } 0.000039 \quad \text{against} \quad \text{Poisson } 0.000045 \]

Verify: confirm the conditions the book gives for expecting agreement

Why: The book's own justification is that the approximation should be good because n is large, greater than 20, and p is small, less than 0.05 — and here n is 200 and p is 0.0102, so both hold comfortably. Note also what close means at this scale: the two differ by about fifteen percent of each other, but both round to zero at any practical number of decimal places, so the difference cannot affect a conclusion.

OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, p. 254

49. Approximate, or not?

Sorting

Check n at least 20 and p at most 0.05.

Sort into buckets

Sort each binomial.

Poisson approximation is appropriate
n = 500, p = 0.004; n = 1000, p = 0.002; n = 50, p = 0.02
Compute the binomial, or use another tool
n = 200, p = 0.4; n = 10, p = 0.01
yes
Many trials with a small success probability, so np is a stable mean and q is close to one.
no
Either p is too large for the variances to agree, or n is too small for the rule of thumb.

Item (d) fails only on n, and mildly: with ten trials at one percent the exact binomial is trivial to compute anyway, so there is nothing to gain. The approximation earns its place when the exact computation is awkward, which needs n to be genuinely large.

50. Worked example: when the conditions fail

Worked example

The same n with a much larger p.

\[ n = 200, \; p = 0.4 \]

Check n

Why: Two hundred trials.

Check p

Why: Four tenths.

\[ \text{far above } 0.05 \]

Say what goes wrong

Why: The binomial's variance is npq, not np.

Compare

Why: npq is 48 against a Poisson variance of 80.

Figure (svg): The solution to Worked example when the conditions fail shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ npq = 48 \quad \text{against} \quad \mu = 80 \]

Verify: confirm why small p is the essential condition

Why: The binomial variance is npq and the Poisson variance is np, so the two agree only when q is close to one — that is, when p is close to zero. At p equal to 0.004 the value of q is 0.996 and the variances differ by under half a percent; at p equal to 0.4 they differ by forty percent. The condition on p is doing the real work, and the condition on n merely ensures the mean is not too small for the shape to settle.

OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 253-254

51. Trap: approximating whenever n is large

Trap

The trap

\[ n = 200, \; p = 0.4 \;\to\; \text{use } P(80) \]

Apply the approximation because n is large

Why: Two hundred trials certainly counts as many.

\[ npq = 48 \quad \text{but} \quad \sigma^2_{\text{Poisson}} = 80 \]

The Poisson has no q to shrink its variance, so it overstates the spread by two thirds whenever p is not small.

The fix

\[ \text{check BOTH: } n \ge 20 \text{ and } p \le 0.05 \]

Verify the small-p condition, which is the binding one

Why: The variances agree only when q is near one.

Both conditions must hold, and it is p that usually decides. A binomial with two hundred trials at p equal to 0.4 is a well-behaved distribution that should simply be computed as a binomial, or handled by the normal approximation of chapter 6 — which is the tool for exactly the case where p is not small.

52. Find the approximating mean

Faded example

A binomial with 400 trials and a success probability of 0.01.

Fill in the blanks

\mu = np = 400 \times 0.01 = 4

Why: The approximating Poisson uses mu equal to np, which is 4. Both conditions hold comfortably here, so the two families agree to several decimal places.

53. One of these is false

Two truths and a lie

All three concern the approximation.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The approximating Poisson uses mu equal to np
  • C. The variances agree only when q is close to one
  • B. A large n alone is enough to justify the approximation

Survives elimination: B

Why: The survivor is false. With n equal to 200 and p equal to 0.4 the trials are plentiful but the variances differ by forty percent, since npq is 48 against the Poisson's 80. Both conditions are needed and the one on p does the real work.

54. Which way does the spread go?

Prediction

Commit before reasoning.

Predict first

When a Poisson approximates a binomial, whose variance is larger?

  • The Poisson's, since npq is smaller than np
  • The binomial's, since it has an extra factor
  • They are exactly equal
  • It depends on n

Correct: The Poisson's.

Why: The binomial variance npq is smaller than np whenever q is below one, which is always. So the Poisson is slightly wider, and the gap is negligible when p is small and substantial when it is not. Knowing the direction is useful: a Poisson approximation errs by spreading the probability a little too far into the tails.

55. The four discrete families

Comparison

Fill the blanks. The chapter's whole map on one table.

Comparison matrix

FamilyWhat X countsParameters
Binomialsuccesses in n independent trialsn and p
Geometrictrials until the first successp alone
Hypergeometricitems from the group of interest, drawn without replacementr, b and n
Poissonoccurrences in an intervalmu alone

Three of the four count successes among trials, and the Poisson counts occurrences with no trials at all. That is why it is the one family whose recognition question is different: instead of asking which binomial characteristic fails, ask whether there are any trials to speak of.

56. Solving a Poisson problem, in order

Pattern

Six steps, and the second is where the errors hide.

  1. Confirm there are no identifiable trials, only a rate of occurrences over an interval.
  2. Read the interval of interest out of the QUESTION, before looking at the stated rate.
  3. Rescale the rate onto that interval by multiplying by the ratio of lengths, and check the direction.
  4. Write X as P(mu), and list which values of x the question is asking about.
  5. Compute exact values directly, and any more-than question as one minus a cumulative value.
  6. Compute sigma as the square root of mu, and use it to sanity-check the answer's plausibility.

If the problem does give a fixed n and a small p, both this family and the binomial apply. Check n at least 20 and p at most 0.05, and use whichever is less work.

OpenStax Introductory Business Statistics 2e, §4.4 Poisson Distribution §4.4 Poisson Distribution

57. Check yourself 1 of 3

Check

Identify the family.

Check your understanding

Which of these is a Poisson experiment?

  • A. The number of emails arriving in an hour (correct)
  • B. The number of heads in 30 coin flips
  • C. The number of cards drawn until the first ace
  • D. The number of women on a committee of 6

Answer: A

Why: Emails arrive at a rate over an interval with no identifiable trials, and the count has no upper limit.

Why B tempts people
Thirty flips are thirty trials, each succeeding or failing: binomial.
Why C tempts people
Drawing until a first success with no fixed n is geometric.
Why D tempts people
A committee is drawn without replacement from two groups: hypergeometric.

58. Check yourself 2 of 3

Check

Rescaling the mean.

Check your understanding

A shop serves an average of 24 customers in a 3-hour period. What is mu for a 30-minute interval?

  • A. 4 (correct)
  • B. 24
  • C. 8
  • D. 144

Answer: A

Why: Thirty minutes is a sixth of three hours, so the mean is 24 divided by 6, which is 4.

Why B tempts people
That is the mean over the stated three-hour interval, not the half hour asked about.
Why C tempts people
That is the hourly rate, which is the mean for a 60-minute interval rather than a 30-minute one.
Why D tempts people
That multiplies where it should divide, and gives a mean larger than the three-hour one.

59. Check yourself 3 of 3

Check

The standard deviation.

Check your understanding

X follows P(64). What is the standard deviation?

  • A. 8 (correct)
  • B. 64
  • C. 4096
  • D. It cannot be found without a second parameter

Answer: A

Why: The variance equals the mean of 64, so the standard deviation is its square root, 8.

Why B tempts people
That is the variance, not the standard deviation. Section 2.7's distinction applies here.
Why C tempts people
That squares the mean rather than taking its root.
Why D tempts people
The Poisson has only one parameter, and it determines both the centre and the spread.

60. Where this shows up outside the textbook

Real world

A hospital emergency department admits an average of 9 patients between midnight and 6 a.m. The night staffing plan can handle up to 4 admissions in any two-hour block without calling in support, and the manager wants to know how often support will be needed.

Discussion prompt

Model the admissions, compute the probability that a given two-hour block exceeds the plan, and say what the model assumes that a real emergency department might violate.

Hint: Rescale first, then use a complement.

Answer:

Rescale onto the two-hour block. Nine admissions over six hours is a rate of 1.5 per hour, so a two-hour block has a mean of 3. The distribution is X following P(3), and the standard deviation is the square root of three, about 1.73.

The probability of exceeding the plan is about 0.185. More than four admissions means one minus the cumulative probability at four, which is one minus about 0.8153. So roughly one two-hour block in five needs support — across three blocks a night, support is needed on most nights at least once.

\[ \mu = 9 \times \tfrac{2}{6} = 3, \qquad P(x > 4) = 1 - P(x \le 4) \approx 0.185 \]

The assumptions are the interesting part, and the first characteristic fails twice over here. It requires a known average rate, and emergency admissions are not uniform through the night — the hours right after midnight are typically busier than those before dawn. Modelling the night as three blocks with a single mean of 3 therefore understates the risk early and overstates it late.

It also requires independence from the time since the last event, and a multi-casualty incident brings several admissions at once, which is precisely a violation. Both failures push in the same direction: real admission counts would show a variance above their mean, so the true probability of exceeding the plan is higher than 0.185. The honest conclusion is that the Poisson gives a floor on how often support is needed, not an estimate — and the way to check it is to compare the observed variance of past blocks against their mean.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

Why must a more-than question about a Poisson variable be answered with a complement?

  • Because the formula only computes cumulative values
  • Because the upper tail is infinite: X has no largest value to sum up to
  • Because complements are always faster
  • Because mu may not be a whole number

Correct: Because the upper tail is infinite.

\[ P(x > k) = 1 - P(x \le k) = 1 - \sum_{i=0}^{k} \frac{\mu^i e^{-\mu}}{i!} \]

Why: A Poisson variable takes every whole number from zero upward, so there is no last value at which a tail sum could stop. Subtracting a finite cumulative probability from one is the only exact route. That is a real difference from the binomial, where the tail is finite and could be added directly if one chose to.

62. Explain it to someone a year behind you

Explain it

They got 0.983 for Leah's break by using a mean of six, and cannot see what went wrong.

Discussion prompt

In three sentences or fewer, locate the error.

Hint: Ask them what interval the six refers to.

Answer:

Ask what period the six calls covers: two hours, while the question asks about a fifteen-minute break.

Point out that their answer says Leah is interrupted more than once in almost every break, which cannot be right for someone taking six calls across a whole two hours.

Rescaling gives a mean of 0.75 for the quarter hour, and the answer drops to 0.1734 — about one break in six, which matches the intuition.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Recognising a Poisson from the absence of trials
  • Rescaling the rate onto the interval of interest
  • Setting up a complement for a more-than question
  • Deciding whether a Poisson may approximate a binomial

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For recognition, ask whether you can point at a trial. For rescaling, write the rate as a fraction with units and check the direction before computing. For complements, write out which values are in and which are out. For the approximation, check n at least 20 and p at most 0.05, remembering that p is the binding condition. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the chapter on one page

Connect it up

Paper. Twenty minutes. This one closes the chapter, so make it a map.

Draw it

Down the left, list the four families: binomial, geometric, hypergeometric and Poisson. Beside each, write what X counts, what parameters it needs, and one phrase in a problem that signals it. Draw an arrow from the binomial to each of the other three, and label each arrow with what it relaxes. In the middle of the page, work Example 4.29 completely: six calls in two hours, rescaled to a fifteen-minute interval, with the mean, the six exact probabilities from zero to five, and P(x > 1) as a complement. Draw the stems to scale and shade the four that make up the answer. Below that, write the mean and standard deviation formulas for all four families side by side and circle the one that needs only a single number. Finish with two approximation rules: hypergeometric to binomial when the sample is under five percent of the population, and binomial to Poisson when n is at least twenty and p at most 0.05 — and beside each, write which quantity is being matched and which is only approximately preserved.

Check your six probabilities by adding them: they should total about 0.9999 rather than exactly one, because the tail above five is small but not empty. That shortfall is itself worth a sentence — it is the clearest possible reminder that a Poisson variable has no largest value.

65. What you can do now

Recap

Five things, and the chapter's whole map besides.

If you seeThen
A rate over time, distance, area or volumePoisson: X counts occurrences
No n anywhere in the problemPoisson rather than a trial-counting family
A rate quoted over a different intervalRescale by the ratio of lengths first
More than, or greater thanOne minus a cumulative value: the tail is infinite
A request for the standard deviationThe square root of mu, not mu
A variance far above the mean in real dataThe Poisson does not fit: look for clumping
n at least 20 and p at most 0.05A Poisson with mu equal to np will serve

That closes the discrete families. Chapter 5 changes the question entirely: when a variable can take any value in an interval rather than isolated whole numbers, single values have probability zero and probability becomes area under a curve. The uniform and exponential distributions are where that idea is built, and the normal distribution of chapter 6 is where it pays off.

OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution §4.6, pp. 250-254 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §4.6 Poisson Distribution — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 250-254
  2. OpenStax Introductory Business Statistics 2e, §4.4 Poisson Distribution — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108