4.4 Geometric Distribution

The binomial's first characteristic relaxed. A geometric experiment repeats a trial until the first success occurs and then stops, so the number of trials is the random variable rather than the number of successes, and it takes the values one, two, three and upward with no bound. The probability of a first success on trial x is q to the power x minus one times p, with no combination factor because exactly one sequence produces it. The mean is one over p, matching the intuition that a one-in-three chance takes about three tries. The distribution is also the only discrete one that is memoryless: previous failures carry no information about what comes next, which is a genuine property of independent trials and a poor model for anything that learns.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 4.4 Geometric Distribution

Title

Statistics · Chapter 4 — Discrete Random Variables

Geometric Distribution

2. By the end of this lesson you can

Objectives

Five outcomes. The second is what makes the whole family different from section 4.3's.

OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 242-247 — the section these objectives are drawn from

3. What you already have

Warm-up

Section 4.3's binomial required a fixed number of trials, and that was the first of its three characteristics.

Discussion prompt

You throw darts at a bullseye until you hit it. What is the number of trials here, and can the binomial formula describe this?

Hint: Ask whether you could write down n before starting.

Answer:

The number of trials is not fixed and cannot be known in advance: it might take one throw or twenty, and in theory it could go on forever. So the first binomial characteristic fails immediately.

What is fixed instead is the number of successes — exactly one, since you stop as soon as you hit. So the two quantities have swapped roles: the successes are now fixed and the trials are random.

That swap is the whole of this section. The book calls the result a geometric distribution, and the random variable X is the number of the trial on which the first success occurs — which is a waiting time rather than a count of successes.

4. Count the trials, not the successes

Concept

In a geometric experiment a trial is repeated until a success occurs, and then you stop. The trials are independent and the probability of a success is the same on each. The random variable X is the number of the trial on which the first success occurs, so it takes the values one, two, three and upward.

geometric experiment — A sequence of independent trials with constant success probability p, repeated until the first success. X is the number of the trial on which that success occurs, and X takes the values 1, 2, 3 and so on without upper bound. The book describes the sequence as failure, failure, failure, success, STOP.

\[ X \sim G(p), \qquad x = 1, 2, 3, \ldots \]

The book notes that the variable can be defined either way — as the number of trials until a success, or as the number of failures before one — and that each form changes the formula and the mean slightly. This lesson uses the first, which is the book's own convention in its examples and the more common one: X counts the trial on which the success lands, so its smallest value is one rather than zero.

Figure (svg): Two columns contrasting the binomial, which fixes the number of trials, with the geometric, which fixes the number of successes

The two are mirror images of one question. Which quantity is fixed in advance and which is random is the whole distinction.

OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, p. 242

5. The four characteristics

Section

Section 1

6. Repeat until a success, then stop

Concept

A trial is repeated until a success occurs. The repeated trials are independent of each other, and the probability p of a success and q of a failure are the same for each trial. The random variable X is the number of the trial on which the first success occurs.

the four characteristics — Repeat until a success and stop; the trials are independent; p and q are constant; and X is the number of the trial on which the first success occurs. In theory the number of trials could go on forever.

\[ \text{F, F, F, F, S, STOP} \;\Longrightarrow\; x = 5 \]

Only the first characteristic differs from the binomial's. Independence and a constant p are shared, which is why the same sampling caution applies: drawing without replacement from a small population breaks the third characteristic here exactly as it broke the binomial's. The book's own illustration is a die rolled until a three appears, where p stays a sixth no matter how many rolls have gone before.

Figure (svg): The four characteristics of a geometric experiment listed as a numbered procedure

In theory the number of trials could go on forever, so X takes the values 1, 2, 3 and upward without bound.

OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, p. 242 — the four characteristics and the die example

7. The checklist

Picture it

Four conditions, of which two are the binomial's.

Figure (svg): The four characteristics of a geometric experiment listed as a numbered procedure

In theory the number of trials could go on forever, so X takes the values 1, 2, 3 and upward without bound.

The fourth characteristic is really a definition rather than a condition, and it is the one that changes what any answer means. In a binomial problem an answer of 5 would be five successes; here it is five trials, of which four failed. Reading the units of X correctly is what stops a geometric answer being interpreted as a binomial one.

8. Worked example: identifying a geometric experiment

Worked example

Example 4.17. Play until you lose.

\[ \text{You play until you lose; the probability of losing is } p = 0.57. \]

Check the stopping rule

Why: You play until you lose, then stop.

Identify the success

Why: Losing is what ends the sequence.

\[ p = 0.57 \]

Define the variable

Why: The number of games played, including the losing one.

\[ X =\text{ games until the loss} \]

List the values

Why: One game at least, and no upper bound.

\[ 1, 2, 3,... \]

Figure (svg): The solution to Worked example identifying a geometric experiment shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ X \sim G(0.57), \quad \text{find } P(x = 5) \]

Verify: confirm what a success is here

Why: Losing is the success, because the success is whatever the sequence is waiting for — the book defines it that way, and the label carries no approval, exactly as in section 4.3. The variable includes the losing game itself, so x equals 5 means four wins followed by a loss. Reading the definition carefully matters because 'the number of games you play until you lose' could otherwise be taken as the number of wins, which would be one less.

OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, p. 243

9. Binomial or geometric?

Sorting

Ask which quantity is fixed in advance.

Sort into buckets

Sort each question.

Binomial: trials fixed, successes counted
How many heads in 20 flips?; How many defective items in a sample of 50?
Geometric: successes fixed at one, trials counted
How many flips until the first head?; How many reports read before finding the first relevant one?; How many darts thrown until the first bullseye?
bin
The number of trials is stated in advance and the count of successes is what varies.
geo
The experiment stops at the first success, so the number of successes is fixed at one and the number of trials varies.

The phrase 'until' is the reliable signal for a geometric problem, and a stated number of trials is the signal for a binomial one. A question containing both — how many flips until the third head — is neither, and belongs to a family this book does not cover.

10. Worked example: geometric or binomial?

Worked example

The same situation, two different questions.

\[ \text{A die is rolled. } p = \frac{1}{6} \text{ for a three.} \]

Question A: how many threes in ten rolls?

Why: Ten trials fixed, successes counted.

Question B: how many rolls until the first three?

Why: Successes fixed at one, trials counted.

Note what is fixed in each

Why: n in the first, the success count in the second.

Note what is random in each

Why: The count in the first, the number of trials in the second.

Figure (svg): The solution to Worked example geometric or binomial shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{fixed } n \to \text{binomial}; \qquad \text{fixed successes} \to \text{geometric} \]

Verify: confirm the two are genuinely different questions

Why: Question A has eleven possible answers, zero through ten, while question B has infinitely many, one upward. They cannot be the same distribution because their value sets differ. The reliable test is to ask what could be written down before the experiment starts: if n is known in advance the problem is binomial, and if instead you know you will stop at the first success it is geometric.

OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 242-243

11. Trap: reading X as a count of successes

Trap

The trap

\[ P(x = 5) \text{ with } p = 0.57 \]

Interpret x = 5 as five losses

Why: In section 4.3 the variable counted successes, so the habit carries over.

\[ \text{but here } x \text{ counts TRIALS, and there is exactly one success} \]

Five means five games were played, of which the first four were wins and the fifth a loss.

The fix

\[ x = 5 \;\Longrightarrow\; \text{four failures, then the first success on trial } 5 \]

Read the fourth characteristic every time: X is the TRIAL NUMBER

Why: The number of successes is always exactly one.

The two families use the same letters for different things, which is the main source of confusion between them. A useful habit is to write out what x equals 5 means as a sentence — 'the first success came on the fifth trial' — before doing any arithmetic. It also explains why x cannot be zero here: there must be at least one trial for a success to occur on.

12. The smallest value

Fill the middle

What X can be in a geometric experiment.

Fill in the blanks

X \text1 ___, 2, 3, \ldots \text___

Why: One, because at least one trial must occur for a success to happen on it. This is a visible difference from the binomial, whose values start at zero — a binomial count of zero successes is perfectly possible, and a geometric trial number of zero is not.

13. One of these is false

Two truths and a lie

All three concern the characteristics.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The trials are independent with a constant p
  • C. In theory the number of trials could go on forever
  • B. The number of trials is fixed in advance

Survives elimination: B

Why: The survivor is false and it is exactly the characteristic the geometric relaxes. The number of trials is the random variable, determined by when the first success happens, and it cannot be known before the experiment runs.

14. Why no upper bound?

Prediction

Commit before reasoning.

Predict first

Why does X have infinitely many possible values?

  • Because p is small
  • Because a run of failures of any length is possible, however unlikely
  • Because the trials are independent
  • It is a simplification; in practice there is a maximum

Correct: Because a run of failures of any length is possible.

Why: Nothing forbids a hundred consecutive failures — its probability is q to the hundredth, which is minute but not zero — so no value of x can be ruled out. The independence is what makes those long runs possible rather than being the reason for the unboundedness itself, and the probabilities decay fast enough that the infinitely many values still sum to one.

15. The formula

Section

Section 2

16. Failures throughout, then a success

Concept

For the first success to occur on trial x, the first x minus one trials must all fail and the xth must succeed. Because the trials are independent, that sequence has probability q multiplied by itself x minus one times, then p.

the geometric formula — P(x) equals q to the power x minus one, times p. There is no combination factor because exactly one sequence of outcomes produces a first success on trial x: failures throughout, then a success.

\[ P(x) = q^{x-1} p \]

The absence of a combination factor is the clearest structural difference from section 4.3. There, x successes in n trials could be arranged in many orders and the formula counted them; here the ordering is completely determined by the definition — every trial before the last failed, and the last succeeded — so there is exactly one arrangement and nothing to count.

Figure (svg): The geometric probability formula shown as a sequence of failures followed by one success

The exponent is x minus one because the success is the last trial and every earlier one failed.

OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 242-243 — the geometric probability, with the die example

17. One sequence, one product

Picture it

Four failures and a success, multiplied by section 3.3's rule.

Figure (svg): The geometric probability formula shown as a sequence of failures followed by one success

The exponent is x minus one because the success is the last trial and every earlier one failed.

The exponent is x minus one rather than x, and that off-by-one is the commonest slip in the whole section. It follows directly from the picture: on trial 5 there are four failures, not five, because the fifth trial is the success. Counting the F chips before reaching for the formula settles it every time.

18. Worked example: a first success on the fifth trial

Worked example

Example 4.17's question, with the numbers.

\[ X \sim G(0.57); \quad \text{find } P(x = 5) \]

Count the failures needed

Why: Trials one through four must fail.

Write their probability

Why: q is 0.43, four times.

\[ 0.43 ^{4} \]

Multiply by the success

Why: The fifth trial succeeds.

\[ \times 0.57 \]

Evaluate

Why: The product.

\[ \text{about } 0.0195 \]

Figure (svg): The solution to Worked example a first success on the fifth trial shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(5) = (0.43)^4(0.57) \approx 0.0195 \]

Verify: confirm the exponent by counting the trials

Why: Four failures plus one success is five trials, matching x. Using an exponent of 5 would describe six trials and answer a different question. The check is the same one section 4.3 used for the binomial's two exponents: the trials accounted for must total x, and here that means x minus one failures and exactly one success.

OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, p. 243

19. Substitute into the formula

Faded example

The probability of a defective steel rod is 0.01. Find the probability the first defect is the ninth rod.

Fill in the blanks

P(9) = (0.99)^8}(0.01) \approx 0.0092

Why: Eight good rods then a defective one, giving about 0.0092. The exponent is one less than x, because the ninth rod is the success rather than one of the failures.

20. Worked example: a cumulative geometric probability

Worked example

Example 4.18's second question, which has a shortcut.

\[ X \sim G(0.35); \quad \text{find } P(x \ge 3) \]

Say what at least three means

Why: The first two reports were not relevant.

\[ \text{trials } 1\text{ and } 2\text{ failed} \]

Write that probability

Why: q twice.

\[ 0.65 ^{2} \]

Evaluate

Why: The square.

\[ 0.4225 \]

Note that no summation was needed

Why: The event is just a run of failures.

Figure (svg): The solution to Worked example a cumulative geometric probability shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(x \ge 3) = q^2 = (0.65)^2 = 0.4225 \]

Verify: confirm against the direct summation

Why: Adding P(3), P(4), P(5) and so on forever would give the same answer, since the tail is a geometric series summing to q squared. The shortcut works because 'at least three trials needed' says exactly that the first two failed and nothing about what happens afterwards. In general P(x is greater than k) equals q to the power k, which converts every upper-tail geometric question into one exponentiation.

OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 243-244

21. Trap: the wrong exponent

Trap

The trap

\[ P(5) = (0.43)^5(0.57) \]

Raise q to the power x

Why: Five is the value of x, so it looks like the exponent.

\[ \text{that describes six trials} \quad \text{(five failures and a success)} \]

On trial five there are four earlier trials, so q appears four times, not five.

The fix

\[ P(5) = (0.43)^4(0.57) = q^{x-1}p \]

Count the trials BEFORE the success, which is x minus one

Why: The success is the xth trial, so it is not among the failures.

Writing the sequence out for a small x settles it permanently: x equals 1 is just a success with probability p, and the formula gives q to the zero times p, which is p. That single check confirms the exponent, and it is worth doing once rather than re-deriving under pressure. It also shows why the formula cannot use x: it would make P(1) equal to qp, which is wrong.

22. One of these is false

Two truths and a lie

All three concern the formula.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. There is no combination factor
  • C. P(1) equals p
  • B. The probabilities increase as x increases

Survives elimination: B

Why: The survivor is false. Each probability is q times the one before it, and q is less than one, so the stems decay steadily. The largest probability is always at x equals 1 — the first success is more likely to come immediately than at any later single trial, however small p is.

23. The upper tail shortcut

Faded example

For a geometric distribution with success probability p.

Fill in the blanks

P(x > k) = q^k}, \textfail ___ \text___ ___

Why: More than k trials are needed exactly when the first k all fail, so the probability is q to the power k. This converts every upper-tail geometric question into a single exponentiation rather than an infinite sum.

24. Which value is most likely?

Prediction

Commit before reasoning.

Predict first

For a geometric distribution with p = 0.1, which single value of x has the largest probability?

  • x = 10
  • x = 1
  • x = 5
  • They are all equally likely

Correct: x = 1.

Why: Each probability is q times the previous one, so the sequence decreases from the start and the maximum is always at x equals 1, whatever p is. That sits oddly beside the mean of 10, and both are correct: the single most likely trial is the first, while the average wait is ten trials because the long tail pulls the mean out. It is section 2.6's mode-against-mean distinction on a strongly right-skewed distribution.

25. The mean and standard deviation

Section

Section 3

26. One over p, and why that is obvious

Concept

The mean of a geometric distribution is one over p, and the standard deviation is the square root of q divided by p squared. Both follow from p alone, since p is the distribution's only parameter.

the geometric mean — Mu equals one over p: the expected number of trials until the first success. A success probability of one in three gives a mean of three trials, and one in a hundred gives a hundred.

\[ \mu = \frac{1}{p}, \qquad \sigma = \sqrt{\frac{q}{p^2}} = \frac{\sqrt{q}}{p} \]

The mean is the rare formula that could have been guessed. If a success happens one time in three, then across many attempts about a third succeed, so the average wait between successes is three attempts. What section 4.2's machinery adds is the proof and the standard deviation, which is not guessable and which turns out to be almost as large as the mean when p is small.

Figure (svg): The geometric mean and standard deviation formulas with a worked instance

The formula matches the intuition exactly: the rarer the success, the longer the wait, in proportion.

OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 243-247 — the geometric mean and standard deviation

27. The reciprocal, and its spread

Picture it

Both summaries from the single parameter.

Figure (svg): The geometric mean and standard deviation formulas with a worked instance

The formula matches the intuition exactly: the rarer the success, the longer the wait, in proportion.

The relationship between the two is worth noticing: for small p, sigma is approximately one over p as well, so the standard deviation is about the same size as the mean. That is the signature of a strongly right-skewed distribution and it means the average is a poor guide to any individual wait — a point the transfer problem takes up.

28. Worked example: how many reports to expect

Worked example

Example 4.18's first question.

\[ 35\% \text{ of accidents are caused by failure to follow instructions; reports are read until one is found} \]

Identify p

Why: The proportion of relevant reports.

\[ p = 0.35 \]

Apply the mean formula

Why: One over p.

\[ \frac{1}{0.35} \]

Evaluate

Why: The reciprocal.

\[ \text{about } 2.86 \]

Interpret

Why: The long-run average number of reports read.

Figure (svg): The solution to Worked example how many reports to expect shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \mu = \frac{1}{0.35} \approx 2.86 \]

Verify: confirm the answer against the intuition

Why: A little over a third of reports are relevant, so finding one should take a little under three reads — and 2.86 is exactly that. The check is worth running because the reciprocal is easy to invert by mistake: an answer of 0.35 reports would be nonsense, since at least one report must always be read. Any geometric mean below one indicates the formula was used upside down.

OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 243-244

29. Both summaries

Faded example

A success probability of 0.02.

Fill in the blanks

\mu = \frac502450 = ___, \qquad \sigma^2 = \frac______ = ___

Why: A mean of 50 trials and a variance of 2450, whose square root is 49.5 — almost exactly the mean, which is the usual pattern for a small p. These are the book's own numbers from its worked variance calculation.

30. Worked example: a rare event's waiting time

Worked example

Example 4.21. The lifetime risk of developing cancer is about one in 67.

\[ p = 0.015: \text{ how many people must be asked before one says they have cancer?} \]

Compute the mean

Why: One over p.

\[ \text{about } 66.67\text{ people} \]

Compute the standard deviation

Why: Root q over p.

\[ \text{about } 66.16 \]

Compare them

Why: Nearly equal.

Say what that implies

Why: The distribution is enormously spread.

Figure (svg): The solution to Worked example a rare event's waiting time shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \mu = \frac{1}{0.015} \approx 66.7, \qquad \sigma \approx 66.2 \]

Verify: confirm what a sigma this large means in practice

Why: A standard deviation almost equal to the mean says the waiting time is extremely variable: it could easily take ten people or two hundred. The mean of 66.7 is the long-run average across many repetitions of the whole search and is a poor prediction of any single one. This is a general feature of geometric distributions with small p, and it is the reason the mean alone should never be quoted as an expected wait without the spread beside it.

OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, p. 247

31. Error analysis: four attempts at a geometric mean

Error analysis

A success probability of 0.35, and the question is how many trials are expected until the first success.

Annotate

On: \( \begin{aligned} &(1)\; \mu = 0.35 \\ &(2)\; \mu = np = ? \\ &(3)\; \mu = 1 - 0.35 = 0.65 \\ &(4)\; \mu = \tfrac{1}{0.35} \approx 2.86 \end{aligned} \)

  • (1) reports p itself. A mean below one is impossible here, since at least one trial always occurs, so the answer refutes itself.
  • (2) reaches for the binomial mean, which needs an n — and there is no n in a geometric experiment, since the number of trials is what is being predicted.
  • (3) computes q, which is the probability of a single failure and not a number of trials at all.
  • (4) is correct: the reciprocal of the success probability, giving about 2.86 trials.

Errors (1) and (3) both produce a number below one, which is the giveaway: a geometric mean is always at least one, and it is large exactly when p is small. Error (2) is the more interesting one, because it shows the two families cannot share a formula — the absence of an n is precisely what defines a geometric problem.

32. How long a wait?

Estimation

A machine produces a defect with probability 0.004.

Predict first

About how many items will be produced before the first defect?

  • About 4
  • About 25
  • About 250
  • About 2,500

Correct: About 250.

Why: The mean is one over 0.004, which is 250. The reciprocal relationship makes these estimates quick: a one-in-a-thousand event takes about a thousand trials on average, and a one-in-four event takes about four. The standard deviation here is also about 250, so the actual wait could easily be fifty items or six hundred.

33. One of these is false

Two truths and a lie

All three concern the geometric summaries.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The mean is always at least one
  • C. For small p, sigma is nearly as large as mu
  • B. The mean is the most likely number of trials

Survives elimination: B

Why: The survivor is false. The most likely value is always x equals 1, whatever p is, because the probabilities decay from the start. For p equal to 0.1 the mean is 10 and the mode is 1, which is section 2.6's mean-mode gap on a strongly right-skewed distribution.

34. Halve p, and what happens?

Prediction

Commit before reasoning.

Predict first

A success probability is halved. What happens to the expected number of trials?

  • It halves
  • It doubles
  • It stays the same
  • It falls by half a trial

Correct: It doubles.

Why: The mean is one over p, so halving p doubles the reciprocal: a one-in-ten event takes about ten trials and a one-in-twenty event takes about twenty. The inverse relationship is the whole content of the formula, and it is why rare events have long waits in exact proportion to their rarity.

35. Memorylessness

Section

Section 4

36. Previous failures carry no information

Concept

The geometric distribution is memoryless, and it is the only distribution with a discrete random variable that is. All events prior to the present are irrelevant: the probability of a success on the next trial begins anew each time, regardless of how many failures have gone before.

memoryless — A distribution is memoryless when the probability of waiting a further n trials, given that k have already passed without success, equals the unconditional probability of waiting n trials. The geometric is the only discrete distribution with this property.

\[ P(X > k + n \mid X > k) = P(X > n) \]

The book's illustration is a baseball player who hits with probability 0.20 and has failed his last ten times at bat. His probability of a hit next time is still 0.20, because the answer ignores his ten previous failures entirely. That follows directly from the second characteristic: the trials are independent, so nothing about the past can inform the future.

Figure (svg): A diagram illustrating that a geometric distribution has no memory of previous failures

Past failures carry no information about the future, which is a real property of independent trials and a poor description of anything that learns.

OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 242-243 — memorylessness and the baseball example

37. Ten failures, and no effect

Picture it

The property stated, and its formal form.

Figure (svg): A diagram illustrating that a geometric distribution has no memory of previous failures

Past failures carry no information about the future, which is a real property of independent trials and a poor description of anything that learns.

The book draws out a consequence worth noticing: testing manufactured parts for defects, the geometric distribution begins with a clean slate each time, with no consideration of previous test results. That is a real property of genuinely independent trials and a poor description of a process that wears out or improves — which is the next idea.

38. Worked example: the batter's next attempt

Worked example

The book's own example, worked through the definition.

\[ p = 0.20; \text{ the last ten attempts all failed. What is the probability of a hit next time?} \]

Recall the second characteristic

Why: The trials are independent.

Ask what the ten failures tell you

Why: Nothing about the eleventh trial.

State the probability

Why: The same as on any trial.

\[ 0.20 \]

Note the general form

Why: Waiting resets each time.

\[ P(X > k + n | X > k) = P(X > n) \]

Figure (svg): The solution to Worked example the batter's next attempt shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(\text{hit next}) = p = 0.20 \]

Verify: confirm this is not the gambler's fallacy in reverse

Why: Section 3.1 warned against expecting a tail after a run of heads, and this is the same principle applied honestly. The failures neither raise the next probability, as a 'due for a hit' argument would claim, nor lower it. Memorylessness says the probability is exactly unchanged, which is the only position consistent with independence — and it is worth stating explicitly, because both errors are common in the other direction.

OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, p. 243

39. Is memorylessness reasonable here?

Discrimination

Ask whether the success probability could change over the sequence.

Sort into buckets

Sort each situation.

Memorylessness is reasonable
rolling a die until a six appears; flipping a coin until the first head; drawing lottery numbers until a win
p probably drifts: a poor fit
a novice throwing darts until the first bullseye; testing machine parts as the machine wears out
ok
The mechanism has no memory and no capacity to change, so the success probability is genuinely constant across trials.
no
Skill improves or equipment degrades over the sequence, so p is not the same on the hundredth trial as on the first.

40. Worked example: the property stated formally

Worked example

Checking the formal statement against the formula.

\[ \text{Show } P(X > k+n \mid X > k) = P(X > n) \]

Write the upper tail

Why: More than m trials means the first m failed.

\[ P(X > m) = q ^{m} \]

Write the conditional

Why: Section 3.1's definition.

\[ P(X > k + n AND X > k) / P(X > k) \]

Simplify the numerator

Why: Exceeding k+n already implies exceeding k.

\[ q ^{k + n} / q ^{k} \]

Cancel

Why: The powers subtract.

\[ q ^{n} = P(X > n) \]

Figure (svg): The solution to Worked example the property stated formally shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \frac{q^{k+n}}{q^k} = q^n = P(X > n) \]

Verify: confirm which property of the exponents did the work

Why: The cancellation works because the upper tail is a pure power of q, so dividing subtracts exponents and the k vanishes. No other discrete distribution has tails of that form, which is why the geometric is the only discrete memoryless one — the book states this and the algebra shows why. Section 5.3's exponential distribution has an exponential tail for the same reason and is the continuous counterpart.

OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 242-243

41. Trap: applying memorylessness to a process that learns

Trap

The trap

\[ \text{a dart thrower hits the centre with probability } 0.17 \]

Model the throws as geometric across a long practice session

Why: The throws look like independent trials with a fixed success rate.

\[ \text{but the thrower improves with practice} \quad (p \text{ is not constant}) \]

The third characteristic fails: a later throw has a higher success probability than an early one.

The fix

\[ \text{geometric requires a constant } p \text{, so no learning} \]

Ask whether the success probability could drift over the sequence

Why: Skill, wear and fatigue all break the constant-p assumption.

The book raises this directly: in experiments requiring skill, such as hitting a baseball or throwing a dart, one might consider that learning during the experiment would alter the probability of a success, and the geometric distribution cannot capture learning. The historical success probability is assumed constant. That is a modelling assumption to be stated rather than a fact, and it is exactly where a geometric model of a human activity is most vulnerable.

42. After twenty failures

Prediction

Commit before reasoning.

Predict first

A geometric process has p = 0.05 and twenty trials have failed. What is the expected number of FURTHER trials until a success?

  • Fewer than 20, since a success is overdue
  • Exactly 20, matching what has passed
  • 20, the same as at the start
  • It cannot be determined

Correct: 20, the same as at the start.

Why: The mean is one over 0.05, which is 20, and memorylessness says the process resets: the expected further wait is the same 20 trials it was before any of the failures occurred. Nothing is overdue, which is section 3.1's gambler's fallacy again — and the coincidence that twenty trials have passed is irrelevant rather than significant.

43. One of these is false

Two truths and a lie

All three concern memorylessness.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The geometric is the only discrete distribution that is memoryless
  • C. Memorylessness follows from the trials being independent
  • B. Memorylessness means a success becomes more likely after a long run of failures

Survives elimination: B

Why: The survivor is false and is the gambler's fallacy. Memorylessness means the probability is exactly UNCHANGED — neither raised nor lowered. The run of failures is information about the past only, and the process begins anew each trial.

44. Where memorylessness fails usefully

Socratic

Many real waiting times are not memoryless.

Discussion prompt

Give a real waiting time that is clearly NOT memoryless, and say which direction the conditioning goes.

Hint: Think about something that wears out, or something that is nearly due.

Answer:

A mechanical component that wears out is the standard case: having run for ten years, it is MORE likely to fail in the next year than a new one was, because the failure rate rises with age. Conditioning on survival makes the remaining wait shorter.

The opposite also occurs. Manufactured electronics often have a high early failure rate, so a component that has survived its first month is LESS likely to fail soon than a fresh one — conditioning on survival makes the remaining wait longer.

The geometric and its continuous cousin the exponential sit exactly between these, with a failure rate that never changes. That makes them the right model for genuinely random events like radioactive decay and the wrong model for anything that ages, which is why reliability engineering uses other families for wear-out.

45. Reading and modelling

Section

Section 5

46. Setting up a geometric problem

Concept

Setting up means answering the same questions as for a binomial with one change: there is no n. Identify one trial, decide what counts as the success, read off p, and note that X is the number of the trial on which that success occurs.

setting up a geometric problem — One trial, a success and its probability p, and a variable counting the trial number of the first success. The absence of a fixed n is the signal, usually carried by the word until.

\[ X \sim G(p): \quad P(x) = q^{x-1}p, \; \mu = \frac{1}{p}, \; \sigma = \frac{\sqrt{q}}{p} \]

The book's Example 4.18 is worth noticing for one detail in its setup: the accident reports are selected randomly AND REPLACED IN THE PILE after reading. That clause exists to keep p constant, satisfying the third characteristic. Without replacement the probability would shift as reports were removed, and the geometric model would not apply — the same replacement issue section 3.2 and section 4.3 both turned on.

Figure (svg): Two columns contrasting the binomial, which fixes the number of trials, with the geometric, which fixes the number of successes

The two are mirror images of one question. Which quantity is fixed in advance and which is random is the whole distinction.

OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 243-247 — the setup of Examples 4.18 and 4.21

47. Two families, one comparison

Picture it

What each fixes and what each counts.

Figure (svg): Two columns contrasting the binomial, which fixes the number of trials, with the geometric, which fixes the number of successes

The two are mirror images of one question. Which quantity is fixed in advance and which is random is the whole distinction.

The rows to carry are the second and third. A binomial answer is a count of successes between zero and n; a geometric answer is a trial number of at least one, with no ceiling. If an answer's units do not match the question's, the wrong family was used, and that check catches the confusion before the arithmetic does.

48. Worked example: a full setup

Worked example

Example 4.18 from the top.

\[ 35\% \text{ of accidents come from failure to follow instructions; reports are read (and replaced) until one is found} \]

What is one trial?

Why: Reading one accident report.

What is the success?

Why: Finding one caused by failure to follow instructions.

\[ p = 0.35 \]

Why does replacement matter?

Why: It keeps p the same on every trial.

Define X and its values

Why: The report on which the first is found.

\[ x = 1, 2, 3,... \]

Figure (svg): The solution to Worked example a full setup shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ X \sim G(0.35), \quad \mu \approx 2.86, \quad P(x \ge 3) = 0.4225 \]

Verify: confirm the replacement clause is doing real work

Why: Without replacement, reading a non-relevant report would remove it from the pile and slightly raise the proportion of relevant ones remaining, so p would creep upward. Over a few reports the drift is tiny, but the clause makes the model exact rather than approximate. It is the same reason section 4.3's binomial needed replacement, and it is why textbook problems specify it so often — the specification is what licenses the formula.

OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 243-244

49. Which family, and why?

Sorting

Look for a fixed n, or for the word until.

Sort into buckets

Sort each problem.

Binomial: an n is given
20 workers sampled; how many have the qualification?; 50 students; how many finished homework?
Geometric: only p is given
reports read until the first relevant one; people asked until one says yes; darts thrown until the first bullseye
bin
A fixed number of trials is stated, so the count of successes is the random variable.
geo
No number of trials is given and the word until appears, so the trial number of the first success is the random variable.

The absence of an n is the reliable structural signal, and 'until' is the reliable verbal one. When both point the same way the identification is safe; when a problem gives an n and also says until, read it again, because it is probably asking two questions.

50. Worked example: a rare-event setup

Worked example

Example 4.21. The lifetime risk of developing cancer is about one in 67.

\[ X = \text{ the number of people asked until one says they have cancer} \]

Read off p

Why: One in 67, about 1.5 percent.

\[ p = 0.015 \]

Compute a single probability

Why: The tenth person is the first to say yes.

\[ (0.985) ^{9}(0.015) \]

Evaluate

Why: About one and a third percent.

\[ 0.0131 \]

Compute the summaries

Why: Reciprocal and the spread formula.

\[ \mu 66.7, \sigma 66.2 \]

Figure (svg): The solution to Worked example a rare-event setup shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(10) = (0.985)^9(0.015) \approx 0.0131 \]

Verify: confirm why the individual probabilities are all small

Why: With a mean of 66.7 the probability is spread across many values, so no single one carries much — the largest, at x equals 1, is only 0.015. That is characteristic of a geometric distribution with small p: the probability is thinly spread over a long tail. It also means questions about ranges are usually more informative than questions about single values, which is why the upper-tail shortcut q to the k is so useful here.

OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, p. 247

51. Trap: looking for an n that is not there

Trap

The trap

\[ \text{'how many rolls until the first six?'} \;\to\; \text{find } n \]

Search the problem for a number of trials

Why: Every binomial problem had one, so its absence looks like missing information.

\[ \text{there is no } n \text{, and its absence is the point} \]

The number of trials is what the question is asking about, so it cannot also be given.

The fix

\[ \text{a geometric problem gives only } p \]

Treat a missing n as the signal for a geometric problem, not as an omission

Why: One parameter is all the family needs.

This is worth naming because a missing quantity usually means a problem is under-specified, and here it means the opposite: the geometric distribution has exactly one parameter, and p determines the probabilities, the mean and the standard deviation together. The word 'until' in the problem is the positive signal, and the absence of an n is the confirming one.

52. Formula to family

Matching

Each formula belongs to one of the two families.

Match the pairs

  • l1. P(x) = C(n,x) p^x q^(n-x)
  • l2. P(x) = q^(x-1) p
  • l3. mu = np
  • l4. mu = 1/p
  • r1. binomial probability
  • r2. geometric probability
  • r3. binomial mean
  • r4. geometric mean

Why: The presence of n is the tell in every case. A geometric formula cannot contain n, because the number of trials is the random variable rather than a parameter — so any formula with an n in it belongs to the other family.

53. One of these is false

Two truths and a lie

All three concern setting up a geometric problem.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The distribution has exactly one parameter
  • C. A replacement clause in the problem exists to keep p constant
  • B. A geometric problem must state the number of trials

Survives elimination: B

Why: The survivor is false and inverts the situation. The number of trials is the random variable, so stating it would answer the question. The absence of an n is the structural signal that a problem is geometric rather than binomial.

54. Explain the swap

Explain it

A classmate cannot see the difference between the two families.

Discussion prompt

In three sentences or fewer, give them one experiment and two questions that separate the families cleanly.

Hint: Use one die and change only the question.

Answer:

Give them a die and ask two questions: how many threes in ten rolls, and how many rolls until the first three.

The first fixes ten trials and lets the count of successes vary, so it is binomial with eleven possible answers; the second fixes one success and lets the number of trials vary, so it is geometric with infinitely many.

Same die, same p — what changed is which quantity was decided in advance, and that is the only thing separating the two families.

55. Binomial against geometric

Comparison

Fill the blanks. One experiment can supply both, depending on the question.

Comparison matrix

FeatureBinomialGeometric
Fixed in advancethe number of trials, nthe number of successes: exactly one
The random variablethe count of successesthe trial number of the first success
Values0 through n1, 2, 3, ... with no upper bound
Meannp1/p
Combination factoryes: many orderings give x successesno: only one ordering works

The two answer opposite halves of the same question, and telling them apart is a matter of reading which quantity the problem decided in advance. A formula containing n is binomial; a formula containing only p is geometric.

56. Solving a geometric problem, in order

Pattern

Six steps, and the first is a family identification rather than a calculation.

  1. Confirm the experiment repeats until a first success and then stops, so no n is given.
  2. Check independence and a constant p, watching for replacement clauses and for skill or wear that would make p drift.
  3. Identify one trial, name the success, and read off p and q.
  4. Define X as the trial number of the first success, taking the values 1, 2, 3 and upward.
  5. For a single value use q to the power x minus one times p; for an upper tail use q to the power k directly.
  6. Check the answer against mu = 1/p, remembering that sigma is nearly as large as mu when p is small.

The exponent is x minus one, not x. Writing out the sequence for x equal to 1 confirms it: a success on the first trial has probability p, and the formula must give that.

OpenStax Introductory Business Statistics 2e, §4.3 Geometric Distribution §4.3 Geometric Distribution

57. Check yourself 1 of 3

Check

Identify the family.

Check your understanding

Which of these is a geometric experiment?

  • A. Rolling a die until the first six appears (correct)
  • B. Rolling a die ten times and counting sixes
  • C. Drawing 5 cards and counting hearts
  • D. Sampling 50 items and counting defects

Answer: A

Why: The experiment repeats until the first success and then stops, so the number of trials is the random variable and no n is given.

Why B tempts people
Ten trials are fixed in advance and successes are counted, which is binomial.
Why C tempts people
Five draws are fixed, and without replacement this is in fact hypergeometric rather than either.
Why D tempts people
Fifty trials are fixed and successes counted: binomial.

58. Check yourself 2 of 3

Check

Apply the formula.

Check your understanding

For a geometric distribution with p = 0.2, what is P(x = 4)?

  • A. (0.8)^3 (0.2) (correct)
  • B. (0.8)^4 (0.2)
  • C. (0.2)^3 (0.8)
  • D. C(4,1)(0.2)(0.8)^3

Answer: A

Why: Three failures then a success on the fourth trial, so q is raised to x minus one and multiplied by p.

Why B tempts people
That uses an exponent of x rather than x minus one, describing five trials rather than four.
Why C tempts people
That swaps p and q, describing three successes followed by a failure.
Why D tempts people
That includes a combination factor, but only one ordering gives a FIRST success on trial four.

59. Check yourself 3 of 3

Check

The mean.

Check your understanding

A process succeeds with probability 0.04. About how many trials until the first success?

  • A. About 25 (correct)
  • B. About 4
  • C. About 0.04
  • D. About 96

Answer: A

Why: The mean is one over p, which is one over 0.04, giving 25 trials.

Why B tempts people
That reads the percentage as the answer rather than taking its reciprocal.
Why C tempts people
That reports p itself. A mean below one is impossible, since at least one trial always occurs.
Why D tempts people
That is roughly q as a percentage, which is a probability rather than a number of trials.

60. Where this shows up outside the textbook

Real world

A recruiter says that about 4 percent of applications lead to an interview, so a candidate should expect an interview after about 25 applications. A candidate has sent 60 applications and had none, and concludes something must be wrong with their CV.

Discussion prompt

Model the situation, say whether 60 without success is surprising, and identify the modelling assumption most likely to be wrong.

Hint: Compute the mean and the upper-tail probability, then ask whether p is really constant across candidates.

Answer:

The model is geometric with p = 0.04. The mean is one over 0.04, which is 25 applications, so the recruiter's figure is right as a long-run average. But sigma is the square root of 0.96 over 0.0016, about 24.5 — almost as large as the mean, so the spread is enormous.

Sixty without success is not surprising. The probability of more than 60 failures is 0.96 to the power 60, which is about 0.086 — roughly one candidate in twelve. On a distribution this skewed, waits far above the mean are routine rather than remarkable.

The assumption most likely wrong is the constant p. The geometric model treats every candidate as having the same 4 percent rate, when in reality the rate varies enormously between candidates and between applications. The 4 percent is an average ACROSS applicants, and no individual's rate need be near it.

\[ \mu = 25, \quad \sigma \approx 24.5, \quad P(X > 60) = (0.96)^{60} \approx 0.086 \]

So the candidate's conclusion may still be right, but the sixty failures alone are weak evidence for it — one in twelve people with a perfectly ordinary CV would see the same run. What would be better evidence is comparison with similar candidates, which is a question about differing p values rather than about one geometric distribution. The general lesson is that a mean quoted without its spread invites exactly this error, and for a geometric distribution the spread is always about as large as the mean.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

Why does the geometric formula have no combination factor?

  • Because there is only one trial
  • Because exactly one sequence produces a first success on trial x: failures throughout, then a success
  • Because p is constant
  • Because the trials are independent

Correct: Because exactly one sequence produces it.

\[ \text{binomial: many orderings} \qquad \text{geometric: exactly one} \]

Why: The definition of the event fixes the ordering completely — every trial before the xth failed and the xth succeeded — so there is nothing to count. The binomial needs its combination factor precisely because x successes among n trials can be arranged in many ways. Independence and a constant p are what let the single sequence's probability be a product, but they do not affect the counting.

62. Explain it to someone a year behind you

Explain it

They have written P(x = 5) as q to the fifth times p.

Discussion prompt

In three sentences or fewer, show them the exponent using a small case.

Hint: Ask them what the formula gives for x = 1.

Answer:

Ask them what their version gives for x equal to 1: it says q times p, but a success on the very first trial should just be p.

The exponent counts the trials BEFORE the success, and on trial five there are four of them, so it is q to the fourth times p.

Writing the sequence out — failure, failure, failure, failure, success — makes the four visible and settles it permanently.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Telling a geometric problem from a binomial one
  • Getting the exponent right in the formula
  • Computing and interpreting the mean 1/p
  • Saying what memorylessness does and does not mean

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the family, look for a fixed n or for the word until. For the exponent, check that the formula gives p when x is 1. For the mean, remember it is at least one and is large when p is small. For memorylessness, it means the probability is unchanged — neither raised nor lowered — by past failures. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Fifteen minutes.

Draw it

At the top, write the four characteristics of a geometric experiment and mark which two it shares with the binomial. Beside them, draw a row of five boxes labelled F, F, F, F, S and write underneath the probability of that exact sequence, then generalise it to the formula. Check your formula by setting x equal to 1 and confirming it gives p. In the middle, work Example 4.18 completely: 35 percent of accident reports are relevant and reports are read with replacement until one is found. Write the setup in four lines, compute the mean and the standard deviation, and compute P(x at least 3) using the q-to-the-k shortcut rather than a sum. Below that, draw the first eight stems of the distribution to scale and mark which value is most likely, then mark the mean and write one sentence on why the two are different. At the bottom, make a two-column table comparing binomial and geometric on five rows: what is fixed, what X counts, the values X takes, the mean, and whether a combination factor appears. Finish by writing the memoryless property in symbols and one sentence naming a real waiting time that is NOT memoryless.

Check your stem drawing: the tallest stem must be at x = 1 and every later stem must be exactly q times the one before it. If your tallest stem is near the mean, the distribution was drawn as though it were binomial — a geometric distribution always decays from its first value.

65. What you can do now

Recap

Five things, and the whole family follows from one parameter.

If you seeThen
The word until, and no nGeometric: X is the trial number
A stated number of trialsBinomial: X is the count of successes
A request for P(x = k)q to the k minus one, times p
A request for P(x greater than k)q to the k, in one step
A request for the expected wait1/p, always at least one
A run of failures so farIrrelevant: the process is memoryless
Skill, learning or wear in the trialsp is not constant; the model does not apply

Section 4.5 relaxes the other binomial characteristic. When sampling is done without replacement from two groups, the trials are not independent and p changes at every draw — and the count of items from the group of interest has a hypergeometric distribution.

OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 242-247 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 242-247
  2. OpenStax Introductory Business Statistics 2e, §4.3 Geometric Distribution — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108