The binomial's first characteristic relaxed. A geometric experiment repeats a trial until the first success occurs and then stops, so the number of trials is the random variable rather than the number of successes, and it takes the values one, two, three and upward with no bound. The probability of a first success on trial x is q to the power x minus one times p, with no combination factor because exactly one sequence produces it. The mean is one over p, matching the intuition that a one-in-three chance takes about three tries. The distribution is also the only discrete one that is memoryless: previous failures carry no information about what comes next, which is a genuine property of independent trials and a poor model for anything that learns.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 4 — Discrete Random Variables
Geometric Distribution
Objectives
Five outcomes. The second is what makes the whole family different from section 4.3's.
OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 242-247 — the section these objectives are drawn from
Warm-up
Section 4.3's binomial required a fixed number of trials, and that was the first of its three characteristics.
Discussion prompt
You throw darts at a bullseye until you hit it. What is the number of trials here, and can the binomial formula describe this?
Hint: Ask whether you could write down n before starting.
Answer:
The number of trials is not fixed and cannot be known in advance: it might take one throw or twenty, and in theory it could go on forever. So the first binomial characteristic fails immediately.
What is fixed instead is the number of successes — exactly one, since you stop as soon as you hit. So the two quantities have swapped roles: the successes are now fixed and the trials are random.
That swap is the whole of this section. The book calls the result a geometric distribution, and the random variable X is the number of the trial on which the first success occurs — which is a waiting time rather than a count of successes.
Concept
In a geometric experiment a trial is repeated until a success occurs, and then you stop. The trials are independent and the probability of a success is the same on each. The random variable X is the number of the trial on which the first success occurs, so it takes the values one, two, three and upward.
geometric experiment — A sequence of independent trials with constant success probability p, repeated until the first success. X is the number of the trial on which that success occurs, and X takes the values 1, 2, 3 and so on without upper bound. The book describes the sequence as failure, failure, failure, success, STOP.
\[ X \sim G(p), \qquad x = 1, 2, 3, \ldots \]
The book notes that the variable can be defined either way — as the number of trials until a success, or as the number of failures before one — and that each form changes the formula and the mean slightly. This lesson uses the first, which is the book's own convention in its examples and the more common one: X counts the trial on which the success lands, so its smallest value is one rather than zero.
Figure (svg): Two columns contrasting the binomial, which fixes the number of trials, with the geometric, which fixes the number of successes
OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, p. 242
Section
Section 1
Concept
A trial is repeated until a success occurs. The repeated trials are independent of each other, and the probability p of a success and q of a failure are the same for each trial. The random variable X is the number of the trial on which the first success occurs.
the four characteristics — Repeat until a success and stop; the trials are independent; p and q are constant; and X is the number of the trial on which the first success occurs. In theory the number of trials could go on forever.
\[ \text{F, F, F, F, S, STOP} \;\Longrightarrow\; x = 5 \]
Only the first characteristic differs from the binomial's. Independence and a constant p are shared, which is why the same sampling caution applies: drawing without replacement from a small population breaks the third characteristic here exactly as it broke the binomial's. The book's own illustration is a die rolled until a three appears, where p stays a sixth no matter how many rolls have gone before.
Figure (svg): The four characteristics of a geometric experiment listed as a numbered procedure
OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, p. 242 — the four characteristics and the die example
Picture it
Four conditions, of which two are the binomial's.
Figure (svg): The four characteristics of a geometric experiment listed as a numbered procedure
The fourth characteristic is really a definition rather than a condition, and it is the one that changes what any answer means. In a binomial problem an answer of 5 would be five successes; here it is five trials, of which four failed. Reading the units of X correctly is what stops a geometric answer being interpreted as a binomial one.
Worked example
Example 4.17. Play until you lose.
\[ \text{You play until you lose; the probability of losing is } p = 0.57. \]
Check the stopping rule
Why: You play until you lose, then stop.
Identify the success
Why: Losing is what ends the sequence.
\[ p = 0.57 \]
Define the variable
Why: The number of games played, including the losing one.
\[ X =\text{ games until the loss} \]
List the values
Why: One game at least, and no upper bound.
\[ 1, 2, 3,... \]
Figure (svg): The solution to Worked example identifying a geometric experiment shown as a ladder of expressions, one row per legal move
\[ X \sim G(0.57), \quad \text{find } P(x = 5) \]
Verify: confirm what a success is here
Why: Losing is the success, because the success is whatever the sequence is waiting for — the book defines it that way, and the label carries no approval, exactly as in section 4.3. The variable includes the losing game itself, so x equals 5 means four wins followed by a loss. Reading the definition carefully matters because 'the number of games you play until you lose' could otherwise be taken as the number of wins, which would be one less.
OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, p. 243
Sorting
Ask which quantity is fixed in advance.
Sort into buckets
Sort each question.
The phrase 'until' is the reliable signal for a geometric problem, and a stated number of trials is the signal for a binomial one. A question containing both — how many flips until the third head — is neither, and belongs to a family this book does not cover.
Worked example
The same situation, two different questions.
\[ \text{A die is rolled. } p = \frac{1}{6} \text{ for a three.} \]
Question A: how many threes in ten rolls?
Why: Ten trials fixed, successes counted.
Question B: how many rolls until the first three?
Why: Successes fixed at one, trials counted.
Note what is fixed in each
Why: n in the first, the success count in the second.
Note what is random in each
Why: The count in the first, the number of trials in the second.
Figure (svg): The solution to Worked example geometric or binomial shown as a ladder of expressions, one row per legal move
\[ \text{fixed } n \to \text{binomial}; \qquad \text{fixed successes} \to \text{geometric} \]
Verify: confirm the two are genuinely different questions
Why: Question A has eleven possible answers, zero through ten, while question B has infinitely many, one upward. They cannot be the same distribution because their value sets differ. The reliable test is to ask what could be written down before the experiment starts: if n is known in advance the problem is binomial, and if instead you know you will stop at the first success it is geometric.
OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 242-243
Trap
\[ P(x = 5) \text{ with } p = 0.57 \]
Interpret x = 5 as five losses
Why: In section 4.3 the variable counted successes, so the habit carries over.
\[ \text{but here } x \text{ counts TRIALS, and there is exactly one success} \]
Five means five games were played, of which the first four were wins and the fifth a loss.
\[ x = 5 \;\Longrightarrow\; \text{four failures, then the first success on trial } 5 \]
Read the fourth characteristic every time: X is the TRIAL NUMBER
Why: The number of successes is always exactly one.
The two families use the same letters for different things, which is the main source of confusion between them. A useful habit is to write out what x equals 5 means as a sentence — 'the first success came on the fifth trial' — before doing any arithmetic. It also explains why x cannot be zero here: there must be at least one trial for a success to occur on.
Fill the middle
What X can be in a geometric experiment.
Fill in the blanks
X \text1 ___, 2, 3, \ldots \text___
Why: One, because at least one trial must occur for a success to happen on it. This is a visible difference from the binomial, whose values start at zero — a binomial count of zero successes is perfectly possible, and a geometric trial number of zero is not.
Two truths and a lie
All three concern the characteristics.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and it is exactly the characteristic the geometric relaxes. The number of trials is the random variable, determined by when the first success happens, and it cannot be known before the experiment runs.
Prediction
Commit before reasoning.
Predict first
Why does X have infinitely many possible values?
Correct: Because a run of failures of any length is possible.
Why: Nothing forbids a hundred consecutive failures — its probability is q to the hundredth, which is minute but not zero — so no value of x can be ruled out. The independence is what makes those long runs possible rather than being the reason for the unboundedness itself, and the probabilities decay fast enough that the infinitely many values still sum to one.
Section
Section 2
Concept
For the first success to occur on trial x, the first x minus one trials must all fail and the xth must succeed. Because the trials are independent, that sequence has probability q multiplied by itself x minus one times, then p.
the geometric formula — P(x) equals q to the power x minus one, times p. There is no combination factor because exactly one sequence of outcomes produces a first success on trial x: failures throughout, then a success.
\[ P(x) = q^{x-1} p \]
The absence of a combination factor is the clearest structural difference from section 4.3. There, x successes in n trials could be arranged in many orders and the formula counted them; here the ordering is completely determined by the definition — every trial before the last failed, and the last succeeded — so there is exactly one arrangement and nothing to count.
Figure (svg): The geometric probability formula shown as a sequence of failures followed by one success
OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 242-243 — the geometric probability, with the die example
Picture it
Four failures and a success, multiplied by section 3.3's rule.
Figure (svg): The geometric probability formula shown as a sequence of failures followed by one success
The exponent is x minus one rather than x, and that off-by-one is the commonest slip in the whole section. It follows directly from the picture: on trial 5 there are four failures, not five, because the fifth trial is the success. Counting the F chips before reaching for the formula settles it every time.
Worked example
Example 4.17's question, with the numbers.
\[ X \sim G(0.57); \quad \text{find } P(x = 5) \]
Count the failures needed
Why: Trials one through four must fail.
Write their probability
Why: q is 0.43, four times.
\[ 0.43 ^{4} \]
Multiply by the success
Why: The fifth trial succeeds.
\[ \times 0.57 \]
Evaluate
Why: The product.
\[ \text{about } 0.0195 \]
Figure (svg): The solution to Worked example a first success on the fifth trial shown as a ladder of expressions, one row per legal move
\[ P(5) = (0.43)^4(0.57) \approx 0.0195 \]
Verify: confirm the exponent by counting the trials
Why: Four failures plus one success is five trials, matching x. Using an exponent of 5 would describe six trials and answer a different question. The check is the same one section 4.3 used for the binomial's two exponents: the trials accounted for must total x, and here that means x minus one failures and exactly one success.
OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, p. 243
Faded example
The probability of a defective steel rod is 0.01. Find the probability the first defect is the ninth rod.
Fill in the blanks
P(9) = (0.99)^8}(0.01) \approx 0.0092
Why: Eight good rods then a defective one, giving about 0.0092. The exponent is one less than x, because the ninth rod is the success rather than one of the failures.
Worked example
Example 4.18's second question, which has a shortcut.
\[ X \sim G(0.35); \quad \text{find } P(x \ge 3) \]
Say what at least three means
Why: The first two reports were not relevant.
\[ \text{trials } 1\text{ and } 2\text{ failed} \]
Write that probability
Why: q twice.
\[ 0.65 ^{2} \]
Evaluate
Why: The square.
\[ 0.4225 \]
Note that no summation was needed
Why: The event is just a run of failures.
Figure (svg): The solution to Worked example a cumulative geometric probability shown as a ladder of expressions, one row per legal move
\[ P(x \ge 3) = q^2 = (0.65)^2 = 0.4225 \]
Verify: confirm against the direct summation
Why: Adding P(3), P(4), P(5) and so on forever would give the same answer, since the tail is a geometric series summing to q squared. The shortcut works because 'at least three trials needed' says exactly that the first two failed and nothing about what happens afterwards. In general P(x is greater than k) equals q to the power k, which converts every upper-tail geometric question into one exponentiation.
OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 243-244
Trap
\[ P(5) = (0.43)^5(0.57) \]
Raise q to the power x
Why: Five is the value of x, so it looks like the exponent.
\[ \text{that describes six trials} \quad \text{(five failures and a success)} \]
On trial five there are four earlier trials, so q appears four times, not five.
\[ P(5) = (0.43)^4(0.57) = q^{x-1}p \]
Count the trials BEFORE the success, which is x minus one
Why: The success is the xth trial, so it is not among the failures.
Writing the sequence out for a small x settles it permanently: x equals 1 is just a success with probability p, and the formula gives q to the zero times p, which is p. That single check confirms the exponent, and it is worth doing once rather than re-deriving under pressure. It also shows why the formula cannot use x: it would make P(1) equal to qp, which is wrong.
Two truths and a lie
All three concern the formula.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. Each probability is q times the one before it, and q is less than one, so the stems decay steadily. The largest probability is always at x equals 1 — the first success is more likely to come immediately than at any later single trial, however small p is.
Faded example
For a geometric distribution with success probability p.
Fill in the blanks
P(x > k) = q^k}, \textfail ___ \text___ ___
Why: More than k trials are needed exactly when the first k all fail, so the probability is q to the power k. This converts every upper-tail geometric question into a single exponentiation rather than an infinite sum.
Prediction
Commit before reasoning.
Predict first
For a geometric distribution with p = 0.1, which single value of x has the largest probability?
Correct: x = 1.
Why: Each probability is q times the previous one, so the sequence decreases from the start and the maximum is always at x equals 1, whatever p is. That sits oddly beside the mean of 10, and both are correct: the single most likely trial is the first, while the average wait is ten trials because the long tail pulls the mean out. It is section 2.6's mode-against-mean distinction on a strongly right-skewed distribution.
Section
Section 3
Concept
The mean of a geometric distribution is one over p, and the standard deviation is the square root of q divided by p squared. Both follow from p alone, since p is the distribution's only parameter.
the geometric mean — Mu equals one over p: the expected number of trials until the first success. A success probability of one in three gives a mean of three trials, and one in a hundred gives a hundred.
\[ \mu = \frac{1}{p}, \qquad \sigma = \sqrt{\frac{q}{p^2}} = \frac{\sqrt{q}}{p} \]
The mean is the rare formula that could have been guessed. If a success happens one time in three, then across many attempts about a third succeed, so the average wait between successes is three attempts. What section 4.2's machinery adds is the proof and the standard deviation, which is not guessable and which turns out to be almost as large as the mean when p is small.
Figure (svg): The geometric mean and standard deviation formulas with a worked instance
OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 243-247 — the geometric mean and standard deviation
Picture it
Both summaries from the single parameter.
Figure (svg): The geometric mean and standard deviation formulas with a worked instance
The relationship between the two is worth noticing: for small p, sigma is approximately one over p as well, so the standard deviation is about the same size as the mean. That is the signature of a strongly right-skewed distribution and it means the average is a poor guide to any individual wait — a point the transfer problem takes up.
Worked example
Example 4.18's first question.
\[ 35\% \text{ of accidents are caused by failure to follow instructions; reports are read until one is found} \]
Identify p
Why: The proportion of relevant reports.
\[ p = 0.35 \]
Apply the mean formula
Why: One over p.
\[ \frac{1}{0.35} \]
Evaluate
Why: The reciprocal.
\[ \text{about } 2.86 \]
Interpret
Why: The long-run average number of reports read.
Figure (svg): The solution to Worked example how many reports to expect shown as a ladder of expressions, one row per legal move
\[ \mu = \frac{1}{0.35} \approx 2.86 \]
Verify: confirm the answer against the intuition
Why: A little over a third of reports are relevant, so finding one should take a little under three reads — and 2.86 is exactly that. The check is worth running because the reciprocal is easy to invert by mistake: an answer of 0.35 reports would be nonsense, since at least one report must always be read. Any geometric mean below one indicates the formula was used upside down.
OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 243-244
Faded example
A success probability of 0.02.
Fill in the blanks
\mu = \frac502450 = ___, \qquad \sigma^2 = \frac______ = ___
Why: A mean of 50 trials and a variance of 2450, whose square root is 49.5 — almost exactly the mean, which is the usual pattern for a small p. These are the book's own numbers from its worked variance calculation.
Worked example
Example 4.21. The lifetime risk of developing cancer is about one in 67.
\[ p = 0.015: \text{ how many people must be asked before one says they have cancer?} \]
Compute the mean
Why: One over p.
\[ \text{about } 66.67\text{ people} \]
Compute the standard deviation
Why: Root q over p.
\[ \text{about } 66.16 \]
Compare them
Why: Nearly equal.
Say what that implies
Why: The distribution is enormously spread.
Figure (svg): The solution to Worked example a rare event's waiting time shown as a ladder of expressions, one row per legal move
\[ \mu = \frac{1}{0.015} \approx 66.7, \qquad \sigma \approx 66.2 \]
Verify: confirm what a sigma this large means in practice
Why: A standard deviation almost equal to the mean says the waiting time is extremely variable: it could easily take ten people or two hundred. The mean of 66.7 is the long-run average across many repetitions of the whole search and is a poor prediction of any single one. This is a general feature of geometric distributions with small p, and it is the reason the mean alone should never be quoted as an expected wait without the spread beside it.
OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, p. 247
Error analysis
A success probability of 0.35, and the question is how many trials are expected until the first success.
Annotate
On: \( \begin{aligned} &(1)\; \mu = 0.35 \\ &(2)\; \mu = np = ? \\ &(3)\; \mu = 1 - 0.35 = 0.65 \\ &(4)\; \mu = \tfrac{1}{0.35} \approx 2.86 \end{aligned} \)
Errors (1) and (3) both produce a number below one, which is the giveaway: a geometric mean is always at least one, and it is large exactly when p is small. Error (2) is the more interesting one, because it shows the two families cannot share a formula — the absence of an n is precisely what defines a geometric problem.
Estimation
A machine produces a defect with probability 0.004.
Predict first
About how many items will be produced before the first defect?
Correct: About 250.
Why: The mean is one over 0.004, which is 250. The reciprocal relationship makes these estimates quick: a one-in-a-thousand event takes about a thousand trials on average, and a one-in-four event takes about four. The standard deviation here is also about 250, so the actual wait could easily be fifty items or six hundred.
Two truths and a lie
All three concern the geometric summaries.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The most likely value is always x equals 1, whatever p is, because the probabilities decay from the start. For p equal to 0.1 the mean is 10 and the mode is 1, which is section 2.6's mean-mode gap on a strongly right-skewed distribution.
Prediction
Commit before reasoning.
Predict first
A success probability is halved. What happens to the expected number of trials?
Correct: It doubles.
Why: The mean is one over p, so halving p doubles the reciprocal: a one-in-ten event takes about ten trials and a one-in-twenty event takes about twenty. The inverse relationship is the whole content of the formula, and it is why rare events have long waits in exact proportion to their rarity.
Section
Section 4
Concept
The geometric distribution is memoryless, and it is the only distribution with a discrete random variable that is. All events prior to the present are irrelevant: the probability of a success on the next trial begins anew each time, regardless of how many failures have gone before.
memoryless — A distribution is memoryless when the probability of waiting a further n trials, given that k have already passed without success, equals the unconditional probability of waiting n trials. The geometric is the only discrete distribution with this property.
\[ P(X > k + n \mid X > k) = P(X > n) \]
The book's illustration is a baseball player who hits with probability 0.20 and has failed his last ten times at bat. His probability of a hit next time is still 0.20, because the answer ignores his ten previous failures entirely. That follows directly from the second characteristic: the trials are independent, so nothing about the past can inform the future.
Figure (svg): A diagram illustrating that a geometric distribution has no memory of previous failures
OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 242-243 — memorylessness and the baseball example
Picture it
The property stated, and its formal form.
Figure (svg): A diagram illustrating that a geometric distribution has no memory of previous failures
The book draws out a consequence worth noticing: testing manufactured parts for defects, the geometric distribution begins with a clean slate each time, with no consideration of previous test results. That is a real property of genuinely independent trials and a poor description of a process that wears out or improves — which is the next idea.
Worked example
The book's own example, worked through the definition.
\[ p = 0.20; \text{ the last ten attempts all failed. What is the probability of a hit next time?} \]
Recall the second characteristic
Why: The trials are independent.
Ask what the ten failures tell you
Why: Nothing about the eleventh trial.
State the probability
Why: The same as on any trial.
\[ 0.20 \]
Note the general form
Why: Waiting resets each time.
\[ P(X > k + n | X > k) = P(X > n) \]
Figure (svg): The solution to Worked example the batter's next attempt shown as a ladder of expressions, one row per legal move
\[ P(\text{hit next}) = p = 0.20 \]
Verify: confirm this is not the gambler's fallacy in reverse
Why: Section 3.1 warned against expecting a tail after a run of heads, and this is the same principle applied honestly. The failures neither raise the next probability, as a 'due for a hit' argument would claim, nor lower it. Memorylessness says the probability is exactly unchanged, which is the only position consistent with independence — and it is worth stating explicitly, because both errors are common in the other direction.
OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, p. 243
Discrimination
Ask whether the success probability could change over the sequence.
Sort into buckets
Sort each situation.
Worked example
Checking the formal statement against the formula.
\[ \text{Show } P(X > k+n \mid X > k) = P(X > n) \]
Write the upper tail
Why: More than m trials means the first m failed.
\[ P(X > m) = q ^{m} \]
Write the conditional
Why: Section 3.1's definition.
\[ P(X > k + n AND X > k) / P(X > k) \]
Simplify the numerator
Why: Exceeding k+n already implies exceeding k.
\[ q ^{k + n} / q ^{k} \]
Cancel
Why: The powers subtract.
\[ q ^{n} = P(X > n) \]
Figure (svg): The solution to Worked example the property stated formally shown as a ladder of expressions, one row per legal move
\[ \frac{q^{k+n}}{q^k} = q^n = P(X > n) \]
Verify: confirm which property of the exponents did the work
Why: The cancellation works because the upper tail is a pure power of q, so dividing subtracts exponents and the k vanishes. No other discrete distribution has tails of that form, which is why the geometric is the only discrete memoryless one — the book states this and the algebra shows why. Section 5.3's exponential distribution has an exponential tail for the same reason and is the continuous counterpart.
OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 242-243
Trap
\[ \text{a dart thrower hits the centre with probability } 0.17 \]
Model the throws as geometric across a long practice session
Why: The throws look like independent trials with a fixed success rate.
\[ \text{but the thrower improves with practice} \quad (p \text{ is not constant}) \]
The third characteristic fails: a later throw has a higher success probability than an early one.
\[ \text{geometric requires a constant } p \text{, so no learning} \]
Ask whether the success probability could drift over the sequence
Why: Skill, wear and fatigue all break the constant-p assumption.
The book raises this directly: in experiments requiring skill, such as hitting a baseball or throwing a dart, one might consider that learning during the experiment would alter the probability of a success, and the geometric distribution cannot capture learning. The historical success probability is assumed constant. That is a modelling assumption to be stated rather than a fact, and it is exactly where a geometric model of a human activity is most vulnerable.
Prediction
Commit before reasoning.
Predict first
A geometric process has p = 0.05 and twenty trials have failed. What is the expected number of FURTHER trials until a success?
Correct: 20, the same as at the start.
Why: The mean is one over 0.05, which is 20, and memorylessness says the process resets: the expected further wait is the same 20 trials it was before any of the failures occurred. Nothing is overdue, which is section 3.1's gambler's fallacy again — and the coincidence that twenty trials have passed is irrelevant rather than significant.
Two truths and a lie
All three concern memorylessness.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and is the gambler's fallacy. Memorylessness means the probability is exactly UNCHANGED — neither raised nor lowered. The run of failures is information about the past only, and the process begins anew each trial.
Socratic
Many real waiting times are not memoryless.
Discussion prompt
Give a real waiting time that is clearly NOT memoryless, and say which direction the conditioning goes.
Hint: Think about something that wears out, or something that is nearly due.
Answer:
A mechanical component that wears out is the standard case: having run for ten years, it is MORE likely to fail in the next year than a new one was, because the failure rate rises with age. Conditioning on survival makes the remaining wait shorter.
The opposite also occurs. Manufactured electronics often have a high early failure rate, so a component that has survived its first month is LESS likely to fail soon than a fresh one — conditioning on survival makes the remaining wait longer.
The geometric and its continuous cousin the exponential sit exactly between these, with a failure rate that never changes. That makes them the right model for genuinely random events like radioactive decay and the wrong model for anything that ages, which is why reliability engineering uses other families for wear-out.
Section
Section 5
Concept
Setting up means answering the same questions as for a binomial with one change: there is no n. Identify one trial, decide what counts as the success, read off p, and note that X is the number of the trial on which that success occurs.
setting up a geometric problem — One trial, a success and its probability p, and a variable counting the trial number of the first success. The absence of a fixed n is the signal, usually carried by the word until.
\[ X \sim G(p): \quad P(x) = q^{x-1}p, \; \mu = \frac{1}{p}, \; \sigma = \frac{\sqrt{q}}{p} \]
The book's Example 4.18 is worth noticing for one detail in its setup: the accident reports are selected randomly AND REPLACED IN THE PILE after reading. That clause exists to keep p constant, satisfying the third characteristic. Without replacement the probability would shift as reports were removed, and the geometric model would not apply — the same replacement issue section 3.2 and section 4.3 both turned on.
Figure (svg): Two columns contrasting the binomial, which fixes the number of trials, with the geometric, which fixes the number of successes
OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 243-247 — the setup of Examples 4.18 and 4.21
Picture it
What each fixes and what each counts.
Figure (svg): Two columns contrasting the binomial, which fixes the number of trials, with the geometric, which fixes the number of successes
The rows to carry are the second and third. A binomial answer is a count of successes between zero and n; a geometric answer is a trial number of at least one, with no ceiling. If an answer's units do not match the question's, the wrong family was used, and that check catches the confusion before the arithmetic does.
Worked example
Example 4.18 from the top.
\[ 35\% \text{ of accidents come from failure to follow instructions; reports are read (and replaced) until one is found} \]
What is one trial?
Why: Reading one accident report.
What is the success?
Why: Finding one caused by failure to follow instructions.
\[ p = 0.35 \]
Why does replacement matter?
Why: It keeps p the same on every trial.
Define X and its values
Why: The report on which the first is found.
\[ x = 1, 2, 3,... \]
Figure (svg): The solution to Worked example a full setup shown as a ladder of expressions, one row per legal move
\[ X \sim G(0.35), \quad \mu \approx 2.86, \quad P(x \ge 3) = 0.4225 \]
Verify: confirm the replacement clause is doing real work
Why: Without replacement, reading a non-relevant report would remove it from the pile and slightly raise the proportion of relevant ones remaining, so p would creep upward. Over a few reports the drift is tiny, but the clause makes the model exact rather than approximate. It is the same reason section 4.3's binomial needed replacement, and it is why textbook problems specify it so often — the specification is what licenses the formula.
OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 243-244
Sorting
Look for a fixed n, or for the word until.
Sort into buckets
Sort each problem.
The absence of an n is the reliable structural signal, and 'until' is the reliable verbal one. When both point the same way the identification is safe; when a problem gives an n and also says until, read it again, because it is probably asking two questions.
Worked example
Example 4.21. The lifetime risk of developing cancer is about one in 67.
\[ X = \text{ the number of people asked until one says they have cancer} \]
Read off p
Why: One in 67, about 1.5 percent.
\[ p = 0.015 \]
Compute a single probability
Why: The tenth person is the first to say yes.
\[ (0.985) ^{9}(0.015) \]
Evaluate
Why: About one and a third percent.
\[ 0.0131 \]
Compute the summaries
Why: Reciprocal and the spread formula.
\[ \mu 66.7, \sigma 66.2 \]
Figure (svg): The solution to Worked example a rare-event setup shown as a ladder of expressions, one row per legal move
\[ P(10) = (0.985)^9(0.015) \approx 0.0131 \]
Verify: confirm why the individual probabilities are all small
Why: With a mean of 66.7 the probability is spread across many values, so no single one carries much — the largest, at x equals 1, is only 0.015. That is characteristic of a geometric distribution with small p: the probability is thinly spread over a long tail. It also means questions about ranges are usually more informative than questions about single values, which is why the upper-tail shortcut q to the k is so useful here.
OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, p. 247
Trap
\[ \text{'how many rolls until the first six?'} \;\to\; \text{find } n \]
Search the problem for a number of trials
Why: Every binomial problem had one, so its absence looks like missing information.
\[ \text{there is no } n \text{, and its absence is the point} \]
The number of trials is what the question is asking about, so it cannot also be given.
\[ \text{a geometric problem gives only } p \]
Treat a missing n as the signal for a geometric problem, not as an omission
Why: One parameter is all the family needs.
This is worth naming because a missing quantity usually means a problem is under-specified, and here it means the opposite: the geometric distribution has exactly one parameter, and p determines the probabilities, the mean and the standard deviation together. The word 'until' in the problem is the positive signal, and the absence of an n is the confirming one.
Matching
Each formula belongs to one of the two families.
Match the pairs
Why: The presence of n is the tell in every case. A geometric formula cannot contain n, because the number of trials is the random variable rather than a parameter — so any formula with an n in it belongs to the other family.
Two truths and a lie
All three concern setting up a geometric problem.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and inverts the situation. The number of trials is the random variable, so stating it would answer the question. The absence of an n is the structural signal that a problem is geometric rather than binomial.
Explain it
A classmate cannot see the difference between the two families.
Discussion prompt
In three sentences or fewer, give them one experiment and two questions that separate the families cleanly.
Hint: Use one die and change only the question.
Answer:
Give them a die and ask two questions: how many threes in ten rolls, and how many rolls until the first three.
The first fixes ten trials and lets the count of successes vary, so it is binomial with eleven possible answers; the second fixes one success and lets the number of trials vary, so it is geometric with infinitely many.
Same die, same p — what changed is which quantity was decided in advance, and that is the only thing separating the two families.
Comparison
Fill the blanks. One experiment can supply both, depending on the question.
Comparison matrix
| Feature | Binomial | Geometric |
|---|---|---|
| Fixed in advance | the number of trials, n | the number of successes: exactly one |
| The random variable | the count of successes | the trial number of the first success |
| Values | 0 through n | 1, 2, 3, ... with no upper bound |
| Mean | np | 1/p |
| Combination factor | yes: many orderings give x successes | no: only one ordering works |
The two answer opposite halves of the same question, and telling them apart is a matter of reading which quantity the problem decided in advance. A formula containing n is binomial; a formula containing only p is geometric.
Pattern
Six steps, and the first is a family identification rather than a calculation.
The exponent is x minus one, not x. Writing out the sequence for x equal to 1 confirms it: a success on the first trial has probability p, and the formula must give that.
OpenStax Introductory Business Statistics 2e, §4.3 Geometric Distribution §4.3 Geometric Distribution
Check
Identify the family.
Check your understanding
Which of these is a geometric experiment?
Answer: A
Why: The experiment repeats until the first success and then stops, so the number of trials is the random variable and no n is given.
Check
Apply the formula.
Check your understanding
For a geometric distribution with p = 0.2, what is P(x = 4)?
Answer: A
Why: Three failures then a success on the fourth trial, so q is raised to x minus one and multiplied by p.
Check
The mean.
Check your understanding
A process succeeds with probability 0.04. About how many trials until the first success?
Answer: A
Why: The mean is one over p, which is one over 0.04, giving 25 trials.
Real world
A recruiter says that about 4 percent of applications lead to an interview, so a candidate should expect an interview after about 25 applications. A candidate has sent 60 applications and had none, and concludes something must be wrong with their CV.
Discussion prompt
Model the situation, say whether 60 without success is surprising, and identify the modelling assumption most likely to be wrong.
Hint: Compute the mean and the upper-tail probability, then ask whether p is really constant across candidates.
Answer:
The model is geometric with p = 0.04. The mean is one over 0.04, which is 25 applications, so the recruiter's figure is right as a long-run average. But sigma is the square root of 0.96 over 0.0016, about 24.5 — almost as large as the mean, so the spread is enormous.
Sixty without success is not surprising. The probability of more than 60 failures is 0.96 to the power 60, which is about 0.086 — roughly one candidate in twelve. On a distribution this skewed, waits far above the mean are routine rather than remarkable.
The assumption most likely wrong is the constant p. The geometric model treats every candidate as having the same 4 percent rate, when in reality the rate varies enormously between candidates and between applications. The 4 percent is an average ACROSS applicants, and no individual's rate need be near it.
\[ \mu = 25, \quad \sigma \approx 24.5, \quad P(X > 60) = (0.96)^{60} \approx 0.086 \]
So the candidate's conclusion may still be right, but the sixty failures alone are weak evidence for it — one in twelve people with a perfectly ordinary CV would see the same run. What would be better evidence is comparison with similar candidates, which is a question about differing p values rather than about one geometric distribution. The general lesson is that a mean quoted without its spread invites exactly this error, and for a geometric distribution the spread is always about as large as the mean.
Commit first
Answer, then rate your confidence honestly.
Predict first
Why does the geometric formula have no combination factor?
Correct: Because exactly one sequence produces it.
\[ \text{binomial: many orderings} \qquad \text{geometric: exactly one} \]
Why: The definition of the event fixes the ordering completely — every trial before the xth failed and the xth succeeded — so there is nothing to count. The binomial needs its combination factor precisely because x successes among n trials can be arranged in many ways. Independence and a constant p are what let the single sequence's probability be a product, but they do not affect the counting.
Explain it
They have written P(x = 5) as q to the fifth times p.
Discussion prompt
In three sentences or fewer, show them the exponent using a small case.
Hint: Ask them what the formula gives for x = 1.
Answer:
Ask them what their version gives for x equal to 1: it says q times p, but a success on the very first trial should just be p.
The exponent counts the trials BEFORE the success, and on trial five there are four of them, so it is q to the fourth times p.
Writing the sequence out — failure, failure, failure, failure, success — makes the four visible and settles it permanently.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the family, look for a fixed n or for the word until. For the exponent, check that the formula gives p when x is 1. For the mean, remember it is at least one and is large when p is small. For memorylessness, it means the probability is unchanged — neither raised nor lowered — by past failures. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top, write the four characteristics of a geometric experiment and mark which two it shares with the binomial. Beside them, draw a row of five boxes labelled F, F, F, F, S and write underneath the probability of that exact sequence, then generalise it to the formula. Check your formula by setting x equal to 1 and confirming it gives p. In the middle, work Example 4.18 completely: 35 percent of accident reports are relevant and reports are read with replacement until one is found. Write the setup in four lines, compute the mean and the standard deviation, and compute P(x at least 3) using the q-to-the-k shortcut rather than a sum. Below that, draw the first eight stems of the distribution to scale and mark which value is most likely, then mark the mean and write one sentence on why the two are different. At the bottom, make a two-column table comparing binomial and geometric on five rows: what is fixed, what X counts, the values X takes, the mean, and whether a combination factor appears. Finish by writing the memoryless property in symbols and one sentence naming a real waiting time that is NOT memoryless.
Check your stem drawing: the tallest stem must be at x = 1 and every later stem must be exactly q times the one before it. If your tallest stem is near the mean, the distribution was drawn as though it were binomial — a geometric distribution always decays from its first value.
Recap
Five things, and the whole family follows from one parameter.
| If you see | Then |
|---|---|
| The word until, and no n | Geometric: X is the trial number |
| A stated number of trials | Binomial: X is the count of successes |
| A request for P(x = k) | q to the k minus one, times p |
| A request for P(x greater than k) | q to the k, in one step |
| A request for the expected wait | 1/p, always at least one |
| A run of failures so far | Irrelevant: the process is memoryless |
| Skill, learning or wear in the trials | p is not constant; the model does not apply |
Section 4.5 relaxes the other binomial characteristic. When sampling is done without replacement from two groups, the trials are not independent and p changes at every draw — and the count of items from the group of interest has a hypergeometric distribution.
OpenStax Introductory Statistics 2e, §4.4 Geometric Distribution §4.4, pp. 242-247 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.