The first named family of distributions. A binomial experiment has three characteristics: a fixed number of trials n, exactly two outcomes on each trial with success probability p and failure probability q summing to one, and independent trials repeated under identical conditions so that p never changes. The random variable X counting the successes then has a binomial distribution, written X follows B of n and p, and its probabilities come from a formula whose three factors count the orderings and give the probability of one of them. Its mean is np and its standard deviation is the square root of npq, so two numbers determine the whole distribution — which is what makes a named family worth having.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 4 — Discrete Random Variables
Binomial Distribution
Objectives
Five outcomes. The first is a check, and everything else depends on it passing.
OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, pp. 237-241 — the section these objectives are drawn from
Warm-up
Section 4.2 computed a mean and standard deviation from a distribution table, row by row.
Discussion prompt
Suppose you flip a fair coin ten times and count the heads. Building the whole table means computing eleven probabilities. Is there anything about this situation that might let you skip the table?
Hint: Ask what is the same about every one of the ten flips.
Answer:
Every flip is identical: the same two outcomes, the same probability, and no flip affects any other. So the whole situation is described by just two numbers — ten flips and a success probability of a half.
If two numbers determine the situation, they ought to determine the distribution, and they do. There is a formula for each probability, and formulas for the mean and standard deviation that need no table at all.
That is what a named distribution is: a shape of problem common enough to be worth solving once in general. This section covers the commonest of them, and the price of using it is checking that the situation really has the three characteristics it assumes.
Concept
There are three characteristics of a binomial experiment. There are a fixed number of trials, denoted n. There are only two possible outcomes, called success and failure, with probabilities p and q which sum to one. And the n trials are independent and repeated under identical conditions, so that p and q remain the same for every trial.
binomial experiment — An experiment with a fixed number n of independent trials, each having exactly two outcomes with the same success probability p on every trial. The random variable X is the number of successes in the n trials, and it takes the values 0 through n.
\[ X \sim B(n, p), \qquad p + q = 1 \]
The word success carries no approval: it names whichever outcome is being counted. The book's Example 4.9 makes the point by defining a success as a student WITHDRAWING from a physics course, because that is the quantity of interest. Choosing which outcome to call a success is free, and it determines p — swapping the labels replaces p by q and counts the other thing, which is often the easier route.
Figure (svg): The three characteristics of a binomial experiment listed as a numbered procedure
OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 237
Section
Section 1
Concept
A fixed number of trials, exactly two outcomes per trial with probabilities p and q summing to one, and independent trials under identical conditions. Because the trials are independent, the outcome of one does not help predict another, and p and q remain the same throughout.
Bernoulli trial — Any experiment with characteristics two and three and n equal to one: a single independent trial with two outcomes. A binomial experiment counts the successes in one or more Bernoulli trials, and is named after Jacob Bernoulli, who studied them in the late 1600s.
\[ \text{fixed } n, \quad \text{two outcomes}, \quad \text{independent with constant } p \]
All three must hold and each fails in a recognisable way. The number of trials is not fixed when the experiment runs until something happens, which is section 4.4's geometric distribution. There are more than two outcomes when a die is rolled and the face recorded. And the trials are not independent with constant p when sampling without replacement from a small population, which is section 4.5's hypergeometric distribution. Each failure has its own named remedy later in the chapter.
Figure (svg): The three characteristics of a binomial experiment listed as a numbered procedure
OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 237 — the three characteristics and the Bernoulli trial
Picture it
The checklist to run before reaching for any binomial formula.
Figure (svg): The three characteristics of a binomial experiment listed as a numbered procedure
It is worth running the list explicitly rather than by feel, because a situation can look binomial and fail the third condition invisibly. Drawing cards without replacement has a fixed n and two outcomes and is not binomial, since p changes on every draw — and the formula would give a confident wrong answer with nothing to flag it.
Worked example
Example 4.12. The book asks why this is a binomial problem.
\[ 70\% \text{ of } 50 \text{ statistics students do homework on time, independently} \]
Is the number of trials fixed?
Why: Fifty students in the class.
\[ n = 50 \]
Are there exactly two outcomes?
Why: On time, or not on time.
Is the probability constant and are trials independent?
Why: Each student works independently at 0.70.
\[ p = 0.70\text{ throughout} \]
Name the failure and q
Why: Not completing on time.
\[ q = 0.30 \]
Figure (svg): The solution to Worked example checking the three characteristics shown as a ladder of expressions, one row per legal move
\[ X \sim B(50, 0.70), \quad x = 0, 1, \ldots, 50 \]
Verify: confirm the values of X and the meaning of a failure
Why: X takes the values 0 through 50, one for every possible number of on-time students, and a failure is a student who does not complete on time. Writing the failure out in words is worth doing because q is defined by it, and a problem that describes only the success rate leaves q to be computed as one minus p. Here 0.30, and p plus q is one as the second characteristic requires.
OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, pp. 238-239
Sorting
Test each situation against the three characteristics.
Sort into buckets
Sort each experiment.
Items (d) and (c) are not dead ends: they are sections 4.4 and 4.5, which are named distributions for exactly those two failures. Recognising which condition breaks tells you which family to reach for instead.
Worked example
Example 4.9. The label is a choice, and it fixes p.
\[ \text{The withdrawal rate from a physics course is } 30\%. \]
Identify what is being counted
Why: The book counts students who withdraw.
Set p accordingly
Why: The withdrawal rate.
\[ p = 0.30 \]
Set q as the complement
Why: Students who stay the whole term.
\[ q = 0.70 \]
Note the alternative labelling
Why: Counting stayers instead swaps the two.
\[ p = 0.70, q = 0.30 \]
Figure (svg): The solution to Worked example choosing which outcome is the success shown as a ladder of expressions, one row per legal move
\[ X = \text{the number who withdraw}, \quad p = 0.30 \]
Verify: confirm both labellings answer the same questions
Why: Counting withdrawals with p equal to 0.30 and counting stayers with p equal to 0.70 describe the same class, and any question about one can be rephrased about the other — the probability that at least 5 of 20 withdraw equals the probability that at most 15 stay. The labels are free, and the useful habit is to choose whichever makes the question a simple inequality rather than a complement.
OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 237
Trap
\[ \text{draw } 5 \text{ cards from a deck; count the hearts} \]
Treat it as binomial with n = 5 and p = 13/52
Why: There is a fixed number of draws and two outcomes, so two conditions hold.
\[ \text{but } p \text{ changes after each draw} \quad \text{(without replacement)} \]
After a heart is drawn, twelve remain among fifty-one, so the second trial has p equal to twelve fifty-firsts rather than a quarter.
\[ \text{binomial only WITH replacement; otherwise hypergeometric} \]
Check the third condition explicitly, especially for sampling problems
Why: Fixed n and two outcomes are not enough on their own.
Section 3.2 established that sampling with replacement gives independent trials and sampling without does not, and this is where that distinction pays. The binomial assumes replacement; section 4.5's hypergeometric distribution is the version for sampling without it. Section 1.2's remark applies too: when the population is very large relative to the sample, p barely moves and the binomial is an excellent approximation, which is why survey data are analysed with it.
Fill the middle
The second characteristic.
Fill in the blanks
p + q = 1, \text___ q = 1 - p
Why: Exactly one, because success and failure are complementary events on each trial — section 3.1's complement rule applied to a single trial. A problem stating only p leaves q to be computed, and forgetting to do so is a common source of a wrong exponent in the formula.
Two truths and a lie
All three concern the characteristics.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false because it omits the third condition. Drawing five cards without replacement has a fixed n and two outcomes and is not binomial, because p changes on every draw. All three characteristics are required, and the third is the one that fails silently.
Prediction
Commit before reasoning.
Predict first
The binomial formula multiplies p by itself x times. What does that assume?
Correct: That the trials are independent.
Why: Section 3.3's multiplication rule gives a joint probability as a product only when the events are independent; otherwise a conditional is required. Raising p to the power x is exactly that product for a sequence of x successes, so the third characteristic is what licenses the formula's middle factor. Neither the size of p nor of n has anything to do with it.
Section
Section 2
Concept
The notation X follows B of n and p is read as X is a random variable with a binomial distribution, whose parameters are n, the number of trials, and p, the probability of a success on each trial.
the notation B(n, p) — X ~ B(n, p) says X has a binomial distribution with n trials and success probability p. Two numbers specify the distribution completely, which is what makes the family worth naming.
\[ X \sim B(20, 0.41): \quad n = 20, \; p = 0.41, \; q = 0.59, \; x = 0, 1, \ldots, 20 \]
Setting a problem up means answering four questions in order: what is one trial, what counts as a success, how many trials are there, and what is being asked about the count. The book's Example 4.12 walks through exactly these, ending with the translation of 'at least 40' into an inequality — which is the step where most binomial problems are lost, because the arithmetic afterwards is mechanical.
Figure (svg): Two columns translating everyday phrases into the inequalities a binomial probability question needs
OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 239 — the notation, and Example 4.13's setup
Picture it
Six phrases and what each one means for x.
Figure (svg): Two columns translating everyday phrases into the inequalities a binomial probability question needs
The endpoint matters here in a way it will not in chapter 5. On a discrete distribution P(x = 12) is a real positive number, so 'at most 12' and 'fewer than 12' differ by that whole stem — for Example 4.13's distribution they differ by about 0.06, which is a substantial slice of the answer.
Worked example
Four questions, answered in order.
\[ 41\% \text{ of adult workers have a high school diploma and no further education; } 20 \text{ are selected} \]
What is one trial?
Why: Selecting one adult worker.
What is a success?
Why: Having the diploma and no further education.
\[ p = 0.41 \]
How many trials?
Why: Twenty workers selected.
\[ n = 20 \]
What is asked?
Why: At most 12 of them.
\[ P(x \le 12) \]
Figure (svg): The solution to Worked example setting up Example 4.13 shown as a ladder of expressions, one row per legal move
\[ X \sim B(20, 0.41), \quad q = 0.59, \quad \text{find } P(x \le 12) \]
Verify: confirm the setup before computing anything
Why: Every later step depends on these four answers, and all four are read from the problem rather than computed. Note in particular that q is 0.59 rather than being given — the problem states only p, and section 4.1's habit of writing the complement immediately avoids a wrong exponent later. The book gives the answer as 0.9738, which stats-lib reproduces exactly.
OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 239
Translation
Each phrase, and the inequality it names.
Match the pairs
Why: At least and at most include their endpoints; more than and fewer than exclude them. The pair to watch is at least 4 and more than 3, which name the same set — recognising that equivalence often turns an awkward phrase into a familiar one.
Worked example
Example 4.10. The book asks only for the statement, not the answer.
\[ \text{Win } 55\% \text{ of games, play } 20 \text{ times, want to win } 15 \]
Define the variable
Why: The number of wins.
\[ X =\text{ number of wins} \]
List its values
Why: Anything from none to all twenty.
\[ 0, 1,..., 20 \]
Read off the parameters
Why: Twenty games at 55 percent.
\[ n = 20, p = 0.55, q = 0.45 \]
State the question
Why: Exactly fifteen wins.
\[ P(x = 15) \]
Figure (svg): The solution to Worked example stating the question mathematically shown as a ladder of expressions, one row per legal move
\[ X \sim B(20, 0.55), \quad \text{find } P(x = 15) \]
Verify: confirm that 'exactly' means a single value
Why: Fifteen wins exactly is one stem of the distribution, so the answer is a single application of the formula rather than a sum. Contrast Example 4.13's 'at most 12', which covers thirteen stems and needs a cumulative calculation. Recognising which of the two a question asks for, before touching a formula, is what decides whether one term or many are needed.
OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 237
Trap
\[ \text{'more than 3 heads' in five flips} \]
Compute P(x is at least 3), including the value 3
Why: More than three and three or more sound alike in a hurry.
\[ P(3) + P(4) + P(5) \quad \text{(the value 3 does not qualify)} \]
For the altered coin P(3) is 0.0879, which is more than five times the correct answer of 0.0156.
\[ P(x > 3) = P(4) + P(5) = 0.0146 + 0.0010 = 0.0156 \]
Write the list of qualifying values before computing anything
Why: The endpoint is decided once, in the listing.
On a discrete distribution the endpoint carries real probability, so 'more than' and 'at least' give genuinely different answers — here by a factor of six, because the excluded stem is larger than both included ones combined. Section 4.1 made the same point and it becomes more consequential here, since binomial questions are usually phrased in exactly these words.
Faded example
Forty-one percent of twenty selected workers have the qualification.
Fill in the blanks
n = 20, \quad p = 0.41, \quad q = 0.59
Why: Twenty trials with a success probability of 0.41, so q is 0.59. The problem states only p, and computing q immediately is what prevents a wrong exponent in the formula later.
Sorting
Decide whether the question needs one application of the formula or a sum.
Sort into buckets
Sort each question for X following B(20, 0.41).
For the ranges, the complement is often shorter: P(at least 15) covers six stems while its complement covers fifteen, so the direct route wins there — but P(more than 3) covers seventeen stems against four for its complement, and the complement wins easily.
Two truths and a lie
All three concern the setup.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. At most 12 includes the value 12 and fewer than 12 excludes it, and on a discrete distribution that value carries real probability — about 0.06 for B(20, 0.41), which is a substantial part of any answer near there.
Section
Section 3
Concept
The probability of exactly x successes in n trials is the number of ways to choose which x trials succeed, multiplied by the probability of x successes and n minus x failures in one particular arrangement.
the binomial formula — P(x) equals n choose x, times p to the power x, times q to the power n minus x. The first factor counts the orderings; the second and third give the probability of any one of them, which is the same for all.
\[ P(x) = \binom{n}{x} p^x q^{n-x} \]
The structure is worth understanding rather than memorising. Any particular sequence with x successes and n minus x failures has probability p to the x times q to the n minus x, by section 3.3's multiplication rule with independence. Every such sequence has the same probability, and they are mutually exclusive, so the total is that common probability times how many sequences there are — which is the combination count.
Figure (svg): The binomial probability formula with each of its three factors labelled
OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, pp. 237-238 — the binomial distribution and Example 4.11
Picture it
What each part of the formula contributes.
Figure (svg): The binomial probability formula with each of its three factors labelled
The two exponents must total n, and checking that is the fastest test of a substitution: x successes and n minus x failures account for every trial exactly once. An expression with exponents summing to anything else has miscounted the trials, and the error is visible without evaluating anything.
Worked example
Example 4.11's distribution, at a single value.
\[ X \sim B(5, 0.25); \quad \text{find } P(x = 4) \]
Count the orderings
Why: Which four of the five flips are heads.
\[ C(5, 4) = 5 \]
Probability of four successes
Why: A quarter, four times.
\[ 0.25 ^{4} = 0.00390625 \]
Probability of one failure
Why: Three quarters, once.
\[ 0.75 ^{1} = 0.75 \]
Multiply the three
Why: Five times the product.
\[ 0.014648 \]
Figure (svg): The solution to Worked example one binomial probability shown as a ladder of expressions, one row per legal move
\[ P(4) = \binom{5}{4}(0.25)^4(0.75)^1 \approx 0.0146 \]
Verify: confirm the exponents account for every trial
Why: Four successes and one failure is five trials, matching n, so no trial has been counted twice or left out. The combination count of 5 is also checkable directly: the single failure can be any one of the five flips, so there are five arrangements. Both checks are quick and between them they catch the two commonest substitution errors.
OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 238
Faded example
X follows B(5, 0.25). Find P(x = 2).
Fill in the blanks
P(2) = \binom23(0.25)^___}(0.75)^___}
Why: Two successes and three failures, and the exponents total five as they must. The combination count is 10, giving 10 times 0.0625 times 0.421875, which is about 0.2637 — the largest single probability after P(1) in this distribution.
Worked example
Example 4.11. The book develops the full density before answering.
\[ X \sim B(5, 0.25); \quad \text{find } P(x > 3) \]
List the qualifying values
Why: More than three means four or five.
\[ x = 4, 5 \]
Compute P(4)
Why: As above.
\[ 0.014648 \]
Compute P(5)
Why: All five heads: a quarter to the fifth.
\[ 0.000977 \]
Add them
Why: Distinct values are exclusive.
\[ 0.015625 \]
Figure (svg): The solution to Worked example the whole distribution shown as a ladder of expressions, one row per legal move
\[ P(x > 3) = 0.0146 + 0.0010 = 0.015625 \]
Verify: confirm the whole distribution sums to one
Why: The six probabilities are 0.2373, 0.3955, 0.2637, 0.0879, 0.0146 and 0.0010, and they total exactly 1.0000 — section 4.1's second condition, which every binomial distribution satisfies automatically. Computing the full set as the book does is more work than the question needs, but it supplies that check and it makes the shape visible: the mass piles at one head, which is where np sits.
OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 238
Error analysis
Five flips of a coin altered to p = 0.25. Four students compute the probability of exactly four heads.
Annotate
On: \( \begin{aligned} &(1)\; (0.25)^4 = 0.0039 \\ &(2)\; \binom{5}{4}(0.25)^4(0.75)^4 \\ &(3)\; \binom{5}{4}(0.25)^4 = 0.0195 \\ &(4)\; \binom{5}{4}(0.25)^4(0.75)^1 \approx 0.0146 \end{aligned} \)
Errors (2) and (3) are both caught by the exponent check — the two powers must sum to n. Error (1) is subtler because the arithmetic is valid; what it computes is the probability of a specific ordering rather than of the count, which is exactly the distinction the combination factor exists to bridge.
Discrimination
Each expression omits or mishandles one part of the formula.
Sort into buckets
Sort each faulty expression by what it got wrong.
Two truths and a lie
All three concern the formula.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The combination counts orderings and is needed whatever p is: for a fair coin, four heads in five flips still happens five ways, and omitting the factor understates the probability by a factor of five. The value of p affects the other two factors and never the count.
Prediction
Commit before reasoning.
Predict first
In the formula, what is n choose x counting?
Correct: The number of orderings producing exactly x successes.
Why: Four heads in five flips can happen five ways, depending on which flip is the tail, and each way has the same probability. The formula multiplies that common probability by the count of ways, which is exactly section 4.1's observation that several outcomes can map to one value of a random variable — here made quantitative.
Section
Section 4
Concept
The mean of a binomial probability distribution is n times p, and the variance is n times p times q, so the standard deviation is the square root of npq.
binomial mean and standard deviation — For X following B(n, p), the mean mu is np and the variance sigma squared is npq, so sigma is the square root of npq. Both follow from n and p without building a table.
\[ \mu = np, \qquad \sigma^2 = npq, \qquad \sigma = \sqrt{npq} \]
Section 4.2 computed a mean by building a table and summing x times P(x) row by row. For a binomial with fifty trials that would be fifty-one rows. The formula np gives the same answer in one multiplication, and this is the main practical payoff of recognising a named family — the general method still works and is never needed.
Figure (svg): The binomial mean and standard deviation formulas with a worked instance
OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 237 — the binomial mean and variance
Picture it
Both summaries from the two parameters.
Figure (svg): The binomial mean and standard deviation formulas with a worked instance
The formula for mu is worth sanity-checking against intuition: twenty workers with a 41 percent rate should yield about eight, and np gives 8.2. Any binomial mean that is not roughly the trials times the rate is an arithmetic error, and that check takes no time at all.
Worked example
Example 4.13's second question.
\[ X \sim B(20, 0.41): \text{ how many workers are expected to have the qualification?} \]
Identify n and p
Why: From the setup.
\[ n = 20, p = 0.41 \]
Apply the formula
Why: Trials times success probability.
\[ \mu = (20) (0.41) \]
Evaluate
Why: The product.
\[ 8.2 \]
Interpret
Why: The long-run average across many such samples.
\[ \text{about } 8\text{ workers} \]
Figure (svg): The solution to Worked example how many workers to expect shown as a ladder of expressions, one row per legal move
\[ \mu = np = (20)(0.41) = 8.2 \]
Verify: confirm the answer is not a possible count
Why: Eight point two workers cannot be observed in any single sample of twenty, since the count is a whole number — section 4.2's point about expected values, and it applies to every binomial mean with a non-integer np. What 8.2 describes is the average across very many samples of twenty, which the Law of Large Numbers connects to what would actually be observed.
OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 239
Faded example
X follows B(50, 0.70).
Fill in the blanks
\mu = (50)(0.70) = 35, \qquad \sigma = \sqrt3.24 \approx ___
Why: Thirty-five students on average, with a standard deviation of about 3.24. Both come from n and p alone; building the fifty-one row table would give the same answers after a great deal more work.
Worked example
The same two parameters, one more step.
\[ X \sim B(20, 0.41): \text{ find } \sigma \]
Compute the variance
Why: n times p times q.
\[ (20) (0.41) (0.59) \]
Evaluate it
Why: The product.
\[ 4.838 \]
Take the square root
Why: Back to counts.
\[ \text{about } 2.20 \]
Sanity-check the size
Why: Against a range of 0 to 20.
Figure (svg): The solution to Worked example the standard deviation shown as a ladder of expressions, one row per legal move
\[ \sigma = \sqrt{(20)(0.41)(0.59)} \approx 2.20 \]
Verify: confirm sigma is small relative to the range
Why: The count runs from 0 to 20 and sigma is about 2.2, roughly a ninth of the range — which is the usual pattern, since a binomial concentrates its probability within a few standard deviations of np. Section 2.7's rule of thumb applies: nearly all the probability lies within about two sigma of the mean, so counts between roughly 4 and 13 cover almost everything, and the book's P(x at most 12) of 0.9738 is consistent with that.
OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, pp. 237-239
Trap
\[ \sigma = npq = (20)(0.41)(0.59) = 4.838 \]
Report npq as the standard deviation
Why: The formula for the variance is the memorable one.
\[ \text{that is } \sigma^2 \text{, in squared workers} \]
A spread of 4.84 would be more than twice the true one, and it is in the wrong units.
\[ \sigma = \sqrt{npq} = \sqrt{4.838} \approx 2.20 \]
Write npq as the VARIANCE, then take the root
Why: Section 4.2's distinction, in a new formula.
This is section 2.7's variance-for-standard-deviation slip in binomial clothing, and the same guard works: a standard deviation carries the variable's own units, so if the answer cannot be described as a number of workers it is the variance. Checking the size against the range catches it too, since 4.84 is a quarter of the whole range for a distribution this concentrated.
Estimation
X follows B(10, 0.2).
Predict first
Around which value does the distribution peak?
Correct: Around 2.
Why: The mean is np, which is 2, and a binomial distribution peaks at or next to its mean. That gives a free sanity check on any binomial calculation: the largest probabilities should sit near np, so a computed distribution peaking elsewhere signals an error in n or p. The figure comparing three values of p shows the peak tracking np in each case.
Two truths and a lie
All three concern the binomial summaries.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and is exactly backwards. The product pq is largest at p equal to 0.5 and shrinks toward zero as p approaches either extreme, so sigma is greatest for a fair coin and smallest when the outcome is nearly certain — which makes sense, since a nearly certain outcome varies very little.
Prediction
Commit before reasoning.
Predict first
For n = 100, which success probability gives the largest standard deviation?
Correct: p = 0.5.
Why: Sigma is the square root of npq, and pq is maximised at p equal to 0.5, where it is 0.25. At p equal to 0.1 or 0.9 the product is 0.09, so sigma is about 3 rather than 5. The intuition is that an outcome close to certain has little room to vary, while a fifty-fifty outcome is maximally uncertain — and section 8.3's sample size formulas use exactly this fact.
Section
Section 5
Concept
A question about a range of values requires adding the probabilities of every value in the range. Because distinct values are mutually exclusive, the sum needs no correction, and the complement is often the shorter route.
cumulative binomial probability — The probability that X is at most some value, found by summing the individual probabilities up to it. A calculator's binomcdf computes it directly; by hand, the complement is often fewer terms.
\[ P(x \le 12) = \sum_{k=0}^{12} \binom{20}{k}(0.41)^k(0.59)^{20-k} \]
The book's Example 4.13 asks for P(x at most 12) with n equal to twenty, which is thirteen separate applications of the formula by hand — and gives the answer as 0.9738 from a calculator. This is where the named family earns its second keep: the probabilities are tabulated and built into every statistical tool, so the work is in setting the problem up correctly rather than in evaluating it.
Figure (svg): The binomial distribution for five flips of a coin with success probability 0.25, peaking at one head and skewed to the right
OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, pp. 239-240 — cumulative probabilities and the calculator functions
Picture it
Example 4.11's distribution with the two stems above three highlighted.
Figure (svg): The binomial distribution for five flips of a coin with success probability 0.25, peaking at one head and skewed to the right
Two stems here, and thirteen for Example 4.13. Which route is shorter is worth a moment's thought before starting: 'more than 3' out of five values covers two stems directly and four by complement, so the direct route wins; 'at most 12' out of twenty-one covers thirteen directly and eight by complement, so the complement wins.
Worked example
Example 4.13. The book quotes the calculator's answer.
\[ X \sim B(20, 0.41); \quad \text{find } P(x \le 12) \]
List the qualifying values
Why: Zero through twelve.
Note what a hand computation needs
Why: Thirteen applications of the formula.
Use the cumulative function
Why: binomcdf with n, p and 12.
\[ 0.9738 \]
Sanity-check against the mean
Why: The mean is 8.2, well below 12.
Figure (svg): The solution to Worked example a cumulative probability shown as a ladder of expressions, one row per legal move
\[ P(x \le 12) = 0.9738 \]
Verify: confirm the answer is plausible from mu and sigma
Why: The mean is 8.2 and sigma is about 2.2, so 12 sits about 1.7 standard deviations above the mean. Section 2.7's rule of thumb puts most of a distribution within two sigma, so a probability of about 0.97 below that point is exactly what should be expected. Running that check before trusting a calculator output catches a mistyped n or p, which otherwise produces a confident wrong number.
OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 239
Faded example
X follows B(20, 0.41). You want P(x at least 5).
Fill in the blanks
P(x \ge 5) = 1 - P(x \le 4), \text20 0 \text___ ___
Why: The complement stops at 4, so the two ranges partition the values 0 through 20 with no overlap. Subtracting the cumulative to 5 instead would drop the value 5 from both sides and undercount the answer by its probability.
Worked example
The same distribution, a range covering most of it.
\[ X \sim B(20, 0.41); \quad \text{find } P(x \ge 5) \]
Count the direct terms
Why: Five through twenty.
Count the complement's terms
Why: Zero through four.
Choose the complement
Why: Fewer terms.
\[ 1 - P(x \le 4) \]
Evaluate
Why: One minus the cumulative to four.
\[ \text{about } 0.9666 \]
Figure (svg): The solution to Worked example choosing the complement shown as a ladder of expressions, one row per legal move
\[ P(x \ge 5) = 1 - P(x \le 4) \approx 1 - 0.0334 = 0.9666 \]
Verify: confirm the endpoint of the complement
Why: At least five is the complement of at most FOUR, not of at most five — the value 5 belongs to the event, so it must be excluded from the complement. Getting this boundary wrong is the commonest error in complement calculations and it shifts the answer by P(5), which here is about 0.036. Writing both value lists before computing is what settles it.
OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, pp. 239-240
Trap
\[ P(x \ge 5) = 1 - P(x \le 5) \]
Subtract the cumulative probability up to the same value
Why: Both expressions mention 5, so they look complementary.
\[ \text{but the value 5 has now been removed from BOTH sides} \]
At least 5 includes the value 5, so its complement must stop at 4.
\[ P(x \ge 5) = 1 - P(x \le 4) \]
Write out both lists of values and check they partition 0 to n
Why: Together they must cover every value exactly once.
The test is that the two events tile the whole range: zero through four and five through twenty account for all twenty-one values with no overlap and no gap. The wrong version leaves the value 5 in neither, which is why it undercounts by P(5). This partition check works for any complement on a discrete distribution and takes one line.
Sorting
Count the terms each way for X following B(20, 0.41).
Sort into buckets
Sort each question by the shorter route.
Item (e) is the borderline case, covering eleven values against ten for its complement, so the two routes are nearly equal. With a calculator's cumulative function the choice hardly matters; by hand it is the difference between a line and a page.
Two truths and a lie
All three concern cumulative probabilities.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. They differ by P(x = 12), which for B(20, 0.41) is about 0.06 — a real slice of any answer near there. Only for chapter 5's continuous distributions, where a single value has probability zero, do the two coincide.
Prediction
Commit before reasoning.
Predict first
For B(20, 0.41), P(x at most 12) is 0.9738. Why is it so close to one?
Correct: Because 12 is about 1.7 sigma above the mean.
Why: The distribution concentrates within a couple of standard deviations of np, so a cut-off 1.7 sigma above the mean leaves very little probability beyond it. That reasoning generalises to every distribution in the book and it is the sanity check to run on any cumulative answer — being more than half of n is not the relevant fact, since a distribution with a mean of 15 would give a very different answer at the same cut-off.
Comparison
Fill the blanks. Two parameters determine every row.
Comparison matrix
| Quantity | Formula | For B(20, 0.41) |
|---|---|---|
| P(x) | n choose x, times p to the x, times q to the n minus x | one term per value of x |
| Mean | mu = np | 8.2 workers |
| Standard deviation | sigma = square root of npq | about 2.20 workers |
| Shape | symmetric when p = 0.5, skewed otherwise | skewed right, since p is below 0.5 |
Everything in the table follows from n and p. That is what a named distribution buys, and it is why the effort of checking the three characteristics is worth making — they are the price of admission to all of this.
Pattern
Six steps, and the first two are where the problems are won or lost.
If the third characteristic fails, do not force the formula. Sampling without replacement wants section 4.5's hypergeometric distribution, and an unfixed number of trials wants section 4.4's geometric.
OpenStax Introductory Business Statistics 2e, §4.2 Binomial Distribution §4.2 Binomial Distribution
Check
Test the characteristics.
Check your understanding
Which of these is NOT a binomial experiment?
Answer: A
Why: Without replacement the composition of the deck changes after each draw, so p is not constant and the trials are not independent — the third characteristic fails.
Check
Apply the formula.
Check your understanding
For X following B(5, 0.25), what is P(x = 4)?
Answer: A
Why: Five orderings, times 0.25 to the fourth, times 0.75 to the first, giving about 0.0146.
Check
The two summaries.
Check your understanding
For X following B(50, 0.70), what are the mean and standard deviation?
Answer: A
Why: The mean is np, which is 35, and sigma is the square root of npq, the square root of 10.5, about 3.24.
Real world
A factory's process produces defective items 3 percent of the time, independently. A quality inspector samples 100 items from a shift and finds 7 defective. The shift supervisor says this is well within normal variation; the plant manager says the process has drifted.
Discussion prompt
Model the count, compute what is normal, and say who is right.
Hint: Check the three characteristics, then use mu and sigma.
Answer:
The model is binomial. A fixed 100 items, each defective or not, independently and at a constant 3 percent — all three characteristics hold, so the count of defectives follows B(100, 0.03). The batch is presumably large enough that sampling barely changes p.
Normal variation is about 3 plus or minus 1.7. The mean is np, which is 3 defectives, and sigma is the square root of 100 times 0.03 times 0.97, about 1.71. So 7 defectives sits about 2.3 standard deviations above the mean.
The supervisor is on weak ground. Section 2.7's rule of thumb puts about two standard deviations as the borderline for far from the mean, and 7 is beyond it. The exact probability of 7 or more from a stable process is about 0.031 — roughly one shift in thirty, which is unusual but not extraordinary.
\[ \mu = 3, \quad \sigma \approx 1.71, \quad P(x \ge 7) \approx 0.031 \]
Neither party is entitled to certainty from one sample, and the honest reading is that a stable process produces this result about three percent of the time. What the calculation converts is an argument into a number: the manager can say how surprising the result is rather than merely that it is surprising. Deciding what level of surprise should trigger action is chapter 9's hypothesis test, and this is precisely the situation it was built for — with the added practical point that a second sample would settle it far better than more argument about the first.
Commit first
Answer, then rate your confidence honestly.
Predict first
Which characteristic does sampling WITHOUT replacement violate?
Correct: The trials are not independent and p does not stay constant.
\[ \text{without replacement} \;\Longrightarrow\; p \text{ changes} \;\Longrightarrow\; \text{not binomial} \]
Why: Removing an item changes the composition of what remains, so the probability of a success on the second draw depends on the first — exactly section 3.2's point about replacement. The number of draws is still fixed and there are still two outcomes, so the first and second characteristics hold; only the third fails. Section 4.5's hypergeometric distribution is the family for this case.
Explain it
They computed P(x = 4) for five flips as 0.25 to the fourth times 0.75, and cannot see what is missing.
Discussion prompt
In three sentences or fewer, show them what their answer actually computes.
Hint: Ask which four flips their calculation assumed were heads.
Answer:
Ask them which flips their answer assumed were the heads — their number is the probability of heads on the first four flips and a tail on the fifth, one specific sequence.
But four heads can also happen with the tail first, or second, or third, or fourth, so there are five such sequences and each has that same probability.
Multiplying by 5, which is 5 choose 4, gives 0.0146 — and that combination factor is exactly what counts the orderings their calculation left out.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the characteristics, run all three explicitly and suspect the third whenever sampling is involved. For the formula, check that the two exponents sum to n. For the summaries, remember npq is the variance and sigma is its root. For the translation, write out the list of qualifying values before computing anything. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top, write the three characteristics of a binomial experiment as a numbered list, and beside each write one experiment that FAILS it and the name of the distribution that handles that failure instead. Below, set up Example 4.13 in four lines: one trial, the success, n and p and q, and the question as an inequality. Then compute mu and sigma from n and p, and write one sentence saying why mu is not a possible count. In the middle, build the complete distribution for X following B(5, 0.25): apply the formula six times, writing each as a combination times two powers, and check that the six probabilities sum to one. Draw the six stems to scale and shade the two that answer P(x more than 3), writing the total beside them. At the bottom, write a two-column table of six phrases — at least, at most, more than, fewer than, exactly, no more than — against the values of x each one admits for n = 20, and mark which two phrases name the same set.
Check your six binomial probabilities by adding them: they must total exactly one, since the values 0 through 5 are everything that can happen. And check the peak sits at x = 1, which is where np = 1.25 lands — if your largest probability is elsewhere, a substitution has gone wrong.
Recap
Five things, and the first is the gate the other four depend on.
| If you see | Then |
|---|---|
| Fixed trials, two outcomes, constant p | Binomial: X follows B(n, p) |
| Sampling without replacement | Not binomial: hypergeometric, section 4.5 |
| Trials until an event occurs | Not binomial: geometric, section 4.4 |
| A request for exactly x | One application of the formula |
| A request for a range | Sum the stems, or take the complement |
| npq reported as sigma | That is the variance; take the square root |
| An answer to sanity-check | Compare with np, and use the two-sigma rule |
Section 4.4 relaxes the first characteristic. When trials continue until the first success rather than for a fixed number, the count of trials has a geometric distribution — and the number of trials becomes the random variable instead of the number of successes.
OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, pp. 237-241 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.