4.3 Binomial Distribution

The first named family of distributions. A binomial experiment has three characteristics: a fixed number of trials n, exactly two outcomes on each trial with success probability p and failure probability q summing to one, and independent trials repeated under identical conditions so that p never changes. The random variable X counting the successes then has a binomial distribution, written X follows B of n and p, and its probabilities come from a formula whose three factors count the orderings and give the probability of one of them. Its mean is np and its standard deviation is the square root of npq, so two numbers determine the whole distribution — which is what makes a named family worth having.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 4.3 Binomial Distribution

Title

Statistics · Chapter 4 — Discrete Random Variables

Binomial Distribution

2. By the end of this lesson you can

Objectives

Five outcomes. The first is a check, and everything else depends on it passing.

OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, pp. 237-241 — the section these objectives are drawn from

3. What you already have

Warm-up

Section 4.2 computed a mean and standard deviation from a distribution table, row by row.

Discussion prompt

Suppose you flip a fair coin ten times and count the heads. Building the whole table means computing eleven probabilities. Is there anything about this situation that might let you skip the table?

Hint: Ask what is the same about every one of the ten flips.

Answer:

Every flip is identical: the same two outcomes, the same probability, and no flip affects any other. So the whole situation is described by just two numbers — ten flips and a success probability of a half.

If two numbers determine the situation, they ought to determine the distribution, and they do. There is a formula for each probability, and formulas for the mean and standard deviation that need no table at all.

That is what a named distribution is: a shape of problem common enough to be worth solving once in general. This section covers the commonest of them, and the price of using it is checking that the situation really has the three characteristics it assumes.

4. Fixed trials, two outcomes, unchanging probability

Concept

There are three characteristics of a binomial experiment. There are a fixed number of trials, denoted n. There are only two possible outcomes, called success and failure, with probabilities p and q which sum to one. And the n trials are independent and repeated under identical conditions, so that p and q remain the same for every trial.

binomial experiment — An experiment with a fixed number n of independent trials, each having exactly two outcomes with the same success probability p on every trial. The random variable X is the number of successes in the n trials, and it takes the values 0 through n.

\[ X \sim B(n, p), \qquad p + q = 1 \]

The word success carries no approval: it names whichever outcome is being counted. The book's Example 4.9 makes the point by defining a success as a student WITHDRAWING from a physics course, because that is the quantity of interest. Choosing which outcome to call a success is free, and it determines p — swapping the labels replaces p by q and counts the other thing, which is often the easier route.

Figure (svg): The three characteristics of a binomial experiment listed as a numbered procedure

The random variable X is then the number of successes in the n trials, and it takes the values 0 through n.

OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 237

5. The three characteristics

Section

Section 1

6. What must be true before the formula applies

Concept

A fixed number of trials, exactly two outcomes per trial with probabilities p and q summing to one, and independent trials under identical conditions. Because the trials are independent, the outcome of one does not help predict another, and p and q remain the same throughout.

Bernoulli trial — Any experiment with characteristics two and three and n equal to one: a single independent trial with two outcomes. A binomial experiment counts the successes in one or more Bernoulli trials, and is named after Jacob Bernoulli, who studied them in the late 1600s.

\[ \text{fixed } n, \quad \text{two outcomes}, \quad \text{independent with constant } p \]

All three must hold and each fails in a recognisable way. The number of trials is not fixed when the experiment runs until something happens, which is section 4.4's geometric distribution. There are more than two outcomes when a die is rolled and the face recorded. And the trials are not independent with constant p when sampling without replacement from a small population, which is section 4.5's hypergeometric distribution. Each failure has its own named remedy later in the chapter.

Figure (svg): The three characteristics of a binomial experiment listed as a numbered procedure

The random variable X is then the number of successes in the n trials, and it takes the values 0 through n.

OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 237 — the three characteristics and the Bernoulli trial

7. Three conditions, all required

Picture it

The checklist to run before reaching for any binomial formula.

Figure (svg): The three characteristics of a binomial experiment listed as a numbered procedure

The random variable X is then the number of successes in the n trials, and it takes the values 0 through n.

It is worth running the list explicitly rather than by feel, because a situation can look binomial and fail the third condition invisibly. Drawing cards without replacement has a fixed n and two outcomes and is not binomial, since p changes on every draw — and the formula would give a confident wrong answer with nothing to flag it.

8. Worked example: checking the three characteristics

Worked example

Example 4.12. The book asks why this is a binomial problem.

\[ 70\% \text{ of } 50 \text{ statistics students do homework on time, independently} \]

Is the number of trials fixed?

Why: Fifty students in the class.

\[ n = 50 \]

Are there exactly two outcomes?

Why: On time, or not on time.

Is the probability constant and are trials independent?

Why: Each student works independently at 0.70.

\[ p = 0.70\text{ throughout} \]

Name the failure and q

Why: Not completing on time.

\[ q = 0.30 \]

Figure (svg): The solution to Worked example checking the three characteristics shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ X \sim B(50, 0.70), \quad x = 0, 1, \ldots, 50 \]

Verify: confirm the values of X and the meaning of a failure

Why: X takes the values 0 through 50, one for every possible number of on-time students, and a failure is a student who does not complete on time. Writing the failure out in words is worth doing because q is defined by it, and a problem that describes only the success rate leaves q to be computed as one minus p. Here 0.30, and p plus q is one as the second characteristic requires.

OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, pp. 238-239

9. Which condition fails?

Sorting

Test each situation against the three characteristics.

Sort into buckets

Sort each experiment.

Binomial: all three hold
flip a coin 20 times and count heads; ask 50 independent students whether they finished homework
Fails: more than two outcomes
roll a die 10 times and record each face
Fails: p changes across trials
draw 5 cards without replacement and count hearts
Fails: n is not fixed
flip a coin until the first head appears
ok
A fixed number of independent trials, two outcomes each, and a success probability that never changes.
two
Recording the face gives six outcomes rather than two. Counting sixes instead would make it binomial.
indep
Sampling without replacement changes the composition, so p differs on each draw.
fixed
The experiment runs until an event occurs, so the number of trials is a random variable rather than a fixed n.

Items (d) and (c) are not dead ends: they are sections 4.4 and 4.5, which are named distributions for exactly those two failures. Recognising which condition breaks tells you which family to reach for instead.

10. Worked example: choosing which outcome is the success

Worked example

Example 4.9. The label is a choice, and it fixes p.

\[ \text{The withdrawal rate from a physics course is } 30\%. \]

Identify what is being counted

Why: The book counts students who withdraw.

Set p accordingly

Why: The withdrawal rate.

\[ p = 0.30 \]

Set q as the complement

Why: Students who stay the whole term.

\[ q = 0.70 \]

Note the alternative labelling

Why: Counting stayers instead swaps the two.

\[ p = 0.70, q = 0.30 \]

Figure (svg): The solution to Worked example choosing which outcome is the success shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ X = \text{the number who withdraw}, \quad p = 0.30 \]

Verify: confirm both labellings answer the same questions

Why: Counting withdrawals with p equal to 0.30 and counting stayers with p equal to 0.70 describe the same class, and any question about one can be rephrased about the other — the probability that at least 5 of 20 withdraw equals the probability that at most 15 stay. The labels are free, and the useful habit is to choose whichever makes the question a simple inequality rather than a complement.

OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 237

11. Trap: applying the formula when p changes

Trap

The trap

\[ \text{draw } 5 \text{ cards from a deck; count the hearts} \]

Treat it as binomial with n = 5 and p = 13/52

Why: There is a fixed number of draws and two outcomes, so two conditions hold.

\[ \text{but } p \text{ changes after each draw} \quad \text{(without replacement)} \]

After a heart is drawn, twelve remain among fifty-one, so the second trial has p equal to twelve fifty-firsts rather than a quarter.

The fix

\[ \text{binomial only WITH replacement; otherwise hypergeometric} \]

Check the third condition explicitly, especially for sampling problems

Why: Fixed n and two outcomes are not enough on their own.

Section 3.2 established that sampling with replacement gives independent trials and sampling without does not, and this is where that distinction pays. The binomial assumes replacement; section 4.5's hypergeometric distribution is the version for sampling without it. Section 1.2's remark applies too: when the population is very large relative to the sample, p barely moves and the binomial is an excellent approximation, which is why survey data are analysed with it.

12. The complement

Fill the middle

The second characteristic.

Fill in the blanks

p + q = 1, \text___ q = 1 - p

Why: Exactly one, because success and failure are complementary events on each trial — section 3.1's complement rule applied to a single trial. A problem stating only p leaves q to be computed, and forgetting to do so is a common source of a wrong exponent in the formula.

13. One of these is false

Two truths and a lie

All three concern the characteristics.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Which outcome is called the success is a free choice
  • C. A Bernoulli trial is a binomial experiment with n equal to one
  • B. A fixed number of trials with two outcomes is enough to make an experiment binomial

Survives elimination: B

Why: The survivor is false because it omits the third condition. Drawing five cards without replacement has a fixed n and two outcomes and is not binomial, because p changes on every draw. All three characteristics are required, and the third is the one that fails silently.

14. Why does independence matter?

Prediction

Commit before reasoning.

Predict first

The binomial formula multiplies p by itself x times. What does that assume?

  • That p is small
  • That the trials are independent, so the joint probability of a sequence is the product of its trial probabilities
  • That n is large
  • That the outcomes are equally likely

Correct: That the trials are independent.

Why: Section 3.3's multiplication rule gives a joint probability as a product only when the events are independent; otherwise a conditional is required. Raising p to the power x is exactly that product for a sequence of x successes, so the third characteristic is what licenses the formula's middle factor. Neither the size of p nor of n has anything to do with it.

15. Notation and setup

Section

Section 2

16. Identify n, p, q and the question

Concept

The notation X follows B of n and p is read as X is a random variable with a binomial distribution, whose parameters are n, the number of trials, and p, the probability of a success on each trial.

the notation B(n, p) — X ~ B(n, p) says X has a binomial distribution with n trials and success probability p. Two numbers specify the distribution completely, which is what makes the family worth naming.

\[ X \sim B(20, 0.41): \quad n = 20, \; p = 0.41, \; q = 0.59, \; x = 0, 1, \ldots, 20 \]

Setting a problem up means answering four questions in order: what is one trial, what counts as a success, how many trials are there, and what is being asked about the count. The book's Example 4.12 walks through exactly these, ending with the translation of 'at least 40' into an inequality — which is the step where most binomial problems are lost, because the arithmetic afterwards is mechanical.

Figure (svg): Two columns translating everyday phrases into the inequalities a binomial probability question needs

On a discrete distribution the endpoint genuinely matters, because P(x = 12) is a real positive number. Getting the inequality wrong shifts the answer by a whole stem.

OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 239 — the notation, and Example 4.13's setup

17. The words and the inequality

Picture it

Six phrases and what each one means for x.

Figure (svg): Two columns translating everyday phrases into the inequalities a binomial probability question needs

On a discrete distribution the endpoint genuinely matters, because P(x = 12) is a real positive number. Getting the inequality wrong shifts the answer by a whole stem.

The endpoint matters here in a way it will not in chapter 5. On a discrete distribution P(x = 12) is a real positive number, so 'at most 12' and 'fewer than 12' differ by that whole stem — for Example 4.13's distribution they differ by about 0.06, which is a substantial slice of the answer.

18. Worked example: setting up Example 4.13

Worked example

Four questions, answered in order.

\[ 41\% \text{ of adult workers have a high school diploma and no further education; } 20 \text{ are selected} \]

What is one trial?

Why: Selecting one adult worker.

What is a success?

Why: Having the diploma and no further education.

\[ p = 0.41 \]

How many trials?

Why: Twenty workers selected.

\[ n = 20 \]

What is asked?

Why: At most 12 of them.

\[ P(x \le 12) \]

Figure (svg): The solution to Worked example setting up Example 4.13 shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ X \sim B(20, 0.41), \quad q = 0.59, \quad \text{find } P(x \le 12) \]

Verify: confirm the setup before computing anything

Why: Every later step depends on these four answers, and all four are read from the problem rather than computed. Note in particular that q is 0.59 rather than being given — the problem states only p, and section 4.1's habit of writing the complement immediately avoids a wrong exponent later. The book gives the answer as 0.9738, which stats-lib reproduces exactly.

OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 239

19. Words to inequality

Translation

Each phrase, and the inequality it names.

Match the pairs

  • l1. at least 40
  • l2. at most 12
  • l3. more than 3
  • l4. exactly 15
  • r1. x is 40 or more
  • r2. x is 12 or fewer
  • r3. x is 4 or more
  • r4. x equals 15 alone

Why: At least and at most include their endpoints; more than and fewer than exclude them. The pair to watch is at least 4 and more than 3, which name the same set — recognising that equivalence often turns an awkward phrase into a familiar one.

20. Worked example: stating the question mathematically

Worked example

Example 4.10. The book asks only for the statement, not the answer.

\[ \text{Win } 55\% \text{ of games, play } 20 \text{ times, want to win } 15 \]

Define the variable

Why: The number of wins.

\[ X =\text{ number of wins} \]

List its values

Why: Anything from none to all twenty.

\[ 0, 1,..., 20 \]

Read off the parameters

Why: Twenty games at 55 percent.

\[ n = 20, p = 0.55, q = 0.45 \]

State the question

Why: Exactly fifteen wins.

\[ P(x = 15) \]

Figure (svg): The solution to Worked example stating the question mathematically shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ X \sim B(20, 0.55), \quad \text{find } P(x = 15) \]

Verify: confirm that 'exactly' means a single value

Why: Fifteen wins exactly is one stem of the distribution, so the answer is a single application of the formula rather than a sum. Contrast Example 4.13's 'at most 12', which covers thirteen stems and needs a cumulative calculation. Recognising which of the two a question asks for, before touching a formula, is what decides whether one term or many are needed.

OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 237

21. Trap: mistranslating the inequality

Trap

The trap

\[ \text{'more than 3 heads' in five flips} \]

Compute P(x is at least 3), including the value 3

Why: More than three and three or more sound alike in a hurry.

\[ P(3) + P(4) + P(5) \quad \text{(the value 3 does not qualify)} \]

For the altered coin P(3) is 0.0879, which is more than five times the correct answer of 0.0156.

The fix

\[ P(x > 3) = P(4) + P(5) = 0.0146 + 0.0010 = 0.0156 \]

Write the list of qualifying values before computing anything

Why: The endpoint is decided once, in the listing.

On a discrete distribution the endpoint carries real probability, so 'more than' and 'at least' give genuinely different answers — here by a factor of six, because the excluded stem is larger than both included ones combined. Section 4.1 made the same point and it becomes more consequential here, since binomial questions are usually phrased in exactly these words.

22. Read off the parameters

Faded example

Forty-one percent of twenty selected workers have the qualification.

Fill in the blanks

n = 20, \quad p = 0.41, \quad q = 0.59

Why: Twenty trials with a success probability of 0.41, so q is 0.59. The problem states only p, and computing q immediately is what prevents a wrong exponent in the formula later.

23. One value, or many?

Sorting

Decide whether the question needs one application of the formula or a sum.

Sort into buckets

Sort each question for X following B(20, 0.41).

One term
P(exactly 8); P(x = 0)
A sum of several terms
P(at most 12); P(at least 15); P(more than 3)
one
The question names a single value of x, so the formula is applied once.
many
The question names a range, so the probabilities of every qualifying value must be added.

For the ranges, the complement is often shorter: P(at least 15) covers six stems while its complement covers fifteen, so the direct route wins there — but P(more than 3) covers seventeen stems against four for its complement, and the complement wins easily.

24. One of these is false

Two truths and a lie

All three concern the setup.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Two numbers, n and p, specify the whole distribution
  • C. X takes the values 0 through n
  • B. At most 12 and fewer than 12 name the same set of values

Survives elimination: B

Why: The survivor is false. At most 12 includes the value 12 and fewer than 12 excludes it, and on a discrete distribution that value carries real probability — about 0.06 for B(20, 0.41), which is a substantial part of any answer near there.

25. The binomial formula

Section

Section 3

26. Count the orderings, times the probability of one

Concept

The probability of exactly x successes in n trials is the number of ways to choose which x trials succeed, multiplied by the probability of x successes and n minus x failures in one particular arrangement.

the binomial formula — P(x) equals n choose x, times p to the power x, times q to the power n minus x. The first factor counts the orderings; the second and third give the probability of any one of them, which is the same for all.

\[ P(x) = \binom{n}{x} p^x q^{n-x} \]

The structure is worth understanding rather than memorising. Any particular sequence with x successes and n minus x failures has probability p to the x times q to the n minus x, by section 3.3's multiplication rule with independence. Every such sequence has the same probability, and they are mutually exclusive, so the total is that common probability times how many sequences there are — which is the combination count.

Figure (svg): The binomial probability formula with each of its three factors labelled

The formula is the multiplication rule for one sequence, multiplied by the number of sequences.

OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, pp. 237-238 — the binomial distribution and Example 4.11

27. Three factors, three jobs

Picture it

What each part of the formula contributes.

Figure (svg): The binomial probability formula with each of its three factors labelled

The formula is the multiplication rule for one sequence, multiplied by the number of sequences.

The two exponents must total n, and checking that is the fastest test of a substitution: x successes and n minus x failures account for every trial exactly once. An expression with exponents summing to anything else has miscounted the trials, and the error is visible without evaluating anything.

28. Worked example: one binomial probability

Worked example

Example 4.11's distribution, at a single value.

\[ X \sim B(5, 0.25); \quad \text{find } P(x = 4) \]

Count the orderings

Why: Which four of the five flips are heads.

\[ C(5, 4) = 5 \]

Probability of four successes

Why: A quarter, four times.

\[ 0.25 ^{4} = 0.00390625 \]

Probability of one failure

Why: Three quarters, once.

\[ 0.75 ^{1} = 0.75 \]

Multiply the three

Why: Five times the product.

\[ 0.014648 \]

Figure (svg): The solution to Worked example one binomial probability shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(4) = \binom{5}{4}(0.25)^4(0.75)^1 \approx 0.0146 \]

Verify: confirm the exponents account for every trial

Why: Four successes and one failure is five trials, matching n, so no trial has been counted twice or left out. The combination count of 5 is also checkable directly: the single failure can be any one of the five flips, so there are five arrangements. Both checks are quick and between them they catch the two commonest substitution errors.

OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 238

29. Substitute into the formula

Faded example

X follows B(5, 0.25). Find P(x = 2).

Fill in the blanks

P(2) = \binom23(0.25)^___}(0.75)^___}

Why: Two successes and three failures, and the exponents total five as they must. The combination count is 10, giving 10 times 0.0625 times 0.421875, which is about 0.2637 — the largest single probability after P(1) in this distribution.

30. Worked example: the whole distribution

Worked example

Example 4.11. The book develops the full density before answering.

\[ X \sim B(5, 0.25); \quad \text{find } P(x > 3) \]

List the qualifying values

Why: More than three means four or five.

\[ x = 4, 5 \]

Compute P(4)

Why: As above.

\[ 0.014648 \]

Compute P(5)

Why: All five heads: a quarter to the fifth.

\[ 0.000977 \]

Add them

Why: Distinct values are exclusive.

\[ 0.015625 \]

Figure (svg): The solution to Worked example the whole distribution shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(x > 3) = 0.0146 + 0.0010 = 0.015625 \]

Verify: confirm the whole distribution sums to one

Why: The six probabilities are 0.2373, 0.3955, 0.2637, 0.0879, 0.0146 and 0.0010, and they total exactly 1.0000 — section 4.1's second condition, which every binomial distribution satisfies automatically. Computing the full set as the book does is more work than the question needs, but it supplies that check and it makes the shape visible: the mass piles at one head, which is where np sits.

OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 238

31. Error analysis: four attempts at P(x = 4) for B(5, 0.25)

Error analysis

Five flips of a coin altered to p = 0.25. Four students compute the probability of exactly four heads.

Annotate

On: \( \begin{aligned} &(1)\; (0.25)^4 = 0.0039 \\ &(2)\; \binom{5}{4}(0.25)^4(0.75)^4 \\ &(3)\; \binom{5}{4}(0.25)^4 = 0.0195 \\ &(4)\; \binom{5}{4}(0.25)^4(0.75)^1 \approx 0.0146 \end{aligned} \)

  • (1) gives the probability of ONE particular sequence — say heads on the first four flips — and forgets that four heads can arise five ways.
  • (2) uses the wrong second exponent. Four successes and four failures is eight trials, but there are only five, so this substitution is impossible on its face.
  • (3) omits the failure factor entirely, as if the fifth flip did not happen. Every trial must appear in one exponent or the other.
  • (4) is correct: the count of orderings, times p to the fourth, times q to the first.

Errors (2) and (3) are both caught by the exponent check — the two powers must sum to n. Error (1) is subtler because the arithmetic is valid; what it computes is the probability of a specific ordering rather than of the count, which is exactly the distinction the combination factor exists to bridge.

32. Which factor is missing?

Discrimination

Each expression omits or mishandles one part of the formula.

Sort into buckets

Sort each faulty expression by what it got wrong.

The ordering count is wrong or missing
(0.25)^4 (0.75)^1, with no combination factor; C(5,1)(0.25)^4(0.75)^1
An exponent is wrong or missing
C(5,4)(0.25)^4, with no failure factor; C(5,4)(0.25)^4(0.75)^4
count
The combination factor is absent or computed for the wrong number, so the arrangements are miscounted.
expo
An exponent is missing or does not pair with the other to total n, so the trials are not all accounted for.

33. One of these is false

Two truths and a lie

All three concern the formula.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The two exponents must sum to n
  • C. Every sequence with x successes has the same probability
  • B. The combination factor is needed only when p is not 0.5

Survives elimination: B

Why: The survivor is false. The combination counts orderings and is needed whatever p is: for a fair coin, four heads in five flips still happens five ways, and omitting the factor understates the probability by a factor of five. The value of p affects the other two factors and never the count.

34. What does the combination count?

Prediction

Commit before reasoning.

Predict first

In the formula, what is n choose x counting?

  • The number of successes
  • The number of different orderings of the trials that produce exactly x successes
  • The number of trials
  • The probability of one ordering

Correct: The number of orderings producing exactly x successes.

Why: Four heads in five flips can happen five ways, depending on which flip is the tail, and each way has the same probability. The formula multiplies that common probability by the count of ways, which is exactly section 4.1's observation that several outcomes can map to one value of a random variable — here made quantitative.

35. The mean and standard deviation

Section

Section 4

36. Two formulas replace the whole table

Concept

The mean of a binomial probability distribution is n times p, and the variance is n times p times q, so the standard deviation is the square root of npq.

binomial mean and standard deviation — For X following B(n, p), the mean mu is np and the variance sigma squared is npq, so sigma is the square root of npq. Both follow from n and p without building a table.

\[ \mu = np, \qquad \sigma^2 = npq, \qquad \sigma = \sqrt{npq} \]

Section 4.2 computed a mean by building a table and summing x times P(x) row by row. For a binomial with fifty trials that would be fifty-one rows. The formula np gives the same answer in one multiplication, and this is the main practical payoff of recognising a named family — the general method still works and is never needed.

Figure (svg): The binomial mean and standard deviation formulas with a worked instance

Section 4.2 built mu and sigma from a table row by row. For a binomial, n and p give both directly.

OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 237 — the binomial mean and variance

37. n and p, and nothing else

Picture it

Both summaries from the two parameters.

Figure (svg): The binomial mean and standard deviation formulas with a worked instance

Section 4.2 built mu and sigma from a table row by row. For a binomial, n and p give both directly.

The formula for mu is worth sanity-checking against intuition: twenty workers with a 41 percent rate should yield about eight, and np gives 8.2. Any binomial mean that is not roughly the trials times the rate is an arithmetic error, and that check takes no time at all.

38. Worked example: how many workers to expect

Worked example

Example 4.13's second question.

\[ X \sim B(20, 0.41): \text{ how many workers are expected to have the qualification?} \]

Identify n and p

Why: From the setup.

\[ n = 20, p = 0.41 \]

Apply the formula

Why: Trials times success probability.

\[ \mu = (20) (0.41) \]

Evaluate

Why: The product.

\[ 8.2 \]

Interpret

Why: The long-run average across many such samples.

\[ \text{about } 8\text{ workers} \]

Figure (svg): The solution to Worked example how many workers to expect shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \mu = np = (20)(0.41) = 8.2 \]

Verify: confirm the answer is not a possible count

Why: Eight point two workers cannot be observed in any single sample of twenty, since the count is a whole number — section 4.2's point about expected values, and it applies to every binomial mean with a non-integer np. What 8.2 describes is the average across very many samples of twenty, which the Law of Large Numbers connects to what would actually be observed.

OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 239

39. Both summaries

Faded example

X follows B(50, 0.70).

Fill in the blanks

\mu = (50)(0.70) = 35, \qquad \sigma = \sqrt3.24 \approx ___

Why: Thirty-five students on average, with a standard deviation of about 3.24. Both come from n and p alone; building the fifty-one row table would give the same answers after a great deal more work.

40. Worked example: the standard deviation

Worked example

The same two parameters, one more step.

\[ X \sim B(20, 0.41): \text{ find } \sigma \]

Compute the variance

Why: n times p times q.

\[ (20) (0.41) (0.59) \]

Evaluate it

Why: The product.

\[ 4.838 \]

Take the square root

Why: Back to counts.

\[ \text{about } 2.20 \]

Sanity-check the size

Why: Against a range of 0 to 20.

Figure (svg): The solution to Worked example the standard deviation shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \sigma = \sqrt{(20)(0.41)(0.59)} \approx 2.20 \]

Verify: confirm sigma is small relative to the range

Why: The count runs from 0 to 20 and sigma is about 2.2, roughly a ninth of the range — which is the usual pattern, since a binomial concentrates its probability within a few standard deviations of np. Section 2.7's rule of thumb applies: nearly all the probability lies within about two sigma of the mean, so counts between roughly 4 and 13 cover almost everything, and the book's P(x at most 12) of 0.9738 is consistent with that.

OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, pp. 237-239

41. Trap: forgetting the square root

Trap

The trap

\[ \sigma = npq = (20)(0.41)(0.59) = 4.838 \]

Report npq as the standard deviation

Why: The formula for the variance is the memorable one.

\[ \text{that is } \sigma^2 \text{, in squared workers} \]

A spread of 4.84 would be more than twice the true one, and it is in the wrong units.

The fix

\[ \sigma = \sqrt{npq} = \sqrt{4.838} \approx 2.20 \]

Write npq as the VARIANCE, then take the root

Why: Section 4.2's distinction, in a new formula.

This is section 2.7's variance-for-standard-deviation slip in binomial clothing, and the same guard works: a standard deviation carries the variable's own units, so if the answer cannot be described as a number of workers it is the variance. Checking the size against the range catches it too, since 4.84 is a quarter of the whole range for a distribution this concentrated.

42. Where is the peak?

Estimation

X follows B(10, 0.2).

Predict first

Around which value does the distribution peak?

  • Around 2
  • Around 5
  • Around 8
  • Around 10

Correct: Around 2.

Why: The mean is np, which is 2, and a binomial distribution peaks at or next to its mean. That gives a free sanity check on any binomial calculation: the largest probabilities should sit near np, so a computed distribution peaking elsewhere signals an error in n or p. The figure comparing three values of p shows the peak tracking np in each case.

43. One of these is false

Two truths and a lie

All three concern the binomial summaries.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The mean np need not be a possible value of X
  • C. The variance is npq and sigma is its square root
  • B. Sigma is largest when p is close to 0 or 1

Survives elimination: B

Why: The survivor is false and is exactly backwards. The product pq is largest at p equal to 0.5 and shrinks toward zero as p approaches either extreme, so sigma is greatest for a fair coin and smallest when the outcome is nearly certain — which makes sense, since a nearly certain outcome varies very little.

44. Which p gives the most spread?

Prediction

Commit before reasoning.

Predict first

For n = 100, which success probability gives the largest standard deviation?

  • p = 0.1
  • p = 0.5
  • p = 0.9
  • They are all equal

Correct: p = 0.5.

Why: Sigma is the square root of npq, and pq is maximised at p equal to 0.5, where it is 0.25. At p equal to 0.1 or 0.9 the product is 0.09, so sigma is about 3 rather than 5. The intuition is that an outcome close to certain has little room to vary, while a fifty-fifty outcome is maximally uncertain — and section 8.3's sample size formulas use exactly this fact.

45. Cumulative probabilities

Section

Section 5

46. Adding stems, and when to take the complement

Concept

A question about a range of values requires adding the probabilities of every value in the range. Because distinct values are mutually exclusive, the sum needs no correction, and the complement is often the shorter route.

cumulative binomial probability — The probability that X is at most some value, found by summing the individual probabilities up to it. A calculator's binomcdf computes it directly; by hand, the complement is often fewer terms.

\[ P(x \le 12) = \sum_{k=0}^{12} \binom{20}{k}(0.41)^k(0.59)^{20-k} \]

The book's Example 4.13 asks for P(x at most 12) with n equal to twenty, which is thirteen separate applications of the formula by hand — and gives the answer as 0.9738 from a calculator. This is where the named family earns its second keep: the probabilities are tabulated and built into every statistical tool, so the work is in setting the problem up correctly rather than in evaluating it.

Figure (svg): The binomial distribution for five flips of a coin with success probability 0.25, peaking at one head and skewed to the right

The book notes the skew: with p = 0.25 rather than 0.5 the distribution leans right, its mass piled at the low counts.

OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, pp. 239-240 — cumulative probabilities and the calculator functions

47. Adding the qualifying stems

Picture it

Example 4.11's distribution with the two stems above three highlighted.

Figure (svg): The binomial distribution for five flips of a coin with success probability 0.25, peaking at one head and skewed to the right

The book notes the skew: with p = 0.25 rather than 0.5 the distribution leans right, its mass piled at the low counts.

Two stems here, and thirteen for Example 4.13. Which route is shorter is worth a moment's thought before starting: 'more than 3' out of five values covers two stems directly and four by complement, so the direct route wins; 'at most 12' out of twenty-one covers thirteen directly and eight by complement, so the complement wins.

48. Worked example: a cumulative probability

Worked example

Example 4.13. The book quotes the calculator's answer.

\[ X \sim B(20, 0.41); \quad \text{find } P(x \le 12) \]

List the qualifying values

Why: Zero through twelve.

Note what a hand computation needs

Why: Thirteen applications of the formula.

Use the cumulative function

Why: binomcdf with n, p and 12.

\[ 0.9738 \]

Sanity-check against the mean

Why: The mean is 8.2, well below 12.

Figure (svg): The solution to Worked example a cumulative probability shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(x \le 12) = 0.9738 \]

Verify: confirm the answer is plausible from mu and sigma

Why: The mean is 8.2 and sigma is about 2.2, so 12 sits about 1.7 standard deviations above the mean. Section 2.7's rule of thumb puts most of a distribution within two sigma, so a probability of about 0.97 below that point is exactly what should be expected. Running that check before trusting a calculator output catches a mistyped n or p, which otherwise produces a confident wrong number.

OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, p. 239

49. Set up the complement

Faded example

X follows B(20, 0.41). You want P(x at least 5).

Fill in the blanks

P(x \ge 5) = 1 - P(x \le 4), \text20 0 \text___ ___

Why: The complement stops at 4, so the two ranges partition the values 0 through 20 with no overlap. Subtracting the cumulative to 5 instead would drop the value 5 from both sides and undercount the answer by its probability.

50. Worked example: choosing the complement

Worked example

The same distribution, a range covering most of it.

\[ X \sim B(20, 0.41); \quad \text{find } P(x \ge 5) \]

Count the direct terms

Why: Five through twenty.

Count the complement's terms

Why: Zero through four.

Choose the complement

Why: Fewer terms.

\[ 1 - P(x \le 4) \]

Evaluate

Why: One minus the cumulative to four.

\[ \text{about } 0.9666 \]

Figure (svg): The solution to Worked example choosing the complement shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(x \ge 5) = 1 - P(x \le 4) \approx 1 - 0.0334 = 0.9666 \]

Verify: confirm the endpoint of the complement

Why: At least five is the complement of at most FOUR, not of at most five — the value 5 belongs to the event, so it must be excluded from the complement. Getting this boundary wrong is the commonest error in complement calculations and it shifts the answer by P(5), which here is about 0.036. Writing both value lists before computing is what settles it.

OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, pp. 239-240

51. Trap: the off-by-one in a complement

Trap

The trap

\[ P(x \ge 5) = 1 - P(x \le 5) \]

Subtract the cumulative probability up to the same value

Why: Both expressions mention 5, so they look complementary.

\[ \text{but the value 5 has now been removed from BOTH sides} \]

At least 5 includes the value 5, so its complement must stop at 4.

The fix

\[ P(x \ge 5) = 1 - P(x \le 4) \]

Write out both lists of values and check they partition 0 to n

Why: Together they must cover every value exactly once.

The test is that the two events tile the whole range: zero through four and five through twenty account for all twenty-one values with no overlap and no gap. The wrong version leaves the value 5 in neither, which is why it undercounts by P(5). This partition check works for any complement on a discrete distribution and takes one line.

52. Direct, or complement?

Sorting

Count the terms each way for X following B(20, 0.41).

Sort into buckets

Sort each question by the shorter route.

Direct: fewer terms
P(x at most 2); P(x at least 18); P(x at most 10)
Complement: fewer terms
P(x at least 3); P(x at most 18)
direct
The range covers ten or fewer of the twenty-one values, so adding them directly is the shorter path.
comp
The range covers more than half the values, so subtracting the smaller complement from one is quicker.

Item (e) is the borderline case, covering eleven values against ten for its complement, so the two routes are nearly equal. With a calculator's cumulative function the choice hardly matters; by hand it is the difference between a line and a page.

53. One of these is false

Two truths and a lie

All three concern cumulative probabilities.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The stems are mutually exclusive, so a range is a plain sum
  • C. The complement of at least 5 is at most 4
  • B. P(x at most 12) and P(x less than 12) are equal

Survives elimination: B

Why: The survivor is false. They differ by P(x = 12), which for B(20, 0.41) is about 0.06 — a real slice of any answer near there. Only for chapter 5's continuous distributions, where a single value has probability zero, do the two coincide.

54. Why is 0.9738 so high?

Prediction

Commit before reasoning.

Predict first

For B(20, 0.41), P(x at most 12) is 0.9738. Why is it so close to one?

  • Because 12 is more than half of 20
  • Because the mean is 8.2 and sigma about 2.2, so 12 lies about 1.7 standard deviations above the mean
  • Because p is less than 0.5
  • Because the distribution is symmetric

Correct: Because 12 is about 1.7 sigma above the mean.

Why: The distribution concentrates within a couple of standard deviations of np, so a cut-off 1.7 sigma above the mean leaves very little probability beyond it. That reasoning generalises to every distribution in the book and it is the sanity check to run on any cumulative answer — being more than half of n is not the relevant fact, since a distribution with a mean of 15 would give a very different answer at the same cut-off.

55. The binomial in one grid

Comparison

Fill the blanks. Two parameters determine every row.

Comparison matrix

QuantityFormulaFor B(20, 0.41)
P(x)n choose x, times p to the x, times q to the n minus xone term per value of x
Meanmu = np8.2 workers
Standard deviationsigma = square root of npqabout 2.20 workers
Shapesymmetric when p = 0.5, skewed otherwiseskewed right, since p is below 0.5

Everything in the table follows from n and p. That is what a named distribution buys, and it is why the effort of checking the three characteristics is worth making — they are the price of admission to all of this.

56. Solving a binomial problem, in order

Pattern

Six steps, and the first two are where the problems are won or lost.

  1. Check the three characteristics: fixed n, two outcomes, independent trials with constant p.
  2. Identify one trial, decide which outcome is the success, and read off n, p and q.
  3. Write the distribution as X follows B(n, p) and list the values 0 through n.
  4. Translate the question into an inequality, writing out the qualifying values explicitly.
  5. For a single value use the formula; for a range, add the qualifying probabilities or take the complement if it has fewer terms.
  6. Check the answer against mu = np and sigma = the square root of npq, using the two-sigma rule of thumb.

If the third characteristic fails, do not force the formula. Sampling without replacement wants section 4.5's hypergeometric distribution, and an unfixed number of trials wants section 4.4's geometric.

OpenStax Introductory Business Statistics 2e, §4.2 Binomial Distribution §4.2 Binomial Distribution

57. Check yourself 1 of 3

Check

Test the characteristics.

Check your understanding

Which of these is NOT a binomial experiment?

  • A. Drawing 5 cards without replacement and counting hearts (correct)
  • B. Flipping a coin 20 times and counting heads
  • C. Asking 50 independent students whether they finished homework
  • D. Rolling a die 10 times and counting sixes

Answer: A

Why: Without replacement the composition of the deck changes after each draw, so p is not constant and the trials are not independent — the third characteristic fails.

Why B tempts people
Fixed n, two outcomes, and a constant p of one half: all three hold.
Why C tempts people
Fixed n, two outcomes, independent trials with a constant p: binomial.
Why D tempts people
Counting sixes gives two outcomes, six or not six, with p a sixth on every independent roll.

58. Check yourself 2 of 3

Check

Apply the formula.

Check your understanding

For X following B(5, 0.25), what is P(x = 4)?

  • A. About 0.0146 (correct)
  • B. About 0.0039
  • C. About 0.0195
  • D. About 0.25

Answer: A

Why: Five orderings, times 0.25 to the fourth, times 0.75 to the first, giving about 0.0146.

Why B tempts people
That is 0.25 to the fourth alone: the probability of one particular ordering, without counting the five ways.
Why C tempts people
That omits the failure factor of 0.75, as though the fifth flip did not occur.
Why D tempts people
That is p itself, which is the probability of a success on one trial rather than of four successes in five.

59. Check yourself 3 of 3

Check

The two summaries.

Check your understanding

For X following B(50, 0.70), what are the mean and standard deviation?

  • A. mu = 35 and sigma about 3.24 (correct)
  • B. mu = 35 and sigma = 10.5
  • C. mu = 0.70 and sigma about 3.24
  • D. mu = 15 and sigma about 3.24

Answer: A

Why: The mean is np, which is 35, and sigma is the square root of npq, the square root of 10.5, about 3.24.

Why B tempts people
That reports npq, the variance, as the standard deviation without taking the root.
Why C tempts people
That reports p as the mean. The mean counts successes, so it must be scaled by n.
Why D tempts people
That computes nq rather than np, counting the failures instead of the successes.

60. Where this shows up outside the textbook

Real world

A factory's process produces defective items 3 percent of the time, independently. A quality inspector samples 100 items from a shift and finds 7 defective. The shift supervisor says this is well within normal variation; the plant manager says the process has drifted.

Discussion prompt

Model the count, compute what is normal, and say who is right.

Hint: Check the three characteristics, then use mu and sigma.

Answer:

The model is binomial. A fixed 100 items, each defective or not, independently and at a constant 3 percent — all three characteristics hold, so the count of defectives follows B(100, 0.03). The batch is presumably large enough that sampling barely changes p.

Normal variation is about 3 plus or minus 1.7. The mean is np, which is 3 defectives, and sigma is the square root of 100 times 0.03 times 0.97, about 1.71. So 7 defectives sits about 2.3 standard deviations above the mean.

The supervisor is on weak ground. Section 2.7's rule of thumb puts about two standard deviations as the borderline for far from the mean, and 7 is beyond it. The exact probability of 7 or more from a stable process is about 0.031 — roughly one shift in thirty, which is unusual but not extraordinary.

\[ \mu = 3, \quad \sigma \approx 1.71, \quad P(x \ge 7) \approx 0.031 \]

Neither party is entitled to certainty from one sample, and the honest reading is that a stable process produces this result about three percent of the time. What the calculation converts is an argument into a number: the manager can say how surprising the result is rather than merely that it is surprising. Deciding what level of surprise should trigger action is chapter 9's hypothesis test, and this is precisely the situation it was built for — with the added practical point that a second sample would settle it far better than more argument about the first.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

Which characteristic does sampling WITHOUT replacement violate?

  • The number of trials is not fixed
  • The trials are not independent and p does not stay constant
  • There are more than two outcomes
  • None; it is still binomial

Correct: The trials are not independent and p does not stay constant.

\[ \text{without replacement} \;\Longrightarrow\; p \text{ changes} \;\Longrightarrow\; \text{not binomial} \]

Why: Removing an item changes the composition of what remains, so the probability of a success on the second draw depends on the first — exactly section 3.2's point about replacement. The number of draws is still fixed and there are still two outcomes, so the first and second characteristics hold; only the third fails. Section 4.5's hypergeometric distribution is the family for this case.

62. Explain it to someone a year behind you

Explain it

They computed P(x = 4) for five flips as 0.25 to the fourth times 0.75, and cannot see what is missing.

Discussion prompt

In three sentences or fewer, show them what their answer actually computes.

Hint: Ask which four flips their calculation assumed were heads.

Answer:

Ask them which flips their answer assumed were the heads — their number is the probability of heads on the first four flips and a tail on the fifth, one specific sequence.

But four heads can also happen with the tail first, or second, or third, or fourth, so there are five such sequences and each has that same probability.

Multiplying by 5, which is 5 choose 4, gives 0.0146 — and that combination factor is exactly what counts the orderings their calculation left out.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Checking the three characteristics before using the formula
  • Substituting correctly into the binomial formula
  • Computing mu and sigma from n and p
  • Translating at least and at most into inequalities

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the characteristics, run all three explicitly and suspect the third whenever sampling is involved. For the formula, check that the two exponents sum to n. For the summaries, remember npq is the variance and sigma is its root. For the translation, write out the list of qualifying values before computing anything. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Fifteen minutes.

Draw it

At the top, write the three characteristics of a binomial experiment as a numbered list, and beside each write one experiment that FAILS it and the name of the distribution that handles that failure instead. Below, set up Example 4.13 in four lines: one trial, the success, n and p and q, and the question as an inequality. Then compute mu and sigma from n and p, and write one sentence saying why mu is not a possible count. In the middle, build the complete distribution for X following B(5, 0.25): apply the formula six times, writing each as a combination times two powers, and check that the six probabilities sum to one. Draw the six stems to scale and shade the two that answer P(x more than 3), writing the total beside them. At the bottom, write a two-column table of six phrases — at least, at most, more than, fewer than, exactly, no more than — against the values of x each one admits for n = 20, and mark which two phrases name the same set.

Check your six binomial probabilities by adding them: they must total exactly one, since the values 0 through 5 are everything that can happen. And check the peak sits at x = 1, which is where np = 1.25 lands — if your largest probability is elsewhere, a substitution has gone wrong.

65. What you can do now

Recap

Five things, and the first is the gate the other four depend on.

If you seeThen
Fixed trials, two outcomes, constant pBinomial: X follows B(n, p)
Sampling without replacementNot binomial: hypergeometric, section 4.5
Trials until an event occursNot binomial: geometric, section 4.4
A request for exactly xOne application of the formula
A request for a rangeSum the stems, or take the complement
npq reported as sigmaThat is the variance; take the square root
An answer to sanity-checkCompare with np, and use the two-sigma rule

Section 4.4 relaxes the first characteristic. When trials continue until the first success rather than for a fixed number, the count of trials has a geometric distribution — and the number of trials becomes the random variable instead of the number of successes.

OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution §4.3, pp. 237-241 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §4.3 Binomial Distribution — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 237-241
  2. OpenStax Introductory Business Statistics 2e, §4.2 Binomial Distribution — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108