4.1 Probability Distribution Function (PDF) for a Discrete Random Variable

The chapter's opening move: instead of asking for the probability of one event, ask how probability is distributed across every value a quantity can take. A random variable attaches a number to each outcome of an experiment, and a discrete probability distribution function lists each value the variable can take beside the probability of taking it. Such a function has exactly two characteristics — each probability lies between zero and one inclusive, and the probabilities sum to one — and a table failing either is not a distribution at all. Because distinct values are mutually exclusive, every range question is a plain sum, and because the variable takes isolated values the distribution is drawn as separated stems rather than a curve.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 4.1 Probability Distribution Function (PDF) for a Discrete Random Variable

Title

Statistics · Chapter 4 — Discrete Random Variables

Probability Distribution Function (PDF) for a Discrete Random Variable

2. By the end of this lesson you can

Objectives

Five outcomes. The second is a two-line check that decides whether anything else is worth doing.

OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229 — the section these objectives are drawn from

3. What you already have

Warm-up

Section 1.3 built relative frequency tables from data, and section 3.1 computed the probability of individual events.

Discussion prompt

Toss a coin ten times and count the heads. What values can that count take, and could you write down the probability of each before tossing anything?

Hint: List the possible counts first, then ask whether they are equally likely.

Answer:

The count can be any whole number from zero to ten, so there are eleven possible values. Nothing in between is possible — 3.5 heads is not an outcome.

And yes, each value's probability can be written down in advance, though not by simple counting: five heads is far more likely than zero, because there are many orderings giving five heads and only one giving none.

A table of every possible value beside its probability is what this chapter calls a probability distribution, and building one shifts the question from 'what is the chance of this event' to 'how is the probability spread across everything that could happen'. That shift is what the rest of the book is built on.

4. A distribution assigns probability to every value at once

Concept

A discrete probability distribution function has two characteristics: each probability is between zero and one inclusive, and the sum of the probabilities is one. Given a random variable X, the notation P(x) means the probability that X takes on the value x.

discrete probability distribution function — A listing of every value a discrete random variable can take, each with its probability. It must satisfy two conditions: every probability lies between zero and one inclusive, and the probabilities sum to one.

\[ 0 \le P(x) \le 1 \quad \text{for every } x, \qquad \sum_x P(x) = 1 \]

The two conditions look modest and they do a great deal of work. The first is inherited from section 3.1, where probabilities were defined to lie between zero and one. The second is the genuinely new demand: it insists that the list is COMPLETE, so that every value the variable could possibly take appears on it. A table whose probabilities sum to 0.9 is not describing a distribution with a gap; it is describing a situation whose possibilities have not all been written down.

Figure (svg): The two conditions for a discrete probability distribution function, each shown alongside a table that violates it

Both conditions are checkable in seconds, and a table failing either is not a probability distribution at all.

OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, p. 228

5. Random variables

Section

Section 1

6. A number attached to every outcome

Concept

A random variable assigns a number to each outcome of an experiment. Where section 3.1 asked whether an event occurred, a random variable records a quantity, and the probability distribution records how likely each of its values is.

random variable — A rule assigning a numerical value to each outcome of an experiment, conventionally written with a capital letter such as X. The lower case x denotes a particular value it can take, and P(x) is the probability that X takes on that value.

\[ X = \text{the number of days per week Nancy attends class} \]

The notation is worth settling now because the whole chapter depends on it. A capital X names the variable — the rule — and a lower case x names one of its values. So P(x = 3) asks for the probability that the rule returns 3, and the four values 0, 1, 2 and 3 are the complete list of what it can return. Section 1.1 introduced variables as characteristics measured on each member; a random variable is the same idea applied to the outcomes of an experiment.

Figure (svg): A diagram showing four outcomes of tossing two coins mapped to the number of heads, illustrating a random variable

A random variable is a rule assigning a number to every outcome, and its distribution says how probability piles up on those numbers.

OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, p. 228 — the notation P(x) for a random variable

7. Outcomes in, numbers out

Picture it

Two coins, four outcomes, three possible values of the count.

Figure (svg): A diagram showing four outcomes of tossing two coins mapped to the number of heads, illustrating a random variable

A random variable is a rule assigning a number to every outcome, and its distribution says how probability piles up on those numbers.

The interesting feature is that the map is many-to-one: HT and TH both send to the value 1, which is why that value carries twice the probability of the others. A random variable deliberately throws information away — it records how many heads, not which coins — and the distribution records what probability survives that compression.

8. Worked example: naming the variable and its values

Worked example

Example 4.2a and b. The description says what to count.

\[ \text{Nancy has classes three days a week and attends some of them.} \]

Find the quantity being counted

Why: How many of her class days she attends.

Write it as a variable of ONE week

Why: A single randomly chosen week.

\[ X =\text{ days attended that week} \]

List the possible values

Why: She has three class days, so she can attend none through all three.

\[ 0, 1, 2, 3 \]

Check nothing between them is possible

Why: Days are counted whole.

Figure (svg): The solution to Worked example naming the variable and its values shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ X \text{ takes the values } 0, 1, 2, 3 \]

Verify: confirm the value list is complete and cannot be extended

Why: Nancy has classes three days a week, so four is impossible and negative attendance is meaningless — the list of four values is exhaustive. Completeness matters because the second condition demands that the probabilities sum to one, and a missing value would make that impossible to satisfy honestly. Listing the values before any probability is written down is the reliable order.

OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, p. 229

9. Which values can it take?

Sorting

List the possible values before doing anything else.

Sort into buckets

Sort each random variable by the values it can take.

Three values
number of heads in two coin tosses; number of practices Jeremiah attends, of two
Four values
number of days Nancy attends, of three class days
Eleven or more
number of heads in ten coin tosses; number of children in a household
three
The count runs from zero to two, giving three possible values.
four
The count runs from zero to three, giving four possible values.
many
The count runs to ten or has no fixed upper bound, so the list is long or unbounded.

Item (d) has no natural maximum, which is allowed: a discrete variable may take infinitely many values, as the geometric and Poisson distributions later in this chapter do. What matters is that the values are isolated, not that they are few.

10. Worked example: two variables on one experiment

Worked example

The same experiment can carry different random variables.

\[ \text{Toss two coins. } X = \text{ number of heads}; \; Y = \text{ number of tails}. \]

List the sample space

Why: Four equally likely outcomes.

Apply X to each

Why: Counting heads.

\[ 2, 1, 1, 0 \]

Apply Y to each

Why: Counting tails.

\[ 0, 1, 1, 2 \]

Note the relationship

Why: Two tosses in all.

\[ Y = 2 - X \]

Figure (svg): The solution to Worked example two variables on one experiment shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ X + Y = 2 \text{ on every outcome} \]

Verify: confirm the two distributions are mirror images

Why: X takes the value 2 with probability one quarter and Y takes the value 0 with the same probability, and so on down: the two distributions are reflections of each other. That is expected, since Y is just 2 minus X. It shows that a random variable is a CHOICE about what to record from an experiment, not a property the experiment has on its own — and that different choices can carry the same information in different arrangements.

OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229

11. Trap: confusing the variable with one of its values

Trap

The trap

\[ X = 3 \]

Write the definition of the variable as an equation

Why: The variable and its value are both being written, so one line looks enough.

\[ \text{that is not a definition; it is one particular case} \]

Writing X equals 3 says the variable took the value 3 on some occasion, which is an event rather than a definition.

The fix

\[ X = \text{the number of days Nancy attends class per week} \]

Define the variable in words, then list its values separately

Why: The capital names the rule; the lower case names an outcome of the rule.

The distinction is exactly section 1.1's between a variable and its data, and the notation preserves it: X is what is being measured and x is what came out. It matters here because P(x = 3) is a sensible expression and P(X) is not — the probability attaches to a value, never to the rule itself.

12. The notation

Fill the middle

What the capital and the lower case each name.

Fill in the blanks

X \textrandom variable ___, \text___ x \text___

Why: The capital names the random variable — the rule assigning a number to each outcome — and the lower case names a particular value. So P(x = 3) is well formed and asks for the probability that the rule returns three.

13. One of these is false

Two truths and a lie

All three concern random variables.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Two different random variables can be defined on the same experiment
  • C. Several outcomes may map to the same value
  • B. A random variable must take a different value for each outcome

Survives elimination: B

Why: The survivor is false. HT and TH both give one head, and that many-to-one mapping is exactly what makes the value 1 twice as likely as the value 2. A random variable deliberately compresses the sample space, and the distribution records what probability accumulates on each value.

14. What does the variable discard?

Prediction

Commit before reasoning.

Predict first

X is the number of heads in two tosses. What information about the experiment does X not record?

  • Nothing; X captures the whole outcome
  • Which coin gave which face, since HT and TH both give X = 1
  • The number of tosses
  • Whether the coins were fair

Correct: Which coin gave which face.

Why: The value 1 is returned by two different outcomes, so knowing X equals 1 does not say which toss was the head. That compression is deliberate and is what makes a distribution manageable: eleven values summarise 1,024 outcomes for ten tosses. The fairness of the coins is an assumption behind the probabilities rather than something the variable records.

15. The two characteristics

Section

Section 2

16. Between zero and one, and summing to one

Concept

A discrete probability distribution function has two characteristics. First, each probability is between zero and one, inclusive. Second, the sum of the probabilities is one.

the two characteristics — Every P(x) satisfies zero is less than or equal to P(x), which is less than or equal to one. And the sum of P(x) over all values x is exactly one. A table violating either is not a probability distribution function.

\[ 0 \le P(x) \le 1 \qquad \text{and} \qquad \sum P(x) = 1 \]

The two conditions fail in different ways and it is worth knowing which is which. A probability outside the range is an arithmetic impossibility — nothing can happen more often than always. A sum that misses one is a bookkeeping failure: either a value has been left off the list, or a probability has been miscomputed. The second is the commoner error and the more instructive one, because it usually points at a possibility the modeller forgot.

Figure (svg): The two conditions for a discrete probability distribution function, each shown alongside a table that violates it

Both conditions are checkable in seconds, and a table failing either is not a probability distribution at all.

OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229 — the two characteristics, and the check on Example 4.1

17. Two conditions, two failures

Picture it

Each condition with a table that breaks it.

Figure (svg): The two conditions for a discrete probability distribution function, each shown alongside a table that violates it

Both conditions are checkable in seconds, and a table failing either is not a probability distribution at all.

The book checks both conditions explicitly on its first example, and the habit is worth copying: before computing anything from a distribution, add the probabilities. It takes seconds and it catches both a missing value and an arithmetic slip, and every calculation later in this chapter assumes the check has passed.

18. Worked example: verifying a distribution

Worked example

Example 4.2c. The book asks what the P(x) column sums to.

\[ P(0) = 0.01, \; P(1) = 0.04, \; P(2) = 0.15, \; P(3) = 0.80 \]

Check each probability's range

Why: All four lie between zero and one.

Add them

Why: One hundredth plus four hundredths plus 0.15 plus 0.80.

\[ 0.01 + 0.04 + 0.15 + 0.80 \]

Evaluate

Why: The total.

\[ 1.00 \]

Conclude

Why: Both conditions are satisfied.

Figure (svg): The solution to Worked example verifying a distribution shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \sum P(x) = 0.01 + 0.04 + 0.15 + 0.80 = 1.00 \]

Verify: confirm the list of values is what makes the sum work

Why: The four values 0 through 3 are every possibility, because Nancy has exactly three class days. Had the value 3 been omitted, the remaining probabilities would total 0.20 and the second condition would fail — correctly, because the table would then be silent about the commonest case. So the sum is not merely an arithmetic check; it is a test that the list of possibilities is complete.

OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, p. 229

19. Valid distribution, or not?

Discrimination

Check both conditions on each.

Sort into buckets

Sort each proposed table.

A valid distribution
0.01, 0.04, 0.15, 0.80; 0.25, 0.25, 0.25, 0.25; 0, 0.5, 0.5
Fails a condition
0.5, 0.3, 0.4; 0.6, 0.5, -0.1
ok
Every probability lies between zero and one inclusive, and the total is exactly one.
no
Either a probability lies outside the permitted range, or the probabilities do not total one.

20. Worked example: finding a missing probability

Worked example

The second condition used to solve rather than to check.

\[ P(0) = 0.01, \; P(1) = 0.04, \; P(2) = 0.15, \; P(3) = \text{?} \]

Write the condition

Why: The four must total one.

\[ \sum\text{ of } P(x) = 1 \]

Add the known probabilities

Why: The first three.

\[ 0.01 + 0.04 + 0.15 = 0.20 \]

Subtract from one

Why: What is left for the last value.

\[ 1 - 0.20 \]

Read the answer

Why: The missing probability.

\[ 0.80 \]

Figure (svg): The solution to Worked example finding a missing probability shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(3) = 1 - (0.01 + 0.04 + 0.15) = 0.80 \]

Verify: confirm the answer is a legal probability

Why: The result 0.80 lies between zero and one, so the first condition is satisfied too. That check matters because the method can produce an illegal answer: if the known probabilities had summed to more than one, the subtraction would give a negative number, and that would be evidence of an error in the given values rather than a strange distribution. Solving with the second condition should always be followed by testing against the first.

OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229

21. Error analysis: four proposed distributions

Error analysis

A student proposes four tables for a variable taking the values 0, 1, 2 and 3.

Annotate

On: \( \begin{aligned} &(1)\; 0.2, \;0.4, \;0.5, \;0.1 \\ &(2)\; 0.3, \;0.4, \;-0.1, \;0.4 \\ &(3)\; 0.2, \;0.3, \;0.2, \;0.1 \\ &(4)\; 0.1, \;0.2, \;0.3, \;0.4 \end{aligned} \)

  • (1) has every entry in range but they total 1.2. Some probability has been double counted, or a value's probability has been overstated.
  • (2) contains a negative entry, which fails the first condition outright — no event can occur less often than never — and its total is a coincidence rather than a rescue.
  • (3) has every entry in range but totals 0.8. Either a fifth value is possible and was omitted, or one of the four is understated.
  • (4) satisfies both conditions: every entry lies in range and the four total exactly one.

Errors (1) and (3) are the same failure in opposite directions and are far commoner than (2), because nothing about the individual numbers looks wrong. That is precisely why the sum must be computed rather than eyeballed.

22. Solve with the second condition

Faded example

A variable takes values 0, 1, 2 with P(0) = 0.35 and P(1) = 0.45.

Fill in the blanks

P(2) = 1 - (0.35 + 0.45) = 1 - 0.8 = 0.2

Why: The three must total one, so the missing probability is 0.20. Check it against the first condition as well: 0.20 lies between zero and one, so the answer is legal. Had the two known probabilities summed past one, the method would have returned a negative number and revealed an error in the data.

23. One of these is false

Two truths and a lie

All three concern the two conditions.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. A probability of exactly zero is permitted
  • C. A total below one usually means a value was left off the list
  • B. A table whose probabilities are all in range is a valid distribution

Survives elimination: B

Why: The survivor is false because it checks only the first condition. The table 0.2, 0.3, 0.2, 0.1 has every entry in range and totals 0.8, so it is not a distribution. Both conditions must hold, and the second is the one that fails most often and least visibly.

24. What does a shortfall mean?

Prediction

Commit before reasoning.

Predict first

A proposed distribution's probabilities total 0.85. What is the most likely explanation?

  • The distribution is valid but incomplete in a harmless way
  • A possible value has been left off the list, or a probability is understated
  • The variable is continuous rather than discrete
  • The probabilities should be rescaled to sum to one

Correct: A value was left off, or a probability is wrong.

Why: Probability must be fully accounted for across everything that can happen, so a shortfall of 0.15 means something that can happen is not on the table. The right response is to find the missing possibility, not to rescale — rescaling would silently redistribute the missing 0.15 across the listed values and assert probabilities nobody computed.

25. Building a PDF table

Section

Section 3

26. From a description to a table

Concept

A probability distribution table, or PDF table, has two columns: one listing every value the random variable can take, and one giving the probability of each. Building it from a verbal description means identifying the variable, listing its values, and reading a probability for each.

PDF table — A two-column table with the values of the random variable in the first column and their probabilities in the second. The book's Example 4.1 and 4.2 both construct one, and the P(x) column must sum to one.

\[ \begin{array}{c|c} x & P(x) \\ \hline 0 & 0.01 \\ 1 & 0.04 \\ 2 & 0.15 \\ 3 & 0.80 \end{array} \]

The table is section 1.3's relative frequency table with theoretical probabilities in place of observed proportions, and the parallel goes further than the shape. There the relative frequency column had to sum to one because every observation fell in exactly one row; here the probability column must sum to one because every outcome produces exactly one value. The same partition argument underlies both.

Figure (svg): A probability distribution table for Nancy's class attendance with a total row summing to one

This is section 1.3's relative frequency table with probabilities in place of observed proportions.

OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, p. 229 — constructing a PDF table for Example 4.2

27. Nancy's four weeks in a hundred

Picture it

The distribution as a table, with the running total made explicit.

Figure (svg): A probability distribution table for Nancy's class attendance with a total row summing to one

This is section 1.3's relative frequency table with probabilities in place of observed proportions.

The running total column is not part of the standard table but it is worth writing while learning, because it turns the second condition into something checked as you go rather than at the end. It is also section 1.3's cumulative relative frequency in a new setting, and it will reappear in chapter 5 as the cumulative distribution function.

28. Worked example: building Nancy's table

Worked example

Example 4.2c. The percentages in the description become the probabilities.

\[ \text{three days } 80\%, \text{ two days } 15\%, \text{ one day } 4\%, \text{ no days } 1\% \]

List the values in order

Why: From none to all three.

\[ 0, 1, 2, 3 \]

Attach each stated percentage

Why: Reading the description onto the right row.

\[ 0.01, 0.04, 0.15, 0.80 \]

Convert percentages to decimals

Why: Dividing each by a hundred.

Check the total

Why: The second condition.

\[ 1.00 \]

Figure (svg): The solution to Worked example building Nancy's table shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(0) = 0.01, \; P(1) = 0.04, \; P(2) = 0.15, \; P(3) = 0.80 \]

Verify: confirm the description was read onto the right rows

Why: The description gives the percentages in decreasing order of days — three, two, one, none — while the table lists values in increasing order, so the four numbers must be reversed as they are transferred. That reversal is where errors occur, and the check is that the largest probability, 0.80, sits against three days, which is what the description says she does most often. Reading the finished table back into a sentence catches a transposition immediately.

OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, p. 229

29. Complete the table

Faded example

Jeremiah attends both of two practices 90 percent of the time, one 8 percent, neither 2 percent.

Fill in the blanks

P(0) = 0.02, \quad P(1) = 0.08, \quad P(2) = 0.90

Why: The description runs from most to fewest and the table runs from fewest to most, so the percentages reverse as they are transferred. The three total 1.00, and the largest probability sits against two practices, which matches what the description says he usually does.

30. Worked example: a table from equally likely outcomes

Worked example

When the outcomes are equally likely, section 3.1's counting builds the table.

\[ X = \text{the number of heads in two tosses of a fair coin} \]

List the sample space

Why: Four equally likely outcomes.

Apply the variable to each

Why: Counting heads.

\[ 2, 1, 1, 0 \]

Count the outcomes for each value

Why: One, two and one respectively.

\[ 1, 2, 1 \]

Divide by four

Why: Each count over the sample space size.

\[ 0.25, 0.50, 0.25 \]

Figure (svg): The solution to Worked example a table from equally likely outcomes shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(0) = \frac{1}{4}, \; P(1) = \frac{2}{4}, \; P(2) = \frac{1}{4} \]

Verify: confirm the total and the shape

Why: The three probabilities sum to one, as required. And the middle value carries twice the probability of either end, because two of the four outcomes produce it — which is the many-to-one mapping the first idea drew. That shape, higher in the middle and lower at the extremes, is the beginning of the binomial distribution that section 4.3 will build for any number of tosses.

OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229

31. Trap: listing outcomes instead of values

Trap

The trap

\[ x: \; HH, \;HT, \;TH, \;TT \]

Put the sample space in the x column

Why: Those are the four things that can happen, so they look like the values.

\[ \text{but a random variable takes NUMBERS} \]

The x column holds the values of X, which are the counts 0, 1 and 2 — three rows, not four.

The fix

\[ x: \; 0, \;1, \;2 \quad \text{with } P(x): \; 0.25, \;0.50, \;0.25 \]

Apply the variable to the outcomes first, then list the distinct results

Why: Outcomes are what happens; values are what X reports about it.

The compression from four outcomes to three values is the whole point of a random variable, and it is why the middle probability is 0.50 rather than 0.25. A table with one row per outcome would have four equal probabilities and would describe the sample space rather than the distribution of X — which is a different and less useful object once the sample space is large.

32. Order the steps

Ranking

Building a PDF table from a verbal description.

Put in order

  1. identify the quantity being counted and name it X
  2. list every value X can take, in order
  3. attach a probability to each value
  4. check that the probabilities sum to one

Why: Naming the variable first prevents the commonest error of tabulating outcomes instead of values. Listing the values before attaching probabilities makes it obvious if one is missing — and the final check is what catches an omission if the listing missed one anyway.

33. One of these is false

Two truths and a lie

All three concern building the table.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The x column lists numbers, not outcomes
  • C. A PDF table is the theoretical counterpart of a relative frequency table
  • B. The table must have one row per outcome of the experiment

Survives elimination: B

Why: The survivor is false. Two coins have four outcomes and three values, and ten coins have 1,024 outcomes and eleven values. That compression is what makes a distribution useful, and it is why the number of rows is the number of distinct values rather than the size of the sample space.

34. How many rows?

Prediction

Commit before reasoning.

Predict first

X is the number of heads in ten tosses of a coin. How many rows does its PDF table have?

  • Two
  • Ten
  • Eleven
  • 1,024

Correct: Eleven.

Why: The count can be any whole number from zero to ten inclusive, which is eleven values. The 1,024 is the number of outcomes in the sample space, and the variable compresses those onto eleven values — which is precisely why the distribution is worth tabulating and the sample space is not.

35. Reading a distribution

Section

Section 4

36. A range is a sum, because the values are exclusive

Concept

Once a distribution is tabulated, the probability of any range of values is found by adding the probabilities of the values in that range. No correction is needed, because a random variable takes exactly one value on any occasion, so distinct values are mutually exclusive.

reading a range — The probability that X falls in a set of values is the sum of the probabilities of those values. Distinct values of a random variable are mutually exclusive, so section 3.3's addition rule applies with its correction term equal to zero.

\[ P(X \le 2) = P(0) + P(1) + P(2) \]

This is section 3.3's addition rule with the subtraction guaranteed to vanish. X cannot be both 1 and 2 on the same occasion, so the events are mutually exclusive by construction and the plain sum is always correct. It is one of the quiet conveniences of working with a distribution rather than with arbitrary events, and it holds for every distribution in this chapter and the next.

Figure (svg): A discrete probability distribution with the stems for zero, one and two days highlighted and their probabilities summed to 0.20

Every range question on a discrete distribution is an addition, because distinct values cannot both occur.

OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229 — reading probabilities from a distribution

37. At most two days

Picture it

Three stems highlighted, and their heights added.

Figure (svg): A discrete probability distribution with the stems for zero, one and two days highlighted and their probabilities summed to 0.20

Every range question on a discrete distribution is an addition, because distinct values cannot both occur.

The answer of 0.20 is small because almost all of Nancy's probability sits on the single value 3. That is worth noticing as a habit: before computing a range, look at where the distribution's mass is, because a range excluding the tall stem will be small however many values it contains.

38. Worked example: a range from the table

Worked example

Adding the stems that qualify.

\[ P(X \le 2) \text{ for Nancy's attendance} \]

Identify the qualifying values

Why: At most two means 0, 1 or 2.

Read their probabilities

Why: From the table.

\[ 0.01, 0.04, 0.15 \]

Add them

Why: The values are mutually exclusive.

\[ 0.20 \]

Sanity-check against the remaining value

Why: One minus P(3).

\[ 1 - 0.80 = 0.20 \]

Figure (svg): The solution to Worked example a range from the table shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(X \le 2) = 0.01 + 0.04 + 0.15 = 0.20 \]

Verify: confirm the complement gives the same answer

Why: At most two and exactly three are complementary events, since those four values exhaust the possibilities, so one minus 0.80 must give the same 0.20 — and it does. On a distribution with many values the complement is often far the shorter route, and section 3.1's advice about 'at least one' events applies here: whenever a range covers most of the values, compute the complement instead.

OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, p. 229

39. Which values qualify?

Sorting

Nancy's variable takes the values 0, 1, 2, 3.

Sort into buckets

Sort each range by how many values it covers.

One value
X is exactly 3
Two values
X is less than 2; X is more than 1
Three values
X is at most 2; X is at least 1
one
The range names a single value exactly, so its probability is one table entry.
two
Two of the four values satisfy the description.
three
Three of the four values satisfy it, so the complement is a single value and is usually the shorter route.

Items (a) and (b) differ by exactly one value and by 0.15 of probability, which on this distribution is three times the whole answer to (b). Listing the qualifying values in writing before adding is what prevents the slip.

40. Worked example: which route is shorter

Worked example

The same distribution, a range covering most of it.

\[ P(X \ge 1) \text{ for Nancy's attendance} \]

Count the qualifying values

Why: One, two and three.

Count the complement's values

Why: Only zero.

Choose the complement

Why: Fewer terms.

\[ 1 - P(0) \]

Compute

Why: One minus one hundredth.

\[ 0.99 \]

Figure (svg): The solution to Worked example which route is shorter shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(X \ge 1) = 1 - P(0) = 1 - 0.01 = 0.99 \]

Verify: confirm the direct route agrees

Why: Adding 0.04, 0.15 and 0.80 gives 0.99, the same answer by three additions instead of one subtraction. On four values the saving is trivial; on the eleven values of ten coin tosses, or the unbounded list of a Poisson variable in section 4.6, it is the difference between a line and a page. The rule of thumb is to count the terms each way and take the shorter.

OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229

41. Trap: dropping or including the endpoint

Trap

The trap

\[ P(X < 2) = 0.01 + 0.04 + 0.15 \]

Include the value 2 in a strictly-less-than range

Why: At most two and less than two look similar in a hurry.

\[ = 0.20 \quad \text{(that is } P(X \le 2)\text{)} \]

Strictly less than two means the values 0 and 1 only, giving 0.05.

The fix

\[ P(X < 2) = 0.01 + 0.04 = 0.05 \]

Write out the list of qualifying values before adding anything

Why: The endpoint is decided once, in the listing, rather than repeatedly in the arithmetic.

On a discrete distribution the endpoint genuinely matters, because P(X = 2) is a real positive number — here 0.15, which is three times the whole of P(X < 2). That is a sharp contrast with chapter 5's continuous distributions, where any single value has probability zero and 'less than' and 'at most' give identical answers. Carrying the discrete habit into chapter 5 is harmless; carrying the continuous habit back here is not.

42. Take the shorter route

Faded example

Nancy's distribution: 0.01, 0.04, 0.15, 0.80 for the values 0 to 3.

Fill in the blanks

P(X \ge 2) = 1 - P(0) - P(1) = 1 - 0.01 - 0.04 = 0.95

Why: Two subtractions rather than two additions makes little difference here, but the direct route of 0.15 plus 0.80 gives the same 0.95, which is the check. Whichever route is used, the two must agree, and that agreement is the cheapest verification available on a range question.

43. One of these is false

Two truths and a lie

All three concern reading ranges.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Distinct values of a random variable are mutually exclusive
  • C. The complement is often the shorter route for a wide range
  • B. P(X < 2) and P(X is at most 2) are the same

Survives elimination: B

Why: The survivor is false for a discrete variable. The two differ by P(X = 2), which is 0.15 on Nancy's distribution — larger than the whole of P(X < 2). Only for the continuous variables of chapter 5, where any single value has probability zero, do the two coincide.

44. Why no subtraction?

Prediction

Commit before reasoning.

Predict first

Section 3.3's addition rule subtracts the intersection. Why is no subtraction needed when adding P(1) and P(2)?

  • Because the probabilities are small
  • Because X cannot take two values at once, so the intersection is empty
  • Because the values are consecutive
  • Because the distribution sums to one

Correct: Because X cannot take two values at once.

Why: A random variable returns exactly one number on each occasion, so the events X equals 1 and X equals 2 are mutually exclusive by construction and their intersection has probability zero. The addition rule still applies; its correction term is simply guaranteed to vanish, which is why range questions on a distribution are always plain sums.

45. Why the stems are separated

Section

Section 5

46. Isolated values, and nothing in between

Concept

A discrete random variable takes isolated values, so its distribution is drawn as separated stems with a dot on top rather than as a joined curve. The height of each stem is the probability of that value, and the space between stems represents values the variable cannot take.

drawing a discrete distribution — Separated vertical stems, one per value, with height equal to that value's probability. The stems are not joined because the variable takes no values between them, and the height is the probability itself rather than an area.

\[ \text{height of the stem at } x \;=\; P(x) \]

This is section 2.1's line-graph rule in a new setting. There, joining points across an unordered axis asserted a trend that relabelling could reverse; here, joining stems would assert that the variable can take values between them. For Nancy's attendance there is no such thing as 2.4 days, so a line drawn from the stem at 2 to the stem at 3 passes through territory that does not exist.

Figure (svg): Two columns contrasting a discrete distribution drawn as separated stems with a continuous curve

Joining the stems of a discrete distribution would assert values the variable cannot take. Chapter 5's continuous distributions are drawn as curves for exactly the opposite reason.

OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229 — the discrete distribution and its values

47. Stems against a curve

Picture it

What each convention asserts about the variable.

Figure (svg): Two columns contrasting a discrete distribution drawn as separated stems with a continuous curve

Joining the stems of a discrete distribution would assert values the variable cannot take. Chapter 5's continuous distributions are drawn as curves for exactly the opposite reason.

The row worth carrying into chapter 5 is the third. On a discrete distribution the HEIGHT is the probability, so a stem of 0.80 means that value occurs eighty percent of the time. On a continuous distribution the height is a density and the probability is an AREA, so a single point has probability zero. Reading a continuous curve as if its height were a probability is the commonest error in chapter 5, and it starts here with the difference in how the two are drawn.

48. Worked example: why 2.4 days is not a value

Worked example

The variable's definition decides which numbers can appear.

\[ X = \text{the number of days Nancy attends class per week} \]

Ask what the variable counts

Why: Whole days of attendance.

Ask whether a fraction is possible

Why: She attends a day or she does not.

Conclude about the values

Why: Only whole numbers from 0 to 3.

Draw accordingly

Why: Separated stems, unjoined.

Figure (svg): The solution to Worked example why 2.4 days is not a value shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ x \in \{0, 1, 2, 3\}, \text{ with nothing between} \]

Verify: confirm the same test against a continuous quantity

Why: The number of HOURS Nancy spends in class would be different: 2.4 hours is perfectly possible, and so is 2.41, so the values fill a range rather than sitting apart. That variable would be continuous and would belong to chapter 5, drawn as a curve. The same situation supplies both kinds of variable depending on what is recorded, which is section 1.2's discrete-against-continuous distinction reappearing exactly where it is needed.

OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, p. 228

49. Stems or a curve?

Discrimination

Ask whether the variable can take values between the ones you would plot.

Sort into buckets

Sort each random variable.

Discrete: separated stems
number of heads in ten tosses; number of children in a household; number of defective items in a batch
Continuous: a curve
time to run a mile, in minutes; weight of a package in kilograms
stems
The variable counts whole things, so its values are isolated and nothing lies between them.
curve
The variable measures, so its values fill a range and a finer instrument would give more decimal places.

50. Worked example: reading a height

Worked example

On a discrete plot the height is the answer.

\[ \text{the stem at } x = 3 \text{ has height } 0.80 \]

Read the height

Why: Eight tenths.

\[ 0.80 \]

State what it means

Why: The probability that X equals exactly 3.

\[ P(X = 3) = 0.80 \]

Note what it does not mean

Why: It is not an area and needs no width.

Check against the total

Why: All four heights sum to one.

Figure (svg): The solution to Worked example reading a height shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(3) = 0.80 \text{, the height of the stem} \]

Verify: confirm the heights must sum to one

Why: The four stem heights are 0.01, 0.04, 0.15 and 0.80, totalling exactly one — which is the second condition drawn. On a discrete plot that total is a property of the HEIGHTS, and it is worth noticing now because on chapter 5's continuous plots the corresponding statement is that the total AREA under the curve is one. The same requirement, expressed differently because the drawing convention differs.

OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229

51. Trap: joining the tops of the stems

Trap

The trap

\[ \text{a smooth line drawn through the four dots} \]

Join the stems to show the shape of the distribution

Why: A joined line makes the pattern easier to see.

\[ \text{the line passes through } x = 2.4, \text{ which is not a possible value} \]

Reading a height off that line at 2.4 would give a probability for something that cannot occur.

The fix

\[ \text{separated stems, each with a dot on top} \]

Leave the gaps, because the gaps are information

Why: The empty space says the variable takes no values there.

Section 2.2's frequency polygon joined its points legitimately because the underlying variable was continuous and had been grouped into classes. Here there is no underlying continuum to interpolate across, so the join would be a claim rather than a convenience. The same principle governed section 2.1's rule against line graphs over categories, and it comes down to the same question: does the space between the plotted points contain anything?

52. What the height means

Fill the middle

On a discrete probability plot.

Fill in the blanks

\textprobability x \text___ ___ \text___ X \text___ x

Why: The height is the probability itself, so a stem of 0.80 means that value occurs eighty percent of the time. Chapter 5's continuous distributions reverse this: there the height is a density and the probability is the area beneath the curve.

53. One of these is false

Two truths and a lie

All three concern drawing a discrete distribution.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The stem heights sum to one
  • C. The gaps between stems represent values the variable cannot take
  • B. A discrete distribution should be drawn with touching bars, like a histogram

Survives elimination: B

Why: The survivor is false for a probability distribution. Section 2.2 did draw discrete DATA as a histogram with touching bars centred on each value, which is a display of observed frequencies over a numeric axis. A probability distribution over isolated values is conventionally drawn as separated stems, precisely to keep the isolation visible.

54. What would a curve assert?

Prediction

Commit before reasoning.

Predict first

Someone draws a smooth curve through the tops of Nancy's four stems. What does the curve claim at x = 2.4?

  • Nothing; it is only a visual aid
  • That there is a probability of Nancy attending 2.4 days, which is not a possible value
  • That the distribution is continuous
  • That the probabilities do not sum to one

Correct: That there is a probability of attending 2.4 days.

Why: A curve has a height at every point of its domain, so drawing one asserts a value at 2.4 — and Nancy cannot attend 2.4 days. Section 2.2's frequency polygon was allowed to join its points because the underlying variable was continuous and merely grouped; here there is no continuum underneath, so the interpolation has nothing to interpolate.

55. Discrete distributions against what came before

Comparison

Fill the blanks. The left column is chapter 3 and the right is this chapter.

Comparison matrix

QuestionChapter 3 answeredChapter 4 answers
What is being asked about?one event: did A happen?a quantity: what value did X take?
What is the answer?a single probabilitya probability for every possible value
What must sum to one?an event and its complementthe probabilities of all the values
How is a range handled?the addition rule, minus the intersectiona plain sum: the values are exclusive

The shift is from asking about one event to describing a whole quantity, and the payoff arrives in the next five sections: some quantities recur so often that their distributions have names and formulas, so the table never has to be built by hand again.

56. Building and using a discrete distribution, in order

Pattern

Six steps, and the last is a check that should never be skipped.

  1. Identify the quantity being counted and define X in a full sentence about ONE trial.
  2. List every value X can take, in increasing order, and confirm the list is exhaustive.
  3. Attach a probability to each value, from the description or by counting equally likely outcomes.
  4. Check the first condition: every probability lies between zero and one inclusive.
  5. Check the second condition: the probabilities sum to exactly one.
  6. Answer range questions by adding the qualifying probabilities, or by subtracting the complement when that is shorter.

If the probabilities fall short of one, a possible value has been omitted rather than a probability being slightly wrong. Go back to step two before adjusting anything.

OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229

57. Check yourself 1 of 3

Check

Test both conditions.

Check your understanding

Is the table with probabilities 0.2, 0.3, 0.4 and 0.2 a valid discrete PDF?

  • A. No: the probabilities sum to 1.1 (correct)
  • B. Yes: every probability is between 0 and 1
  • C. No: a probability is negative
  • D. Yes: there are four values

Answer: A

Why: Every entry lies in range, so the first condition holds, but the four total 1.1 rather than 1, so the second fails and the table is not a distribution.

Why B tempts people
That checks only the first condition. Both must hold.
Why C tempts people
No entry is negative here; the failure is in the total.
Why D tempts people
The number of values is not one of the conditions.

58. Check yourself 2 of 3

Check

Read a range.

Check your understanding

Nancy's distribution is 0.01, 0.04, 0.15, 0.80 for x = 0, 1, 2, 3. What is P(X < 2)?

  • A. 0.05 (correct)
  • B. 0.20
  • C. 0.15
  • D. 0.95

Answer: A

Why: Strictly less than 2 means the values 0 and 1, giving 0.01 plus 0.04, which is 0.05.

Why B tempts people
That is P(X is at most 2), which includes the value 2 and its probability of 0.15.
Why C tempts people
That is P(X = 2) alone, which the range excludes.
Why D tempts people
That is P(X is at least 2), the complement of the correct answer.

59. Check yourself 3 of 3

Check

Count the values.

Check your understanding

X is the number of heads in three tosses of a coin. How many rows does its PDF table have?

  • A. Four (correct)
  • B. Three
  • C. Eight
  • D. Six

Answer: A

Why: The count can be 0, 1, 2 or 3, which is four values, so the table has four rows.

Why B tempts people
That would be the count if zero heads were impossible, which it is not.
Why C tempts people
Eight is the number of outcomes in the sample space, which the variable compresses onto four values.
Why D tempts people
Six corresponds to nothing here; three tosses give four possible counts.

60. Where this shows up outside the textbook

Real world

A support team models the number of tickets arriving in an hour and publishes a distribution: 0 tickets with probability 0.10, 1 with 0.25, 2 with 0.30, 3 with 0.20, and 4 with 0.10. A manager uses it to plan staffing for 'up to three tickets an hour'.

Discussion prompt

Check the distribution, say what the manager's plan actually covers, and identify what the model is assuming.

Hint: Add the probabilities first, then read the range the manager described.

Answer:

The distribution does not sum to one. The five probabilities total 0.95, so five percent of the probability is unaccounted for — and since each listed value is in range, the failure is the second condition. Some possibility has been left off the table.

What is missing is the tail. Five or more tickets in an hour is clearly possible and simply was not listed, which is the commonest way this condition fails: the modeller enumerated the ordinary cases and stopped. The missing 0.05 belongs to those busy hours.

The manager's plan covers less than it appears to. Up to three tickets is 0.10 plus 0.25 plus 0.30 plus 0.20, which is 0.85 — so the plan fails about one hour in seven, not one in twenty as the incomplete table might suggest.

\[ \sum P(x) = 0.95 \ne 1 \;\Longrightarrow\; \text{a value is missing} \]

The general lesson is that the sum-to-one condition is a completeness test rather than an arithmetic formality, and a shortfall almost always points at an omitted tail. It matters most for exactly the planning question being asked here, because the omitted values are the extreme ones — the busy hours a staffing model exists to survive. Section 4.6's Poisson distribution is the standard model for counts of this kind, and it assigns a probability to every non-negative whole number precisely so that no tail can be forgotten.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

A proposed discrete distribution has all its probabilities between 0 and 1 but they sum to 0.9. What is wrong?

  • Nothing; that is within rounding tolerance
  • The second condition fails, which usually means a possible value has been left off the list
  • One of the probabilities must be negative
  • The variable must be continuous

Correct: The second condition fails, and a value is probably missing.

\[ \sum P(x) = 1 \text{ is a COMPLETENESS condition, not an arithmetic one} \]

Why: Probability must be fully accounted for across everything the variable can do, so a shortfall of 0.1 means something possible is not on the table. It is not a rounding matter and it is not fixed by rescaling, which would silently redistribute the missing probability across the listed values. The right move is to go back and find the omitted possibility, which is usually in the tail.

62. Explain it to someone a year behind you

Explain it

They have made a table with one row per outcome of tossing two coins, and four equal probabilities of 0.25.

Discussion prompt

In three sentences or fewer, show them what a distribution of the count would look like instead.

Hint: Ask them what value X takes on HT and on TH.

Answer:

Ask them what the number of heads is for HT and for TH: both give one, so those two outcomes collapse onto a single value of the variable.

So the table has three rows rather than four — the values 0, 1 and 2 — and the middle one carries 0.50 because two of the four outcomes produce it.

Their table describes the sample space, which is fine for two coins and hopeless for ten, where 1,024 outcomes compress onto just eleven values.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Defining the random variable and listing its values
  • Testing a table against both conditions
  • Building a distribution by counting equally likely outcomes
  • Handling endpoints in a range question

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the definition, write X as a full sentence about one trial and list the values before any probability. For the conditions, check the range of each entry and then add them. For counting, apply the variable to every outcome first and then group. For endpoints, write the qualifying values out before adding, since P(X = 2) is a real positive number here. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Fifteen minutes.

Draw it

At the top, write the two conditions for a discrete PDF in a box, and beside each write a four-row table that violates it and one word saying what has gone wrong. Below, build Nancy's distribution from this description: she has classes three days a week and attends all three 80 percent of weeks, two days 15 percent, one day 4 percent and none 1 percent. Write the variable as a full sentence, list the values, build the two-column table, add a running-total column, and check the total. Draw the distribution as separated stems with a dot on each, label the height of the tallest, and write one sentence saying why you must not join them. In the middle, answer four questions from your table, writing the qualifying values before adding: P(X = 2), P(X is at most 2), P(X is less than 2), and P(X is at least 1) — and for the last one, do it both directly and by the complement. At the bottom, build the distribution of the number of heads in two coin tosses by listing the four outcomes, applying the variable to each, and grouping; explain in one sentence why the table has three rows rather than four.

Check your two answers to P(X at most 2) and P(X less than 2): they must differ by exactly 0.15, the probability of the value 2. If they came out equal, the endpoint was handled the continuous way, which belongs to chapter 5 and not here.

65. What you can do now

Recap

Five things, and the two conditions are the ones you will check for the rest of the book.

If you seeThen
A quantity counted on each trialA discrete random variable
A probability outside 0 to 1Not a distribution: the first condition fails
Probabilities summing to less than oneA value has been left off the list
A range questionAdd the qualifying values; no correction needed
A range covering most of the valuesUse the complement instead
Isolated whole-number valuesDraw separated stems, height equals probability
Values filling a rangeContinuous: chapter 5, and area rather than height

Section 4.2 asks what a distribution's centre and spread are. The expected value is the long-term average of the variable, computed by multiplying each value by its probability and adding — which is section 2.5's weighted mean with probabilities in place of relative frequencies.

OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 228-229

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108