The chapter's opening move: instead of asking for the probability of one event, ask how probability is distributed across every value a quantity can take. A random variable attaches a number to each outcome of an experiment, and a discrete probability distribution function lists each value the variable can take beside the probability of taking it. Such a function has exactly two characteristics — each probability lies between zero and one inclusive, and the probabilities sum to one — and a table failing either is not a distribution at all. Because distinct values are mutually exclusive, every range question is a plain sum, and because the variable takes isolated values the distribution is drawn as separated stems rather than a curve.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 4 — Discrete Random Variables
Probability Distribution Function (PDF) for a Discrete Random Variable
Objectives
Five outcomes. The second is a two-line check that decides whether anything else is worth doing.
OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229 — the section these objectives are drawn from
Warm-up
Section 1.3 built relative frequency tables from data, and section 3.1 computed the probability of individual events.
Discussion prompt
Toss a coin ten times and count the heads. What values can that count take, and could you write down the probability of each before tossing anything?
Hint: List the possible counts first, then ask whether they are equally likely.
Answer:
The count can be any whole number from zero to ten, so there are eleven possible values. Nothing in between is possible — 3.5 heads is not an outcome.
And yes, each value's probability can be written down in advance, though not by simple counting: five heads is far more likely than zero, because there are many orderings giving five heads and only one giving none.
A table of every possible value beside its probability is what this chapter calls a probability distribution, and building one shifts the question from 'what is the chance of this event' to 'how is the probability spread across everything that could happen'. That shift is what the rest of the book is built on.
Concept
A discrete probability distribution function has two characteristics: each probability is between zero and one inclusive, and the sum of the probabilities is one. Given a random variable X, the notation P(x) means the probability that X takes on the value x.
discrete probability distribution function — A listing of every value a discrete random variable can take, each with its probability. It must satisfy two conditions: every probability lies between zero and one inclusive, and the probabilities sum to one.
\[ 0 \le P(x) \le 1 \quad \text{for every } x, \qquad \sum_x P(x) = 1 \]
The two conditions look modest and they do a great deal of work. The first is inherited from section 3.1, where probabilities were defined to lie between zero and one. The second is the genuinely new demand: it insists that the list is COMPLETE, so that every value the variable could possibly take appears on it. A table whose probabilities sum to 0.9 is not describing a distribution with a gap; it is describing a situation whose possibilities have not all been written down.
Figure (svg): The two conditions for a discrete probability distribution function, each shown alongside a table that violates it
OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, p. 228
Section
Section 1
Concept
A random variable assigns a number to each outcome of an experiment. Where section 3.1 asked whether an event occurred, a random variable records a quantity, and the probability distribution records how likely each of its values is.
random variable — A rule assigning a numerical value to each outcome of an experiment, conventionally written with a capital letter such as X. The lower case x denotes a particular value it can take, and P(x) is the probability that X takes on that value.
\[ X = \text{the number of days per week Nancy attends class} \]
The notation is worth settling now because the whole chapter depends on it. A capital X names the variable — the rule — and a lower case x names one of its values. So P(x = 3) asks for the probability that the rule returns 3, and the four values 0, 1, 2 and 3 are the complete list of what it can return. Section 1.1 introduced variables as characteristics measured on each member; a random variable is the same idea applied to the outcomes of an experiment.
Figure (svg): A diagram showing four outcomes of tossing two coins mapped to the number of heads, illustrating a random variable
OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, p. 228 — the notation P(x) for a random variable
Picture it
Two coins, four outcomes, three possible values of the count.
Figure (svg): A diagram showing four outcomes of tossing two coins mapped to the number of heads, illustrating a random variable
The interesting feature is that the map is many-to-one: HT and TH both send to the value 1, which is why that value carries twice the probability of the others. A random variable deliberately throws information away — it records how many heads, not which coins — and the distribution records what probability survives that compression.
Worked example
Example 4.2a and b. The description says what to count.
\[ \text{Nancy has classes three days a week and attends some of them.} \]
Find the quantity being counted
Why: How many of her class days she attends.
Write it as a variable of ONE week
Why: A single randomly chosen week.
\[ X =\text{ days attended that week} \]
List the possible values
Why: She has three class days, so she can attend none through all three.
\[ 0, 1, 2, 3 \]
Check nothing between them is possible
Why: Days are counted whole.
Figure (svg): The solution to Worked example naming the variable and its values shown as a ladder of expressions, one row per legal move
\[ X \text{ takes the values } 0, 1, 2, 3 \]
Verify: confirm the value list is complete and cannot be extended
Why: Nancy has classes three days a week, so four is impossible and negative attendance is meaningless — the list of four values is exhaustive. Completeness matters because the second condition demands that the probabilities sum to one, and a missing value would make that impossible to satisfy honestly. Listing the values before any probability is written down is the reliable order.
OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, p. 229
Sorting
List the possible values before doing anything else.
Sort into buckets
Sort each random variable by the values it can take.
Item (d) has no natural maximum, which is allowed: a discrete variable may take infinitely many values, as the geometric and Poisson distributions later in this chapter do. What matters is that the values are isolated, not that they are few.
Worked example
The same experiment can carry different random variables.
\[ \text{Toss two coins. } X = \text{ number of heads}; \; Y = \text{ number of tails}. \]
List the sample space
Why: Four equally likely outcomes.
Apply X to each
Why: Counting heads.
\[ 2, 1, 1, 0 \]
Apply Y to each
Why: Counting tails.
\[ 0, 1, 1, 2 \]
Note the relationship
Why: Two tosses in all.
\[ Y = 2 - X \]
Figure (svg): The solution to Worked example two variables on one experiment shown as a ladder of expressions, one row per legal move
\[ X + Y = 2 \text{ on every outcome} \]
Verify: confirm the two distributions are mirror images
Why: X takes the value 2 with probability one quarter and Y takes the value 0 with the same probability, and so on down: the two distributions are reflections of each other. That is expected, since Y is just 2 minus X. It shows that a random variable is a CHOICE about what to record from an experiment, not a property the experiment has on its own — and that different choices can carry the same information in different arrangements.
OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229
Trap
\[ X = 3 \]
Write the definition of the variable as an equation
Why: The variable and its value are both being written, so one line looks enough.
\[ \text{that is not a definition; it is one particular case} \]
Writing X equals 3 says the variable took the value 3 on some occasion, which is an event rather than a definition.
\[ X = \text{the number of days Nancy attends class per week} \]
Define the variable in words, then list its values separately
Why: The capital names the rule; the lower case names an outcome of the rule.
The distinction is exactly section 1.1's between a variable and its data, and the notation preserves it: X is what is being measured and x is what came out. It matters here because P(x = 3) is a sensible expression and P(X) is not — the probability attaches to a value, never to the rule itself.
Fill the middle
What the capital and the lower case each name.
Fill in the blanks
X \textrandom variable ___, \text___ x \text___
Why: The capital names the random variable — the rule assigning a number to each outcome — and the lower case names a particular value. So P(x = 3) is well formed and asks for the probability that the rule returns three.
Two truths and a lie
All three concern random variables.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. HT and TH both give one head, and that many-to-one mapping is exactly what makes the value 1 twice as likely as the value 2. A random variable deliberately compresses the sample space, and the distribution records what probability accumulates on each value.
Prediction
Commit before reasoning.
Predict first
X is the number of heads in two tosses. What information about the experiment does X not record?
Correct: Which coin gave which face.
Why: The value 1 is returned by two different outcomes, so knowing X equals 1 does not say which toss was the head. That compression is deliberate and is what makes a distribution manageable: eleven values summarise 1,024 outcomes for ten tosses. The fairness of the coins is an assumption behind the probabilities rather than something the variable records.
Section
Section 2
Concept
A discrete probability distribution function has two characteristics. First, each probability is between zero and one, inclusive. Second, the sum of the probabilities is one.
the two characteristics — Every P(x) satisfies zero is less than or equal to P(x), which is less than or equal to one. And the sum of P(x) over all values x is exactly one. A table violating either is not a probability distribution function.
\[ 0 \le P(x) \le 1 \qquad \text{and} \qquad \sum P(x) = 1 \]
The two conditions fail in different ways and it is worth knowing which is which. A probability outside the range is an arithmetic impossibility — nothing can happen more often than always. A sum that misses one is a bookkeeping failure: either a value has been left off the list, or a probability has been miscomputed. The second is the commoner error and the more instructive one, because it usually points at a possibility the modeller forgot.
Figure (svg): The two conditions for a discrete probability distribution function, each shown alongside a table that violates it
OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229 — the two characteristics, and the check on Example 4.1
Picture it
Each condition with a table that breaks it.
Figure (svg): The two conditions for a discrete probability distribution function, each shown alongside a table that violates it
The book checks both conditions explicitly on its first example, and the habit is worth copying: before computing anything from a distribution, add the probabilities. It takes seconds and it catches both a missing value and an arithmetic slip, and every calculation later in this chapter assumes the check has passed.
Worked example
Example 4.2c. The book asks what the P(x) column sums to.
\[ P(0) = 0.01, \; P(1) = 0.04, \; P(2) = 0.15, \; P(3) = 0.80 \]
Check each probability's range
Why: All four lie between zero and one.
Add them
Why: One hundredth plus four hundredths plus 0.15 plus 0.80.
\[ 0.01 + 0.04 + 0.15 + 0.80 \]
Evaluate
Why: The total.
\[ 1.00 \]
Conclude
Why: Both conditions are satisfied.
Figure (svg): The solution to Worked example verifying a distribution shown as a ladder of expressions, one row per legal move
\[ \sum P(x) = 0.01 + 0.04 + 0.15 + 0.80 = 1.00 \]
Verify: confirm the list of values is what makes the sum work
Why: The four values 0 through 3 are every possibility, because Nancy has exactly three class days. Had the value 3 been omitted, the remaining probabilities would total 0.20 and the second condition would fail — correctly, because the table would then be silent about the commonest case. So the sum is not merely an arithmetic check; it is a test that the list of possibilities is complete.
OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, p. 229
Discrimination
Check both conditions on each.
Sort into buckets
Sort each proposed table.
Worked example
The second condition used to solve rather than to check.
\[ P(0) = 0.01, \; P(1) = 0.04, \; P(2) = 0.15, \; P(3) = \text{?} \]
Write the condition
Why: The four must total one.
\[ \sum\text{ of } P(x) = 1 \]
Add the known probabilities
Why: The first three.
\[ 0.01 + 0.04 + 0.15 = 0.20 \]
Subtract from one
Why: What is left for the last value.
\[ 1 - 0.20 \]
Read the answer
Why: The missing probability.
\[ 0.80 \]
Figure (svg): The solution to Worked example finding a missing probability shown as a ladder of expressions, one row per legal move
\[ P(3) = 1 - (0.01 + 0.04 + 0.15) = 0.80 \]
Verify: confirm the answer is a legal probability
Why: The result 0.80 lies between zero and one, so the first condition is satisfied too. That check matters because the method can produce an illegal answer: if the known probabilities had summed to more than one, the subtraction would give a negative number, and that would be evidence of an error in the given values rather than a strange distribution. Solving with the second condition should always be followed by testing against the first.
OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229
Error analysis
A student proposes four tables for a variable taking the values 0, 1, 2 and 3.
Annotate
On: \( \begin{aligned} &(1)\; 0.2, \;0.4, \;0.5, \;0.1 \\ &(2)\; 0.3, \;0.4, \;-0.1, \;0.4 \\ &(3)\; 0.2, \;0.3, \;0.2, \;0.1 \\ &(4)\; 0.1, \;0.2, \;0.3, \;0.4 \end{aligned} \)
Errors (1) and (3) are the same failure in opposite directions and are far commoner than (2), because nothing about the individual numbers looks wrong. That is precisely why the sum must be computed rather than eyeballed.
Faded example
A variable takes values 0, 1, 2 with P(0) = 0.35 and P(1) = 0.45.
Fill in the blanks
P(2) = 1 - (0.35 + 0.45) = 1 - 0.8 = 0.2
Why: The three must total one, so the missing probability is 0.20. Check it against the first condition as well: 0.20 lies between zero and one, so the answer is legal. Had the two known probabilities summed past one, the method would have returned a negative number and revealed an error in the data.
Two truths and a lie
All three concern the two conditions.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false because it checks only the first condition. The table 0.2, 0.3, 0.2, 0.1 has every entry in range and totals 0.8, so it is not a distribution. Both conditions must hold, and the second is the one that fails most often and least visibly.
Prediction
Commit before reasoning.
Predict first
A proposed distribution's probabilities total 0.85. What is the most likely explanation?
Correct: A value was left off, or a probability is wrong.
Why: Probability must be fully accounted for across everything that can happen, so a shortfall of 0.15 means something that can happen is not on the table. The right response is to find the missing possibility, not to rescale — rescaling would silently redistribute the missing 0.15 across the listed values and assert probabilities nobody computed.
Section
Section 3
Concept
A probability distribution table, or PDF table, has two columns: one listing every value the random variable can take, and one giving the probability of each. Building it from a verbal description means identifying the variable, listing its values, and reading a probability for each.
PDF table — A two-column table with the values of the random variable in the first column and their probabilities in the second. The book's Example 4.1 and 4.2 both construct one, and the P(x) column must sum to one.
\[ \begin{array}{c|c} x & P(x) \\ \hline 0 & 0.01 \\ 1 & 0.04 \\ 2 & 0.15 \\ 3 & 0.80 \end{array} \]
The table is section 1.3's relative frequency table with theoretical probabilities in place of observed proportions, and the parallel goes further than the shape. There the relative frequency column had to sum to one because every observation fell in exactly one row; here the probability column must sum to one because every outcome produces exactly one value. The same partition argument underlies both.
Figure (svg): A probability distribution table for Nancy's class attendance with a total row summing to one
OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, p. 229 — constructing a PDF table for Example 4.2
Picture it
The distribution as a table, with the running total made explicit.
Figure (svg): A probability distribution table for Nancy's class attendance with a total row summing to one
The running total column is not part of the standard table but it is worth writing while learning, because it turns the second condition into something checked as you go rather than at the end. It is also section 1.3's cumulative relative frequency in a new setting, and it will reappear in chapter 5 as the cumulative distribution function.
Worked example
Example 4.2c. The percentages in the description become the probabilities.
\[ \text{three days } 80\%, \text{ two days } 15\%, \text{ one day } 4\%, \text{ no days } 1\% \]
List the values in order
Why: From none to all three.
\[ 0, 1, 2, 3 \]
Attach each stated percentage
Why: Reading the description onto the right row.
\[ 0.01, 0.04, 0.15, 0.80 \]
Convert percentages to decimals
Why: Dividing each by a hundred.
Check the total
Why: The second condition.
\[ 1.00 \]
Figure (svg): The solution to Worked example building Nancy's table shown as a ladder of expressions, one row per legal move
\[ P(0) = 0.01, \; P(1) = 0.04, \; P(2) = 0.15, \; P(3) = 0.80 \]
Verify: confirm the description was read onto the right rows
Why: The description gives the percentages in decreasing order of days — three, two, one, none — while the table lists values in increasing order, so the four numbers must be reversed as they are transferred. That reversal is where errors occur, and the check is that the largest probability, 0.80, sits against three days, which is what the description says she does most often. Reading the finished table back into a sentence catches a transposition immediately.
OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, p. 229
Faded example
Jeremiah attends both of two practices 90 percent of the time, one 8 percent, neither 2 percent.
Fill in the blanks
P(0) = 0.02, \quad P(1) = 0.08, \quad P(2) = 0.90
Why: The description runs from most to fewest and the table runs from fewest to most, so the percentages reverse as they are transferred. The three total 1.00, and the largest probability sits against two practices, which matches what the description says he usually does.
Worked example
When the outcomes are equally likely, section 3.1's counting builds the table.
\[ X = \text{the number of heads in two tosses of a fair coin} \]
List the sample space
Why: Four equally likely outcomes.
Apply the variable to each
Why: Counting heads.
\[ 2, 1, 1, 0 \]
Count the outcomes for each value
Why: One, two and one respectively.
\[ 1, 2, 1 \]
Divide by four
Why: Each count over the sample space size.
\[ 0.25, 0.50, 0.25 \]
Figure (svg): The solution to Worked example a table from equally likely outcomes shown as a ladder of expressions, one row per legal move
\[ P(0) = \frac{1}{4}, \; P(1) = \frac{2}{4}, \; P(2) = \frac{1}{4} \]
Verify: confirm the total and the shape
Why: The three probabilities sum to one, as required. And the middle value carries twice the probability of either end, because two of the four outcomes produce it — which is the many-to-one mapping the first idea drew. That shape, higher in the middle and lower at the extremes, is the beginning of the binomial distribution that section 4.3 will build for any number of tosses.
OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229
Trap
\[ x: \; HH, \;HT, \;TH, \;TT \]
Put the sample space in the x column
Why: Those are the four things that can happen, so they look like the values.
\[ \text{but a random variable takes NUMBERS} \]
The x column holds the values of X, which are the counts 0, 1 and 2 — three rows, not four.
\[ x: \; 0, \;1, \;2 \quad \text{with } P(x): \; 0.25, \;0.50, \;0.25 \]
Apply the variable to the outcomes first, then list the distinct results
Why: Outcomes are what happens; values are what X reports about it.
The compression from four outcomes to three values is the whole point of a random variable, and it is why the middle probability is 0.50 rather than 0.25. A table with one row per outcome would have four equal probabilities and would describe the sample space rather than the distribution of X — which is a different and less useful object once the sample space is large.
Ranking
Building a PDF table from a verbal description.
Put in order
Why: Naming the variable first prevents the commonest error of tabulating outcomes instead of values. Listing the values before attaching probabilities makes it obvious if one is missing — and the final check is what catches an omission if the listing missed one anyway.
Two truths and a lie
All three concern building the table.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. Two coins have four outcomes and three values, and ten coins have 1,024 outcomes and eleven values. That compression is what makes a distribution useful, and it is why the number of rows is the number of distinct values rather than the size of the sample space.
Prediction
Commit before reasoning.
Predict first
X is the number of heads in ten tosses of a coin. How many rows does its PDF table have?
Correct: Eleven.
Why: The count can be any whole number from zero to ten inclusive, which is eleven values. The 1,024 is the number of outcomes in the sample space, and the variable compresses those onto eleven values — which is precisely why the distribution is worth tabulating and the sample space is not.
Section
Section 4
Concept
Once a distribution is tabulated, the probability of any range of values is found by adding the probabilities of the values in that range. No correction is needed, because a random variable takes exactly one value on any occasion, so distinct values are mutually exclusive.
reading a range — The probability that X falls in a set of values is the sum of the probabilities of those values. Distinct values of a random variable are mutually exclusive, so section 3.3's addition rule applies with its correction term equal to zero.
\[ P(X \le 2) = P(0) + P(1) + P(2) \]
This is section 3.3's addition rule with the subtraction guaranteed to vanish. X cannot be both 1 and 2 on the same occasion, so the events are mutually exclusive by construction and the plain sum is always correct. It is one of the quiet conveniences of working with a distribution rather than with arbitrary events, and it holds for every distribution in this chapter and the next.
Figure (svg): A discrete probability distribution with the stems for zero, one and two days highlighted and their probabilities summed to 0.20
OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229 — reading probabilities from a distribution
Picture it
Three stems highlighted, and their heights added.
Figure (svg): A discrete probability distribution with the stems for zero, one and two days highlighted and their probabilities summed to 0.20
The answer of 0.20 is small because almost all of Nancy's probability sits on the single value 3. That is worth noticing as a habit: before computing a range, look at where the distribution's mass is, because a range excluding the tall stem will be small however many values it contains.
Worked example
Adding the stems that qualify.
\[ P(X \le 2) \text{ for Nancy's attendance} \]
Identify the qualifying values
Why: At most two means 0, 1 or 2.
Read their probabilities
Why: From the table.
\[ 0.01, 0.04, 0.15 \]
Add them
Why: The values are mutually exclusive.
\[ 0.20 \]
Sanity-check against the remaining value
Why: One minus P(3).
\[ 1 - 0.80 = 0.20 \]
Figure (svg): The solution to Worked example a range from the table shown as a ladder of expressions, one row per legal move
\[ P(X \le 2) = 0.01 + 0.04 + 0.15 = 0.20 \]
Verify: confirm the complement gives the same answer
Why: At most two and exactly three are complementary events, since those four values exhaust the possibilities, so one minus 0.80 must give the same 0.20 — and it does. On a distribution with many values the complement is often far the shorter route, and section 3.1's advice about 'at least one' events applies here: whenever a range covers most of the values, compute the complement instead.
OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, p. 229
Sorting
Nancy's variable takes the values 0, 1, 2, 3.
Sort into buckets
Sort each range by how many values it covers.
Items (a) and (b) differ by exactly one value and by 0.15 of probability, which on this distribution is three times the whole answer to (b). Listing the qualifying values in writing before adding is what prevents the slip.
Worked example
The same distribution, a range covering most of it.
\[ P(X \ge 1) \text{ for Nancy's attendance} \]
Count the qualifying values
Why: One, two and three.
Count the complement's values
Why: Only zero.
Choose the complement
Why: Fewer terms.
\[ 1 - P(0) \]
Compute
Why: One minus one hundredth.
\[ 0.99 \]
Figure (svg): The solution to Worked example which route is shorter shown as a ladder of expressions, one row per legal move
\[ P(X \ge 1) = 1 - P(0) = 1 - 0.01 = 0.99 \]
Verify: confirm the direct route agrees
Why: Adding 0.04, 0.15 and 0.80 gives 0.99, the same answer by three additions instead of one subtraction. On four values the saving is trivial; on the eleven values of ten coin tosses, or the unbounded list of a Poisson variable in section 4.6, it is the difference between a line and a page. The rule of thumb is to count the terms each way and take the shorter.
OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229
Trap
\[ P(X < 2) = 0.01 + 0.04 + 0.15 \]
Include the value 2 in a strictly-less-than range
Why: At most two and less than two look similar in a hurry.
\[ = 0.20 \quad \text{(that is } P(X \le 2)\text{)} \]
Strictly less than two means the values 0 and 1 only, giving 0.05.
\[ P(X < 2) = 0.01 + 0.04 = 0.05 \]
Write out the list of qualifying values before adding anything
Why: The endpoint is decided once, in the listing, rather than repeatedly in the arithmetic.
On a discrete distribution the endpoint genuinely matters, because P(X = 2) is a real positive number — here 0.15, which is three times the whole of P(X < 2). That is a sharp contrast with chapter 5's continuous distributions, where any single value has probability zero and 'less than' and 'at most' give identical answers. Carrying the discrete habit into chapter 5 is harmless; carrying the continuous habit back here is not.
Faded example
Nancy's distribution: 0.01, 0.04, 0.15, 0.80 for the values 0 to 3.
Fill in the blanks
P(X \ge 2) = 1 - P(0) - P(1) = 1 - 0.01 - 0.04 = 0.95
Why: Two subtractions rather than two additions makes little difference here, but the direct route of 0.15 plus 0.80 gives the same 0.95, which is the check. Whichever route is used, the two must agree, and that agreement is the cheapest verification available on a range question.
Two truths and a lie
All three concern reading ranges.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false for a discrete variable. The two differ by P(X = 2), which is 0.15 on Nancy's distribution — larger than the whole of P(X < 2). Only for the continuous variables of chapter 5, where any single value has probability zero, do the two coincide.
Prediction
Commit before reasoning.
Predict first
Section 3.3's addition rule subtracts the intersection. Why is no subtraction needed when adding P(1) and P(2)?
Correct: Because X cannot take two values at once.
Why: A random variable returns exactly one number on each occasion, so the events X equals 1 and X equals 2 are mutually exclusive by construction and their intersection has probability zero. The addition rule still applies; its correction term is simply guaranteed to vanish, which is why range questions on a distribution are always plain sums.
Section
Section 5
Concept
A discrete random variable takes isolated values, so its distribution is drawn as separated stems with a dot on top rather than as a joined curve. The height of each stem is the probability of that value, and the space between stems represents values the variable cannot take.
drawing a discrete distribution — Separated vertical stems, one per value, with height equal to that value's probability. The stems are not joined because the variable takes no values between them, and the height is the probability itself rather than an area.
\[ \text{height of the stem at } x \;=\; P(x) \]
This is section 2.1's line-graph rule in a new setting. There, joining points across an unordered axis asserted a trend that relabelling could reverse; here, joining stems would assert that the variable can take values between them. For Nancy's attendance there is no such thing as 2.4 days, so a line drawn from the stem at 2 to the stem at 3 passes through territory that does not exist.
Figure (svg): Two columns contrasting a discrete distribution drawn as separated stems with a continuous curve
OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229 — the discrete distribution and its values
Picture it
What each convention asserts about the variable.
Figure (svg): Two columns contrasting a discrete distribution drawn as separated stems with a continuous curve
The row worth carrying into chapter 5 is the third. On a discrete distribution the HEIGHT is the probability, so a stem of 0.80 means that value occurs eighty percent of the time. On a continuous distribution the height is a density and the probability is an AREA, so a single point has probability zero. Reading a continuous curve as if its height were a probability is the commonest error in chapter 5, and it starts here with the difference in how the two are drawn.
Worked example
The variable's definition decides which numbers can appear.
\[ X = \text{the number of days Nancy attends class per week} \]
Ask what the variable counts
Why: Whole days of attendance.
Ask whether a fraction is possible
Why: She attends a day or she does not.
Conclude about the values
Why: Only whole numbers from 0 to 3.
Draw accordingly
Why: Separated stems, unjoined.
Figure (svg): The solution to Worked example why 2.4 days is not a value shown as a ladder of expressions, one row per legal move
\[ x \in \{0, 1, 2, 3\}, \text{ with nothing between} \]
Verify: confirm the same test against a continuous quantity
Why: The number of HOURS Nancy spends in class would be different: 2.4 hours is perfectly possible, and so is 2.41, so the values fill a range rather than sitting apart. That variable would be continuous and would belong to chapter 5, drawn as a curve. The same situation supplies both kinds of variable depending on what is recorded, which is section 1.2's discrete-against-continuous distinction reappearing exactly where it is needed.
OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, p. 228
Discrimination
Ask whether the variable can take values between the ones you would plot.
Sort into buckets
Sort each random variable.
Worked example
On a discrete plot the height is the answer.
\[ \text{the stem at } x = 3 \text{ has height } 0.80 \]
Read the height
Why: Eight tenths.
\[ 0.80 \]
State what it means
Why: The probability that X equals exactly 3.
\[ P(X = 3) = 0.80 \]
Note what it does not mean
Why: It is not an area and needs no width.
Check against the total
Why: All four heights sum to one.
Figure (svg): The solution to Worked example reading a height shown as a ladder of expressions, one row per legal move
\[ P(3) = 0.80 \text{, the height of the stem} \]
Verify: confirm the heights must sum to one
Why: The four stem heights are 0.01, 0.04, 0.15 and 0.80, totalling exactly one — which is the second condition drawn. On a discrete plot that total is a property of the HEIGHTS, and it is worth noticing now because on chapter 5's continuous plots the corresponding statement is that the total AREA under the curve is one. The same requirement, expressed differently because the drawing convention differs.
OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229
Trap
\[ \text{a smooth line drawn through the four dots} \]
Join the stems to show the shape of the distribution
Why: A joined line makes the pattern easier to see.
\[ \text{the line passes through } x = 2.4, \text{ which is not a possible value} \]
Reading a height off that line at 2.4 would give a probability for something that cannot occur.
\[ \text{separated stems, each with a dot on top} \]
Leave the gaps, because the gaps are information
Why: The empty space says the variable takes no values there.
Section 2.2's frequency polygon joined its points legitimately because the underlying variable was continuous and had been grouped into classes. Here there is no underlying continuum to interpolate across, so the join would be a claim rather than a convenience. The same principle governed section 2.1's rule against line graphs over categories, and it comes down to the same question: does the space between the plotted points contain anything?
Fill the middle
On a discrete probability plot.
Fill in the blanks
\textprobability x \text___ ___ \text___ X \text___ x
Why: The height is the probability itself, so a stem of 0.80 means that value occurs eighty percent of the time. Chapter 5's continuous distributions reverse this: there the height is a density and the probability is the area beneath the curve.
Two truths and a lie
All three concern drawing a discrete distribution.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false for a probability distribution. Section 2.2 did draw discrete DATA as a histogram with touching bars centred on each value, which is a display of observed frequencies over a numeric axis. A probability distribution over isolated values is conventionally drawn as separated stems, precisely to keep the isolation visible.
Prediction
Commit before reasoning.
Predict first
Someone draws a smooth curve through the tops of Nancy's four stems. What does the curve claim at x = 2.4?
Correct: That there is a probability of attending 2.4 days.
Why: A curve has a height at every point of its domain, so drawing one asserts a value at 2.4 — and Nancy cannot attend 2.4 days. Section 2.2's frequency polygon was allowed to join its points because the underlying variable was continuous and merely grouped; here there is no continuum underneath, so the interpolation has nothing to interpolate.
Comparison
Fill the blanks. The left column is chapter 3 and the right is this chapter.
Comparison matrix
| Question | Chapter 3 answered | Chapter 4 answers |
|---|---|---|
| What is being asked about? | one event: did A happen? | a quantity: what value did X take? |
| What is the answer? | a single probability | a probability for every possible value |
| What must sum to one? | an event and its complement | the probabilities of all the values |
| How is a range handled? | the addition rule, minus the intersection | a plain sum: the values are exclusive |
The shift is from asking about one event to describing a whole quantity, and the payoff arrives in the next five sections: some quantities recur so often that their distributions have names and formulas, so the table never has to be built by hand again.
Pattern
Six steps, and the last is a check that should never be skipped.
If the probabilities fall short of one, a possible value has been omitted rather than a probability being slightly wrong. Go back to step two before adjusting anything.
OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229
Check
Test both conditions.
Check your understanding
Is the table with probabilities 0.2, 0.3, 0.4 and 0.2 a valid discrete PDF?
Answer: A
Why: Every entry lies in range, so the first condition holds, but the four total 1.1 rather than 1, so the second fails and the table is not a distribution.
Check
Read a range.
Check your understanding
Nancy's distribution is 0.01, 0.04, 0.15, 0.80 for x = 0, 1, 2, 3. What is P(X < 2)?
Answer: A
Why: Strictly less than 2 means the values 0 and 1, giving 0.01 plus 0.04, which is 0.05.
Check
Count the values.
Check your understanding
X is the number of heads in three tosses of a coin. How many rows does its PDF table have?
Answer: A
Why: The count can be 0, 1, 2 or 3, which is four values, so the table has four rows.
Real world
A support team models the number of tickets arriving in an hour and publishes a distribution: 0 tickets with probability 0.10, 1 with 0.25, 2 with 0.30, 3 with 0.20, and 4 with 0.10. A manager uses it to plan staffing for 'up to three tickets an hour'.
Discussion prompt
Check the distribution, say what the manager's plan actually covers, and identify what the model is assuming.
Hint: Add the probabilities first, then read the range the manager described.
Answer:
The distribution does not sum to one. The five probabilities total 0.95, so five percent of the probability is unaccounted for — and since each listed value is in range, the failure is the second condition. Some possibility has been left off the table.
What is missing is the tail. Five or more tickets in an hour is clearly possible and simply was not listed, which is the commonest way this condition fails: the modeller enumerated the ordinary cases and stopped. The missing 0.05 belongs to those busy hours.
The manager's plan covers less than it appears to. Up to three tickets is 0.10 plus 0.25 plus 0.30 plus 0.20, which is 0.85 — so the plan fails about one hour in seven, not one in twenty as the incomplete table might suggest.
\[ \sum P(x) = 0.95 \ne 1 \;\Longrightarrow\; \text{a value is missing} \]
The general lesson is that the sum-to-one condition is a completeness test rather than an arithmetic formality, and a shortfall almost always points at an omitted tail. It matters most for exactly the planning question being asked here, because the omitted values are the extreme ones — the busy hours a staffing model exists to survive. Section 4.6's Poisson distribution is the standard model for counts of this kind, and it assigns a probability to every non-negative whole number precisely so that no tail can be forgotten.
Commit first
Answer, then rate your confidence honestly.
Predict first
A proposed discrete distribution has all its probabilities between 0 and 1 but they sum to 0.9. What is wrong?
Correct: The second condition fails, and a value is probably missing.
\[ \sum P(x) = 1 \text{ is a COMPLETENESS condition, not an arithmetic one} \]
Why: Probability must be fully accounted for across everything the variable can do, so a shortfall of 0.1 means something possible is not on the table. It is not a rounding matter and it is not fixed by rescaling, which would silently redistribute the missing probability across the listed values. The right move is to go back and find the omitted possibility, which is usually in the tail.
Explain it
They have made a table with one row per outcome of tossing two coins, and four equal probabilities of 0.25.
Discussion prompt
In three sentences or fewer, show them what a distribution of the count would look like instead.
Hint: Ask them what value X takes on HT and on TH.
Answer:
Ask them what the number of heads is for HT and for TH: both give one, so those two outcomes collapse onto a single value of the variable.
So the table has three rows rather than four — the values 0, 1 and 2 — and the middle one carries 0.50 because two of the four outcomes produce it.
Their table describes the sample space, which is fine for two coins and hopeless for ten, where 1,024 outcomes compress onto just eleven values.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the definition, write X as a full sentence about one trial and list the values before any probability. For the conditions, check the range of each entry and then add them. For counting, apply the variable to every outcome first and then group. For endpoints, write the qualifying values out before adding, since P(X = 2) is a real positive number here. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top, write the two conditions for a discrete PDF in a box, and beside each write a four-row table that violates it and one word saying what has gone wrong. Below, build Nancy's distribution from this description: she has classes three days a week and attends all three 80 percent of weeks, two days 15 percent, one day 4 percent and none 1 percent. Write the variable as a full sentence, list the values, build the two-column table, add a running-total column, and check the total. Draw the distribution as separated stems with a dot on each, label the height of the tallest, and write one sentence saying why you must not join them. In the middle, answer four questions from your table, writing the qualifying values before adding: P(X = 2), P(X is at most 2), P(X is less than 2), and P(X is at least 1) — and for the last one, do it both directly and by the complement. At the bottom, build the distribution of the number of heads in two coin tosses by listing the four outcomes, applying the variable to each, and grouping; explain in one sentence why the table has three rows rather than four.
Check your two answers to P(X at most 2) and P(X less than 2): they must differ by exactly 0.15, the probability of the value 2. If they came out equal, the endpoint was handled the continuous way, which belongs to chapter 5 and not here.
Recap
Five things, and the two conditions are the ones you will check for the rest of the book.
| If you see | Then |
|---|---|
| A quantity counted on each trial | A discrete random variable |
| A probability outside 0 to 1 | Not a distribution: the first condition fails |
| Probabilities summing to less than one | A value has been left off the list |
| A range question | Add the qualifying values; no correction needed |
| A range covering most of the values | Use the complement instead |
| Isolated whole-number values | Draw separated stems, height equals probability |
| Values filling a range | Continuous: chapter 5, and area rather than height |
Section 4.2 asks what a distribution's centre and spread are. The expected value is the long-term average of the variable, computed by multiplying each value by its probability and adding — which is section 2.5's weighted mean with probabilities in place of relative frequencies.
OpenStax Introductory Statistics 2e, §4.1 Probability Distribution Function (PDF) for a Discrete Random Variable §4.1, pp. 228-229 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.