The centre and spread of a distribution rather than of a data set. The expected value is the long-term average of a random variable, found by multiplying each value by its probability and adding the products, and it is denoted by the Greek letter mu because it is a parameter of a process rather than a statistic computed from data. It need not be one of the values the variable can take. The standard deviation is the square root of the sum of squared deviations weighted by their probabilities, with no divisor because the probabilities already sum to one. Both are computed from a single expected value table, and the Law of Large Numbers is what connects them to what actually happens over many trials.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 4 — Discrete Random Variables
Mean or Expected Value and Standard Deviation
Objectives
Five outcomes. The first is one line of arithmetic and the fifth is what it does and does not promise.
OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-236 — the section these objectives are drawn from
Warm-up
Section 2.5 computed a mean from a frequency table by weighting each value by how often it occurred and dividing by the total count.
Discussion prompt
A soccer team plays zero days a week 20 percent of the time, one day 50 percent, and two days 30 percent. What is the average number of days per week, and what changed from the section 2.5 calculation?
Hint: Try the weighted mean with the percentages as the weights.
Answer:
Multiply and add: zero times 0.2, plus one times 0.5, plus two times 0.3, giving 0 plus 0.5 plus 0.6, which is 1.1 days per week.
What changed is where the weights came from. In section 2.5 they were counts of observations divided by n; here they are probabilities, and no division is needed because probabilities already sum to one.
So the arithmetic is identical and the meaning is different: that 1.1 is not the average of any data collected but a property of the process itself. The book calls it the expected value and denotes it mu, the same Greek letter section 1.1 reserved for a population parameter.
Concept
The expected value is often referred to as the long-term average or mean: over the long term of doing an experiment over and over, you would expect this average. To find it, multiply each value of the random variable by its probability and add the products. It is denoted by the Greek letter mu.
expected value — The long-term average of a random variable, denoted mu. Computed as the sum over all values of x times P(x). It is the mean of the distribution, and it need not be one of the values the variable can take.
\[ \mu = \sum x \cdot P(x) \]
The name is unfortunate and worth disarming immediately: the expected value is very often a value you should not expect at all. The soccer team's 1.1 days is impossible in any single week, since the team plays a whole number of days. What the number describes is the average over many weeks, which the Law of Large Numbers says the observed average will approach. Reading 'expected' as 'long-run average' rather than 'likely' removes almost all the confusion the term causes.
Figure (svg): An expected value table for the soccer team with columns for the value, its probability and their product, summing to 1.1
OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-231
Section
Section 1
Concept
To find the expected value or long-term average, simply multiply each value of the random variable by its probability and add the products. Constructing a table with a third column for x times P(x) organises the work.
the expected value table — A PDF table with a third column holding x times P(x) for each row. The total of that column is the expected value. The book calls it an expected value table and uses it throughout the chapter.
\[ \mu = \sum x P(x) = x_1P(x_1) + x_2P(x_2) + \cdots \]
The table earns its keep by keeping two checks visible at once. The P(x) column must total one, which is section 4.1's second condition and confirms the distribution before anything is computed from it; and the x times P(x) column totals the expected value. Doing both in one table means an invalid distribution is caught before its mean is computed and reported, which is the order that matters.
Figure (svg): An expected value table for the soccer team with columns for the value, its probability and their product, summing to 1.1
OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-231 — the expected value and its table
Picture it
The soccer team's distribution with the product column added.
Figure (svg): An expected value table for the soccer team with columns for the value, its probability and their product, summing to 1.1
The value 1.1 is the sum of the third column, and it lies between the smallest and largest possible values as any weighted average must. Notice it sits nearer 1 than 2 because the value 1 carries the most probability — the expected value is pulled toward wherever the probability is concentrated, exactly as section 2.5's mean was pulled toward wherever the data were.
Worked example
Example 4.3, with the table the book builds.
\[ P(0) = 0.2, \quad P(1) = 0.5, \quad P(2) = 0.3 \]
Check the distribution first
Why: The three probabilities.
\[ 0.2 + 0.5 + 0.3 = 1 \]
Multiply each value by its probability
Why: Zero, one and two in turn.
\[ 0, 0.5, 0.6 \]
Add the products
Why: The third column's total.
\[ 0 + 0.5 + 0.6 \]
Read the expected value
Why: The sum.
\[ 1.1 \]
Figure (svg): The solution to Worked example the soccer team shown as a ladder of expressions, one row per legal move
\[ \mu = (0)(0.2) + (1)(0.5) + (2)(0.3) = 1.1 \]
Verify: confirm the answer lies inside the range of values
Why: The team plays between zero and two days, and 1.1 falls between them — which any weighted average must, since it is a mixture of the values themselves. An expected value outside the range of possible values is proof of an arithmetic error, and it is the fastest check available. Here it also sits just above one, which matches the distribution's concentration on that value.
OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, p. 231
Faded example
The soccer team: values 0, 1, 2 with probabilities 0.2, 0.5, 0.3.
Fill in the blanks
\mu = (0)(0.2) + (1)(0.5) + (2)(0.3) = 1.1
Why: The three products are 0, 0.5 and 0.6, totalling 1.1. Notice the first term contributes nothing because the value is zero — a value of zero always drops out of an expected value however likely it is, which is worth remembering when a distribution has many zeros.
Worked example
Example 4.4, a longer table with the same procedure.
\[ P = 0.04, \;0.22, \;0.46, \;0.18, \;0.08, \;0.02 \text{ for } x = 0 \text{ through } 5 \]
Verify the distribution
Why: The six probabilities.
\[ \text{they total } 1.00 \]
Form each product
Why: Value times probability.
\[ 0, 0.22, 0.92, 0.54, 0.32, 0.10 \]
Add them
Why: The expected value.
\[ 2.10 \]
Interpret
Why: The long-run weekly average.
\[ 2.1\text{ wakings per week} \]
Figure (svg): The solution to Worked example the night-waking distribution shown as a ladder of expressions, one row per legal move
\[ \mu = 0(0.04) + 1(0.22) + 2(0.46) + 3(0.18) + 4(0.08) + 5(0.02) = 2.1 \]
Verify: confirm the expected value is not a possible outcome
Why: A baby wakes its mother a whole number of times in a week, so 2.1 will never be observed on any single week. The book's own caption puts it as expecting a newborn to wake its parents 2.1 times per week ON THE AVERAGE, and the qualifier is doing all the work. The number is a property of the long run, and treating it as a prediction for next week is the commonest misreading of the whole chapter.
OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 231-232
Trap
\[ x = 0, 1, 2 \text{, so the average is } \frac{0 + 1 + 2}{3} \]
Average the three possible values
Why: Three numbers are listed, so their mean looks like the answer.
\[ = 1 \quad \text{(wrong: it ignores the probabilities)} \]
That treats playing zero days and playing one day as equally likely, when one day happens two and a half times as often.
\[ \mu = (0)(0.2) + (1)(0.5) + (2)(0.3) = 1.1 \]
Weight every value by its own probability
Why: The probabilities are what distinguish the values.
This is section 2.5's averaging-the-averages error in a new setting, and it gives the right answer only when every value is equally likely — which is why it survives on dice and coins and fails everywhere else. On the night-waking distribution the unweighted average of 0 through 5 is 2.5 against a true expected value of 2.1, and the gap grows with the asymmetry of the distribution.
Estimation
A variable takes the values 3, 7 and 12 with some probabilities.
Predict first
What range must its expected value lie in?
Correct: Between 3 and 12.
Why: An expected value is a weighted average of the values, so it lies between the smallest and largest of them — reaching an endpoint only when all the probability sits on that value. It need not be near the middle value, and it need not be 7: a distribution concentrated on 3 gives an expected value just above 3. This range check catches most arithmetic slips instantly.
Two truths and a lie
All three concern the expected value.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false: the most likely value is the mode. For the night-waking distribution the mode is 2, with probability 0.46, while the expected value is 2.1 — close here, but on a skewed distribution they can be far apart. The soccer team's mode is 1 and its expected value 1.1, and a distribution with a long tail separates them much further.
Prediction
Commit before reasoning.
Predict first
In an expected value calculation, what does the row for the value zero contribute?
Correct: Nothing, whatever its probability.
Why: The product is zero times its probability, which is zero however likely the value is. The row still matters for checking that the probabilities sum to one, but it adds nothing to the mean — so a distribution with a great deal of probability on zero has an expected value dragged down not by that row's contribution but by the smaller weights left for the others.
Section
Section 2
Concept
The standard deviation of a probability distribution is the square root of its variance, and the variance is the sum of the squared deviations weighted by their probabilities. For each value, multiply the square of its deviation from mu by its probability, then add and take the root.
the standard deviation of a distribution — Sigma, the square root of the sum over all values of the squared deviation times the probability. It measures how far the variable's values fall from mu, exactly as section 2.7's s measured how far data fell from x-bar.
\[ \sigma = \sqrt{\sum (x - \mu)^2 P(x)} \]
Section 2.7's computation is recognisable here with one change: there the squared deviations were divided by n minus 1, and here they are weighted by probabilities and not divided at all. The reason is that the probabilities already sum to one, so the weighted sum IS an average. There is no sample and so no n, and no correction for estimating a mean from data, because mu is known rather than estimated.
Figure (svg): A four-column table computing the expected value and standard deviation for the night-waking distribution
OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-232 — the standard deviation of a PDF
Picture it
Example 4.4 with all four columns.
Figure (svg): A four-column table computing the expected value and standard deviation for the night-waking distribution
The fourth column cannot be started until the third has been totalled, because every entry needs mu. That dependency is the only thing making this a two-pass calculation, and it is why the book computes the expected value first and then says to use mu to complete the table.
Worked example
Example 4.4. The fourth column, then a square root.
\[ \mu = 2.1, \quad P = 0.04, 0.22, 0.46, 0.18, 0.08, 0.02 \]
Form each squared deviation
Why: Value minus mu, squared.
\[ 4.41, 1.21, 0.01, 0.81, 3.61, 8.41 \]
Weight each by its probability
Why: The fourth column.
\[ 0.1764, 0.2662, 0.0046, 0.1458, 0.2888, 0.1682 \]
Add them for the variance
Why: The column total.
\[ 1.05 \]
Take the square root
Why: Back to the variable's units.
\[ \text{about } 1.0247 \]
Figure (svg): The solution to Worked example sigma for the night wakings shown as a ladder of expressions, one row per legal move
\[ \sigma = \sqrt{1.05} \approx 1.0247 \]
Verify: confirm sigma is a plausible size for this distribution
Why: The values run from 0 to 5, a range of five, and sigma is about one — roughly a fifth of the range, which is the usual order of magnitude. A sigma larger than the range would be impossible, and one very much smaller would suggest the probability is concentrated on a single value. Notice also which rows contribute most: the value 2 sits almost exactly at mu and contributes only 0.0046, while the value 5 contributes 0.1682 despite having a probability of just 0.02, because its deviation is squared.
OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, p. 232
Faded example
The weighted squared deviations for the night-waking distribution total 1.05.
Fill in the blanks
\sigma = \sqrt1.05} \approx 1.0247
Why: No divisor appears, because the probabilities weighting the terms already sum to one. The square root returns the answer to the variable's own units — wakings per week rather than squared wakings.
Worked example
Comparing the formula with section 2.7's.
\[ s = \sqrt{\frac{\sum (x - \bar{x})^2}{n-1}} \quad \text{against} \quad \sigma = \sqrt{\sum (x - \mu)^2 P(x)} \]
Note what section 2.7 divided by
Why: The number of observations, less one.
\[ n - 1 \]
Note what weights the terms here
Why: Each term carries its own probability.
\[ P(x) \]
Ask what the probabilities sum to
Why: The second condition of section 4.1.
Conclude
Why: Weighting by numbers summing to one IS averaging.
Figure (svg): The solution to Worked example why there is no divisor shown as a ladder of expressions, one row per legal move
\[ \sum P(x) = 1 \;\Longrightarrow\; \sum (x-\mu)^2 P(x) \text{ is already an average} \]
Verify: confirm the two agree when the values are equally likely
Why: Take three values each with probability one third. The distribution formula gives the sum of squared deviations times one third, which is the sum divided by three — exactly the POPULATION standard deviation of those three numbers. The book notes this: when all outcomes are equally likely, these formulas coincide with the mean and standard deviation of the set of possible outcomes. Note it matches the population form with N rather than the sample form with n minus 1, which is right, because mu here is known rather than estimated from data.
OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, p. 230
Error analysis
For a distribution with mu = 2.1, the weighted squared deviations total 1.05. Four students report sigma.
Annotate
On: \( \begin{aligned} &(1)\; \sigma = 1.05 \\ &(2)\; \sigma = \sqrt{\tfrac{1.05}{6}} \approx 0.42 \\ &(3)\; \sigma = \sqrt{\tfrac{1.05}{5}} \approx 0.46 \\ &(4)\; \sigma = \sqrt{1.05} \approx 1.0247 \end{aligned} \)
Errors (2) and (3) are the same instinct: every standard deviation so far has had a divisor, so one is supplied from habit. The guard is to ask what the weights sum to — when they sum to one, the averaging is done.
Discrimination
mu is 2.1, and each row contributes its squared deviation times its probability.
Sort into buckets
Sort each row of the night-waking distribution by its contribution.
Two truths and a lie
All three concern sigma for a distribution.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor describes the VARIANCE rather than the standard deviation. Sigma is its square root, which is what returns the answer to the variable's own units. Reporting 1.05 rather than 1.0247 here is the commonest slip, and it is the same one section 2.7 warned about for data.
Prediction
Commit before reasoning.
Predict first
Two distributions have the same expected value, but one has a much smaller sigma. What does that say?
Correct: Its values are more tightly concentrated near mu.
Why: Sigma measures how far the variable's values fall from its mean, exactly as section 2.7's s did for data. A small sigma means the probability sits close to mu, so an individual outcome is likely to be near the long-run average. The number of possible values is irrelevant — a distribution with many values packed near mu can have a smaller sigma than one with three values spread widely.
Section
Section 3
Concept
An expected value table can be built for money as easily as for counts. Let X be the amount of profit, list the possible profits with their probabilities, and the expected value is the average gain or loss per play over the long term.
expected profit — The expected value of a random variable representing money gained. A negative expected profit means the game loses money on average over many plays, however dramatic the individual outcomes.
\[ \mu = \sum (\text{profit}) \cdot P(\text{profit}) \]
The book's Example 4.5 is worth working slowly because the setup is where the difficulty lies. Five digits are chosen from zero to nine with replacement, so the chance of matching all five in order is a tenth to the fifth power, which is 0.00001. The values of X are the PROFITS — negative two dollars and one hundred thousand dollars — not the digits, and getting that right is most of the problem.
Figure (svg): A comparison showing that a single play of the game yields either a loss of two dollars or a profit of one hundred thousand, while the long-run average is a loss of about one dollar
OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, p. 233 — Example 4.5, the five-digit game
Picture it
Two possible outcomes, and one long-run average.
Figure (svg): A comparison showing that a single play of the game yields either a loss of two dollars or a profit of one hundred thousand, while the long-run average is a loss of about one dollar
The expected loss of about a dollar is never what happens on a play: every play loses two dollars or gains a hundred thousand. What the number says is that a player doing this repeatedly loses about a dollar a go, and the Law of Large Numbers is what guarantees the observed average approaches it. That is exactly why the game is offered.
Worked example
Example 4.5. The values are profits, not digits.
\[ \text{Pay } 2 \text{ dollars; match five digits in order to profit } 100\,000. \]
Find the probability of winning
Why: One digit in ten, five times, with replacement.
Find the probability of losing
Why: The complement.
\[ 0.99999 \]
List the profits, not the digits
Why: Lose two dollars, or profit one hundred thousand.
\[ -2\text{ and } 100 000 \]
Build and total the table
Why: Each profit times its probability.
\[ -1.99998 + 1 \]
Figure (svg): The solution to Worked example the five-digit game shown as a ladder of expressions, one row per legal move
\[ \mu = (-2)(0.99999) + (100000)(0.00001) = -0.99998 \]
Verify: confirm the two terms are nearly balanced and which one wins
Why: The loss term contributes about minus two dollars and the win term contributes exactly one dollar, so the game returns about half of what it takes. The win term is exactly 1 because a hundred thousand times a hundred-thousandth is one — a useful check on the probability, since if the prize and the odds were fair the two terms would cancel. They do not, and the shortfall is the operator's margin.
OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, p. 233
Faded example
The prize is 100,000 dollars and the probability of winning is 0.00001.
Fill in the blanks
(100000)(0.00001) = 1, \quad \text-1.99998 ___
Why: The win contributes exactly one dollar per play and the loss about two, so the expected profit is about minus one. The neatness of the win term is a coincidence of the numbers chosen, and it makes the comparison unusually easy to see.
Worked example
Running the calculation backwards.
\[ \text{What prize would make the expected profit zero?} \]
Write the condition
Why: Expected profit equals zero.
\[ (-2) (0.99999) + W(0.00001) = 0 \]
Isolate the win term
Why: Move the loss across.
\[ W(0.00001) = 1.99998 \]
Divide
Why: Solve for the prize.
\[ W = 199 998 \]
Compare with the actual prize
Why: One hundred thousand offered.
Figure (svg): The solution to Worked example what prize would make it fair shown as a ladder of expressions, one row per legal move
\[ W = \frac{1.99998}{0.00001} = 199\,998 \]
Verify: confirm the interpretation of a fair game
Why: A fair game has an expected profit of zero, so a player neither gains nor loses in the long run — and the operator makes nothing either. The prize offered is about half the fair value, which is why the expected profit is about minus one dollar on a two dollar stake. Every commercial game of chance is priced this way, and computing the fair prize is the standard way to see by how much.
OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, p. 233
Trap
\[ x = 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 \]
Take the values of X to be the digits
Why: Those are the numbers in the problem.
\[ \text{an expected value of } 4.5 \text{ digits} \quad \text{(answering nothing)} \]
The question asks about profit, so the variable must be the money and its values are the two possible profits.
\[ X = \text{the amount of money you profit}; \quad x = -2 \text{ or } 100\,000 \]
Define X as the quantity the question is about, then find that quantity's values
Why: The digits generate the probabilities; the money is what is being averaged.
The book is explicit that the values of x are not the digits zero through nine but the two profits, and the distinction is the whole setup. It generalises: in any decision problem the random variable is the CONSEQUENCE, and the underlying experiment merely supplies the probabilities. Writing the definition of X as a full sentence before listing any values is what keeps the two apart.
Sorting
A game is fair when the expected profit is zero.
Sort into buckets
Sort each game by its expected profit.
Item (d) pays two dollars for a one-dollar stake on a half chance, giving an expected profit of half times one minus half times one, which is zero. Item (e) pays three, giving plus 0.50 — and no commercial operator offers it.
Two truths and a lie
All three concern expected value and money.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. On any given play of the five-digit game you either lose two dollars or gain a hundred thousand — you never lose the expected one dollar. The expected value describes the long run and says nothing about the next play, which is exactly why the games remain attractive to play once.
Socratic
The expected profit is about minus one dollar per play.
Discussion prompt
If the expected value is negative, why does anyone play? Give a reason that does not require the player to be mistaken about the arithmetic.
Hint: Ask whether money is the only thing being bought.
Answer:
One honest reason is that the player is not buying an expected value but a small chance of a life-changing sum, which cannot be obtained any other way. Two dollars for a one-in-a-hundred-thousand shot at a hundred thousand is a trade some people rationally prefer to keeping the two dollars.
Another is that the entertainment is worth something. If the experience of playing is worth a dollar to the player, the game is roughly break-even in the terms that matter to them.
What the expected value does establish is what the money side of the transaction costs — about a dollar a play — so a player who believes they will come out ahead financially is mistaken, whatever else the game may be worth. Separating the financial question from the others is what the calculation is for.
Section
Section 4
Concept
The Law of Large Numbers states that as the number of trials in a probability experiment increases, the difference between the theoretical probability of an event and the relative frequency approaches zero. Applied to an average, it says the observed mean of many trials approaches the expected value.
the Law of Large Numbers — As the number of trials increases, the observed relative frequency approaches the theoretical probability, and the observed average approaches the expected value. It is what makes an expected value a claim about the world rather than only about a table.
\[ \text{observed average of } n \text{ trials} \;\longrightarrow\; \mu \quad \text{as } n \text{ grows} \]
The book illustrates it with Karl Pearson's 24,000 coin tosses, the same experiment section 1.1 and section 3.1 both cited — it recurs because it is the cleanest demonstration available. Without this law an expected value would be a property of a table and nothing more. With it, the table makes a checkable prediction about long runs, which is what allows insurers, casinos and manufacturers to plan on expected values while being unable to predict any individual case.
Figure (svg): A comparison showing that a single play of the game yields either a loss of two dollars or a profit of one hundred thousand, while the long-run average is a loss of about one dollar
OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-231 — the Law of Large Numbers and the expected value
Picture it
What happens once, and what happens over many plays.
Figure (svg): A comparison showing that a single play of the game yields either a loss of two dollars or a profit of one hundred thousand, while the long-run average is a loss of about one dollar
The asymmetry between the two panels is the whole content of the law. The left panel is unpredictable and always will be; the right panel becomes more predictable the more plays there are. An operator running the game a million times faces essentially no uncertainty about the total, which is why the business exists at all.
Worked example
The expected value is 1.1 days per week for the soccer team.
\[ \text{Is the team certain to average } 1.1 \text{ days over a season?} \]
What one week gives
Why: Zero, one or two days.
\[ \text{never } 1.1 \]
What ten weeks might give
Why: An average anywhere from 0 to 2.
What a thousand weeks gives
Why: An average very close to 1.1.
State the promise carefully
Why: Approaches, not equals.
Figure (svg): The solution to Worked example what the law does and does not say shown as a ladder of expressions, one row per legal move
\[ \text{observed average} \to 1.1 \text{, in the long run} \]
Verify: confirm the law is about the average and not about individual weeks
Why: Nothing in the law makes any particular week more likely to be near 1.1 — a week is still zero, one or two days with the same probabilities as ever. What settles down is the AVERAGE, because dividing by a growing number of weeks dilutes any individual week's influence. This is the same distinction section 3.1 drew when refuting the gambler's fallacy: the long run is achieved by dilution, not by correction.
OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-231
Prediction
Commit before reasoning.
Predict first
The expected profit of a game is minus one dollar. After how many plays is a player's average loss reliably close to a dollar?
Correct: Only after very many plays.
Why: For this game in particular the convergence is slow, because almost every play loses exactly two dollars and the average is dragged to minus one only by the rare enormous win. A player could play ten thousand times and average minus two, having never won. The law promises convergence in the long run and says nothing about how long, which for a highly skewed distribution can be very long indeed.
Worked example
Why an insurer can plan on an expected value it can never observe.
\[ \text{A policy has an expected annual claim of } 200 \text{ dollars.} \]
What one policy does
Why: Most years nothing, occasionally a large claim.
What the expected value says
Why: The long-run average per policy-year.
\[ 200\text{ dollars} \]
What a million policies do
Why: The average claim per policy settles near 200.
What the insurer must charge
Why: More than 200, to cover costs and risk.
Figure (svg): The solution to Worked example the law behind an insurance business shown as a ladder of expressions, one row per legal move
\[ \text{total claims} \approx 1\,000\,000 \times 200 \text{ dollars} \]
Verify: confirm what would break the argument
Why: The law assumes the policies are independent, which section 3.2 warned is exactly what fails under a common cause. A flood claiming on a hundred thousand policies at once destroys the averaging, because the outcomes are no longer independent draws from the distribution. That is why insurers reinsure catastrophe risk separately and why the independence assumption behind an expected value calculation is worth naming explicitly.
OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-231
Trap
\[ \mu = 2.1 \text{ wakings per week} \]
Predict that this week the baby will wake its mother about twice
Why: The word 'expected' invites a prediction about the next trial.
\[ \text{a reasonable guess, but not what the number says} \]
The distribution says 0.46 for two wakings and 0.22 for one, so the single most likely outcome is two — which happens to be near mu here, and need not be.
\[ \mu \text{ is the average over many weeks, not a prediction for one} \]
Read 'expected value' as 'long-run average' every time
Why: The word is a technical term and not the ordinary English one.
The two come apart most sharply when mu is impossible, as with the soccer team's 1.1 days or an expected 2.4 children. They also come apart on a skewed distribution: a lottery's expected winnings are a few pence and the overwhelmingly most likely outcome is nothing at all. If a prediction for a single trial is wanted, the mode is the relevant summary and the distribution itself is better than either.
Two truths and a lie
All three concern the Law of Large Numbers.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The law is about approach rather than arrival: the observed average gets arbitrarily close and is never guaranteed to hit mu exactly, particularly since mu is often not even an achievable average of whole-number outcomes. Section 3.1's version of the law carried the same hedge.
Fill the middle
What happens as the number of trials increases.
Fill in the blanks
\textexpected ___ \text___
Why: The expected value, computed from the distribution. The law is the bridge between the theoretical table and observable long runs, and without it a mu computed on paper would be a fact about the paper.
Explain it
A classmate objects that calling 1.1 days the 'expected' value is absurd, since the team never plays 1.1 days.
Discussion prompt
In three sentences or fewer, agree with the objection and rescue the concept.
Hint: Concede the word and defend the quantity.
Answer:
Tell them they are right about the word: 'expected' is a poor name, because nobody should expect 1.1 days in any week and the value is often impossible.
What the number means is the long-run average — play a thousand weeks, total the days, divide by a thousand, and you get very close to 1.1.
Reading it as 'long-run average value' every time removes the absurdity and leaves a useful quantity, which is exactly what an insurer or a casino plans on.
Section
Section 5
Concept
The expected value and standard deviation of a distribution are the theoretical counterparts of the mean and standard deviation of a data set. The arithmetic is the same weighted average in both cases; what differs is whether the weights are observed relative frequencies or theoretical probabilities.
parameter against statistic, again — Mu and sigma computed from a distribution are parameters: fixed properties of the process. X-bar and s computed from data are statistics: they vary from sample to sample. Section 1.1's distinction, now with the parameters actually computable because the distribution is known.
\[ \bar{x} = \sum x \cdot \frac{f}{n} \qquad \text{against} \qquad \mu = \sum x \cdot P(x) \]
Section 1.1 introduced parameters as fixed numbers that are almost never known. Here they are known, because the distribution is given rather than estimated — which is what a probability model buys. Chapters 7 and 8 close the circle by using a known distribution to say how far a computed x-bar is likely to fall from an unknown mu, and that argument needs both halves of this table.
Figure (svg): Two columns contrasting the mean and standard deviation of a data set with the expected value and standard deviation of a distribution
OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-232 — mu and sigma as the distribution's mean and standard deviation
Picture it
Where the weights come from is the whole difference.
Figure (svg): Two columns contrasting the mean and standard deviation of a data set with the expected value and standard deviation of a distribution
The last row is the one that matters for the rest of the book. A statistic changes when you collect more data; a parameter does not change at all, because it is a property of the process rather than of any sample. That is why chapter 8 estimates mu from x-bar and never the reverse, and why the notation has kept them apart since section 1.1.
Worked example
A frequency table and a distribution with identical weights.
\[ \text{50 weeks observed: } 2, 11, 23, 9, 4, 1 \text{ weeks with } 0 \text{ to } 5 \text{ wakings} \]
Compute the relative frequencies
Why: Each count over fifty.
\[ 0.04, 0.22, 0.46, 0.18, 0.08, 0.02 \]
Compute the data's mean
Why: Weighted by relative frequency.
\[ x - b a r = 2.1 \]
Compare with the distribution's mu
Why: Weighted by probability.
\[ \mu = 2.1 \]
Say why they agree
Why: The weights are numerically identical.
Figure (svg): The solution to Worked example the same numbers, two ways shown as a ladder of expressions, one row per legal move
\[ \bar{x} = 2.1 = \mu, \text{ in this case only} \]
Verify: confirm they would not generally agree
Why: The agreement here is engineered: the observed relative frequencies happen to equal the probabilities exactly. Observe another fifty weeks and the counts will differ, x-bar will move, and mu will not — because mu is a property of the process and x-bar is a property of a sample. That divergence is sampling variability, and measuring it is what chapters 7 and 8 do.
OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-232
Sorting
Ask whether the number came from data or from a model.
Sort into buckets
Sort each quantity.
Item (e) is the one the Law of Large Numbers connects to item (b): the recorded average is a statistic that approaches the parameter as the plays accumulate. That connection is what makes a model checkable against data.
Worked example
Three standard deviation formulas, and when each applies.
\[ \text{data as a sample; data as a population; a distribution} \]
Data treated as a sample
Why: Squared deviations over n minus 1.
Data treated as a population
Why: Squared deviations over N.
A probability distribution
Why: Squared deviations weighted by P(x).
Note what unifies them
Why: All three are averages of squared deviations.
Figure (svg): The solution to Worked example which divisor, and why shown as a ladder of expressions, one row per legal move
\[ \frac{1}{n-1}, \qquad \frac{1}{N}, \qquad P(x) \]
Verify: confirm the distribution case matches the population case for equally likely values
Why: Give k values each probability one over k, and the distribution formula becomes the sum of squared deviations divided by k — the population formula exactly. That is the book's own remark that the formulas coincide when all outcomes are equally likely, and it explains why there is no n minus 1 here: the correction in the sample formula compensates for estimating a mean from data, and nothing is being estimated when the distribution is given.
OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, p. 230
Trap
\[ \sigma = \sqrt{\frac{\sum (x - \mu)^2 P(x)}{n - 1}} \]
Divide by n minus 1, as section 2.7 did
Why: Every standard deviation so far has had that divisor.
\[ \text{but there is no sample here, and no n} \]
The n minus 1 corrects for estimating a mean from data. Here mu is given, and the probabilities have already averaged the terms.
\[ \sigma = \sqrt{\sum (x - \mu)^2 P(x)} \]
Ask what the weights sum to before adding any divisor
Why: Weights summing to one have already done the averaging.
The question 'what am I averaging over' settles it every time. Over n observations, divide by n; over n observations with a mean estimated from them, divide by n minus 1; over a distribution whose probabilities sum to one, divide by nothing. The three formulas look different and are the same operation with three sources of weight.
Matching
Four averages of squared deviations.
Match the pairs
Why: The fourth row is what unifies them and is worth holding rather than the three separately: every standard deviation in this book averages squared deviations and takes a root, and the only question is where the weights come from. Answering that question settles the divisor without memorising three formulas.
Two truths and a lie
All three concern the distinction.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and it is the core of the distinction. Mu is computed from the probabilities and does not move whatever data arrive. What moves is x-bar, the observed average, which the Law of Large Numbers says drifts toward the unmoving mu.
Prediction
Commit before reasoning.
Predict first
The distribution formula for sigma has no denominator. Why not?
Correct: The probabilities already sum to one.
Why: Averaging means weighting by numbers that total one, and that is exactly what the probabilities do. Section 2.7's divisors existed because its weights were raw counts that had to be normalised. The absence of a sample size is a symptom rather than the reason — a distribution with equally likely values effectively divides by the number of values, and it does so through the probabilities.
Comparison
Fill the blanks. The arithmetic is the same; the weights are not.
Comparison matrix
| Quantity | For data (chapter 2) | For a distribution (chapter 4) |
|---|---|---|
| Centre | x-bar, weighted by relative frequency | mu, weighted by probability |
| Spread | s, dividing by n minus 1 | sigma, weighted by P(x), no divisor |
| What it is | a statistic: it varies with the sample | a parameter: fixed by the process |
| What connects them | the Law of Large Numbers | the observed average approaches mu |
The bottom row is the whole point of having both columns. A distribution says what a process does; data say what it did; and the law is the promise that the second approaches the first, which is what makes a model testable and what chapters 7 to 13 are built on.
Pattern
Six steps, and the fourth cannot begin before the third is finished.
For a money problem, define X as the PROFIT before anything else, and list its values as payoffs rather than as the underlying outcomes. The experiment supplies the probabilities; the payoffs are what is being averaged.
OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-236
Check
Compute an expected value.
Check your understanding
A variable takes the values 0, 1 and 2 with probabilities 0.2, 0.5 and 0.3. What is mu?
Answer: A
Why: Multiplying and adding gives 0 plus 0.5 plus 0.6, which is 1.1.
Check
Finish a sigma.
Check your understanding
For a distribution, the squared deviations weighted by their probabilities total 1.05. What is sigma?
Answer: A
Why: The weighted total is already the variance, so sigma is its square root, about 1.0247.
Check
Interpret the expected value.
Check your understanding
A game has an expected profit of minus one dollar. What does that mean?
Answer: A
Why: An expected value is a long-run average, so it describes what happens across many plays rather than on any single one.
Real world
A retailer offers a two-year extended warranty for 80 dollars on a device costing 400. Their records show that 6 percent of devices fail within two years and that a failure costs them 400 dollars to replace. A customer asks whether the warranty is worth buying.
Discussion prompt
Compute the expected value from both sides, say what the warranty costs the customer on average, and give a reason a customer might buy it anyway.
Hint: Build one expected value table for the retailer and one for the customer.
Answer:
For the retailer, the expected payout is 24 dollars. A failure costs 400 with probability 0.06 and nothing with probability 0.94, giving 400 times 0.06, which is 24. They charge 80, so the expected profit per warranty is 56 dollars.
For the customer, the warranty costs 56 dollars on average. They pay 80 for certain and receive an expected 24 back, so buying it has an expected value of minus 56 — the mirror image of the retailer's gain, as it must be.
A customer may still buy it rationally. They are not buying an expected value but the removal of a 400-dollar risk they may be unable to absorb. If losing 400 unexpectedly would cause real difficulty and 80 would not, paying 56 in expectation to convert an uncertain large loss into a certain small one can be a sensible trade.
\[ \text{expected payout} = (400)(0.06) + (0)(0.94) = 24 \text{ dollars} \]
This is the structure of all insurance: the seller profits on the expected value and the buyer purchases certainty. What the calculation establishes is the price of that certainty, which is 56 dollars here — and knowing it is what lets a customer decide whether the certainty is worth that much to them. The one thing the arithmetic rules out is the belief that the warranty is financially advantageous on average, which it cannot be, since the retailer would not otherwise offer it.
Commit first
Answer, then rate your confidence honestly.
Predict first
Why does the standard deviation of a distribution have no divisor?
Correct: Because the probabilities already sum to one.
\[ \sum P(x) = 1 \;\Longrightarrow\; \sum (x-\mu)^2 P(x) \text{ is already an average} \]
Why: Averaging means weighting by numbers totalling one, and that is exactly what a probability distribution supplies. Section 2.7 needed a divisor because its weights were raw counts requiring normalisation. Being a parameter is true but is not the reason, and discreteness is irrelevant — chapter 5's continuous distributions integrate against a density that also integrates to one, for the same reason.
Explain it
They say an expected value of 2.4 children per family is nonsense because nobody has 0.4 of a child.
Discussion prompt
In three sentences or fewer, agree with the observation and explain what the number is for.
Hint: Concede the impossibility and shift to what is being averaged.
Answer:
Tell them they are right that no family has 2.4 children, and that the word 'expected' is a poor name for the quantity.
What it means is that if you total the children across a very large number of families and divide by the number of families, you get about 2.4.
That is exactly the number a government needs for planning school places, even though it describes no family at all — which is why the quantity is worth having despite the name.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the table, check the P(x) column totals one before computing anything from it. For sigma, ask what the weights sum to — if one, no divisor. For a money problem, define X as the profit in a full sentence before listing values. For interpretation, read 'expected value' as 'long-run average' every time. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top, build the complete four-column table for this distribution: x from 0 to 5 with probabilities 0.04, 0.22, 0.46, 0.18, 0.08 and 0.02. Check the probability column totals one, compute mu from the third column, then use that mu to fill the fourth column and compute sigma. Write both answers with their units. Beside the table, draw the distribution as stems and mark mu with a triangle underneath, noting in one sentence why the mark falls where no stem does. In the middle, build the expected value table for the five-digit game: pay 2 dollars, profit 100,000 with probability 0.00001. Write the definition of X as a full sentence first, then the two values, then the two products, then the total. Below it, solve for the prize that would make the game fair and compare it with the prize offered. At the bottom, make a two-column table headed 'data' and 'distribution' with four rows: what the weights are, what the symbols are, what divisor is used, and whether the quantity varies when you collect more observations.
Check your sigma against the range: the values span 0 to 5 and your sigma should be around one, roughly a fifth of the range. If it came out near 0.42 or 0.46 you divided by the number of values, which the probabilities had already done.
Recap
Five things, and the last one is what keeps the first four honest.
| If you see | Then |
|---|---|
| A distribution and a request for a mean | Sum of x times P(x) |
| An expected value outside the range of values | An arithmetic error |
| A request for sigma from a distribution | Weight squared deviations by P(x); no divisor |
| A weighted total reported as sigma | That is the variance; take the root |
| A money problem | Let X be the profit, not the underlying outcome |
| A negative expected profit | Unfavourable over many plays; says nothing about one |
| An impossible expected value | Normal: it is a long-run average |
Section 4.3 gives the first named distribution. When a fixed number of independent trials each succeed with the same probability, the count of successes has a binomial distribution — and its expected value and standard deviation come from formulas rather than from a table, which is what makes named families worth having.
OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-236 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.