4.2 Mean or Expected Value and Standard Deviation

The centre and spread of a distribution rather than of a data set. The expected value is the long-term average of a random variable, found by multiplying each value by its probability and adding the products, and it is denoted by the Greek letter mu because it is a parameter of a process rather than a statistic computed from data. It need not be one of the values the variable can take. The standard deviation is the square root of the sum of squared deviations weighted by their probabilities, with no divisor because the probabilities already sum to one. Both are computed from a single expected value table, and the Law of Large Numbers is what connects them to what actually happens over many trials.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 4.2 Mean or Expected Value and Standard Deviation

Title

Statistics · Chapter 4 — Discrete Random Variables

Mean or Expected Value and Standard Deviation

2. By the end of this lesson you can

Objectives

Five outcomes. The first is one line of arithmetic and the fifth is what it does and does not promise.

OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-236 — the section these objectives are drawn from

3. What you already have

Warm-up

Section 2.5 computed a mean from a frequency table by weighting each value by how often it occurred and dividing by the total count.

Discussion prompt

A soccer team plays zero days a week 20 percent of the time, one day 50 percent, and two days 30 percent. What is the average number of days per week, and what changed from the section 2.5 calculation?

Hint: Try the weighted mean with the percentages as the weights.

Answer:

Multiply and add: zero times 0.2, plus one times 0.5, plus two times 0.3, giving 0 plus 0.5 plus 0.6, which is 1.1 days per week.

What changed is where the weights came from. In section 2.5 they were counts of observations divided by n; here they are probabilities, and no division is needed because probabilities already sum to one.

So the arithmetic is identical and the meaning is different: that 1.1 is not the average of any data collected but a property of the process itself. The book calls it the expected value and denotes it mu, the same Greek letter section 1.1 reserved for a population parameter.

4. The long-term average of a random variable

Concept

The expected value is often referred to as the long-term average or mean: over the long term of doing an experiment over and over, you would expect this average. To find it, multiply each value of the random variable by its probability and add the products. It is denoted by the Greek letter mu.

expected value — The long-term average of a random variable, denoted mu. Computed as the sum over all values of x times P(x). It is the mean of the distribution, and it need not be one of the values the variable can take.

\[ \mu = \sum x \cdot P(x) \]

The name is unfortunate and worth disarming immediately: the expected value is very often a value you should not expect at all. The soccer team's 1.1 days is impossible in any single week, since the team plays a whole number of days. What the number describes is the average over many weeks, which the Law of Large Numbers says the observed average will approach. Reading 'expected' as 'long-run average' rather than 'likely' removes almost all the confusion the term causes.

Figure (svg): An expected value table for the soccer team with columns for the value, its probability and their product, summing to 1.1

The soccer team would on average expect to play 1.1 days per week over the long term.

OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-231

5. Computing an expected value

Section

Section 1

6. Multiply each value by its probability and add

Concept

To find the expected value or long-term average, simply multiply each value of the random variable by its probability and add the products. Constructing a table with a third column for x times P(x) organises the work.

the expected value table — A PDF table with a third column holding x times P(x) for each row. The total of that column is the expected value. The book calls it an expected value table and uses it throughout the chapter.

\[ \mu = \sum x P(x) = x_1P(x_1) + x_2P(x_2) + \cdots \]

The table earns its keep by keeping two checks visible at once. The P(x) column must total one, which is section 4.1's second condition and confirms the distribution before anything is computed from it; and the x times P(x) column totals the expected value. Doing both in one table means an invalid distribution is caught before its mean is computed and reported, which is the order that matters.

Figure (svg): An expected value table for the soccer team with columns for the value, its probability and their product, summing to 1.1

The soccer team would on average expect to play 1.1 days per week over the long term.

OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-231 — the expected value and its table

7. Three rows and two totals

Picture it

The soccer team's distribution with the product column added.

Figure (svg): An expected value table for the soccer team with columns for the value, its probability and their product, summing to 1.1

The soccer team would on average expect to play 1.1 days per week over the long term.

The value 1.1 is the sum of the third column, and it lies between the smallest and largest possible values as any weighted average must. Notice it sits nearer 1 than 2 because the value 1 carries the most probability — the expected value is pulled toward wherever the probability is concentrated, exactly as section 2.5's mean was pulled toward wherever the data were.

8. Worked example: the soccer team

Worked example

Example 4.3, with the table the book builds.

\[ P(0) = 0.2, \quad P(1) = 0.5, \quad P(2) = 0.3 \]

Check the distribution first

Why: The three probabilities.

\[ 0.2 + 0.5 + 0.3 = 1 \]

Multiply each value by its probability

Why: Zero, one and two in turn.

\[ 0, 0.5, 0.6 \]

Add the products

Why: The third column's total.

\[ 0 + 0.5 + 0.6 \]

Read the expected value

Why: The sum.

\[ 1.1 \]

Figure (svg): The solution to Worked example the soccer team shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \mu = (0)(0.2) + (1)(0.5) + (2)(0.3) = 1.1 \]

Verify: confirm the answer lies inside the range of values

Why: The team plays between zero and two days, and 1.1 falls between them — which any weighted average must, since it is a mixture of the values themselves. An expected value outside the range of possible values is proof of an arithmetic error, and it is the fastest check available. Here it also sits just above one, which matches the distribution's concentration on that value.

OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, p. 231

9. Build the product column

Faded example

The soccer team: values 0, 1, 2 with probabilities 0.2, 0.5, 0.3.

Fill in the blanks

\mu = (0)(0.2) + (1)(0.5) + (2)(0.3) = 1.1

Why: The three products are 0, 0.5 and 0.6, totalling 1.1. Notice the first term contributes nothing because the value is zero — a value of zero always drops out of an expected value however likely it is, which is worth remembering when a distribution has many zeros.

10. Worked example: the night-waking distribution

Worked example

Example 4.4, a longer table with the same procedure.

\[ P = 0.04, \;0.22, \;0.46, \;0.18, \;0.08, \;0.02 \text{ for } x = 0 \text{ through } 5 \]

Verify the distribution

Why: The six probabilities.

\[ \text{they total } 1.00 \]

Form each product

Why: Value times probability.

\[ 0, 0.22, 0.92, 0.54, 0.32, 0.10 \]

Add them

Why: The expected value.

\[ 2.10 \]

Interpret

Why: The long-run weekly average.

\[ 2.1\text{ wakings per week} \]

Figure (svg): The solution to Worked example the night-waking distribution shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \mu = 0(0.04) + 1(0.22) + 2(0.46) + 3(0.18) + 4(0.08) + 5(0.02) = 2.1 \]

Verify: confirm the expected value is not a possible outcome

Why: A baby wakes its mother a whole number of times in a week, so 2.1 will never be observed on any single week. The book's own caption puts it as expecting a newborn to wake its parents 2.1 times per week ON THE AVERAGE, and the qualifier is doing all the work. The number is a property of the long run, and treating it as a prediction for next week is the commonest misreading of the whole chapter.

OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 231-232

11. Trap: averaging the values instead of weighting them

Trap

The trap

\[ x = 0, 1, 2 \text{, so the average is } \frac{0 + 1 + 2}{3} \]

Average the three possible values

Why: Three numbers are listed, so their mean looks like the answer.

\[ = 1 \quad \text{(wrong: it ignores the probabilities)} \]

That treats playing zero days and playing one day as equally likely, when one day happens two and a half times as often.

The fix

\[ \mu = (0)(0.2) + (1)(0.5) + (2)(0.3) = 1.1 \]

Weight every value by its own probability

Why: The probabilities are what distinguish the values.

This is section 2.5's averaging-the-averages error in a new setting, and it gives the right answer only when every value is equally likely — which is why it survives on dice and coins and fails everywhere else. On the night-waking distribution the unweighted average of 0 through 5 is 2.5 against a true expected value of 2.1, and the gap grows with the asymmetry of the distribution.

12. Where must the answer lie?

Estimation

A variable takes the values 3, 7 and 12 with some probabilities.

Predict first

What range must its expected value lie in?

  • Between 0 and 12
  • Between 3 and 12
  • Exactly 7
  • Between 7 and 12

Correct: Between 3 and 12.

Why: An expected value is a weighted average of the values, so it lies between the smallest and largest of them — reaching an endpoint only when all the probability sits on that value. It need not be near the middle value, and it need not be 7: a distribution concentrated on 3 gives an expected value just above 3. This range check catches most arithmetic slips instantly.

13. One of these is false

Two truths and a lie

All three concern the expected value.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The expected value need not be a value the variable can take
  • C. The expected value lies between the smallest and largest possible values
  • B. The expected value is the most likely value

Survives elimination: B

Why: The survivor is false: the most likely value is the mode. For the night-waking distribution the mode is 2, with probability 0.46, while the expected value is 2.1 — close here, but on a skewed distribution they can be far apart. The soccer team's mode is 1 and its expected value 1.1, and a distribution with a long tail separates them much further.

14. What does a zero value contribute?

Prediction

Commit before reasoning.

Predict first

In an expected value calculation, what does the row for the value zero contribute?

  • Nothing, whatever its probability
  • Its probability
  • One minus its probability
  • It depends on the other values

Correct: Nothing, whatever its probability.

Why: The product is zero times its probability, which is zero however likely the value is. The row still matters for checking that the probabilities sum to one, but it adds nothing to the mean — so a distribution with a great deal of probability on zero has an expected value dragged down not by that row's contribution but by the smaller weights left for the others.

15. The standard deviation of a distribution

Section

Section 2

16. The same table, one column further

Concept

The standard deviation of a probability distribution is the square root of its variance, and the variance is the sum of the squared deviations weighted by their probabilities. For each value, multiply the square of its deviation from mu by its probability, then add and take the root.

the standard deviation of a distribution — Sigma, the square root of the sum over all values of the squared deviation times the probability. It measures how far the variable's values fall from mu, exactly as section 2.7's s measured how far data fell from x-bar.

\[ \sigma = \sqrt{\sum (x - \mu)^2 P(x)} \]

Section 2.7's computation is recognisable here with one change: there the squared deviations were divided by n minus 1, and here they are weighted by probabilities and not divided at all. The reason is that the probabilities already sum to one, so the weighted sum IS an average. There is no sample and so no n, and no correction for estimating a mean from data, because mu is known rather than estimated.

Figure (svg): A four-column table computing the expected value and standard deviation for the night-waking distribution

One table gives both. The fourth column cannot be started until the third has been totalled, since it needs mu.

OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-232 — the standard deviation of a PDF

17. One table, both answers

Picture it

Example 4.4 with all four columns.

Figure (svg): A four-column table computing the expected value and standard deviation for the night-waking distribution

One table gives both. The fourth column cannot be started until the third has been totalled, since it needs mu.

The fourth column cannot be started until the third has been totalled, because every entry needs mu. That dependency is the only thing making this a two-pass calculation, and it is why the book computes the expected value first and then says to use mu to complete the table.

18. Worked example: sigma for the night wakings

Worked example

Example 4.4. The fourth column, then a square root.

\[ \mu = 2.1, \quad P = 0.04, 0.22, 0.46, 0.18, 0.08, 0.02 \]

Form each squared deviation

Why: Value minus mu, squared.

\[ 4.41, 1.21, 0.01, 0.81, 3.61, 8.41 \]

Weight each by its probability

Why: The fourth column.

\[ 0.1764, 0.2662, 0.0046, 0.1458, 0.2888, 0.1682 \]

Add them for the variance

Why: The column total.

\[ 1.05 \]

Take the square root

Why: Back to the variable's units.

\[ \text{about } 1.0247 \]

Figure (svg): The solution to Worked example sigma for the night wakings shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \sigma = \sqrt{1.05} \approx 1.0247 \]

Verify: confirm sigma is a plausible size for this distribution

Why: The values run from 0 to 5, a range of five, and sigma is about one — roughly a fifth of the range, which is the usual order of magnitude. A sigma larger than the range would be impossible, and one very much smaller would suggest the probability is concentrated on a single value. Notice also which rows contribute most: the value 2 sits almost exactly at mu and contributes only 0.0046, while the value 5 contributes 0.1682 despite having a probability of just 0.02, because its deviation is squared.

OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, p. 232

19. Finish the calculation

Faded example

The weighted squared deviations for the night-waking distribution total 1.05.

Fill in the blanks

\sigma = \sqrt1.05} \approx 1.0247

Why: No divisor appears, because the probabilities weighting the terms already sum to one. The square root returns the answer to the variable's own units — wakings per week rather than squared wakings.

20. Worked example: why there is no divisor

Worked example

Comparing the formula with section 2.7's.

\[ s = \sqrt{\frac{\sum (x - \bar{x})^2}{n-1}} \quad \text{against} \quad \sigma = \sqrt{\sum (x - \mu)^2 P(x)} \]

Note what section 2.7 divided by

Why: The number of observations, less one.

\[ n - 1 \]

Note what weights the terms here

Why: Each term carries its own probability.

\[ P(x) \]

Ask what the probabilities sum to

Why: The second condition of section 4.1.

Conclude

Why: Weighting by numbers summing to one IS averaging.

Figure (svg): The solution to Worked example why there is no divisor shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \sum P(x) = 1 \;\Longrightarrow\; \sum (x-\mu)^2 P(x) \text{ is already an average} \]

Verify: confirm the two agree when the values are equally likely

Why: Take three values each with probability one third. The distribution formula gives the sum of squared deviations times one third, which is the sum divided by three — exactly the POPULATION standard deviation of those three numbers. The book notes this: when all outcomes are equally likely, these formulas coincide with the mean and standard deviation of the set of possible outcomes. Note it matches the population form with N rather than the sample form with n minus 1, which is right, because mu here is known rather than estimated from data.

OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, p. 230

21. Error analysis: four attempts at sigma

Error analysis

For a distribution with mu = 2.1, the weighted squared deviations total 1.05. Four students report sigma.

Annotate

On: \( \begin{aligned} &(1)\; \sigma = 1.05 \\ &(2)\; \sigma = \sqrt{\tfrac{1.05}{6}} \approx 0.42 \\ &(3)\; \sigma = \sqrt{\tfrac{1.05}{5}} \approx 0.46 \\ &(4)\; \sigma = \sqrt{1.05} \approx 1.0247 \end{aligned} \)

  • (1) reports the variance and calls it the standard deviation. The two differ by a square root, and only sigma is in the variable's own units.
  • (2) divides by the number of values, importing section 2.7's population divisor. The probabilities have already done that averaging, so dividing again shrinks sigma wrongly.
  • (3) divides by one less than the number of values, importing the sample divisor. There is no sample here at all — this is a distribution, not data.
  • (4) is correct: the weighted sum is already the variance, and sigma is its square root.

Errors (2) and (3) are the same instinct: every standard deviation so far has had a divisor, so one is supplied from habit. The guard is to ask what the weights sum to — when they sum to one, the averaging is done.

22. Which rows dominate sigma?

Discrimination

mu is 2.1, and each row contributes its squared deviation times its probability.

Sort into buckets

Sort each row of the night-waking distribution by its contribution.

Contributes a lot despite low probability
x = 5, P = 0.02, contributing 0.1682; x = 0, P = 0.04, contributing 0.1764
Contributes little despite high probability
x = 2, P = 0.46, contributing 0.0046
Middling on both counts
x = 3, P = 0.18, contributing 0.1458
big
The value sits far from mu, and squaring that distance outweighs a small probability.
small
The value sits almost exactly at mu, so its squared deviation is tiny however likely it is.
mid
A moderate deviation with a moderate probability, contributing an unremarkable amount.

23. One of these is false

Two truths and a lie

All three concern sigma for a distribution.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The formula has no divisor because the probabilities sum to one
  • C. For equally likely values the formula agrees with the population standard deviation of those values
  • B. Sigma is the sum of the squared deviations weighted by their probabilities

Survives elimination: B

Why: The survivor describes the VARIANCE rather than the standard deviation. Sigma is its square root, which is what returns the answer to the variable's own units. Reporting 1.05 rather than 1.0247 here is the commonest slip, and it is the same one section 2.7 warned about for data.

24. What does a small sigma mean?

Prediction

Commit before reasoning.

Predict first

Two distributions have the same expected value, but one has a much smaller sigma. What does that say?

  • Its values are more tightly concentrated near mu
  • It has fewer possible values
  • Its probabilities are smaller
  • Its expected value is more accurate

Correct: Its values are more tightly concentrated near mu.

Why: Sigma measures how far the variable's values fall from its mean, exactly as section 2.7's s did for data. A small sigma means the probability sits close to mu, so an individual outcome is likely to be near the long-run average. The number of possible values is irrelevant — a distribution with many values packed near mu can have a smaller sigma than one with three values spread widely.

25. Expected value and decisions

Section

Section 3

26. What a gamble is worth per play

Concept

An expected value table can be built for money as easily as for counts. Let X be the amount of profit, list the possible profits with their probabilities, and the expected value is the average gain or loss per play over the long term.

expected profit — The expected value of a random variable representing money gained. A negative expected profit means the game loses money on average over many plays, however dramatic the individual outcomes.

\[ \mu = \sum (\text{profit}) \cdot P(\text{profit}) \]

The book's Example 4.5 is worth working slowly because the setup is where the difficulty lies. Five digits are chosen from zero to nine with replacement, so the chance of matching all five in order is a tenth to the fifth power, which is 0.00001. The values of X are the PROFITS — negative two dollars and one hundred thousand dollars — not the digits, and getting that right is most of the problem.

Figure (svg): A comparison showing that a single play of the game yields either a loss of two dollars or a profit of one hundred thousand, while the long-run average is a loss of about one dollar

Example 4.5: the expected value describes the long run and never the next play.

OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, p. 233 — Example 4.5, the five-digit game

27. One play against many

Picture it

Two possible outcomes, and one long-run average.

Figure (svg): A comparison showing that a single play of the game yields either a loss of two dollars or a profit of one hundred thousand, while the long-run average is a loss of about one dollar

Example 4.5: the expected value describes the long run and never the next play.

The expected loss of about a dollar is never what happens on a play: every play loses two dollars or gains a hundred thousand. What the number says is that a player doing this repeatedly loses about a dollar a go, and the Law of Large Numbers is what guarantees the observed average approaches it. That is exactly why the game is offered.

28. Worked example: the five-digit game

Worked example

Example 4.5. The values are profits, not digits.

\[ \text{Pay } 2 \text{ dollars; match five digits in order to profit } 100\,000. \]

Find the probability of winning

Why: One digit in ten, five times, with replacement.

Find the probability of losing

Why: The complement.

\[ 0.99999 \]

List the profits, not the digits

Why: Lose two dollars, or profit one hundred thousand.

\[ -2\text{ and } 100 000 \]

Build and total the table

Why: Each profit times its probability.

\[ -1.99998 + 1 \]

Figure (svg): The solution to Worked example the five-digit game shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \mu = (-2)(0.99999) + (100000)(0.00001) = -0.99998 \]

Verify: confirm the two terms are nearly balanced and which one wins

Why: The loss term contributes about minus two dollars and the win term contributes exactly one dollar, so the game returns about half of what it takes. The win term is exactly 1 because a hundred thousand times a hundred-thousandth is one — a useful check on the probability, since if the prize and the odds were fair the two terms would cancel. They do not, and the shortfall is the operator's margin.

OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, p. 233

29. The win term

Faded example

The prize is 100,000 dollars and the probability of winning is 0.00001.

Fill in the blanks

(100000)(0.00001) = 1, \quad \text-1.99998 ___

Why: The win contributes exactly one dollar per play and the loss about two, so the expected profit is about minus one. The neatness of the win term is a coincidence of the numbers chosen, and it makes the comparison unusually easy to see.

30. Worked example: what prize would make it fair

Worked example

Running the calculation backwards.

\[ \text{What prize would make the expected profit zero?} \]

Write the condition

Why: Expected profit equals zero.

\[ (-2) (0.99999) + W(0.00001) = 0 \]

Isolate the win term

Why: Move the loss across.

\[ W(0.00001) = 1.99998 \]

Divide

Why: Solve for the prize.

\[ W = 199 998 \]

Compare with the actual prize

Why: One hundred thousand offered.

Figure (svg): The solution to Worked example what prize would make it fair shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ W = \frac{1.99998}{0.00001} = 199\,998 \]

Verify: confirm the interpretation of a fair game

Why: A fair game has an expected profit of zero, so a player neither gains nor loses in the long run — and the operator makes nothing either. The prize offered is about half the fair value, which is why the expected profit is about minus one dollar on a two dollar stake. Every commercial game of chance is priced this way, and computing the fair prize is the standard way to see by how much.

OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, p. 233

31. Trap: using the outcomes rather than the payoffs

Trap

The trap

\[ x = 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 \]

Take the values of X to be the digits

Why: Those are the numbers in the problem.

\[ \text{an expected value of } 4.5 \text{ digits} \quad \text{(answering nothing)} \]

The question asks about profit, so the variable must be the money and its values are the two possible profits.

The fix

\[ X = \text{the amount of money you profit}; \quad x = -2 \text{ or } 100\,000 \]

Define X as the quantity the question is about, then find that quantity's values

Why: The digits generate the probabilities; the money is what is being averaged.

The book is explicit that the values of x are not the digits zero through nine but the two profits, and the distinction is the whole setup. It generalises: in any decision problem the random variable is the CONSEQUENCE, and the underlying experiment merely supplies the probabilities. Writing the definition of X as a full sentence before listing any values is what keeps the two apart.

32. Fair, favourable or unfavourable?

Sorting

A game is fair when the expected profit is zero.

Sort into buckets

Sort each game by its expected profit.

Unfavourable to the player
expected profit of -0.99998 dollars
Fair
expected profit of exactly 0; pay 1 dollar, win 2 dollars on a fair coin
Favourable to the player
expected profit of +0.50 dollars; pay 1 dollar, win 3 dollars on a fair coin
bad
The expected profit is negative, so a player loses money on average over many plays.
fair
The expected profit is zero, so neither the player nor the operator gains in the long run.
good
The expected profit is positive, so the player gains on average — which is why such games are rarely offered.

Item (d) pays two dollars for a one-dollar stake on a half chance, giving an expected profit of half times one minus half times one, which is zero. Item (e) pays three, giving plus 0.50 — and no commercial operator offers it.

33. One of these is false

Two truths and a lie

All three concern expected value and money.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. A negative expected profit means the game loses money on average
  • C. The values of X are the payoffs, not the underlying outcomes
  • B. A negative expected profit means you will lose on any given play

Survives elimination: B

Why: The survivor is false. On any given play of the five-digit game you either lose two dollars or gain a hundred thousand — you never lose the expected one dollar. The expected value describes the long run and says nothing about the next play, which is exactly why the games remain attractive to play once.

34. Why do people play?

Socratic

The expected profit is about minus one dollar per play.

Discussion prompt

If the expected value is negative, why does anyone play? Give a reason that does not require the player to be mistaken about the arithmetic.

Hint: Ask whether money is the only thing being bought.

Answer:

One honest reason is that the player is not buying an expected value but a small chance of a life-changing sum, which cannot be obtained any other way. Two dollars for a one-in-a-hundred-thousand shot at a hundred thousand is a trade some people rationally prefer to keeping the two dollars.

Another is that the entertainment is worth something. If the experience of playing is worth a dollar to the player, the game is roughly break-even in the terms that matter to them.

What the expected value does establish is what the money side of the transaction costs — about a dollar a play — so a player who believes they will come out ahead financially is mistaken, whatever else the game may be worth. Separating the financial question from the others is what the calculation is for.

35. The Law of Large Numbers

Section

Section 4

36. What connects the expected value to what happens

Concept

The Law of Large Numbers states that as the number of trials in a probability experiment increases, the difference between the theoretical probability of an event and the relative frequency approaches zero. Applied to an average, it says the observed mean of many trials approaches the expected value.

the Law of Large Numbers — As the number of trials increases, the observed relative frequency approaches the theoretical probability, and the observed average approaches the expected value. It is what makes an expected value a claim about the world rather than only about a table.

\[ \text{observed average of } n \text{ trials} \;\longrightarrow\; \mu \quad \text{as } n \text{ grows} \]

The book illustrates it with Karl Pearson's 24,000 coin tosses, the same experiment section 1.1 and section 3.1 both cited — it recurs because it is the cleanest demonstration available. Without this law an expected value would be a property of a table and nothing more. With it, the table makes a checkable prediction about long runs, which is what allows insurers, casinos and manufacturers to plan on expected values while being unable to predict any individual case.

Figure (svg): A comparison showing that a single play of the game yields either a loss of two dollars or a profit of one hundred thousand, while the long-run average is a loss of about one dollar

Example 4.5: the expected value describes the long run and never the next play.

OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-231 — the Law of Large Numbers and the expected value

37. The long run is the only claim

Picture it

What happens once, and what happens over many plays.

Figure (svg): A comparison showing that a single play of the game yields either a loss of two dollars or a profit of one hundred thousand, while the long-run average is a loss of about one dollar

Example 4.5: the expected value describes the long run and never the next play.

The asymmetry between the two panels is the whole content of the law. The left panel is unpredictable and always will be; the right panel becomes more predictable the more plays there are. An operator running the game a million times faces essentially no uncertainty about the total, which is why the business exists at all.

38. Worked example: what the law does and does not say

Worked example

The expected value is 1.1 days per week for the soccer team.

\[ \text{Is the team certain to average } 1.1 \text{ days over a season?} \]

What one week gives

Why: Zero, one or two days.

\[ \text{never } 1.1 \]

What ten weeks might give

Why: An average anywhere from 0 to 2.

What a thousand weeks gives

Why: An average very close to 1.1.

State the promise carefully

Why: Approaches, not equals.

Figure (svg): The solution to Worked example what the law does and does not say shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{observed average} \to 1.1 \text{, in the long run} \]

Verify: confirm the law is about the average and not about individual weeks

Why: Nothing in the law makes any particular week more likely to be near 1.1 — a week is still zero, one or two days with the same probabilities as ever. What settles down is the AVERAGE, because dividing by a growing number of weeks dilutes any individual week's influence. This is the same distinction section 3.1 drew when refuting the gambler's fallacy: the long run is achieved by dilution, not by correction.

OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-231

39. How many trials?

Prediction

Commit before reasoning.

Predict first

The expected profit of a game is minus one dollar. After how many plays is a player's average loss reliably close to a dollar?

  • After one play
  • After ten plays
  • Only after very many plays, and the more the closer
  • Never; the law says nothing about averages

Correct: Only after very many plays.

Why: For this game in particular the convergence is slow, because almost every play loses exactly two dollars and the average is dragged to minus one only by the rare enormous win. A player could play ten thousand times and average minus two, having never won. The law promises convergence in the long run and says nothing about how long, which for a highly skewed distribution can be very long indeed.

40. Worked example: the law behind an insurance business

Worked example

Why an insurer can plan on an expected value it can never observe.

\[ \text{A policy has an expected annual claim of } 200 \text{ dollars.} \]

What one policy does

Why: Most years nothing, occasionally a large claim.

What the expected value says

Why: The long-run average per policy-year.

\[ 200\text{ dollars} \]

What a million policies do

Why: The average claim per policy settles near 200.

What the insurer must charge

Why: More than 200, to cover costs and risk.

Figure (svg): The solution to Worked example the law behind an insurance business shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{total claims} \approx 1\,000\,000 \times 200 \text{ dollars} \]

Verify: confirm what would break the argument

Why: The law assumes the policies are independent, which section 3.2 warned is exactly what fails under a common cause. A flood claiming on a hundred thousand policies at once destroys the averaging, because the outcomes are no longer independent draws from the distribution. That is why insurers reinsure catastrophe risk separately and why the independence assumption behind an expected value calculation is worth naming explicitly.

OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-231

41. Trap: expecting the expected value

Trap

The trap

\[ \mu = 2.1 \text{ wakings per week} \]

Predict that this week the baby will wake its mother about twice

Why: The word 'expected' invites a prediction about the next trial.

\[ \text{a reasonable guess, but not what the number says} \]

The distribution says 0.46 for two wakings and 0.22 for one, so the single most likely outcome is two — which happens to be near mu here, and need not be.

The fix

\[ \mu \text{ is the average over many weeks, not a prediction for one} \]

Read 'expected value' as 'long-run average' every time

Why: The word is a technical term and not the ordinary English one.

The two come apart most sharply when mu is impossible, as with the soccer team's 1.1 days or an expected 2.4 children. They also come apart on a skewed distribution: a lottery's expected winnings are a few pence and the overwhelmingly most likely outcome is nothing at all. If a prediction for a single trial is wanted, the mode is the relevant summary and the distribution itself is better than either.

42. One of these is false

Two truths and a lie

All three concern the Law of Large Numbers.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. It says the observed average approaches the expected value as trials accumulate
  • C. It requires the trials to be independent draws from the same distribution
  • B. It guarantees that after enough trials the average will equal mu exactly

Survives elimination: B

Why: The survivor is false. The law is about approach rather than arrival: the observed average gets arbitrarily close and is never guaranteed to hit mu exactly, particularly since mu is often not even an achievable average of whole-number outcomes. Section 3.1's version of the law carried the same hedge.

43. State it

Fill the middle

What happens as the number of trials increases.

Fill in the blanks

\textexpected ___ \text___

Why: The expected value, computed from the distribution. The law is the bridge between the theoretical table and observable long runs, and without it a mu computed on paper would be a fact about the paper.

44. Explain the word

Explain it

A classmate objects that calling 1.1 days the 'expected' value is absurd, since the team never plays 1.1 days.

Discussion prompt

In three sentences or fewer, agree with the objection and rescue the concept.

Hint: Concede the word and defend the quantity.

Answer:

Tell them they are right about the word: 'expected' is a poor name, because nobody should expect 1.1 days in any week and the value is often impossible.

What the number means is the long-run average — play a thousand weeks, total the days, divide by a thousand, and you get very close to 1.1.

Reading it as 'long-run average value' every time removes the absurdity and leaves a useful quantity, which is exactly what an insurer or a casino plans on.

45. Distributions against data

Section

Section 5

46. The same arithmetic, different weights

Concept

The expected value and standard deviation of a distribution are the theoretical counterparts of the mean and standard deviation of a data set. The arithmetic is the same weighted average in both cases; what differs is whether the weights are observed relative frequencies or theoretical probabilities.

parameter against statistic, again — Mu and sigma computed from a distribution are parameters: fixed properties of the process. X-bar and s computed from data are statistics: they vary from sample to sample. Section 1.1's distinction, now with the parameters actually computable because the distribution is known.

\[ \bar{x} = \sum x \cdot \frac{f}{n} \qquad \text{against} \qquad \mu = \sum x \cdot P(x) \]

Section 1.1 introduced parameters as fixed numbers that are almost never known. Here they are known, because the distribution is given rather than estimated — which is what a probability model buys. Chapters 7 and 8 close the circle by using a known distribution to say how far a computed x-bar is likely to fall from an unknown mu, and that argument needs both halves of this table.

Figure (svg): Two columns contrasting the mean and standard deviation of a data set with the expected value and standard deviation of a distribution

The arithmetic is the same weighted average in both columns. What differs is where the weights come from, and that is the difference between a statistic and a parameter.

OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-232 — mu and sigma as the distribution's mean and standard deviation

47. Two columns, one arithmetic

Picture it

Where the weights come from is the whole difference.

Figure (svg): Two columns contrasting the mean and standard deviation of a data set with the expected value and standard deviation of a distribution

The arithmetic is the same weighted average in both columns. What differs is where the weights come from, and that is the difference between a statistic and a parameter.

The last row is the one that matters for the rest of the book. A statistic changes when you collect more data; a parameter does not change at all, because it is a property of the process rather than of any sample. That is why chapter 8 estimates mu from x-bar and never the reverse, and why the notation has kept them apart since section 1.1.

48. Worked example: the same numbers, two ways

Worked example

A frequency table and a distribution with identical weights.

\[ \text{50 weeks observed: } 2, 11, 23, 9, 4, 1 \text{ weeks with } 0 \text{ to } 5 \text{ wakings} \]

Compute the relative frequencies

Why: Each count over fifty.

\[ 0.04, 0.22, 0.46, 0.18, 0.08, 0.02 \]

Compute the data's mean

Why: Weighted by relative frequency.

\[ x - b a r = 2.1 \]

Compare with the distribution's mu

Why: Weighted by probability.

\[ \mu = 2.1 \]

Say why they agree

Why: The weights are numerically identical.

Figure (svg): The solution to Worked example the same numbers, two ways shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \bar{x} = 2.1 = \mu, \text{ in this case only} \]

Verify: confirm they would not generally agree

Why: The agreement here is engineered: the observed relative frequencies happen to equal the probabilities exactly. Observe another fifty weeks and the counts will differ, x-bar will move, and mu will not — because mu is a property of the process and x-bar is a property of a sample. That divergence is sampling variability, and measuring it is what chapters 7 and 8 do.

OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-232

49. Statistic or parameter?

Sorting

Ask whether the number came from data or from a model.

Sort into buckets

Sort each quantity.

A statistic: from data, and it varies
the mean of 50 observed weeks; s, computed from a sample; the average profit over 1,000 recorded plays
A parameter: from a model, and it is fixed
the expected value of a given distribution; sigma of a probability distribution
stat
The number is computed from observations, so collecting different observations would give a different value.
par
The number is computed from the distribution, so it is a property of the process and does not vary with any sample.

Item (e) is the one the Law of Large Numbers connects to item (b): the recorded average is a statistic that approaches the parameter as the plays accumulate. That connection is what makes a model checkable against data.

50. Worked example: which divisor, and why

Worked example

Three standard deviation formulas, and when each applies.

\[ \text{data as a sample; data as a population; a distribution} \]

Data treated as a sample

Why: Squared deviations over n minus 1.

Data treated as a population

Why: Squared deviations over N.

A probability distribution

Why: Squared deviations weighted by P(x).

Note what unifies them

Why: All three are averages of squared deviations.

Figure (svg): The solution to Worked example which divisor, and why shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \frac{1}{n-1}, \qquad \frac{1}{N}, \qquad P(x) \]

Verify: confirm the distribution case matches the population case for equally likely values

Why: Give k values each probability one over k, and the distribution formula becomes the sum of squared deviations divided by k — the population formula exactly. That is the book's own remark that the formulas coincide when all outcomes are equally likely, and it explains why there is no n minus 1 here: the correction in the sample formula compensates for estimating a mean from data, and nothing is being estimated when the distribution is given.

OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, p. 230

51. Trap: importing the n minus 1

Trap

The trap

\[ \sigma = \sqrt{\frac{\sum (x - \mu)^2 P(x)}{n - 1}} \]

Divide by n minus 1, as section 2.7 did

Why: Every standard deviation so far has had that divisor.

\[ \text{but there is no sample here, and no n} \]

The n minus 1 corrects for estimating a mean from data. Here mu is given, and the probabilities have already averaged the terms.

The fix

\[ \sigma = \sqrt{\sum (x - \mu)^2 P(x)} \]

Ask what the weights sum to before adding any divisor

Why: Weights summing to one have already done the averaging.

The question 'what am I averaging over' settles it every time. Over n observations, divide by n; over n observations with a mean estimated from them, divide by n minus 1; over a distribution whose probabilities sum to one, divide by nothing. The three formulas look different and are the same operation with three sources of weight.

52. Formula to its setting

Matching

Four averages of squared deviations.

Match the pairs

  • l1. divide by n minus 1
  • l2. divide by N
  • l3. weight by P(x), no divisor
  • l4. all three
  • r1. data treated as a sample, mean estimated from it
  • r2. data treated as the whole population
  • r3. a probability distribution with mu given
  • r4. average the squared deviations, then take a root

Why: The fourth row is what unifies them and is worth holding rather than the three separately: every standard deviation in this book averages squared deviations and takes a root, and the only question is where the weights come from. Answering that question settles the divisor without memorising three formulas.

53. One of these is false

Two truths and a lie

All three concern the distinction.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Mu and sigma of a distribution are parameters
  • C. The arithmetic is the same weighted average in both settings
  • B. A distribution's mu changes as you collect more data

Survives elimination: B

Why: The survivor is false and it is the core of the distinction. Mu is computed from the probabilities and does not move whatever data arrive. What moves is x-bar, the observed average, which the Law of Large Numbers says drifts toward the unmoving mu.

54. Why no divisor?

Prediction

Commit before reasoning.

Predict first

The distribution formula for sigma has no denominator. Why not?

  • Because distributions have no sample size
  • Because the probabilities weighting the terms already sum to one, so the weighted sum is already an average
  • Because sigma is always smaller than s
  • It is an omission in the formula

Correct: The probabilities already sum to one.

Why: Averaging means weighting by numbers that total one, and that is exactly what the probabilities do. Section 2.7's divisors existed because its weights were raw counts that had to be normalised. The absence of a sample size is a symptom rather than the reason — a distribution with equally likely values effectively divides by the number of values, and it does so through the probabilities.

55. Centre and spread, data against distribution

Comparison

Fill the blanks. The arithmetic is the same; the weights are not.

Comparison matrix

QuantityFor data (chapter 2)For a distribution (chapter 4)
Centrex-bar, weighted by relative frequencymu, weighted by probability
Spreads, dividing by n minus 1sigma, weighted by P(x), no divisor
What it isa statistic: it varies with the samplea parameter: fixed by the process
What connects themthe Law of Large Numbersthe observed average approaches mu

The bottom row is the whole point of having both columns. A distribution says what a process does; data say what it did; and the law is the promise that the second approaches the first, which is what makes a model testable and what chapters 7 to 13 are built on.

56. Computing mu and sigma, in order

Pattern

Six steps, and the fourth cannot begin before the third is finished.

  1. Check the distribution: every probability in range, and the probabilities summing to one.
  2. Add a column for x times P(x) and fill it row by row.
  3. Total that column; the result is mu.
  4. Add a column for the squared deviation times the probability, using the mu just computed.
  5. Total that column; the result is the variance.
  6. Take the square root for sigma, and check it against the range of the values.

For a money problem, define X as the PROFIT before anything else, and list its values as payoffs rather than as the underlying outcomes. The experiment supplies the probabilities; the payoffs are what is being averaged.

OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-236

57. Check yourself 1 of 3

Check

Compute an expected value.

Check your understanding

A variable takes the values 0, 1 and 2 with probabilities 0.2, 0.5 and 0.3. What is mu?

  • A. 1.1 (correct)
  • B. 1
  • C. 0.5
  • D. 1.5

Answer: A

Why: Multiplying and adding gives 0 plus 0.5 plus 0.6, which is 1.1.

Why B tempts people
That averages the three values without weighting, which would be right only if all three were equally likely.
Why C tempts people
That is the largest single probability rather than a weighted average of the values.
Why D tempts people
That is the midpoint of the range, which the expected value equals only for a symmetric distribution.

58. Check yourself 2 of 3

Check

Finish a sigma.

Check your understanding

For a distribution, the squared deviations weighted by their probabilities total 1.05. What is sigma?

  • A. About 1.02 (correct)
  • B. 1.05
  • C. About 0.42
  • D. About 0.46

Answer: A

Why: The weighted total is already the variance, so sigma is its square root, about 1.0247.

Why B tempts people
That is the variance, which is in squared units and is not the standard deviation.
Why C tempts people
That divides by the number of values before taking the root, but the probabilities have already averaged the terms.
Why D tempts people
That divides by one less than the number of values, importing a sample correction where there is no sample.

59. Check yourself 3 of 3

Check

Interpret the expected value.

Check your understanding

A game has an expected profit of minus one dollar. What does that mean?

  • A. Over many plays, the average loss is about a dollar a play (correct)
  • B. You lose a dollar every time you play
  • C. You have a one-dollar chance of winning
  • D. The game is fair

Answer: A

Why: An expected value is a long-run average, so it describes what happens across many plays rather than on any single one.

Why B tempts people
On any single play you lose two dollars or win a large prize; the one-dollar figure is never an outcome.
Why C tempts people
A probability is a number between zero and one, not an amount of money.
Why D tempts people
A fair game has an expected profit of zero. A negative expected profit is unfavourable to the player.

60. Where this shows up outside the textbook

Real world

A retailer offers a two-year extended warranty for 80 dollars on a device costing 400. Their records show that 6 percent of devices fail within two years and that a failure costs them 400 dollars to replace. A customer asks whether the warranty is worth buying.

Discussion prompt

Compute the expected value from both sides, say what the warranty costs the customer on average, and give a reason a customer might buy it anyway.

Hint: Build one expected value table for the retailer and one for the customer.

Answer:

For the retailer, the expected payout is 24 dollars. A failure costs 400 with probability 0.06 and nothing with probability 0.94, giving 400 times 0.06, which is 24. They charge 80, so the expected profit per warranty is 56 dollars.

For the customer, the warranty costs 56 dollars on average. They pay 80 for certain and receive an expected 24 back, so buying it has an expected value of minus 56 — the mirror image of the retailer's gain, as it must be.

A customer may still buy it rationally. They are not buying an expected value but the removal of a 400-dollar risk they may be unable to absorb. If losing 400 unexpectedly would cause real difficulty and 80 would not, paying 56 in expectation to convert an uncertain large loss into a certain small one can be a sensible trade.

\[ \text{expected payout} = (400)(0.06) + (0)(0.94) = 24 \text{ dollars} \]

This is the structure of all insurance: the seller profits on the expected value and the buyer purchases certainty. What the calculation establishes is the price of that certainty, which is 56 dollars here — and knowing it is what lets a customer decide whether the certainty is worth that much to them. The one thing the arithmetic rules out is the belief that the warranty is financially advantageous on average, which it cannot be, since the retailer would not otherwise offer it.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

Why does the standard deviation of a distribution have no divisor?

  • Because distributions are always symmetric
  • Because the probabilities weighting the squared deviations already sum to one, so the weighted sum is an average
  • Because sigma is a parameter rather than a statistic
  • Because the values are discrete

Correct: Because the probabilities already sum to one.

\[ \sum P(x) = 1 \;\Longrightarrow\; \sum (x-\mu)^2 P(x) \text{ is already an average} \]

Why: Averaging means weighting by numbers totalling one, and that is exactly what a probability distribution supplies. Section 2.7 needed a divisor because its weights were raw counts requiring normalisation. Being a parameter is true but is not the reason, and discreteness is irrelevant — chapter 5's continuous distributions integrate against a density that also integrates to one, for the same reason.

62. Explain it to someone a year behind you

Explain it

They say an expected value of 2.4 children per family is nonsense because nobody has 0.4 of a child.

Discussion prompt

In three sentences or fewer, agree with the observation and explain what the number is for.

Hint: Concede the impossibility and shift to what is being averaged.

Answer:

Tell them they are right that no family has 2.4 children, and that the word 'expected' is a poor name for the quantity.

What it means is that if you total the children across a very large number of families and divide by the number of families, you get about 2.4.

That is exactly the number a government needs for planning school places, even though it describes no family at all — which is why the quantity is worth having despite the name.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Building the expected value table and totalling it
  • Computing sigma without importing a divisor
  • Setting up a money problem with the right values of X
  • Saying what an expected value does and does not promise

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the table, check the P(x) column totals one before computing anything from it. For sigma, ask what the weights sum to — if one, no divisor. For a money problem, define X as the profit in a full sentence before listing values. For interpretation, read 'expected value' as 'long-run average' every time. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Fifteen minutes.

Draw it

At the top, build the complete four-column table for this distribution: x from 0 to 5 with probabilities 0.04, 0.22, 0.46, 0.18, 0.08 and 0.02. Check the probability column totals one, compute mu from the third column, then use that mu to fill the fourth column and compute sigma. Write both answers with their units. Beside the table, draw the distribution as stems and mark mu with a triangle underneath, noting in one sentence why the mark falls where no stem does. In the middle, build the expected value table for the five-digit game: pay 2 dollars, profit 100,000 with probability 0.00001. Write the definition of X as a full sentence first, then the two values, then the two products, then the total. Below it, solve for the prize that would make the game fair and compare it with the prize offered. At the bottom, make a two-column table headed 'data' and 'distribution' with four rows: what the weights are, what the symbols are, what divisor is used, and whether the quantity varies when you collect more observations.

Check your sigma against the range: the values span 0 to 5 and your sigma should be around one, roughly a fifth of the range. If it came out near 0.42 or 0.46 you divided by the number of values, which the probabilities had already done.

65. What you can do now

Recap

Five things, and the last one is what keeps the first four honest.

If you seeThen
A distribution and a request for a meanSum of x times P(x)
An expected value outside the range of valuesAn arithmetic error
A request for sigma from a distributionWeight squared deviations by P(x); no divisor
A weighted total reported as sigmaThat is the variance; take the root
A money problemLet X be the profit, not the underlying outcome
A negative expected profitUnfavourable over many plays; says nothing about one
An impossible expected valueNormal: it is a long-run average

Section 4.3 gives the first named distribution. When a fixed number of independent trials each succeed with the same probability, the count of successes has a binomial distribution — and its expected value and standard deviation come from formulas rather than from a table, which is what makes named families worth having.

OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation §4.2, pp. 230-236 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §4.2 Mean or Expected Value and Standard Deviation — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 230-236

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108