1.1 Definitions of Statistics, Probability, and Key Terms

Separate a population from a sample, a parameter from a statistic, and numerical from categorical data using a college whose every first-year is visible; then see probability as the proportion long runs of coin tosses settle on.

Subject: Statistics · 64 slides · applied lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Definitions of Statistics, Probability, and Key Terms

Title

Statistics · §1.1

Who a study asks, what it measures, and what chance can promise

2. You will leave able to do these five things

Objectives

  1. Separate everyone studied from those actually measured
  2. Tell a whole-group number from a subgroup's
  3. Sort measurements into amounts and categories
  4. Compute an average and a share, with checks
  5. Read a chance as a long-run share

3. Ten rules from earlier courses carry every step

Concept

Summarising a group of numbers

\[ \text{mean} = \text{total} \div \text{count} \]

Mean (average)

Why: a total shared equally

\[ \text{proportion} = \text{part} \div \text{whole} \]

Proportion (share)

Why: part against whole

\[ u \div v = w \text{ so } w \times v = u \]

Divide back

Why: multiply to check

\[ 3 \times u = u + u + u \]

Multiply

Why: copies of a value

Reading, ordering and counting

\[ u + v + w = w + v + u \]

Any order

Why: totals stay the same

\[ u - v \]

Subtract

Why: measures a gap

\[ 2 \div 3 \approx 0.67 \]

Round

Why: ≈ keeps decimals

\[ 60 \text{ min} = 1 \text{ h} \]

Hours

Why: 60 minutes each

\[ 4 \text{ rows} \times 4 \text{ columns} = 16 \]

Count a grid

Why: rows times columns

\[ 2 < 3 \]

Compare

Why: < points at the smaller

4. Four students first in the bookstore line spent $200 to $300

Prediction

Figure (svg): Four blue dots on a spending axis from $50 to $300: one at 200, two stacked at 250, one at 300; no other first-years are drawn

Predict first

Made-up Pine College has 20 first-years. The 4 first in the bookstore line spent $200, $250, $250, $300.

Can their average stand for all 20?

  • Not yet: these 4 may differ from the rest
  • Yes: an average of real spending is exact
  • Yes: they are first-years, so they speak for all

Correct: Not yet: these 4 may differ from the rest

Why: Their average will be exact, but only for these 4. Being first-years does not make them typical: who stands first in a bookstore line in the first week of term may spend more than most.

5. The four in line average $250, but 16 first-years went unasked

Worked example

Figure (svg): Four blue dots on a spending axis from $50 to $300: one at 200, two stacked at 250, one at 300; no other first-years are drawn

Dot plot: one dot per value above a number line.

\[ \textcolor{#1f5fbf}{200} + \textcolor{#1f5fbf}{250} + \textcolor{#1f5fbf}{250} + \textcolor{#1f5fbf}{300} = \textcolor{#1f5fbf}{1000} \]

Add the four amounts

Why: sharing needs a total first

\[ \textcolor{#1f5fbf}{1000} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{250} \]

Share the total among 4

Why: mean rule: equal portions

Figure (svg): Four blue dots on a spending axis from $50 to $300: one at 200, two stacked at 250, one at 300; no other first-years are drawn; a grey dashed line at their mean, line: 250

\[ 20 - \textcolor{#1f5fbf}{4} = 16 \]

Subtract the 4 asked from 20

Why: the rest could change the answer

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 4 dots filled, the other 16 hollow; a grey dashed line at line: 250

\[ \textcolor{#1f5fbf}{4} \times \textcolor{#6b7280}{250} = \textcolor{#1f5fbf}{1000} \]

Check: multiply the mean back

Why: rebuilds the line's total

6. Everyone, or just some: population and sample

Section

Idea 1 of 4

7. A college of 2,000 first-years needs 5 minutes per interview

Concept

Figure (svg): 2,000 first-years drawn to scale as 40 rows of 50 marks

Made-up college: learn all 2,000 first-years' supply spending.

Discussion prompt

How long would interviewing every one take, and what cheaper plan could still tell us about all 2,000?

Answer:

Ask fewer, chosen to resemble the rest. Next: time both plans.

8. Asking all 2,000 takes about 167 hours; asking 100 takes about 8

Worked example

\[ 2000 \times 5 = 10{,}000 \]

Multiply students by 5 minutes

Why: each interview costs time

\[ 10{,}000 \div 60 \approx 166.7 \]

Divide the minutes by 60

Why: hours show the real workload

\[ 100 \times 5 = 500 \]

Time 100 interviews instead

Why: a cheaper plan to compare

\[ 500 \div 60 \approx 8.3 \]

Convert 500 minutes to hours

Why: compares directly with 166.7 hours

\[ 166.7 \times 60 = 10{,}002 \approx 10{,}000 \]

Check: turn hours back into minutes

Why: rounding lands near 10,000

9. All 20 Pine first-years together spent $3,000

Worked example

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300, with counts above

\[ 3 \times \textcolor{#1f5fbf}{50},\ 5 \times \textcolor{#1f5fbf}{100},\ 5 \times \textcolor{#1f5fbf}{150} = \textcolor{#1f5fbf}{150},\ \textcolor{#1f5fbf}{500},\ \textcolor{#1f5fbf}{750} \]

Total the lowest three stacks

Why: repeated amounts multiply

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300, with counts above; a dashed box around the stacks at 50, 100 and 150

\[ 4 \times \textcolor{#1f5fbf}{200},\ 2 \times \textcolor{#1f5fbf}{250},\ 1 \times \textcolor{#1f5fbf}{300} = \textcolor{#1f5fbf}{800},\ \textcolor{#1f5fbf}{500},\ \textcolor{#1f5fbf}{300} \]

Total the highest three stacks

Why: covers the dollars not yet counted

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300, with counts above; a dashed box around the stacks at 200, 250 and 300

\[ \textcolor{#1f5fbf}{150 + 500 + 750 + 800 + 500 + 300} = \textcolor{#1f5fbf}{3000} \]

Add the six stack totals

Why: so the class total can be shared

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300, with counts above; a dashed box around all six stacks

\[ \textcolor{#1f5fbf}{300 + 500 + 800 + 750 + 500 + 150} = \textcolor{#1f5fbf}{3000} \]

Check: add in reverse

Why: new partial sums, same total

10. All 20 average $150, so the line's $250 misses by $100

Worked example

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300

\( {150 + 500 + 750 + 800 + 500 + 300 = 3000} \)

\[ \textcolor{#1f5fbf}{3000} \div \textcolor{#1f5fbf}{20} = \textcolor{#6b7280}{150} \]

Share $3,000 among 20

Why: equal shares: the benchmark to match

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; a grey dashed line at the mean 150

\[ \textcolor{#6b7280}{250} - \textcolor{#6b7280}{150} = \textcolor{#6b7280}{100} \]

Subtract 150 from the line's 250

Why: measures the line's error in dollars

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 4 dots filled, the other 16 hollow; grey lines at 150 and 250 with an arrow marked misses by 100

Every student in the line spent $200 or more.

\[ \textcolor{#1f5fbf}{20} \times \textcolor{#6b7280}{150} = \textcolor{#1f5fbf}{3000} \]

Check: multiply the mean back

Why: 20 equal shares rebuild the total

Figure (svg): The whole class's $3,000 drawn as one blue bar cut into 20 equal pieces, with a bracket round the first piece labelled one share: $150

11. The whole group is the population; the 4 asked are a sample

Worked example

Figure (svg): Four blue dots on a spending axis from $50 to $300: one at 200, two stacked at 250, one at 300; no other first-years are drawn

Population: everyone a study is about.

\[ \textcolor{#1f5fbf}{\text{all 20 dots}}\text{: population} \]

Name the group asked about

Why: the question covers every first-year

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; all 20 dots filled

Sample: the part measured. Sampling: choosing it.

\[ \textcolor{#1f5fbf}{\text{4 filled dots}}\text{: sample} \]

Name the group measured

Why: only their spending is known

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 4 dots filled, the other 16 hollow

\[ 16 + \textcolor{#1f5fbf}{4} = 20 \]

Check: hollow plus filled

Why: together they rebuild the population

12. All 4 in line spent $200 or more; only 0.35 of the population did

Worked example

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300, with counts above

Representative sample: shaped like its population.

\[ \text{at } \$200\text{+: } \textcolor{#1f5fbf}{4 + 2 + 1} = \textcolor{#1f5fbf}{7} \]

Count the boxed first-years

Why: the population's big spenders, for comparison

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; counts above; the stacks at 200, 250 and 300 boxed

\[ \textcolor{#1f5fbf}{7} \div \textcolor{#1f5fbf}{20} = \textcolor{#6b7280}{0.35} \]

Divide 7 by 20

Why: a share any sample size can match

\[ \textcolor{#1f5fbf}{4} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{1} \]

Divide the line's count by 4

Why: to set the line against 0.35

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 4 dots filled inside the box, the rest hollow

\[ \textcolor{#6b7280}{0.35} \times \textcolor{#1f5fbf}{20} = \textcolor{#1f5fbf}{7} \]

Check: multiply back

Why: returns the 7 boxed dots

13. Eight students from the line average $218.75

Worked example

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's first 4 dots filled, the other 16 hollow

\[ \textcolor{#1f5fbf}{200 + 200 + 200 + 150} = \textcolor{#1f5fbf}{750} \]

Add the next 4 in line

Why: doubles the sample to test size

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 8 dots filled, the other 12 hollow

\[ \textcolor{#1f5fbf}{1000} + \textcolor{#1f5fbf}{750} = \textcolor{#1f5fbf}{1750} \]

Add the first 4's 1000

Why: pools all 8 answers for one mean

\[ \textcolor{#1f5fbf}{1750} \div \textcolor{#1f5fbf}{8} = \textcolor{#6b7280}{218.75} \]

Share $1,750 among 8

Why: equal shares, comparable with 150

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 8 dots filled; grey lines at 150 and 218.75

\[ \textcolor{#1f5fbf}{8} \times \textcolor{#6b7280}{218.75} = \textcolor{#1f5fbf}{1750} \]

Check: multiply the mean back

Why: rebuilds the 8 students' total

14. Four names drawn from a hat average $137.50

Worked example

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300

Hat draw: each first-year equally likely (these draws made up).

\[ \textcolor{#1f5fbf}{50 + 100 + 150 + 250} = \textcolor{#1f5fbf}{550} \]

Add draw A's four amounts

Why: to weigh a fair draw against 150

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw A's 4 dots at 50, 100, 150 and 250 filled, the other 16 hollow

\[ \textcolor{#1f5fbf}{550} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{137.5} \]

Share $550 among the 4 drawn

Why: sets draw A beside 150

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw A's dots filled; grey lines at 150 and 137.5

\[ \textcolor{#1f5fbf}{4} \times \textcolor{#6b7280}{137.5} = \textcolor{#1f5fbf}{550} \]

Check: multiply the mean back

Why: rebuilds draw A's total

15. Doubling the line did not remove its lean

Trap

The trap

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis of dollars with ticks every 50 from 50 to 300, labelled at 50, 150 and 250: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 8 dots filled, the other 12 hollow

\( {\textcolor{#1f5fbf}{1750} \div \textcolor{#1f5fbf}{8} = \textcolor{#6b7280}{218.75}} \)

\[ 4 < 8 \]

Trust the bigger 8

Why: assumes size cancels lean

\[ \textcolor{#6b7280}{218.75} - \textcolor{#6b7280}{150} = \textcolor{#6b7280}{68.75} \]

Measure its miss

Why: fails: still too high

The fix

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis of dollars with ticks every 50 from 50 to 300, labelled at 50, 150 and 250: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw A's 4 dots filled, the other 16 hollow

\( {\textcolor{#1f5fbf}{550} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{137.5}} \)

\[ \text{line: none under } \$150 \]

Read the line's dots

Why: big spenders first

The hat can draw any of the 20.

\[ \textcolor{#6b7280}{137.5} < \textcolor{#6b7280}{150} < \textcolor{#6b7280}{218.75} \]

Check: place the means

Why: lean lands high

16. A poll of 1,500 people speaks for about 150 million voters

Prediction

Predict first

A national poll asks 1,500 of about 150 million voters (illustrative numbers).

What lets so few represent so many?

  • Choosing them so every voter could be picked
  • 1,500 is a large fraction of the voters
  • Asking only people who always vote
  • Any 1,500 answers average out

Correct: Choosing them so every voter could be picked

Why: 1,500 is a tiny fraction of the voters (the next slide computes it). How people are chosen matters: a draw that gives every voter a chance tends to mirror the population, and size alone did not fix the bookstore line.

17. Selection, before size, does the work

Worked example

Figure (svg): Two sample rows on a spending axis from $0 to $300 with a grey dashed line at all 20: 150. Row A: empty. Row line: empty

\[ 1500 \div 150{,}000{,}000 = 0.00001 \]

Divide poll by voters

Why: shows how few are asked

\[ 0.00001 < 0.5 \]

Reject 'a large fraction'

Why: far below even half

\[ \text{always-voters only} \]

Reject 'always vote'

Why: occasional voters go missing

\[ \text{line's 8: } \textcolor{#6b7280}{218.75} \text{ vs } \textcolor{#6b7280}{150} \]

Reject 'average out'

Why: 8 from the line missed

Figure (svg): Two sample rows on a spending axis from $0 to $300 with a grey dashed line at all 20: 150. Row A: dots at 50, 100, 150, 250 and a grey diamond at its mean 137.5. Row line: dots at 150, four at 200, two at 250 and 300, and a grey diamond at its mean 218.75

\[ 0.00001 \times 150{,}000{,}000 = 1500 \]

Check: multiply back

Why: returns the poll's 1,500

18. An inspector opens 40 of today's 50,000 cans

Faded example

Figure (svg): A made-up day's 50,000 cans drawn to scale as 25 rows of 50 marks, 1 mark = 40 cans

Book: canned-drink makers sample cans to check a 16-ounce fill.

Made-up day: 50,000 cans filled; an inspector opens 40.

Fill in the blanks

population: 50000 cans; sample: 40 cans; proportion opened: 0.0008

Why: The population is every can filled today, 50,000. The sample is the 40 opened. The proportion opened is 40 ÷ 50,000 = 0.0008. Opening every can would leave none to sell.

19. The inspector's sample is 0.0008 of the day's cans

Worked example

Figure (svg): A made-up day's 50,000 cans drawn to scale as 25 rows of 50 marks, 1 mark = 40 cans

\[ \text{population} = \textcolor{#1f5fbf}{50{,}000} \text{ cans} \]

Name everything filled today

Why: the fill question covers every can

\[ \text{sample} = \textcolor{#1f5fbf}{40} \text{ cans} \]

Name the cans opened

Why: only these give ounce readings

Figure (svg): A made-up day's 50,000 cans drawn to scale as 25 rows of 50 marks, 1 mark = 40 cans; 40 blue pins, one per opened can, spread evenly across every row (a pin is drawn wider than one can)

\[ \textcolor{#1f5fbf}{40} \div \textcolor{#1f5fbf}{50{,}000} = \textcolor{#6b7280}{0.0008} \]

Divide the sample by the population

Why: sizes the check against the day

\[ \textcolor{#6b7280}{0.0008} \times \textcolor{#1f5fbf}{50{,}000} = \textcolor{#1f5fbf}{40} \]

Check: multiply back

Why: returns the 40 opened cans

20. One group, many samples: parameter and statistic

Section

Idea 2 of 4

21. A second hat draw picks $100, $200, $250 and $300

Concept

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw B's dots at 100, 200, 250 and 300 filled, the other 16 hollow

Draw A averaged $137.50 (the first hat draw).

Discussion prompt

Draw B will average a different amount. Which draw's average, if either, is the mean of all 20?

Answer:

Neither has to be. Next: average draw B and set it beside $150.

22. Draw B averages $212.50, $62.50 above the class mean

Worked example

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw B's dots filled, the other 16 hollow

\[ \textcolor{#1f5fbf}{100 + 200 + 250 + 300} = \textcolor{#1f5fbf}{850} \]

Add draw B's amounts

Why: a total must exist before sharing

\[ \textcolor{#1f5fbf}{850} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{212.5} \]

Share among the 4 drawn

Why: makes B comparable with 150

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw B filled; a grey line at 212.5

\[ \textcolor{#6b7280}{212.5} - \textcolor{#6b7280}{150} = \textcolor{#6b7280}{62.5} \]

Subtract the class mean

Why: the error if B were reported

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; grey lines at 150 and 212.5 with an arrow marked 62.5

\[ \textcolor{#1f5fbf}{4} \times \textcolor{#6b7280}{212.5} = \textcolor{#1f5fbf}{850} \]

Check: multiply back

Why: rebuilds draw B's total

23. The class mean is one parameter; each draw's mean is a statistic

Worked example

Figure (svg): Sample rows on a spending axis from $0 to $300. row A still empty; row B still empty

Parameter: a whole population's number.

Statistic: a sample's number, estimating the parameter.

\[ \text{all 20: } \textcolor{#6b7280}{150} \]

Label the class mean a parameter

Why: it describes the entire population

Figure (svg): Sample rows on a spending axis from $0 to $300. row A still empty; row B still empty; a grey dashed line at the parameter, all 20: 150, through every row

\[ \text{A: } \textcolor{#6b7280}{137.5},\ \ \text{B: } \textcolor{#6b7280}{212.5} \]

Label each draw's mean a statistic

Why: each describes one sample

Figure (svg): Sample rows on a spending axis from $0 to $300. row A: dots at 50, 100, 150, 250 and a grey diamond at its mean 137.5; row B: dots at 100, 200, 250, 300 and a grey diamond at its mean 212.5; a grey dashed line at the parameter, all 20: 150, through every row

\[ \textcolor{#1f5fbf}{20} \times \textcolor{#6b7280}{150} = \textcolor{#1f5fbf}{3000},\ \ \textcolor{#1f5fbf}{4} \times \textcolor{#6b7280}{212.5} = \textcolor{#1f5fbf}{850} \]

Check: rebuild each group's total

Why: each mean shares its own group's dollars

24. Four hat draws give four statistics scattered around $150

Worked example

Figure (svg): Sample rows on a spending axis from $0 to $300. row A: dots at 50, 100, 150, 250 and a grey diamond at its mean 137.5; row B: dots at 100, 200, 250, 300 and a grey diamond at its mean 212.5; a grey dashed line at the parameter, all 20: 150, through every row

\[ \textcolor{#1f5fbf}{50 + 150 + 200 + 200} = \textcolor{#1f5fbf}{600} \]

Add draw C

Why: another fair sample to compare

\[ \textcolor{#1f5fbf}{600} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{150} \]

Divide C's total by 4

Why: places C's statistic on the axis

Figure (svg): Sample rows on a spending axis from $0 to $300. row A: dots at 50, 100, 150, 250 and a grey diamond at its mean 137.5; row B: dots at 100, 200, 250, 300 and a grey diamond at its mean 212.5; row C: dots at 50, 150, 200, 200 and a grey diamond at its mean 150; row D still empty; a grey dashed line at the parameter, all 20: 150, through every row

\[ \textcolor{#1f5fbf}{50 + 100 + 100 + 150} = \textcolor{#1f5fbf}{400} \]

Add draw D

Why: tests whether draws agree

\[ \textcolor{#1f5fbf}{400} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{100} \]

Divide D's total by 4

Why: shows how far a fair draw strays

Figure (svg): Sample rows on a spending axis from $0 to $300. row A: dots at 50, 100, 150, 250 and a grey diamond at its mean 137.5; row B: dots at 100, 200, 250, 300 and a grey diamond at its mean 212.5; row C: dots at 50, 150, 200, 200 and a grey diamond at its mean 150; row D: dots at 50, 100, 100, 150 and a grey diamond at its mean 100; a grey dashed line at the parameter, all 20: 150, through every row

\[ \textcolor{#1f5fbf}{4} \times \textcolor{#6b7280}{150} = \textcolor{#1f5fbf}{600},\ \ \textcolor{#1f5fbf}{4} \times \textcolor{#6b7280}{100} = \textcolor{#1f5fbf}{400} \]

Check: multiply both back

Why: rebuilds both draw totals

25. Describing a sample is certain; claiming the population can miss

Worked example

Problem: the council must print one sentence about all 20 first-years.

Descriptive statistics: summarise those measured. Inferential: claim about the population.

\[ \text{draw A's 4 averaged } \textcolor{#6b7280}{137.5} \]

Describe draw A

Why: certain: all four known

\[ \text{all 20 average about } \textcolor{#6b7280}{137.5} \]

Infer for the whole class

Why: claims 16 unseen students

\[ \textcolor{#6b7280}{150} - \textcolor{#6b7280}{137.5} = \textcolor{#6b7280}{12.5} \]

Subtract the estimate from 150

Why: the inference's error, seen only here

\[ (\textcolor{#6b7280}{137.5} + \textcolor{#6b7280}{212.5} + \textcolor{#6b7280}{150} + \textcolor{#6b7280}{100}) \div 4 = \textcolor{#6b7280}{150} \]

Check: average all four draws' means

Why: these four happen to centre

26. A class of 40 with 22 men has a proportion of men of 0.55

Worked example

Figure (svg): Forty students as a grid of 4 rows of 10 dots, all hollow

Book: a math class of 40: 22 men, 18 women.

\[ \textcolor{#1f5fbf}{22} \div \textcolor{#1f5fbf}{40} = \textcolor{#6b7280}{0.55} \]

Divide the men by 40

Why: a share any class size can match

Figure (svg): Forty students as a grid of 4 rows of 10 dots; the first 22 filled (men), 18 hollow (women), with a grey note 0.55 of the class

\[ \textcolor{#1f5fbf}{18} \div \textcolor{#1f5fbf}{40} = \textcolor{#6b7280}{0.45} \]

Divide the women by 40

Why: the other group's share, for comparison

Book: a sample of all math classes.

\[ \textcolor{#6b7280}{0.55} \times \textcolor{#1f5fbf}{40} = \textcolor{#1f5fbf}{22},\ \ \textcolor{#6b7280}{0.45} \times \textcolor{#1f5fbf}{40} = \textcolor{#1f5fbf}{18} \]

Check: multiply both back

Why: returns the two head counts

27. Treating B's 0.75 as exact invents 8 spenders

Trap

The trap

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis of dollars with ticks every 50 from 50 to 300, labelled at 50, 150 and 250: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw B's 4 dots filled, the other 16 hollow; the stacks at 200, 250 and 300 boxed

\[ \textcolor{#1f5fbf}{3} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{0.75} \]

Divide by 4

Why: B's share, ready to scale

\[ \textcolor{#6b7280}{0.75} \times 20 = 15 \]

Apply 0.75 to all 20

Why: treats it as exact

\[ 15 - 7 = 8 \]

Compare with Pine's 7

Why: fails: 8 invented

The fix

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis of dollars with ticks every 50 from 50 to 300, labelled at 50, 150 and 250: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw A's 4 dots filled, the other 16 hollow; the stacks at 200, 250 and 300 boxed, with 1 filled dot inside

\[ \textcolor{#1f5fbf}{1} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{0.25} \]

Divide A's boxed 1 by 4

Why: B's fair twin

\[ \textcolor{#6b7280}{0.75} - \textcolor{#6b7280}{0.25} = \textcolor{#6b7280}{0.5} \]

Subtract A from B

Why: sizes their disagreement

\[ \textcolor{#1f5fbf}{3} - \textcolor{#1f5fbf}{1} = \textcolor{#1f5fbf}{2} \text{ of } \textcolor{#1f5fbf}{4} \]

Check: recount the boxes

Why: a tally, not a share

28. Four claims about a gym's treadmill times, and one needs inference

Prediction

Predict first

A made-up gym times 30 of its 900 members, chosen at random.

Which claim needs inference?

  • 'Members here average about 21 minutes.'
  • 'The 30 timed averaged 21 minutes.'
  • 'The fastest of the 30 took 12 minutes.'
  • '12 of the 30 finished under 20 minutes.'

Correct: 'Members here average about 21 minutes.'

Why: Only that claim talks about members who were never timed. The other three describe the 30 whose times are on the sheet: descriptive statistics, certain for those 30.

29. Only the claim about all 900 members goes beyond the 30 timed

Worked example

\[ 900 - 30 = 870 \]

Subtract the timed from 900

Why: unseen members are what inference claims

\[ \text{'members here': all } 900 \]

Find whose times it covers

Why: includes the 870 untimed: inference

\[ \text{averaged, fastest, under 20: the } 30 \]

Find whose times the rest cover

Why: only the timed 30: descriptive

\[ \text{on the sheet: 3 claims yes, 1 no} \]

Check: verify each from the 30 times

Why: only the inference cannot be verified

30. Example 1.4: an insurer checks 500 doctors for lawsuits

Faded example

Figure (svg): 500 doctors as a grid of 20 rows of 25 small dots, all hollow

Example 1.4: 500 doctors drawn at random from a professional directory.

Made-up count: 115 had a malpractice lawsuit.

Fill in the blanks

sample: 500 doctors; statistic: 115 ÷ 500 = 0.23; the parameter covers every doctor in the directory

Why: The sample is the 500 doctors checked. Their proportion, 115 ÷ 500 = 0.23, is a statistic. The book's population is all medical doctors listed in the professional directory, so the parameter is the proportion among them, which nobody has checked.

31. 0.23 estimates the directory's proportion

Worked example

Figure (svg): 500 doctors as a grid of 20 rows of 25 small dots, all hollow

\[ \text{sample} = \textcolor{#1f5fbf}{500} \text{ doctors} \]

Name who was checked

Why: only their records are known

\[ \textcolor{#1f5fbf}{115} \div \textcolor{#1f5fbf}{500} = \textcolor{#6b7280}{0.23} \]

Divide sued by 500

Why: a sample's share: statistic

Figure (svg): 500 doctors as a grid of 20 rows of 25 small dots; the first 115 filled, the other 385 hollow, with a grey note 115 sued: 0.23

\[ \text{parameter: the directory's proportion} \]

Name the unknown target

Why: book's population: listed doctors

Book's variable: whether one doctor was in a malpractice suit; data: yes or no.

\[ \textcolor{#6b7280}{0.23} \times \textcolor{#1f5fbf}{500} = \textcolor{#1f5fbf}{115} \]

Check: multiply back

Why: returns the 115 sued

32. Amounts or categories: variables and data

Section

Idea 3 of 4

33. A survey records 3 exam scores and 4 voters' parties

Concept

Figure (svg): Four voters in rows: 1 Democrat, 2 Independent, 3 Republican, 4 Democrat

Book: exam scores 86, 75, 92. Made up: four voters' parties.

Discussion prompt

Which list can be averaged, and what could the other list give instead?

Answer:

Scores can; parties cannot. Next: an attempt to average parties anyway.

34. Number codes give the same 4 voters an average of 1.75 or 2

Worked example

Figure (svg): Four voters in rows: 1 Democrat, 2 Independent, 3 Republican, 4 Democrat

\[ \textcolor{#1f5fbf}{1 + 2 + 3 + 1} = \textcolor{#1f5fbf}{7} \]

Add the code A numbers

Why: treats party labels as amounts

Figure (svg): Four voters in rows: 1 Democrat, 2 Independent, 3 Republican, 4 Democrat, with made-up code A numbers 1, 2, 3, 1

\[ \textcolor{#1f5fbf}{7} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{1.75} \]

Divide by the 4 voters

Why: follows the mean rule anyway

\[ \textcolor{#1f5fbf}{2 + 3 + 1 + 2} = \textcolor{#1f5fbf}{8} \]

Add the code B numbers

Why: same voters, codes reshuffled

Figure (svg): Four voters in rows: 1 Democrat, 2 Independent, 3 Republican, 4 Democrat, with made-up code B numbers 2, 3, 1, 2

\[ \textcolor{#1f5fbf}{8} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{2} \]

Divide by 4 voters again

Why: to test whether the coding decides

\[ \textcolor{#1f5fbf}{4} \times \textcolor{#6b7280}{1.75} = \textcolor{#1f5fbf}{7} \]

Check: multiply code A's average back

Why: 1.75 is arithmetically right

35. Exam scores have equal units, so their mean, 84.3, means something

Worked example

Variable: a trait measured on each member, named X, Y, …

Numerical: equal units (points, hours). Categorical: categories (party).

\[ \textcolor{#1f5fbf}{92} - \textcolor{#1f5fbf}{86} = 6 \text{ points} \]

Subtract two scores

Why: the gap is measured in points

\[ \textcolor{#1f5fbf}{86 + 75 + 92} = \textcolor{#1f5fbf}{253} \]

Add the three scores

Why: to reach one mean worth reading

\[ \textcolor{#1f5fbf}{253} \div \textcolor{#1f5fbf}{3} \approx \textcolor{#6b7280}{84.3} \]

Share among 3 exams

Why: mean rule, kept to one decimal

\[ \textcolor{#1f5fbf}{3} \times \textcolor{#6b7280}{84.3} = \textcolor{#1f5fbf}{252.9} \approx \textcolor{#1f5fbf}{253} \]

Check: multiply back

Why: within rounding of 253

36. Categorical data are summarised by a proportion, not a mean

Worked example

Figure (svg): Two rows of cells. X, exam score: 86, 75, 92, with 75 outlined as one datum and a bracket marking all three as data. Y, party: Dem, Ind, Rep, Dem

A report needs one summary for the party column, and names for its values.

Data: a variable's values. Datum: one value.

\[ \textcolor{#1f5fbf}{\text{Democrat}}\text{: } \textcolor{#1f5fbf}{2} \text{ of } \textcolor{#1f5fbf}{4} \]

Count Y's Democrat data

Why: categories are counted, not added

Figure (svg): Two rows of cells. X, exam score: 86, 75, 92, with 75 outlined as one datum and a bracket marking all three as data. Y, party: Dem, Ind, Rep, Dem, with both Dem cells highlighted

\[ \textcolor{#1f5fbf}{2} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{0.5} \]

Divide by the 4 voters

Why: turns the count into a share

\[ \textcolor{#6b7280}{0.5} \times \textcolor{#1f5fbf}{4} = \textcolor{#1f5fbf}{2} \]

Check: multiply back

Why: returns the two Democrats

37. Example 1.1 asks about all first-years but measures 100

Worked example

Example 1.1: ABC College first-years' mean supply spending, no books.

100 surveyed at random; three spent $150, $200, $225.

\[ \text{population: all ABC first-years this term} \]

Find who the question is about

Why: the asked-for mean covers them

\[ \text{sample: the 100 surveyed} \]

Find who gave answers

Why: only these students were measured

Book's printed sample: one statistics section; we keep the 100.

\[ \text{an ABC sophomore: in neither group} \]

Check: test a student outside

Why: the question names first-years only

38. Example 1.1's parameter and statistic are means of those two groups

Worked example

\( {\text{population: all ABC first-years}},\quad\allowbreak \allowbreak {\text{sample: the 100 surveyed}} \)

\[ \text{parameter: mean spent by all first-years} \]

Name the population's number

Why: unknown: nobody asked everyone

\[ \text{statistic: mean spent by the 100} \]

Name the sample's number

Why: estimates the parameter

\[ X = \text{amount one first-year spent} \]

Define the variable X

Why: one dollar amount each: numerical

\[ \text{data: } \textcolor{#1f5fbf}{\$150,\ \$200,\ \$225,\ \ldots} \]

List the data given

Why: actual values X took

\[ X = \textcolor{#1f5fbf}{\$150} \text{ for one first-year} \]

Check: read one datum as X

Why: a single student's dollar amount fits

39. Example 1.3's crash test leaves four key terms to you

Faded example

Example 1.3: a random sample of 75 cars crashed at 35 miles per hour.

Goal: the proportion of driver dummies with head injuries.

\[ \text{population: all cars with front-seat dummies} \]

Name the group of interest

Why: the goal covers every such car

\[ \text{sample: the 75 cars} \]

Name the cars crashed

Why: only these were measured

Discussion prompt

Now name the parameter, the statistic, the variable X and its data.

Answer:

The next slide gives all four, each with its reason.

40. Example 1.3's variable is yes-or-no, so its statistic is a proportion

Worked example

\[ \text{parameter: proportion injured, all such cars} \]

Name the population's number

Why: the unknown the goal asks for

\[ \text{statistic: proportion injured among the 75} \]

Name the sample's number

Why: computed from the crashed cars

\[ X = \text{whether one driver dummy is injured} \]

Define the variable X

Why: one answer per dummy: categorical

\[ \text{data: } \textcolor{#1f5fbf}{\text{yes}} \text{ or } \textcolor{#1f5fbf}{\text{no}} \]

Describe the data

Why: words, so counted, never averaged

\[ \textcolor{#1f5fbf}{21} \div \textcolor{#1f5fbf}{75} = \textcolor{#6b7280}{0.28} \]

Check: try a made-up 21 injured

Why: yes/no answers do yield a proportion

41. Averaging ZIP codes gives a ZIP with decimals

Trap

The trap

\[ \begin{aligned} &\textcolor{#1f5fbf}{10001 + 10001} \\ &{+}\ \textcolor{#1f5fbf}{60614 + 94103} = \textcolor{#1f5fbf}{174719} \end{aligned} \]

Add 4 made-up ZIP codes

Why: treats place labels as amounts

\[ \textcolor{#1f5fbf}{174719} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{43679.75} \]

Divide by 4 customers

Why: fails: no ZIP code has decimals

The fix

\[ \text{ZIP code} = \text{a place label} \]

Classify ZIP codes

Why: digits name places

\[ \textcolor{#1f5fbf}{10001}: 2,\ \ \textcolor{#1f5fbf}{60614}: 1,\ \ \textcolor{#1f5fbf}{94103}: 1 \]

Tally each ZIP code

Why: labels are counted, not pooled

\[ 2 + 1 + 1 = \textcolor{#1f5fbf}{4} \]

Check: add the tallies

Why: every customer counted once

42. Four survey variables, and only one has equal units

Prediction

Predict first

A student survey records four variables for each respondent.

Which one is numerical?

  • Hours slept last night
  • Student ID number
  • Phone brand
  • Whether they own a car

Correct: Hours slept last night

Why: Hours have equal units: 7 hours is one hour more than 6. An ID number looks numerical but only labels a student; phone brand and car ownership are categories.

43. Only hours slept keeps equal units between values

Worked example

\[ \textcolor{#1f5fbf}{7} - \textcolor{#1f5fbf}{6} = 1 \text{ hour} \]

Subtract two sleep times

Why: equal-sized unit: numerical

\[ \textcolor{#1f5fbf}{20417} - \textcolor{#1f5fbf}{20416} = 1 \text{ of nothing} \]

Subtract two made-up IDs

Why: labels only: categorical

\[ \textcolor{#1f5fbf}{\text{brand A}} - \textcolor{#1f5fbf}{\text{brand B}},\ \ \textcolor{#1f5fbf}{\text{yes}} - \textcolor{#1f5fbf}{\text{no}}\text{: no gap} \]

Subtract each categorical pair

Why: with no unit there is no gap

\[ \textcolor{#1f5fbf}{7} + \textcolor{#1f5fbf}{6} = \textcolor{#1f5fbf}{13} \text{ hours} \]

Add two students' sleep times

Why: equal units allow totals

\[ \textcolor{#1f5fbf}{13} \div \textcolor{#1f5fbf}{2} = \textcolor{#6b7280}{6.5} \text{ hours} \]

Share 13 hours between the 2 students

Why: mean rule, a real amount

\[ \textcolor{#1f5fbf}{6} < \textcolor{#6b7280}{6.5} < \textcolor{#1f5fbf}{7} \]

Check: mean between the two answers

Why: means stay inside their range

44. The book's 14 sleep times leave the mean to you

Faded example

Figure (svg): Dot plot of the book's 14 sleep times on an axis from 5 to 9 hours: 1 at 5, 1 at 5.5, 3 at 6, 4 at 6.5, 2 at 7, 2 at 8, 1 at 9, with counts above

Book's class exercise: 14 students' nightly sleep, in hours.

Fill in the blanks

variable type: numerical; total = 93.5 hours; mean ≈ 6.68 hours

Why: Hours have equal units, so the variable is numerical. The stack totals 5, 5.5, 18, 26, 14, 16 and 9 add to 93.5 hours, and 93.5 ÷ 14 ≈ 6.68 hours.

45. The 14 students sleep about 6.68 hours

Worked example

Figure (svg): Dot plot of the book's 14 sleep times on an axis from 5 to 9 hours: 1 at 5, 1 at 5.5, 3 at 6, 4 at 6.5, 2 at 7, 2 at 8, 1 at 9, with counts above

\[ 3 \times \textcolor{#1f5fbf}{6},\ 4 \times \textcolor{#1f5fbf}{6.5} = \textcolor{#1f5fbf}{18},\ \textcolor{#1f5fbf}{26} \]

Total two stacks

Why: repeated hours multiply

Figure (svg): Dot plot of the book's 14 sleep times on an axis from 5 to 9 hours: 1 at 5, 1 at 5.5, 3 at 6, 4 at 6.5, 2 at 7, 2 at 8, 1 at 9, with counts above; a dashed box around the stacks from 6 to 6.5

\[ 2 \times \textcolor{#1f5fbf}{7},\ 2 \times \textcolor{#1f5fbf}{8} = \textcolor{#1f5fbf}{14},\ \textcolor{#1f5fbf}{16} \]

Total two more

Why: so one addition finishes

Figure (svg): Dot plot of the book's 14 sleep times on an axis from 5 to 9 hours: 1 at 5, 1 at 5.5, 3 at 6, 4 at 6.5, 2 at 7, 2 at 8, 1 at 9, with counts above; a dashed box around the stacks from 7 to 8

\[ \textcolor{#1f5fbf}{5 + 5.5 + 18 + 26 + 14 + 16 + 9} = \textcolor{#1f5fbf}{93.5} \]

Add all stack totals

Why: the mean rule needs it

Figure (svg): Dot plot of the book's 14 sleep times on an axis from 5 to 9 hours: 1 at 5, 1 at 5.5, 3 at 6, 4 at 6.5, 2 at 7, 2 at 8, 1 at 9, with counts above; a dashed box around the stacks from 5 to 9

\[ \textcolor{#1f5fbf}{93.5} \div \textcolor{#1f5fbf}{14} \approx \textcolor{#6b7280}{6.68} \]

Share among 14

Why: mean = total ÷ count

Figure (svg): Dot plot of the book's 14 sleep times on an axis from 5 to 9 hours: 1 at 5, 1 at 5.5, 3 at 6, 4 at 6.5, 2 at 7, 2 at 8, 1 at 9, with counts above; a grey dashed line at the mean 6.68

\[ \textcolor{#1f5fbf}{14} \times \textcolor{#6b7280}{6.68} = \textcolor{#1f5fbf}{93.52} \approx \textcolor{#1f5fbf}{93.5} \]

Check: multiply back

Why: rounds to 93.5

46. Chance, over many tries: probability

Section

Idea 4 of 4

47. A fair coin tossed 4 times 'should' give 2 heads

Concept

Figure (svg): Four empty toss boxes numbered 1 to 4, each marked ?, under the heading one fair coin, 4 tosses

Fair: heads (H), tails (T) equally likely; H share: 1 ÷ 2 = 0.5.

Discussion prompt

Toss it 4 times. How often will you see exactly 2 heads?

Answer:

Less often than it sounds. Next: every possible outcome, listed.

48. Only 6 of the 16 possible four-toss runs give exactly 2 heads

Worked example

Figure (svg): Four empty toss boxes numbered 1 to 4, each marked ?

Run: 4 tosses in order, like HTTH.

\[ 4 \times 4 = 16 \]

Multiply grid rows by columns

Why: every first pair meets every last

Figure (svg): All 16 runs of 4 tosses in a 4 by 4 grid: rows HH, HT, TH, TT for tosses 1–2, columns HH, HT, TH, TT for tosses 3–4

\[ \textcolor{#1f5fbf}{\text{exactly 2 H}}\text{: } \textcolor{#1f5fbf}{6} \text{ runs} \]

Count the runs with two heads

Why: tests how often 'half' happens

Figure (svg): All 16 runs of 4 tosses in a 4 by 4 grid: rows HH, HT, TH, TT for tosses 1–2, columns HH, HT, TH, TT for tosses 3–4; the 6 runs with exactly two heads shaded: HHTT, HTHT, HTTH, THHT, THTH, TTHH

\[ \textcolor{#1f5fbf}{6} \div 16 = \textcolor{#6b7280}{0.375} \]

Divide by all 16 runs

Why: if every run is equally likely

\[ 3 + 2 + 2 + 3 = 10 \text{ unshaded} \]

Check: count each row's unshaded cells

Why: counted, not subtracted

49. Over 2,000 tosses the proportion of heads was 0.498, near 0.5

Worked example

Figure (svg): Proportion of heads on a line from 0 to 1 with a grey dashed line at 0.5; row "4 tosses" empty; row "2,000" empty

\[ 1 \div 4 = \textcolor{#6b7280}{0.25} \]

Divide 1 head by 4

Why: 4 tosses step by 0.25

Figure (svg): Proportion of heads on a line from 0 to 1 with a grey dashed line at 0.5; row "4 tosses" shows grey dots at 0, 0.25, 0.5, 0.75 and 1; row "2,000" empty

\[ \textcolor{#1f5fbf}{996} \div 2000 = \textcolor{#6b7280}{0.498} \]

Divide 996 by 2,000

Why: puts the long run on that scale

Figure (svg): Proportion of heads on a line from 0 to 1 with a grey dashed line at 0.5; row "4 tosses" shows grey dots at 0, 0.25, 0.5, 0.75 and 1; row "2,000" shows a grey dot at 0.498

\[ \textcolor{#6b7280}{0.5} - \textcolor{#6b7280}{0.498} = \textcolor{#6b7280}{0.002} \]

Subtract 0.498 from 0.5

Why: measures how close a long run settles

Probability: the proportion long runs settle near.

\[ \textcolor{#6b7280}{0.498} \times 2000 = \textcolor{#1f5fbf}{996} \]

Check: multiply back

Why: returns 996 heads

50. Probability 0.5 predicts about 1,000 heads in 2,000 tosses, not exactly

Worked example

Figure (svg): A bar of 2,000 tosses with 996 heads filled blue

\[ \textcolor{#6b7280}{0.5} \times 2000 = \textcolor{#6b7280}{1000} \]

Multiply the tosses by 0.5

Why: heads expected over the long run

Figure (svg): A bar of 2,000 tosses with 996 heads filled blue and a grey dashed line at half: 1,000

\[ \textcolor{#6b7280}{1000} - \textcolor{#1f5fbf}{996} = 4 \]

Subtract the author's 996 heads

Why: compares prediction with what happened

4 heads short in 2,000: close, not exact.

\[ 4 \div 2000 = \textcolor{#6b7280}{0.002} \]

Check: shortfall per toss

Why: matches the proportion's miss of 0.002

51. Three heads in a row do not make tails due

Trap

The trap

So far: HHH; one toss left.

\[ 4 \div 2 = 2 \text{ tails 'owed'} \]

Halve 4 tosses

Why: assumes runs split evenly

\[ \text{tails so far} = 0 \]

Count tails in HHH

Why: none yet: tails 'due'

\[ \text{toss 4: surely T} \]

Predict tails

Why: fails: no coin memory

The fix

\[ \text{HHHH},\ \text{HHHT} \]

List the runs starting HHH

Why: two grid cells qualify

\[ 1 \div 2 = \textcolor{#6b7280}{0.5} \]

Divide 1 tails run by 2

Why: equally likely runs, one each

\[ \text{HHHH: 1 of 16} = \text{HHHT: 1 of 16} \]

Check: each run's share of 16

Why: equal shares: tails stays 0.5

52. A weather app gives a 0.3 chance of rain on 10 similar days

Prediction

Predict first

A made-up app gives each of 10 similar days a 0.3 chance of rain.

What does 0.3 predict?

  • About 3 rainy days, maybe 2 or 4
  • Exactly 3 rainy days
  • Rain for 0.3 of each day
  • No rain, since 0.3 is under 0.5

Correct: About 3 rainy days, maybe 2 or 4

Why: 0.3 is a long-run proportion of days like these, so 10 days give about 0.3 × 10 = 3 rainy days, with the same wobble short coin runs show. It says nothing about hours within one day.

53. A 0.3 chance predicts about 3 rainy days in 10, not exactly 3

Worked example

Figure (svg): Ten made-up similar days as a row of 10 cells

\[ \textcolor{#6b7280}{0.3} \times 10 = \textcolor{#1f5fbf}{3} \]

Multiply the days by 0.3

Why: rainy days expected over the run

Figure (svg): Ten made-up similar days as a row of 10 cells; row 3: 3 cells shaded as rainy

\[ \text{10 days: a short run} \]

Reject 'exactly 3'

Why: like 4 tosses, counts wobble

Figure (svg): Ten made-up similar days as a row of 10 cells; row 3: 3 cells shaded as rainy; rows 2 and 4 below: other possible 10-day runs with 2 and 4 rainy days

\[ \textcolor{#6b7280}{0.3}\text{: a proportion of days} \]

Reject 'rain for 0.3 of each day'

Why: it counts days, not hours

\[ 0 < \textcolor{#6b7280}{0.3} \]

Reject 'no rain'

Why: any chance above 0 allows rain

\[ \textcolor{#1f5fbf}{3} \div 10 = \textcolor{#6b7280}{0.3} \]

Check: divide back

Why: 3 of 10 matches the forecast

54. Karl Pearson's 24,000 tosses leave the proportion to you

Faded example

Figure (svg): Zoomed proportion line from 0.49 to 0.51 with a grey dashed line at 0.5; row "2,000" shows a grey dot at 0.498; row "24,000" empty

Book: Karl Pearson tossed a coin 24,000 times and got 12,012 heads.

Fill in the blanks

proportion of heads = 0.5005; miss from 0.5 = 0.0005

Why: 12,012 ÷ 24,000 = 0.5005, and 0.5005 − 0.5 = 0.0005: a smaller miss than the author's 0.002 over 2,000 tosses.

55. Pearson's proportion, 0.5005, misses 0.5 by only 0.0005

Worked example

Figure (svg): Zoomed proportion line from 0.49 to 0.51 with a grey dashed line at 0.5; row "2,000" shows a grey dot at 0.498; row "24,000" empty

\[ 12{,}012 \div 24{,}000 = \textcolor{#6b7280}{0.5005} \]

Divide heads by tosses

Why: puts his run on the 0-to-1 scale

Figure (svg): Zoomed proportion line from 0.49 to 0.51 with a grey dashed line at 0.5; row "2,000" shows a grey dot at 0.498; row "24,000" shows a grey dot at 0.5005

\[ \textcolor{#6b7280}{0.5005} - \textcolor{#6b7280}{0.5} = \textcolor{#6b7280}{0.0005} \]

Subtract 0.5

Why: how far this longer run misses

\[ \textcolor{#6b7280}{0.0005} < \textcolor{#6b7280}{0.002} \]

Compare with the 2,000-toss miss

Why: tests whether a longer run helped

\[ \textcolor{#6b7280}{0.5005} \times 24{,}000 = 12{,}012 \]

Check: multiply back

Why: returns Pearson's 12,012 heads

56. Here, 12 times the tosses gave a miss 4 times smaller

Worked example

Figure (svg): Bars drawn to scale. Tosses in the run, axis 0 to 24 thousand: author 2,000, Pearson 24,000

\[ 24{,}000 \div 2000 = 12 \]

Divide Pearson's tosses by 2,000

Why: counts author-length runs inside his

Figure (svg): Bars drawn to scale. Tosses in the run, axis 0 to 24 thousand: author 2,000, Pearson 24,000, with Pearson's bar divided into 12 pieces each as long as the author's bar

\[ \textcolor{#6b7280}{0.002} \div \textcolor{#6b7280}{0.0005} = 4 \]

Divide the author's miss by Pearson's

Why: counts Pearson-sized misses inside it

Figure (svg): Bars drawn to scale. Tosses in the run, axis 0 to 24 thousand: author 2,000, Pearson 24,000, with Pearson's bar divided into 12 pieces each as long as the author's bar. Miss from 0.5, axis 0 to 0.002: author 0.002 divided into 4 pieces each as long as Pearson's 0.0005 bar

One pair of runs: a hint, not a law.

\[ 4 \times \textcolor{#6b7280}{0.0005} = \textcolor{#6b7280}{0.002} \]

Check: multiply back

Why: rebuilds the author's miss

57. Five questions decode any study's key terms

Pattern

  1. Who is the question about? The population.
  2. Who was actually measured? The sample.
  3. Whose number? Population: parameter. Sample: statistic.
  4. Equal units? Average them. Categories? Use proportions.
  5. A chance? A long-run proportion, not a promise.

58. A classmate labels three families' average as Try It 1.1's statistic

Check

Figure (svg): 100 surveyed families as a 10 by 10 grid of dots; the first 3 filled (amounts listed), the other 97 hollow

Try It 1.1: 100 Knoll families surveyed; three spent $65, $75, $95.

Check your understanding

A classmate calls the three amounts' average the statistic. Which verdict is right?

  • A. Wrong: the statistic uses all 100 families (correct)
  • B. Right: any average of surveyed families is the statistic
  • C. Wrong: that average is the parameter
  • D. Wrong: dollar amounts cannot be averaged

Answer: A

Why: The statistic is the mean uniform spending of the whole sample, all 100 families. Three families are only part of the data, so their average describes just those three.

Why B tempts people
The three were surveyed, but the statistic covers the whole sample, not a part of it.
Why C tempts people
The parameter describes every Knoll Academy family, not three surveyed ones.
Why D tempts people
Dollars have equal units, so spending is numerical and can be averaged.

59. Three amounts are data, not the statistic

Worked example

Figure (svg): 100 surveyed families as a 10 by 10 grid of dots, all hollow

\[ \textcolor{#1f5fbf}{65 + 75 + 95} = \textcolor{#1f5fbf}{235} \]

Add the 3 amounts

Why: tests the classmate

Figure (svg): 100 surveyed families as a 10 by 10 grid of dots; the first 3 filled (amounts listed), the other 97 hollow

\[ \textcolor{#1f5fbf}{235} \div \textcolor{#1f5fbf}{3} \approx \textcolor{#6b7280}{78.33} \]

Share among 3

Why: follows the classmate's method

\[ 100 - \textcolor{#1f5fbf}{3} = 97 \]

Subtract from 100

Why: families 78.33 ignores

Figure (svg): 100 surveyed families as a 10 by 10 grid of dots; the first 3 filled (amounts listed), the other 97 hollow; a grey note 97 unlisted, skipped

\[ \textcolor{#1f5fbf}{\$95} - \textcolor{#1f5fbf}{\$75} = \textcolor{#1f5fbf}{\$20} \]

Subtract two amounts

Why: equal units: averaging allowed

\[ \text{statistic} = \text{total of 100} \div 100 \]

Check the divisor

Why: must be 100, not 3

60. Try It 1.2's athlete survey has one phrase that names the variable

Check

Try It 1.2, one of its six blanks: athletes' heights, in metres.

Check your understanding

Which phrase describes the variable?

  • A. the height of one athlete (correct)
  • B. 1.82, 1.76, 1.69, 1.93
  • C. the average height of athletes in the survey
  • D. all athletes in the university

Answer: A

Why: A variable is a trait measured on each member: one athlete's height. The four numbers are data, the survey's average is the statistic, and all athletes in the university form the population.

Why B tempts people
These are values the variable took: data.
Why C tempts people
An average of the surveyed athletes is a statistic.
Why D tempts people
That is the whole group studied: the population.

61. One athlete's height is the variable; the four numbers are its data

Worked example

\[ \text{height of one athlete} \]

Test: measured on each member?

Why: yes, one value per athlete

\[ \textcolor{#1f5fbf}{1.82, 1.76, 1.69, 1.93} \]

Test the four numbers

Why: values the variable took: data

\[ \textcolor{#1f5fbf}{1.93} - \textcolor{#1f5fbf}{1.82} = \textcolor{#1f5fbf}{0.11} \text{ m} \]

Subtract two heights

Why: equal metre units: numerical

\[ \text{average height in the survey} \]

Test the survey-average phrase

Why: one sample's number: a statistic

\[ \text{all athletes in the university} \]

Test the last phrase

Why: the whole group studied: the population

\[ \textcolor{#1f5fbf}{1.82} \text{ m} = \text{one athlete's height} \]

Check: read one datum as the variable

Why: each value fits the variable's definition

62. The $250 came from an unrepresentative sample

Worked example

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 4 dots filled, the other 16 hollow

\[ \text{population: all 20;\ \ sample: the line's 4} \]

Sort the two groups

Why: the question covers 20

\[ X = \text{one first-year's supply spending} \]

Name the variable

Why: dollars: equal units

\[ \text{statistic } \textcolor{#6b7280}{250},\ \text{parameter } \textcolor{#6b7280}{150} \]

Pair the two means

Why: to size the overstatement

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 4 dots filled; grey lines at 150 and 250

\[ \textcolor{#6b7280}{250} \div \textcolor{#6b7280}{150} \approx 1.67 \]

Divide by the parameter

Why: overstatement as a ratio

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 4 dots filled; grey lines at 150 and 250 with an arrow marked 1.67 times 150

\[ 1.67 \times \textcolor{#6b7280}{150} = 250.5 \approx \textcolor{#6b7280}{250} \]

Check: multiply back

Why: rounds to 250

63. Draw A's $137.50 is honest as an estimate

Worked example

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw A's dots filled, the rest hollow

\[ \textcolor{#6b7280}{150} - \textcolor{#6b7280}{137.5} = \textcolor{#6b7280}{12.5} \]

Measure draw A's miss

Why: visible only in made-up Pine

Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw A's dots filled; grey lines at 150 and 137.5

\[ \textcolor{#6b7280}{212.5} - \textcolor{#6b7280}{100} = \textcolor{#6b7280}{112.5} \]

Subtract draw D's mean from draw B's

Why: sizes what four names can vary

Figure (svg): Sample rows on a spending axis from $0 to $300. row A: dots at 50, 100, 150, 250 and a grey diamond at its mean 137.5; row B: dots at 100, 200, 250, 300 and a grey diamond at its mean 212.5; row C: dots at 50, 150, 200, 200 and a grey diamond at its mean 150; row D: dots at 50, 100, 100, 150 and a grey diamond at its mean 100; a grey dashed line at the parameter, all 20: 150, through every row

\[ \text{report: about } \$137.50 \text{, from 4} \]

Label it an estimate

Why: that spread is why 'about'

\[ \textcolor{#6b7280}{212.5} - \textcolor{#6b7280}{112.5} = \textcolor{#6b7280}{100} \]

Check: take the spread off draw B

Why: lands on a mean already computed

64. You can now read any study's key terms and chances

Recap

OpenStax Introductory Statistics 2e, §1.1 Definitions of Statistics, Probability, and Key Terms §1.1, pp. 5-10 — Examples 1.1–1.4 and the Try Its trace back here

Sources

  1. OpenStax Introductory Statistics 2e, §1.1 Definitions of Statistics, Probability, and Key Terms — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 5-10
  2. OpenStax Introductory Business Statistics 2e, §1.1 Definitions of Statistics, Probability, and Key Terms — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108