Separate a population from a sample, a parameter from a statistic, and numerical from categorical data using a college whose every first-year is visible; then see probability as the proportion long runs of coin tosses settle on.
Subject: Statistics · 64 slides · applied lesson
Open the interactive version of this deck
Title
Statistics · §1.1
Who a study asks, what it measures, and what chance can promise
Objectives
Concept
Summarising a group of numbers
\[ \text{mean} = \text{total} \div \text{count} \]
Mean (average)
Why: a total shared equally
\[ \text{proportion} = \text{part} \div \text{whole} \]
Proportion (share)
Why: part against whole
\[ u \div v = w \text{ so } w \times v = u \]
Divide back
Why: multiply to check
\[ 3 \times u = u + u + u \]
Multiply
Why: copies of a value
Reading, ordering and counting
\[ u + v + w = w + v + u \]
Any order
Why: totals stay the same
\[ u - v \]
Subtract
Why: measures a gap
\[ 2 \div 3 \approx 0.67 \]
Round
Why: ≈ keeps decimals
\[ 60 \text{ min} = 1 \text{ h} \]
Hours
Why: 60 minutes each
\[ 4 \text{ rows} \times 4 \text{ columns} = 16 \]
Count a grid
Why: rows times columns
\[ 2 < 3 \]
Compare
Why: < points at the smaller
Prediction
Figure (svg): Four blue dots on a spending axis from $50 to $300: one at 200, two stacked at 250, one at 300; no other first-years are drawn
Predict first
Made-up Pine College has 20 first-years. The 4 first in the bookstore line spent $200, $250, $250, $300.
Can their average stand for all 20?
Correct: Not yet: these 4 may differ from the rest
Why: Their average will be exact, but only for these 4. Being first-years does not make them typical: who stands first in a bookstore line in the first week of term may spend more than most.
Worked example
Figure (svg): Four blue dots on a spending axis from $50 to $300: one at 200, two stacked at 250, one at 300; no other first-years are drawn
Dot plot: one dot per value above a number line.
\[ \textcolor{#1f5fbf}{200} + \textcolor{#1f5fbf}{250} + \textcolor{#1f5fbf}{250} + \textcolor{#1f5fbf}{300} = \textcolor{#1f5fbf}{1000} \]
Add the four amounts
Why: sharing needs a total first
\[ \textcolor{#1f5fbf}{1000} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{250} \]
Share the total among 4
Why: mean rule: equal portions
Figure (svg): Four blue dots on a spending axis from $50 to $300: one at 200, two stacked at 250, one at 300; no other first-years are drawn; a grey dashed line at their mean, line: 250
\[ 20 - \textcolor{#1f5fbf}{4} = 16 \]
Subtract the 4 asked from 20
Why: the rest could change the answer
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 4 dots filled, the other 16 hollow; a grey dashed line at line: 250
\[ \textcolor{#1f5fbf}{4} \times \textcolor{#6b7280}{250} = \textcolor{#1f5fbf}{1000} \]
Check: multiply the mean back
Why: rebuilds the line's total
Section
Idea 1 of 4
Concept
Figure (svg): 2,000 first-years drawn to scale as 40 rows of 50 marks
Made-up college: learn all 2,000 first-years' supply spending.
Discussion prompt
How long would interviewing every one take, and what cheaper plan could still tell us about all 2,000?
Answer:
Ask fewer, chosen to resemble the rest. Next: time both plans.
Worked example
\[ 2000 \times 5 = 10{,}000 \]
Multiply students by 5 minutes
Why: each interview costs time
\[ 10{,}000 \div 60 \approx 166.7 \]
Divide the minutes by 60
Why: hours show the real workload
\[ 100 \times 5 = 500 \]
Time 100 interviews instead
Why: a cheaper plan to compare
\[ 500 \div 60 \approx 8.3 \]
Convert 500 minutes to hours
Why: compares directly with 166.7 hours
\[ 166.7 \times 60 = 10{,}002 \approx 10{,}000 \]
Check: turn hours back into minutes
Why: rounding lands near 10,000
Worked example
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300, with counts above
\[ 3 \times \textcolor{#1f5fbf}{50},\ 5 \times \textcolor{#1f5fbf}{100},\ 5 \times \textcolor{#1f5fbf}{150} = \textcolor{#1f5fbf}{150},\ \textcolor{#1f5fbf}{500},\ \textcolor{#1f5fbf}{750} \]
Total the lowest three stacks
Why: repeated amounts multiply
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300, with counts above; a dashed box around the stacks at 50, 100 and 150
\[ 4 \times \textcolor{#1f5fbf}{200},\ 2 \times \textcolor{#1f5fbf}{250},\ 1 \times \textcolor{#1f5fbf}{300} = \textcolor{#1f5fbf}{800},\ \textcolor{#1f5fbf}{500},\ \textcolor{#1f5fbf}{300} \]
Total the highest three stacks
Why: covers the dollars not yet counted
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300, with counts above; a dashed box around the stacks at 200, 250 and 300
\[ \textcolor{#1f5fbf}{150 + 500 + 750 + 800 + 500 + 300} = \textcolor{#1f5fbf}{3000} \]
Add the six stack totals
Why: so the class total can be shared
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300, with counts above; a dashed box around all six stacks
\[ \textcolor{#1f5fbf}{300 + 500 + 800 + 750 + 500 + 150} = \textcolor{#1f5fbf}{3000} \]
Check: add in reverse
Why: new partial sums, same total
Worked example
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300
\( {150 + 500 + 750 + 800 + 500 + 300 = 3000} \)
\[ \textcolor{#1f5fbf}{3000} \div \textcolor{#1f5fbf}{20} = \textcolor{#6b7280}{150} \]
Share $3,000 among 20
Why: equal shares: the benchmark to match
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; a grey dashed line at the mean 150
\[ \textcolor{#6b7280}{250} - \textcolor{#6b7280}{150} = \textcolor{#6b7280}{100} \]
Subtract 150 from the line's 250
Why: measures the line's error in dollars
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 4 dots filled, the other 16 hollow; grey lines at 150 and 250 with an arrow marked misses by 100
Every student in the line spent $200 or more.
\[ \textcolor{#1f5fbf}{20} \times \textcolor{#6b7280}{150} = \textcolor{#1f5fbf}{3000} \]
Check: multiply the mean back
Why: 20 equal shares rebuild the total
Figure (svg): The whole class's $3,000 drawn as one blue bar cut into 20 equal pieces, with a bracket round the first piece labelled one share: $150
Worked example
Figure (svg): Four blue dots on a spending axis from $50 to $300: one at 200, two stacked at 250, one at 300; no other first-years are drawn
Population: everyone a study is about.
\[ \textcolor{#1f5fbf}{\text{all 20 dots}}\text{: population} \]
Name the group asked about
Why: the question covers every first-year
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; all 20 dots filled
Sample: the part measured. Sampling: choosing it.
\[ \textcolor{#1f5fbf}{\text{4 filled dots}}\text{: sample} \]
Name the group measured
Why: only their spending is known
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 4 dots filled, the other 16 hollow
\[ 16 + \textcolor{#1f5fbf}{4} = 20 \]
Check: hollow plus filled
Why: together they rebuild the population
Worked example
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300, with counts above
Representative sample: shaped like its population.
\[ \text{at } \$200\text{+: } \textcolor{#1f5fbf}{4 + 2 + 1} = \textcolor{#1f5fbf}{7} \]
Count the boxed first-years
Why: the population's big spenders, for comparison
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; counts above; the stacks at 200, 250 and 300 boxed
\[ \textcolor{#1f5fbf}{7} \div \textcolor{#1f5fbf}{20} = \textcolor{#6b7280}{0.35} \]
Divide 7 by 20
Why: a share any sample size can match
\[ \textcolor{#1f5fbf}{4} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{1} \]
Divide the line's count by 4
Why: to set the line against 0.35
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 4 dots filled inside the box, the rest hollow
\[ \textcolor{#6b7280}{0.35} \times \textcolor{#1f5fbf}{20} = \textcolor{#1f5fbf}{7} \]
Check: multiply back
Why: returns the 7 boxed dots
Worked example
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's first 4 dots filled, the other 16 hollow
\[ \textcolor{#1f5fbf}{200 + 200 + 200 + 150} = \textcolor{#1f5fbf}{750} \]
Add the next 4 in line
Why: doubles the sample to test size
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 8 dots filled, the other 12 hollow
\[ \textcolor{#1f5fbf}{1000} + \textcolor{#1f5fbf}{750} = \textcolor{#1f5fbf}{1750} \]
Add the first 4's 1000
Why: pools all 8 answers for one mean
\[ \textcolor{#1f5fbf}{1750} \div \textcolor{#1f5fbf}{8} = \textcolor{#6b7280}{218.75} \]
Share $1,750 among 8
Why: equal shares, comparable with 150
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 8 dots filled; grey lines at 150 and 218.75
\[ \textcolor{#1f5fbf}{8} \times \textcolor{#6b7280}{218.75} = \textcolor{#1f5fbf}{1750} \]
Check: multiply the mean back
Why: rebuilds the 8 students' total
Worked example
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300
Hat draw: each first-year equally likely (these draws made up).
\[ \textcolor{#1f5fbf}{50 + 100 + 150 + 250} = \textcolor{#1f5fbf}{550} \]
Add draw A's four amounts
Why: to weigh a fair draw against 150
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw A's 4 dots at 50, 100, 150 and 250 filled, the other 16 hollow
\[ \textcolor{#1f5fbf}{550} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{137.5} \]
Share $550 among the 4 drawn
Why: sets draw A beside 150
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw A's dots filled; grey lines at 150 and 137.5
\[ \textcolor{#1f5fbf}{4} \times \textcolor{#6b7280}{137.5} = \textcolor{#1f5fbf}{550} \]
Check: multiply the mean back
Why: rebuilds draw A's total
Trap
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis of dollars with ticks every 50 from 50 to 300, labelled at 50, 150 and 250: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 8 dots filled, the other 12 hollow
\( {\textcolor{#1f5fbf}{1750} \div \textcolor{#1f5fbf}{8} = \textcolor{#6b7280}{218.75}} \)
\[ 4 < 8 \]
Trust the bigger 8
Why: assumes size cancels lean
\[ \textcolor{#6b7280}{218.75} - \textcolor{#6b7280}{150} = \textcolor{#6b7280}{68.75} \]
Measure its miss
Why: fails: still too high
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis of dollars with ticks every 50 from 50 to 300, labelled at 50, 150 and 250: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw A's 4 dots filled, the other 16 hollow
\( {\textcolor{#1f5fbf}{550} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{137.5}} \)
\[ \text{line: none under } \$150 \]
Read the line's dots
Why: big spenders first
The hat can draw any of the 20.
\[ \textcolor{#6b7280}{137.5} < \textcolor{#6b7280}{150} < \textcolor{#6b7280}{218.75} \]
Check: place the means
Why: lean lands high
Prediction
Predict first
A national poll asks 1,500 of about 150 million voters (illustrative numbers).
What lets so few represent so many?
Correct: Choosing them so every voter could be picked
Why: 1,500 is a tiny fraction of the voters (the next slide computes it). How people are chosen matters: a draw that gives every voter a chance tends to mirror the population, and size alone did not fix the bookstore line.
Worked example
Figure (svg): Two sample rows on a spending axis from $0 to $300 with a grey dashed line at all 20: 150. Row A: empty. Row line: empty
\[ 1500 \div 150{,}000{,}000 = 0.00001 \]
Divide poll by voters
Why: shows how few are asked
\[ 0.00001 < 0.5 \]
Reject 'a large fraction'
Why: far below even half
\[ \text{always-voters only} \]
Reject 'always vote'
Why: occasional voters go missing
\[ \text{line's 8: } \textcolor{#6b7280}{218.75} \text{ vs } \textcolor{#6b7280}{150} \]
Reject 'average out'
Why: 8 from the line missed
Figure (svg): Two sample rows on a spending axis from $0 to $300 with a grey dashed line at all 20: 150. Row A: dots at 50, 100, 150, 250 and a grey diamond at its mean 137.5. Row line: dots at 150, four at 200, two at 250 and 300, and a grey diamond at its mean 218.75
\[ 0.00001 \times 150{,}000{,}000 = 1500 \]
Check: multiply back
Why: returns the poll's 1,500
Faded example
Figure (svg): A made-up day's 50,000 cans drawn to scale as 25 rows of 50 marks, 1 mark = 40 cans
Book: canned-drink makers sample cans to check a 16-ounce fill.
Made-up day: 50,000 cans filled; an inspector opens 40.
Fill in the blanks
population: 50000 cans; sample: 40 cans; proportion opened: 0.0008
Why: The population is every can filled today, 50,000. The sample is the 40 opened. The proportion opened is 40 ÷ 50,000 = 0.0008. Opening every can would leave none to sell.
Worked example
Figure (svg): A made-up day's 50,000 cans drawn to scale as 25 rows of 50 marks, 1 mark = 40 cans
\[ \text{population} = \textcolor{#1f5fbf}{50{,}000} \text{ cans} \]
Name everything filled today
Why: the fill question covers every can
\[ \text{sample} = \textcolor{#1f5fbf}{40} \text{ cans} \]
Name the cans opened
Why: only these give ounce readings
Figure (svg): A made-up day's 50,000 cans drawn to scale as 25 rows of 50 marks, 1 mark = 40 cans; 40 blue pins, one per opened can, spread evenly across every row (a pin is drawn wider than one can)
\[ \textcolor{#1f5fbf}{40} \div \textcolor{#1f5fbf}{50{,}000} = \textcolor{#6b7280}{0.0008} \]
Divide the sample by the population
Why: sizes the check against the day
\[ \textcolor{#6b7280}{0.0008} \times \textcolor{#1f5fbf}{50{,}000} = \textcolor{#1f5fbf}{40} \]
Check: multiply back
Why: returns the 40 opened cans
Section
Idea 2 of 4
Concept
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw B's dots at 100, 200, 250 and 300 filled, the other 16 hollow
Draw A averaged $137.50 (the first hat draw).
Discussion prompt
Draw B will average a different amount. Which draw's average, if either, is the mean of all 20?
Answer:
Neither has to be. Next: average draw B and set it beside $150.
Worked example
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw B's dots filled, the other 16 hollow
\[ \textcolor{#1f5fbf}{100 + 200 + 250 + 300} = \textcolor{#1f5fbf}{850} \]
Add draw B's amounts
Why: a total must exist before sharing
\[ \textcolor{#1f5fbf}{850} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{212.5} \]
Share among the 4 drawn
Why: makes B comparable with 150
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw B filled; a grey line at 212.5
\[ \textcolor{#6b7280}{212.5} - \textcolor{#6b7280}{150} = \textcolor{#6b7280}{62.5} \]
Subtract the class mean
Why: the error if B were reported
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; grey lines at 150 and 212.5 with an arrow marked 62.5
\[ \textcolor{#1f5fbf}{4} \times \textcolor{#6b7280}{212.5} = \textcolor{#1f5fbf}{850} \]
Check: multiply back
Why: rebuilds draw B's total
Worked example
Figure (svg): Sample rows on a spending axis from $0 to $300. row A still empty; row B still empty
Parameter: a whole population's number.
Statistic: a sample's number, estimating the parameter.
\[ \text{all 20: } \textcolor{#6b7280}{150} \]
Label the class mean a parameter
Why: it describes the entire population
Figure (svg): Sample rows on a spending axis from $0 to $300. row A still empty; row B still empty; a grey dashed line at the parameter, all 20: 150, through every row
\[ \text{A: } \textcolor{#6b7280}{137.5},\ \ \text{B: } \textcolor{#6b7280}{212.5} \]
Label each draw's mean a statistic
Why: each describes one sample
Figure (svg): Sample rows on a spending axis from $0 to $300. row A: dots at 50, 100, 150, 250 and a grey diamond at its mean 137.5; row B: dots at 100, 200, 250, 300 and a grey diamond at its mean 212.5; a grey dashed line at the parameter, all 20: 150, through every row
\[ \textcolor{#1f5fbf}{20} \times \textcolor{#6b7280}{150} = \textcolor{#1f5fbf}{3000},\ \ \textcolor{#1f5fbf}{4} \times \textcolor{#6b7280}{212.5} = \textcolor{#1f5fbf}{850} \]
Check: rebuild each group's total
Why: each mean shares its own group's dollars
Worked example
Figure (svg): Sample rows on a spending axis from $0 to $300. row A: dots at 50, 100, 150, 250 and a grey diamond at its mean 137.5; row B: dots at 100, 200, 250, 300 and a grey diamond at its mean 212.5; a grey dashed line at the parameter, all 20: 150, through every row
\[ \textcolor{#1f5fbf}{50 + 150 + 200 + 200} = \textcolor{#1f5fbf}{600} \]
Add draw C
Why: another fair sample to compare
\[ \textcolor{#1f5fbf}{600} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{150} \]
Divide C's total by 4
Why: places C's statistic on the axis
Figure (svg): Sample rows on a spending axis from $0 to $300. row A: dots at 50, 100, 150, 250 and a grey diamond at its mean 137.5; row B: dots at 100, 200, 250, 300 and a grey diamond at its mean 212.5; row C: dots at 50, 150, 200, 200 and a grey diamond at its mean 150; row D still empty; a grey dashed line at the parameter, all 20: 150, through every row
\[ \textcolor{#1f5fbf}{50 + 100 + 100 + 150} = \textcolor{#1f5fbf}{400} \]
Add draw D
Why: tests whether draws agree
\[ \textcolor{#1f5fbf}{400} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{100} \]
Divide D's total by 4
Why: shows how far a fair draw strays
Figure (svg): Sample rows on a spending axis from $0 to $300. row A: dots at 50, 100, 150, 250 and a grey diamond at its mean 137.5; row B: dots at 100, 200, 250, 300 and a grey diamond at its mean 212.5; row C: dots at 50, 150, 200, 200 and a grey diamond at its mean 150; row D: dots at 50, 100, 100, 150 and a grey diamond at its mean 100; a grey dashed line at the parameter, all 20: 150, through every row
\[ \textcolor{#1f5fbf}{4} \times \textcolor{#6b7280}{150} = \textcolor{#1f5fbf}{600},\ \ \textcolor{#1f5fbf}{4} \times \textcolor{#6b7280}{100} = \textcolor{#1f5fbf}{400} \]
Check: multiply both back
Why: rebuilds both draw totals
Worked example
Problem: the council must print one sentence about all 20 first-years.
Descriptive statistics: summarise those measured. Inferential: claim about the population.
\[ \text{draw A's 4 averaged } \textcolor{#6b7280}{137.5} \]
Describe draw A
Why: certain: all four known
\[ \text{all 20 average about } \textcolor{#6b7280}{137.5} \]
Infer for the whole class
Why: claims 16 unseen students
\[ \textcolor{#6b7280}{150} - \textcolor{#6b7280}{137.5} = \textcolor{#6b7280}{12.5} \]
Subtract the estimate from 150
Why: the inference's error, seen only here
\[ (\textcolor{#6b7280}{137.5} + \textcolor{#6b7280}{212.5} + \textcolor{#6b7280}{150} + \textcolor{#6b7280}{100}) \div 4 = \textcolor{#6b7280}{150} \]
Check: average all four draws' means
Why: these four happen to centre
Worked example
Figure (svg): Forty students as a grid of 4 rows of 10 dots, all hollow
Book: a math class of 40: 22 men, 18 women.
\[ \textcolor{#1f5fbf}{22} \div \textcolor{#1f5fbf}{40} = \textcolor{#6b7280}{0.55} \]
Divide the men by 40
Why: a share any class size can match
Figure (svg): Forty students as a grid of 4 rows of 10 dots; the first 22 filled (men), 18 hollow (women), with a grey note 0.55 of the class
\[ \textcolor{#1f5fbf}{18} \div \textcolor{#1f5fbf}{40} = \textcolor{#6b7280}{0.45} \]
Divide the women by 40
Why: the other group's share, for comparison
Book: a sample of all math classes.
\[ \textcolor{#6b7280}{0.55} \times \textcolor{#1f5fbf}{40} = \textcolor{#1f5fbf}{22},\ \ \textcolor{#6b7280}{0.45} \times \textcolor{#1f5fbf}{40} = \textcolor{#1f5fbf}{18} \]
Check: multiply both back
Why: returns the two head counts
Trap
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis of dollars with ticks every 50 from 50 to 300, labelled at 50, 150 and 250: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw B's 4 dots filled, the other 16 hollow; the stacks at 200, 250 and 300 boxed
\[ \textcolor{#1f5fbf}{3} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{0.75} \]
Divide by 4
Why: B's share, ready to scale
\[ \textcolor{#6b7280}{0.75} \times 20 = 15 \]
Apply 0.75 to all 20
Why: treats it as exact
\[ 15 - 7 = 8 \]
Compare with Pine's 7
Why: fails: 8 invented
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis of dollars with ticks every 50 from 50 to 300, labelled at 50, 150 and 250: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw A's 4 dots filled, the other 16 hollow; the stacks at 200, 250 and 300 boxed, with 1 filled dot inside
\[ \textcolor{#1f5fbf}{1} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{0.25} \]
Divide A's boxed 1 by 4
Why: B's fair twin
\[ \textcolor{#6b7280}{0.75} - \textcolor{#6b7280}{0.25} = \textcolor{#6b7280}{0.5} \]
Subtract A from B
Why: sizes their disagreement
\[ \textcolor{#1f5fbf}{3} - \textcolor{#1f5fbf}{1} = \textcolor{#1f5fbf}{2} \text{ of } \textcolor{#1f5fbf}{4} \]
Check: recount the boxes
Why: a tally, not a share
Prediction
Predict first
A made-up gym times 30 of its 900 members, chosen at random.
Which claim needs inference?
Correct: 'Members here average about 21 minutes.'
Why: Only that claim talks about members who were never timed. The other three describe the 30 whose times are on the sheet: descriptive statistics, certain for those 30.
Worked example
\[ 900 - 30 = 870 \]
Subtract the timed from 900
Why: unseen members are what inference claims
\[ \text{'members here': all } 900 \]
Find whose times it covers
Why: includes the 870 untimed: inference
\[ \text{averaged, fastest, under 20: the } 30 \]
Find whose times the rest cover
Why: only the timed 30: descriptive
\[ \text{on the sheet: 3 claims yes, 1 no} \]
Check: verify each from the 30 times
Why: only the inference cannot be verified
Faded example
Figure (svg): 500 doctors as a grid of 20 rows of 25 small dots, all hollow
Example 1.4: 500 doctors drawn at random from a professional directory.
Made-up count: 115 had a malpractice lawsuit.
Fill in the blanks
sample: 500 doctors; statistic: 115 ÷ 500 = 0.23; the parameter covers every doctor in the directory
Why: The sample is the 500 doctors checked. Their proportion, 115 ÷ 500 = 0.23, is a statistic. The book's population is all medical doctors listed in the professional directory, so the parameter is the proportion among them, which nobody has checked.
Worked example
Figure (svg): 500 doctors as a grid of 20 rows of 25 small dots, all hollow
\[ \text{sample} = \textcolor{#1f5fbf}{500} \text{ doctors} \]
Name who was checked
Why: only their records are known
\[ \textcolor{#1f5fbf}{115} \div \textcolor{#1f5fbf}{500} = \textcolor{#6b7280}{0.23} \]
Divide sued by 500
Why: a sample's share: statistic
Figure (svg): 500 doctors as a grid of 20 rows of 25 small dots; the first 115 filled, the other 385 hollow, with a grey note 115 sued: 0.23
\[ \text{parameter: the directory's proportion} \]
Name the unknown target
Why: book's population: listed doctors
Book's variable: whether one doctor was in a malpractice suit; data: yes or no.
\[ \textcolor{#6b7280}{0.23} \times \textcolor{#1f5fbf}{500} = \textcolor{#1f5fbf}{115} \]
Check: multiply back
Why: returns the 115 sued
Section
Idea 3 of 4
Concept
Figure (svg): Four voters in rows: 1 Democrat, 2 Independent, 3 Republican, 4 Democrat
Book: exam scores 86, 75, 92. Made up: four voters' parties.
Discussion prompt
Which list can be averaged, and what could the other list give instead?
Answer:
Scores can; parties cannot. Next: an attempt to average parties anyway.
Worked example
Figure (svg): Four voters in rows: 1 Democrat, 2 Independent, 3 Republican, 4 Democrat
\[ \textcolor{#1f5fbf}{1 + 2 + 3 + 1} = \textcolor{#1f5fbf}{7} \]
Add the code A numbers
Why: treats party labels as amounts
Figure (svg): Four voters in rows: 1 Democrat, 2 Independent, 3 Republican, 4 Democrat, with made-up code A numbers 1, 2, 3, 1
\[ \textcolor{#1f5fbf}{7} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{1.75} \]
Divide by the 4 voters
Why: follows the mean rule anyway
\[ \textcolor{#1f5fbf}{2 + 3 + 1 + 2} = \textcolor{#1f5fbf}{8} \]
Add the code B numbers
Why: same voters, codes reshuffled
Figure (svg): Four voters in rows: 1 Democrat, 2 Independent, 3 Republican, 4 Democrat, with made-up code B numbers 2, 3, 1, 2
\[ \textcolor{#1f5fbf}{8} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{2} \]
Divide by 4 voters again
Why: to test whether the coding decides
\[ \textcolor{#1f5fbf}{4} \times \textcolor{#6b7280}{1.75} = \textcolor{#1f5fbf}{7} \]
Check: multiply code A's average back
Why: 1.75 is arithmetically right
Worked example
Variable: a trait measured on each member, named X, Y, …
Numerical: equal units (points, hours). Categorical: categories (party).
\[ \textcolor{#1f5fbf}{92} - \textcolor{#1f5fbf}{86} = 6 \text{ points} \]
Subtract two scores
Why: the gap is measured in points
\[ \textcolor{#1f5fbf}{86 + 75 + 92} = \textcolor{#1f5fbf}{253} \]
Add the three scores
Why: to reach one mean worth reading
\[ \textcolor{#1f5fbf}{253} \div \textcolor{#1f5fbf}{3} \approx \textcolor{#6b7280}{84.3} \]
Share among 3 exams
Why: mean rule, kept to one decimal
\[ \textcolor{#1f5fbf}{3} \times \textcolor{#6b7280}{84.3} = \textcolor{#1f5fbf}{252.9} \approx \textcolor{#1f5fbf}{253} \]
Check: multiply back
Why: within rounding of 253
Worked example
Figure (svg): Two rows of cells. X, exam score: 86, 75, 92, with 75 outlined as one datum and a bracket marking all three as data. Y, party: Dem, Ind, Rep, Dem
A report needs one summary for the party column, and names for its values.
Data: a variable's values. Datum: one value.
\[ \textcolor{#1f5fbf}{\text{Democrat}}\text{: } \textcolor{#1f5fbf}{2} \text{ of } \textcolor{#1f5fbf}{4} \]
Count Y's Democrat data
Why: categories are counted, not added
Figure (svg): Two rows of cells. X, exam score: 86, 75, 92, with 75 outlined as one datum and a bracket marking all three as data. Y, party: Dem, Ind, Rep, Dem, with both Dem cells highlighted
\[ \textcolor{#1f5fbf}{2} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{0.5} \]
Divide by the 4 voters
Why: turns the count into a share
\[ \textcolor{#6b7280}{0.5} \times \textcolor{#1f5fbf}{4} = \textcolor{#1f5fbf}{2} \]
Check: multiply back
Why: returns the two Democrats
Worked example
Example 1.1: ABC College first-years' mean supply spending, no books.
100 surveyed at random; three spent $150, $200, $225.
\[ \text{population: all ABC first-years this term} \]
Find who the question is about
Why: the asked-for mean covers them
\[ \text{sample: the 100 surveyed} \]
Find who gave answers
Why: only these students were measured
Book's printed sample: one statistics section; we keep the 100.
\[ \text{an ABC sophomore: in neither group} \]
Check: test a student outside
Why: the question names first-years only
Worked example
\( {\text{population: all ABC first-years}},\quad\allowbreak \allowbreak {\text{sample: the 100 surveyed}} \)
\[ \text{parameter: mean spent by all first-years} \]
Name the population's number
Why: unknown: nobody asked everyone
\[ \text{statistic: mean spent by the 100} \]
Name the sample's number
Why: estimates the parameter
\[ X = \text{amount one first-year spent} \]
Define the variable X
Why: one dollar amount each: numerical
\[ \text{data: } \textcolor{#1f5fbf}{\$150,\ \$200,\ \$225,\ \ldots} \]
List the data given
Why: actual values X took
\[ X = \textcolor{#1f5fbf}{\$150} \text{ for one first-year} \]
Check: read one datum as X
Why: a single student's dollar amount fits
Faded example
Example 1.3: a random sample of 75 cars crashed at 35 miles per hour.
Goal: the proportion of driver dummies with head injuries.
\[ \text{population: all cars with front-seat dummies} \]
Name the group of interest
Why: the goal covers every such car
\[ \text{sample: the 75 cars} \]
Name the cars crashed
Why: only these were measured
Discussion prompt
Now name the parameter, the statistic, the variable X and its data.
Answer:
The next slide gives all four, each with its reason.
Worked example
\[ \text{parameter: proportion injured, all such cars} \]
Name the population's number
Why: the unknown the goal asks for
\[ \text{statistic: proportion injured among the 75} \]
Name the sample's number
Why: computed from the crashed cars
\[ X = \text{whether one driver dummy is injured} \]
Define the variable X
Why: one answer per dummy: categorical
\[ \text{data: } \textcolor{#1f5fbf}{\text{yes}} \text{ or } \textcolor{#1f5fbf}{\text{no}} \]
Describe the data
Why: words, so counted, never averaged
\[ \textcolor{#1f5fbf}{21} \div \textcolor{#1f5fbf}{75} = \textcolor{#6b7280}{0.28} \]
Check: try a made-up 21 injured
Why: yes/no answers do yield a proportion
Trap
\[ \begin{aligned} &\textcolor{#1f5fbf}{10001 + 10001} \\ &{+}\ \textcolor{#1f5fbf}{60614 + 94103} = \textcolor{#1f5fbf}{174719} \end{aligned} \]
Add 4 made-up ZIP codes
Why: treats place labels as amounts
\[ \textcolor{#1f5fbf}{174719} \div \textcolor{#1f5fbf}{4} = \textcolor{#6b7280}{43679.75} \]
Divide by 4 customers
Why: fails: no ZIP code has decimals
\[ \text{ZIP code} = \text{a place label} \]
Classify ZIP codes
Why: digits name places
\[ \textcolor{#1f5fbf}{10001}: 2,\ \ \textcolor{#1f5fbf}{60614}: 1,\ \ \textcolor{#1f5fbf}{94103}: 1 \]
Tally each ZIP code
Why: labels are counted, not pooled
\[ 2 + 1 + 1 = \textcolor{#1f5fbf}{4} \]
Check: add the tallies
Why: every customer counted once
Prediction
Predict first
A student survey records four variables for each respondent.
Which one is numerical?
Correct: Hours slept last night
Why: Hours have equal units: 7 hours is one hour more than 6. An ID number looks numerical but only labels a student; phone brand and car ownership are categories.
Worked example
\[ \textcolor{#1f5fbf}{7} - \textcolor{#1f5fbf}{6} = 1 \text{ hour} \]
Subtract two sleep times
Why: equal-sized unit: numerical
\[ \textcolor{#1f5fbf}{20417} - \textcolor{#1f5fbf}{20416} = 1 \text{ of nothing} \]
Subtract two made-up IDs
Why: labels only: categorical
\[ \textcolor{#1f5fbf}{\text{brand A}} - \textcolor{#1f5fbf}{\text{brand B}},\ \ \textcolor{#1f5fbf}{\text{yes}} - \textcolor{#1f5fbf}{\text{no}}\text{: no gap} \]
Subtract each categorical pair
Why: with no unit there is no gap
\[ \textcolor{#1f5fbf}{7} + \textcolor{#1f5fbf}{6} = \textcolor{#1f5fbf}{13} \text{ hours} \]
Add two students' sleep times
Why: equal units allow totals
\[ \textcolor{#1f5fbf}{13} \div \textcolor{#1f5fbf}{2} = \textcolor{#6b7280}{6.5} \text{ hours} \]
Share 13 hours between the 2 students
Why: mean rule, a real amount
\[ \textcolor{#1f5fbf}{6} < \textcolor{#6b7280}{6.5} < \textcolor{#1f5fbf}{7} \]
Check: mean between the two answers
Why: means stay inside their range
Faded example
Figure (svg): Dot plot of the book's 14 sleep times on an axis from 5 to 9 hours: 1 at 5, 1 at 5.5, 3 at 6, 4 at 6.5, 2 at 7, 2 at 8, 1 at 9, with counts above
Book's class exercise: 14 students' nightly sleep, in hours.
Fill in the blanks
variable type: numerical; total = 93.5 hours; mean ≈ 6.68 hours
Why: Hours have equal units, so the variable is numerical. The stack totals 5, 5.5, 18, 26, 14, 16 and 9 add to 93.5 hours, and 93.5 ÷ 14 ≈ 6.68 hours.
Worked example
Figure (svg): Dot plot of the book's 14 sleep times on an axis from 5 to 9 hours: 1 at 5, 1 at 5.5, 3 at 6, 4 at 6.5, 2 at 7, 2 at 8, 1 at 9, with counts above
\[ 3 \times \textcolor{#1f5fbf}{6},\ 4 \times \textcolor{#1f5fbf}{6.5} = \textcolor{#1f5fbf}{18},\ \textcolor{#1f5fbf}{26} \]
Total two stacks
Why: repeated hours multiply
Figure (svg): Dot plot of the book's 14 sleep times on an axis from 5 to 9 hours: 1 at 5, 1 at 5.5, 3 at 6, 4 at 6.5, 2 at 7, 2 at 8, 1 at 9, with counts above; a dashed box around the stacks from 6 to 6.5
\[ 2 \times \textcolor{#1f5fbf}{7},\ 2 \times \textcolor{#1f5fbf}{8} = \textcolor{#1f5fbf}{14},\ \textcolor{#1f5fbf}{16} \]
Total two more
Why: so one addition finishes
Figure (svg): Dot plot of the book's 14 sleep times on an axis from 5 to 9 hours: 1 at 5, 1 at 5.5, 3 at 6, 4 at 6.5, 2 at 7, 2 at 8, 1 at 9, with counts above; a dashed box around the stacks from 7 to 8
\[ \textcolor{#1f5fbf}{5 + 5.5 + 18 + 26 + 14 + 16 + 9} = \textcolor{#1f5fbf}{93.5} \]
Add all stack totals
Why: the mean rule needs it
Figure (svg): Dot plot of the book's 14 sleep times on an axis from 5 to 9 hours: 1 at 5, 1 at 5.5, 3 at 6, 4 at 6.5, 2 at 7, 2 at 8, 1 at 9, with counts above; a dashed box around the stacks from 5 to 9
\[ \textcolor{#1f5fbf}{93.5} \div \textcolor{#1f5fbf}{14} \approx \textcolor{#6b7280}{6.68} \]
Share among 14
Why: mean = total ÷ count
Figure (svg): Dot plot of the book's 14 sleep times on an axis from 5 to 9 hours: 1 at 5, 1 at 5.5, 3 at 6, 4 at 6.5, 2 at 7, 2 at 8, 1 at 9, with counts above; a grey dashed line at the mean 6.68
\[ \textcolor{#1f5fbf}{14} \times \textcolor{#6b7280}{6.68} = \textcolor{#1f5fbf}{93.52} \approx \textcolor{#1f5fbf}{93.5} \]
Check: multiply back
Why: rounds to 93.5
Section
Idea 4 of 4
Concept
Figure (svg): Four empty toss boxes numbered 1 to 4, each marked ?, under the heading one fair coin, 4 tosses
Fair: heads (H), tails (T) equally likely; H share: 1 ÷ 2 = 0.5.
Discussion prompt
Toss it 4 times. How often will you see exactly 2 heads?
Answer:
Less often than it sounds. Next: every possible outcome, listed.
Worked example
Figure (svg): Four empty toss boxes numbered 1 to 4, each marked ?
Run: 4 tosses in order, like HTTH.
\[ 4 \times 4 = 16 \]
Multiply grid rows by columns
Why: every first pair meets every last
Figure (svg): All 16 runs of 4 tosses in a 4 by 4 grid: rows HH, HT, TH, TT for tosses 1–2, columns HH, HT, TH, TT for tosses 3–4
\[ \textcolor{#1f5fbf}{\text{exactly 2 H}}\text{: } \textcolor{#1f5fbf}{6} \text{ runs} \]
Count the runs with two heads
Why: tests how often 'half' happens
Figure (svg): All 16 runs of 4 tosses in a 4 by 4 grid: rows HH, HT, TH, TT for tosses 1–2, columns HH, HT, TH, TT for tosses 3–4; the 6 runs with exactly two heads shaded: HHTT, HTHT, HTTH, THHT, THTH, TTHH
\[ \textcolor{#1f5fbf}{6} \div 16 = \textcolor{#6b7280}{0.375} \]
Divide by all 16 runs
Why: if every run is equally likely
\[ 3 + 2 + 2 + 3 = 10 \text{ unshaded} \]
Check: count each row's unshaded cells
Why: counted, not subtracted
Worked example
Figure (svg): Proportion of heads on a line from 0 to 1 with a grey dashed line at 0.5; row "4 tosses" empty; row "2,000" empty
\[ 1 \div 4 = \textcolor{#6b7280}{0.25} \]
Divide 1 head by 4
Why: 4 tosses step by 0.25
Figure (svg): Proportion of heads on a line from 0 to 1 with a grey dashed line at 0.5; row "4 tosses" shows grey dots at 0, 0.25, 0.5, 0.75 and 1; row "2,000" empty
\[ \textcolor{#1f5fbf}{996} \div 2000 = \textcolor{#6b7280}{0.498} \]
Divide 996 by 2,000
Why: puts the long run on that scale
Figure (svg): Proportion of heads on a line from 0 to 1 with a grey dashed line at 0.5; row "4 tosses" shows grey dots at 0, 0.25, 0.5, 0.75 and 1; row "2,000" shows a grey dot at 0.498
\[ \textcolor{#6b7280}{0.5} - \textcolor{#6b7280}{0.498} = \textcolor{#6b7280}{0.002} \]
Subtract 0.498 from 0.5
Why: measures how close a long run settles
Probability: the proportion long runs settle near.
\[ \textcolor{#6b7280}{0.498} \times 2000 = \textcolor{#1f5fbf}{996} \]
Check: multiply back
Why: returns 996 heads
Worked example
Figure (svg): A bar of 2,000 tosses with 996 heads filled blue
\[ \textcolor{#6b7280}{0.5} \times 2000 = \textcolor{#6b7280}{1000} \]
Multiply the tosses by 0.5
Why: heads expected over the long run
Figure (svg): A bar of 2,000 tosses with 996 heads filled blue and a grey dashed line at half: 1,000
\[ \textcolor{#6b7280}{1000} - \textcolor{#1f5fbf}{996} = 4 \]
Subtract the author's 996 heads
Why: compares prediction with what happened
4 heads short in 2,000: close, not exact.
\[ 4 \div 2000 = \textcolor{#6b7280}{0.002} \]
Check: shortfall per toss
Why: matches the proportion's miss of 0.002
Trap
So far: HHH; one toss left.
\[ 4 \div 2 = 2 \text{ tails 'owed'} \]
Halve 4 tosses
Why: assumes runs split evenly
\[ \text{tails so far} = 0 \]
Count tails in HHH
Why: none yet: tails 'due'
\[ \text{toss 4: surely T} \]
Predict tails
Why: fails: no coin memory
\[ \text{HHHH},\ \text{HHHT} \]
List the runs starting HHH
Why: two grid cells qualify
\[ 1 \div 2 = \textcolor{#6b7280}{0.5} \]
Divide 1 tails run by 2
Why: equally likely runs, one each
\[ \text{HHHH: 1 of 16} = \text{HHHT: 1 of 16} \]
Check: each run's share of 16
Why: equal shares: tails stays 0.5
Prediction
Predict first
A made-up app gives each of 10 similar days a 0.3 chance of rain.
What does 0.3 predict?
Correct: About 3 rainy days, maybe 2 or 4
Why: 0.3 is a long-run proportion of days like these, so 10 days give about 0.3 × 10 = 3 rainy days, with the same wobble short coin runs show. It says nothing about hours within one day.
Worked example
Figure (svg): Ten made-up similar days as a row of 10 cells
\[ \textcolor{#6b7280}{0.3} \times 10 = \textcolor{#1f5fbf}{3} \]
Multiply the days by 0.3
Why: rainy days expected over the run
Figure (svg): Ten made-up similar days as a row of 10 cells; row 3: 3 cells shaded as rainy
\[ \text{10 days: a short run} \]
Reject 'exactly 3'
Why: like 4 tosses, counts wobble
Figure (svg): Ten made-up similar days as a row of 10 cells; row 3: 3 cells shaded as rainy; rows 2 and 4 below: other possible 10-day runs with 2 and 4 rainy days
\[ \textcolor{#6b7280}{0.3}\text{: a proportion of days} \]
Reject 'rain for 0.3 of each day'
Why: it counts days, not hours
\[ 0 < \textcolor{#6b7280}{0.3} \]
Reject 'no rain'
Why: any chance above 0 allows rain
\[ \textcolor{#1f5fbf}{3} \div 10 = \textcolor{#6b7280}{0.3} \]
Check: divide back
Why: 3 of 10 matches the forecast
Faded example
Figure (svg): Zoomed proportion line from 0.49 to 0.51 with a grey dashed line at 0.5; row "2,000" shows a grey dot at 0.498; row "24,000" empty
Book: Karl Pearson tossed a coin 24,000 times and got 12,012 heads.
Fill in the blanks
proportion of heads = 0.5005; miss from 0.5 = 0.0005
Why: 12,012 ÷ 24,000 = 0.5005, and 0.5005 − 0.5 = 0.0005: a smaller miss than the author's 0.002 over 2,000 tosses.
Worked example
Figure (svg): Zoomed proportion line from 0.49 to 0.51 with a grey dashed line at 0.5; row "2,000" shows a grey dot at 0.498; row "24,000" empty
\[ 12{,}012 \div 24{,}000 = \textcolor{#6b7280}{0.5005} \]
Divide heads by tosses
Why: puts his run on the 0-to-1 scale
Figure (svg): Zoomed proportion line from 0.49 to 0.51 with a grey dashed line at 0.5; row "2,000" shows a grey dot at 0.498; row "24,000" shows a grey dot at 0.5005
\[ \textcolor{#6b7280}{0.5005} - \textcolor{#6b7280}{0.5} = \textcolor{#6b7280}{0.0005} \]
Subtract 0.5
Why: how far this longer run misses
\[ \textcolor{#6b7280}{0.0005} < \textcolor{#6b7280}{0.002} \]
Compare with the 2,000-toss miss
Why: tests whether a longer run helped
\[ \textcolor{#6b7280}{0.5005} \times 24{,}000 = 12{,}012 \]
Check: multiply back
Why: returns Pearson's 12,012 heads
Worked example
Figure (svg): Bars drawn to scale. Tosses in the run, axis 0 to 24 thousand: author 2,000, Pearson 24,000
\[ 24{,}000 \div 2000 = 12 \]
Divide Pearson's tosses by 2,000
Why: counts author-length runs inside his
Figure (svg): Bars drawn to scale. Tosses in the run, axis 0 to 24 thousand: author 2,000, Pearson 24,000, with Pearson's bar divided into 12 pieces each as long as the author's bar
\[ \textcolor{#6b7280}{0.002} \div \textcolor{#6b7280}{0.0005} = 4 \]
Divide the author's miss by Pearson's
Why: counts Pearson-sized misses inside it
Figure (svg): Bars drawn to scale. Tosses in the run, axis 0 to 24 thousand: author 2,000, Pearson 24,000, with Pearson's bar divided into 12 pieces each as long as the author's bar. Miss from 0.5, axis 0 to 0.002: author 0.002 divided into 4 pieces each as long as Pearson's 0.0005 bar
One pair of runs: a hint, not a law.
\[ 4 \times \textcolor{#6b7280}{0.0005} = \textcolor{#6b7280}{0.002} \]
Check: multiply back
Why: rebuilds the author's miss
Pattern
Check
Figure (svg): 100 surveyed families as a 10 by 10 grid of dots; the first 3 filled (amounts listed), the other 97 hollow
Try It 1.1: 100 Knoll families surveyed; three spent $65, $75, $95.
Check your understanding
A classmate calls the three amounts' average the statistic. Which verdict is right?
Answer: A
Why: The statistic is the mean uniform spending of the whole sample, all 100 families. Three families are only part of the data, so their average describes just those three.
Worked example
Figure (svg): 100 surveyed families as a 10 by 10 grid of dots, all hollow
\[ \textcolor{#1f5fbf}{65 + 75 + 95} = \textcolor{#1f5fbf}{235} \]
Add the 3 amounts
Why: tests the classmate
Figure (svg): 100 surveyed families as a 10 by 10 grid of dots; the first 3 filled (amounts listed), the other 97 hollow
\[ \textcolor{#1f5fbf}{235} \div \textcolor{#1f5fbf}{3} \approx \textcolor{#6b7280}{78.33} \]
Share among 3
Why: follows the classmate's method
\[ 100 - \textcolor{#1f5fbf}{3} = 97 \]
Subtract from 100
Why: families 78.33 ignores
Figure (svg): 100 surveyed families as a 10 by 10 grid of dots; the first 3 filled (amounts listed), the other 97 hollow; a grey note 97 unlisted, skipped
\[ \textcolor{#1f5fbf}{\$95} - \textcolor{#1f5fbf}{\$75} = \textcolor{#1f5fbf}{\$20} \]
Subtract two amounts
Why: equal units: averaging allowed
\[ \text{statistic} = \text{total of 100} \div 100 \]
Check the divisor
Why: must be 100, not 3
Check
Try It 1.2, one of its six blanks: athletes' heights, in metres.
Check your understanding
Which phrase describes the variable?
Answer: A
Why: A variable is a trait measured on each member: one athlete's height. The four numbers are data, the survey's average is the statistic, and all athletes in the university form the population.
Worked example
\[ \text{height of one athlete} \]
Test: measured on each member?
Why: yes, one value per athlete
\[ \textcolor{#1f5fbf}{1.82, 1.76, 1.69, 1.93} \]
Test the four numbers
Why: values the variable took: data
\[ \textcolor{#1f5fbf}{1.93} - \textcolor{#1f5fbf}{1.82} = \textcolor{#1f5fbf}{0.11} \text{ m} \]
Subtract two heights
Why: equal metre units: numerical
\[ \text{average height in the survey} \]
Test the survey-average phrase
Why: one sample's number: a statistic
\[ \text{all athletes in the university} \]
Test the last phrase
Why: the whole group studied: the population
\[ \textcolor{#1f5fbf}{1.82} \text{ m} = \text{one athlete's height} \]
Check: read one datum as the variable
Why: each value fits the variable's definition
Worked example
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 4 dots filled, the other 16 hollow
\[ \text{population: all 20;\ \ sample: the line's 4} \]
Sort the two groups
Why: the question covers 20
\[ X = \text{one first-year's supply spending} \]
Name the variable
Why: dollars: equal units
\[ \text{statistic } \textcolor{#6b7280}{250},\ \text{parameter } \textcolor{#6b7280}{150} \]
Pair the two means
Why: to size the overstatement
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 4 dots filled; grey lines at 150 and 250
\[ \textcolor{#6b7280}{250} \div \textcolor{#6b7280}{150} \approx 1.67 \]
Divide by the parameter
Why: overstatement as a ratio
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; the line's 4 dots filled; grey lines at 150 and 250 with an arrow marked 1.67 times 150
\[ 1.67 \times \textcolor{#6b7280}{150} = 250.5 \approx \textcolor{#6b7280}{250} \]
Check: multiply back
Why: rounds to 250
Worked example
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw A's dots filled, the rest hollow
\[ \textcolor{#6b7280}{150} - \textcolor{#6b7280}{137.5} = \textcolor{#6b7280}{12.5} \]
Measure draw A's miss
Why: visible only in made-up Pine
Figure (svg): Dot plot of Pine College's 20 first-years' supply spending on an axis from $50 to $300: stacks of 3 at 50, 5 at 100, 5 at 150, 4 at 200, 2 at 250, 1 at 300; draw A's dots filled; grey lines at 150 and 137.5
\[ \textcolor{#6b7280}{212.5} - \textcolor{#6b7280}{100} = \textcolor{#6b7280}{112.5} \]
Subtract draw D's mean from draw B's
Why: sizes what four names can vary
Figure (svg): Sample rows on a spending axis from $0 to $300. row A: dots at 50, 100, 150, 250 and a grey diamond at its mean 137.5; row B: dots at 100, 200, 250, 300 and a grey diamond at its mean 212.5; row C: dots at 50, 150, 200, 200 and a grey diamond at its mean 150; row D: dots at 50, 100, 100, 150 and a grey diamond at its mean 100; a grey dashed line at the parameter, all 20: 150, through every row
\[ \text{report: about } \$137.50 \text{, from 4} \]
Label it an estimate
Why: that spread is why 'about'
\[ \textcolor{#6b7280}{212.5} - \textcolor{#6b7280}{112.5} = \textcolor{#6b7280}{100} \]
Check: take the spread off draw B
Why: lands on a mean already computed
Recap
OpenStax Introductory Statistics 2e, §1.1 Definitions of Statistics, Probability, and Key Terms §1.1, pp. 5-10 — Examples 1.1–1.4 and the Try Its trace back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.