The taxonomy of data and the methods for collecting it: qualitative data that name a category against quantitative data that are always numbers, and within quantitative the split between discrete data that are counted and continuous data that are measured; pie charts and bar graphs for qualitative data and the two situations where a pie chart is not merely ugly but wrong; the four random sampling methods — simple random, stratified, cluster and systematic — with convenience sampling as the non-random alternative; and the difference between sampling error, which every sample has and a larger sample reduces, and sampling bias, which no sample size can cure.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 1 — Sampling and Data
Data, Sampling, and Variation in Data and Sampling
Objectives
Five outcomes. The last one is the reason the other four matter.
OpenStax Introductory Statistics 2e, §1.2 Data, Sampling, and Variation in Data and Sampling §1.2, pp. 10-25 — the section these objectives are drawn from
Warm-up
Section 1.1 said a statistic is only as good as the sample it came from, and left the question of how the sample is chosen hanging.
Discussion prompt
A radio programme asks listeners to call in and say whether they support a new road. 4,000 people call, and 78 percent say no. Why should nobody report this as the public's view?
Hint: Ask who ended up in this sample, and who never had a chance of being in it.
Answer:
Nobody chose this sample: it chose itself. Only people listening to that programme, at that hour, who felt strongly enough to call and pay for the call, are in it — and people who are angry about a proposal call in far more often than people who are content with it.
The size is a distraction. Four thousand is a large number, and it makes the result no more trustworthy at all, because every one of the 4,000 came through the same filter. A poll of 400 properly chosen people would be far better evidence.
The book calls this a self-selected sample, and lists it among the things to be evaluated critically. This lesson is about the methods that avoid it, and about the vocabulary for saying exactly what has gone wrong.
Concept
Data may come from a population or from a sample, and most data fall into two categories: qualitative, which describes an attribute, and quantitative, which is always a number. How those data were gathered is a separate question, and it is the one that decides whether the sample can speak for the population at all.
qualitative and quantitative data — Qualitative data result from categorising or describing attributes, and are generally described by words or letters; they are also called categorical. Quantitative data are always numbers, resulting from counting or measuring attributes of a population.
\[ \text{data} \;\longrightarrow\; \begin{cases} \text{qualitative: a category} \\ \text{quantitative: a number} \end{cases} \]
Researchers often prefer quantitative data because it lends itself more easily to mathematical analysis — it makes no sense to find an average hair colour or blood type. But the preference does not license pretending: a categorical variable stays categorical however inconvenient that is, and section 1.1's postal-code example showed what happens when it is treated otherwise. The second half of this lesson turns to sampling, where the stakes are higher still, because a mistake there cannot be detected in the data at all.
Figure (svg): A tree classifying data into qualitative and quantitative, with quantitative splitting further into discrete and continuous
OpenStax Introductory Statistics 2e, §1.2 Data, Sampling, and Variation in Data and Sampling §1.2, p. 10
Section
Section 1
Concept
Qualitative data are the result of categorising or describing attributes of a population: hair colour, blood type, ethnic group, the car a person drives, the street a person lives on. They are generally described by words or letters. Quantitative data are always numbers, and are the result of counting or measuring.
qualitative data — Data that categorise or describe an attribute, generally recorded as words or letters. Also called categorical data. Averaging them is meaningless: there is no average hair colour and no average blood type.
\[ \text{qualitative: } \{ \text{black}, \text{brown}, \text{blonde} \} \qquad \text{quantitative: } \{ 6.2, 7.0, 9.1 \} \]
The book adds a warning worth carrying: you may collect data as numbers and report them categorically. Quiz scores recorded through the term are quantitative, but reported at the end as A, B, C, D or F they have become qualitative — information has been deliberately discarded in exchange for a simpler summary. The category a data set belongs to is a fact about how it is recorded now, not about how it was originally captured.
Figure (svg): A tree classifying data into qualitative and quantitative, with quantitative splitting further into discrete and continuous
OpenStax Introductory Statistics 2e, §1.2 Data, Sampling, and Variation in Data and Sampling §1.2, pp. 10-12 — the two categories, with the book's examples
Picture it
The first question splits everything; the second is asked only of the right-hand branch.
Figure (svg): A tree classifying data into qualitative and quantitative, with quantitative splitting further into discrete and continuous
Drawing this as a tree rather than three boxes in a row matters. Discrete and continuous are subdivisions of quantitative, not alternatives to qualitative, so the question 'is it discrete or continuous?' is meaningless for hair colour. Students who memorise three flat categories reliably try to answer it anyway.
Worked example
Example 1.7. A single situation containing all three types, which is the normal case.
\[ \text{Three cans of soup (19 oz, 14.1 oz, 19 oz), two packages of nuts, four vegetables, two desserts} \]
Find what was counted
Why: Three cans, two packages, four vegetables, two desserts.
Find what was measured
Why: The weights of the soups, to the tenth of an ounce.
Find what was named
Why: Tomato bisque, lentil, walnuts, broccoli, pistachio.
Note that one trip supplied all three
Why: The type belongs to the variable, not to the situation.
Figure (svg): The solution to Worked example one shopping trip, three kinds of data shown as a ladder of expressions, one row per legal move
\[ \text{counts} \to \text{discrete}, \quad \text{weights} \to \text{continuous}, \quad \text{kinds} \to \text{qualitative} \]
Verify: confirm the 14.1 could not have been a count
Why: A count of items can only be a whole number, so 14.1 ounces is decisive evidence that a measurement rather than a count produced it. The reverse check works too: the number of cans could never be 2.6. Whenever a data set contains a value with a fractional part, it cannot be discrete counting data — which makes scanning for a decimal point a fast first test.
OpenStax Introductory Statistics 2e, §1.2 Data, Sampling, and Variation in Data and Sampling §1.2, p. 11
Sorting
First question: amounts or names. Second question, only if amounts: counted or measured.
Sort into buckets
Sort each data set.
Items (b) and (e) come from the same five backpacks and land in different buckets, which is the point: the type belongs to the variable being recorded, never to the people or objects it was recorded from.
Worked example
Example 1.9, worked with the book's own hint in hand.
\[ \text{shoes owned; car driven; distance to the shop; classes taken; calculator used; dog weights; correct answers; IQ} \]
Apply the first question to each
Why: Car driven and calculator used name a make, not an amount.
For the rest, ask counted or measured
Why: Shoes owned, classes taken and correct answers are counts.
Distance and dog weights are measurements
Why: They can take fractional values as precisely as the instrument allows.
Handle the contentious one
Why: IQ scores are treated as continuous, being scaled measurements rather than counts.
Figure (svg): The solution to Worked example classifying eight variables shown as a ladder of expressions, one row per legal move
\[ \text{a, d, g discrete}; \quad \text{c, f, h continuous}; \quad \text{b, e qualitative} \]
Verify: confirm using the book's own hint
Why: The hint is that discrete data often start with the words 'the number of'. Read the three discrete answers back: the number of pairs of shoes, the number of classes, the number of correct answers. All three fit; none of the continuous ones does, because 'the number of distance' is not a phrase. The book flags IQ as likely to cause discussion, and it is worth knowing why: IQ scores are reported as whole numbers, so they look like counts, but nothing is being counted — they are a constructed scale, and the convention is to treat them as continuous.
OpenStax Introductory Statistics 2e, §1.2 Data, Sampling, and Variation in Data and Sampling §1.2, p. 12
Trap
\[ \text{data: } 94062, \; 94085, \; 95014 \quad \text{(postal codes)} \]
Call it quantitative because the values are numbers
Why: Every entry is written in digits, so the student stops there.
\[ \text{quantitative discrete} \quad \text{(wrong)} \]
The same reasoning would make a telephone number, a jersey number and a student identification number quantitative.
\[ \text{qualitative: the digits are a NAME for a place} \]
Ask what the value does, not what it looks like
Why: A quantitative value is an amount; a qualitative value identifies a group.
The decisive test is whether arithmetic on the values means anything. Adding two postal codes and halving gives a number that names no place, and the answer would change if the postal service renumbered — so the operation was never measuring anything. Compare the number of pairs of shoes, where the average is immediately meaningful. Digits are a notation, and notation is not evidence about what is being recorded.
Fill the middle
The book's own rule of thumb for spotting one of the three.
Fill in the blanks
\textdiscrete\; \text___ \;\text___ ___
Why: Counting produces whole numbers only, which is precisely what discrete means. The hint is reliable in the forward direction but not backward: plenty of discrete data are described without the phrase, such as the count of children in a household or a shoe size. Use it to confirm a classification rather than to make one.
Two truths and a lie
All three concern the taxonomy.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and the tree shows why: discrete and continuous are subdivisions of quantitative data only. Hair colour is neither, and asking which of the two it is has no answer. This is the commonest error made by students who memorise three flat categories instead of a two-stage tree.
Discrimination
Every item here is quantitative. Only the second question remains.
Sort into buckets
Sort each variable.
Section
Section 2
Concept
Two graphs used to display qualitative data are pie charts and bar graphs. In a pie chart the categories are wedges of a circle, proportional in size to the percent in each category. In a bar graph the length of each bar is proportional to the number or percent in that category. A Pareto chart is a bar graph whose bars are sorted from largest to smallest.
pie chart and bar graph — A pie chart represents categories as wedges of a circle, proportional to the percent in each. A bar graph represents them as bars whose lengths are proportional to the number or percent. A pie chart may only be used when the categories are exclusive and exhaustive, so that the percentages total exactly 100.
\[ \text{wedge angle} = 360^\circ \times \text{proportion in the category} \]
There are no strict rules about which graph to use, and it is a good idea to look at several. But there are two situations where a pie chart is not a matter of taste. If categories overlap, the percentages sum to more than 100 and the wedges would have to overlap too. If a category has been omitted, the percentages fall short of 100 and the circle cannot be completed honestly. In both cases the book's instruction is the same: use a bar graph.
Figure (svg): Two tables showing percentages that sum to more than 100 percent and to less than 100 percent, each marked as unsuitable for a pie chart
OpenStax Introductory Statistics 2e, §1.2 Data, Sampling, and Variation in Data and Sampling §1.2, pp. 13-17 — displays for qualitative data, and the two failures
Picture it
De Anza has 22,496 students; Foothill has 14,183. The raw counts cannot be compared directly.
Figure (svg): A bar chart comparing the percentage of full-time and part-time students at De Anza College and Foothill College
Displaying percentages alongside the numbers is often helpful, and it is particularly important when comparing sets that do not have the same totals. Foothill's part-time share, at 71.4 percent, is far above De Anza's 59.1 percent — a difference the raw counts actively hide, since De Anza has more part-time students in absolute terms while having proportionally fewer.
Worked example
The book's De Anza characteristics table. Check before drawing, not after.
\[ \text{full-time } 40.9\%, \;\text{ intend to transfer } 48.6\%, \;\text{ under } 25: 61.0\% \]
Add the percentages
Why: The three values as given.
\[ 40.9 + 48.6 + 61.0 = 150.5 \]
Compare with 100
Why: The total overshoots by more than half again.
\[ 150.5\text{ exceeds } 100 \]
Diagnose why
Why: A student may be full-time, intending to transfer AND under 25 at once.
Choose the legal display
Why: Bars may exceed a total; wedges may not.
Figure (svg): The solution to Worked example may this table be a pie chart shown as a ladder of expressions, one row per legal move
\[ \text{total} = 150.5\% > 100\% \;\Longrightarrow\; \text{no pie chart} \]
Verify: confirm the failure is about overlap, not about arithmetic
Why: Nothing here is miscounted: each percentage is correct on its own. What fails is the assumption a pie chart makes, that every member falls into exactly one wedge. These three characteristics are not alternatives to one another, so a single student is counted in up to three of them. A bar graph makes no such assumption — it simply reports three independent percentages side by side — which is why it stays legal.
OpenStax Introductory Statistics 2e, §1.2 Data, Sampling, and Variation in Data and Sampling §1.2, p. 15
Elimination
A survey asks respondents which of five streaming services they subscribe to. Many subscribe to several, and the percentages total 214.
Eliminate the wrong options
Eliminate the three unsuitable options and keep the one that works.
Survives elimination: B
Why: The survivor is the bar graph. Bars carry no requirement that the categories be exclusive or that they sum to anything in particular — each bar simply reports one percentage against a common scale. The rescaling option in choice C is worth dwelling on, because it is the tempting wrong answer: it produces a chart that looks correct and reports numbers that are not true of anybody.
Worked example
The book's ethnicity table, with the Other/Unknown category omitted.
\[ \text{listed frequencies do not add to the total number of students} \]
Add the listed frequencies
Why: Asian, Black, Filipino, Hispanic/Latino, Native American and the rest.
Identify what is missing
Why: Students who declined to respond or fitted no listed category.
Decide on the display
Why: The circle would have a gap of unexplained size.
Note the repair
Why: Restore the missing category and the percentages total 100 again.
Figure (svg): The solution to Worked example percentages that fall short shown as a ladder of expressions, one row per legal move
\[ \sum \text{percentages} < 100\% \;\Longrightarrow\; \text{a category is missing} \]
Verify: confirm this failure is repairable and the previous one is not
Why: These two failures are not the same. A missing category can be put back, and the book's later pie charts do exactly that, including Other/Unknown so the percentages reach 100. Overlapping categories cannot be repaired that way — no amount of adding categories makes 'full-time' and 'under 25' mutually exclusive. So a total below 100 signals an omission to be fixed, while a total above 100 signals that the pie chart was the wrong instrument from the start.
OpenStax Introductory Statistics 2e, §1.2 Data, Sampling, and Variation in Data and Sampling §1.2, pp. 15-17
Error analysis
Using the book's enrolment table: De Anza 9,200 full-time and 13,296 part-time out of 22,496; Foothill 4,059 full-time and 10,124 part-time out of 14,183. Four students draw conclusions.
Annotate
On: \( \begin{aligned} &(1)\; \text{De Anza has more part-time students, so it is the more part-time college} \\ &(2)\; 59.1\% \text{ against } 71.4\% \\ &(3)\; \text{Foothill is smaller, so its percentages are less reliable} \\ &(4)\; \text{the four percentages sum to } 200\%, \text{ so something is wrong} \end{aligned} \)
Errors (1) and (4) are the two ways a percentage table gets misread: comparing across totals using counts, and adding across groups that were never meant to be pooled. Both are avoided by asking, before any comparison, what each percentage is a percentage OF.
Matching
Four displays named in this section.
Match the pairs
Why: The Pareto chart is only a sorting convention rather than a different kind of graph, but sorting is a real gain: it puts the biggest category first and makes the ranking readable without comparing bar heights by eye. The book makes the same point about pie charts, noting that a wedge-sorted pie is more informative than an alphabetical one.
Estimation
A pie chart of four categories has one wedge occupying a right angle.
Predict first
What percent of the total is that category?
Correct: 25 percent.
Why: A wedge's angle is 360 degrees times the proportion, so a right angle of 90 degrees is a quarter of the circle and therefore 25 percent. This works in both directions and is worth being fluent in: a category holding 20 percent takes a 72 degree wedge, and one holding half takes a straight line across the circle. The counts are not needed, which is exactly what a pie chart is for — it shows shares and deliberately discards the totals.
Prediction
Commit before reasoning.
Predict first
Why does the book print percentages beside the counts for the two colleges?
Correct: Because the totals differ, so the counts are not comparable.
Why: De Anza's 22,496 students against Foothill's 14,183 means De Anza has more students in nearly every category simply by being larger. Dividing by each college's own total removes the size difference and leaves the composition, which is what the comparison is about. The other options invert the truth in various ways: percentages are derived from counts rather than easier than them, and they are less informative rather than more accurate, since they discard the totals entirely.
Section
Section 3
Concept
A sample should have the same characteristics as the population it represents, and statisticians use random sampling to achieve that. In each form of random sampling, each member of the population initially has an equal chance of being selected. The four common methods differ in how the chance process is organised.
simple random sample — A sample chosen so that any group of n individuals is equally likely to be selected as any other group of n individuals. It is the strongest of the methods and the one the others are compared against.
\[ \text{SRS: every group of size } n \text{ equally likely} \]
Stratified and cluster sampling are the pair students confuse, and the difference is exactly one word each. Stratified takes SOME members from EVERY group; cluster takes ALL members of SOME groups. Stratifying a college by department means sampling within every department; clustering means picking four departments and surveying everyone in them.
Figure (svg): Four panels showing the same population of forty individuals with different selection patterns for simple random, stratified, cluster and systematic sampling
OpenStax Introductory Statistics 2e, §1.2 Data, Sampling, and Variation in Data and Sampling §1.2, pp. 17-19 — the four methods, with the book's college examples
Picture it
The same forty people, sampled four ways. Look at the pattern of filled circles.
Figure (svg): Four panels showing the same population of forty individuals with different selection patterns for simple random, stratified, cluster and systematic sampling
The cluster panel is the one that should look alarming: two entire rows are taken and two are ignored completely. That is legitimate, and it is much cheaper than the alternatives when the clusters are geographically spread, but it means the sample's quality depends entirely on whether the clusters happen to resemble one another. Stratified sampling is the opposite bet, deliberately reaching into every group so that no stratum can be missed.
Worked example
The book's own example. Lisa needs three people from a class of 31, numbered 00 to 30.
\[ \text{random numbers: } 0.94360; \; 0.99832; \; 0.14669; \; 0.51470; \; 0.40581; \; 0.73381; \; 0.04399 \]
Read the first number in two-digit groups
Why: 94, 43, 36, 60 — all above 30.
Read the second
Why: 99, 98, 83, 32 — again all above 30.
Read the third
Why: 14 is in range and corresponds to Macierz.
Continue, skipping repeats and out-of-range values
Why: The fifth gives 05 for Cuningham and the seventh gives 04 for Cuarismo.
Figure (svg): The solution to Worked example drawing a simple random sample shown as a ladder of expressions, one row per legal move
\[ 14 \to \text{Macierz}, \quad 05 \to \text{Cuningham}, \quad 04 \to \text{Cuarismo} \]
Verify: confirm that discarding numbers does not spoil the randomness
Why: Values above 30 are discarded rather than reduced into range, and that matters. Wrapping 94 around by subtracting 62 would give some identification numbers two ways of being selected and others one, breaking the equal-chance requirement. Discarding is wasteful of random digits and perfectly fair, which is the correct trade. The fourth number also contained 14, and it was skipped as a repeat because each member may be selected only once — sampling without replacement.
OpenStax Introductory Statistics 2e, §1.2 Data, Sampling, and Variation in Data and Sampling §1.2, p. 18
Sorting
Read for what happens after the population is organised.
Sort into buckets
Sort each design.
Items (b) and (f) are the same design at different scales, and both are cluster. The phone-book example in (d) is the book's own, and it is worth noting that systematic sampling is chosen mainly because it is simple to carry out.
Worked example
The same college, the same faculty, two different designs.
\[ \text{Sampling faculty at a college organised into 20 departments} \]
Design A: number the members of every department and draw from each
Why: Some faculty from all 20 departments appear.
Design B: number the 20 departments and draw four of them
Why: All faculty from 4 departments appear; 16 departments contribute nobody.
Ask which groups are represented
Why: A reaches every department; B reaches four.
Ask which is cheaper to run
Why: B requires visiting four departments rather than twenty.
Figure (svg): The solution to Worked example telling stratified from cluster shown as a ladder of expressions, one row per legal move
\[ \text{stratified: some from ALL} \qquad \text{cluster: ALL from some} \]
Verify: confirm by asking what each design cannot detect
Why: The cluster design cannot detect anything peculiar to the sixteen departments it missed: if engineering faculty differ systematically from the rest and engineering was not drawn, nothing in the data will reveal it. The stratified design cannot miss a department, but it costs far more to administer. Neither is simply better — they trade coverage against cost, and naming which trade a study made is the point of being able to tell them apart.
OpenStax Introductory Statistics 2e, §1.2 Data, Sampling, and Variation in Data and Sampling §1.2, p. 19
Trap
\[ \text{four departments were chosen at random and everyone in them surveyed} \]
Call it stratified, because the population was divided into groups
Why: Both methods start by dividing into groups, so the student stops at the shared step.
\[ \text{stratified} \quad \text{(wrong)} \]
Dividing into groups is common to both. What distinguishes them is what happens next, and that step was skipped.
\[ \text{cluster: ALL members of SOME groups} \]
Ask which groups end up represented in the sample
Why: Every group, or only the chosen ones?
If sixteen of the twenty departments contribute nobody, it is a cluster sample. If all twenty contribute somebody, it is stratified. A one-question test that never fails: count the groups with at least one member in the sample, and compare with the number of groups in the population. Equal means stratified; fewer means cluster.
Ranking
Four designs for surveying a college's faculty, from cheapest to run to most expensive.
Put in order
Why: Convenience costs almost nothing and buys almost nothing. Cluster is cheap because only four departments must be visited. Systematic needs a complete list but no organisation beyond it. Stratified is the most work: it needs a complete list broken into twenty groups, and a separate draw from each. The ordering is very nearly the reverse of the quality ordering, which is the honest reason poor designs persist.
Edge cases
A random start and then every nth member sounds unimpeachable.
Discussion prompt
Construct a list where taking every fourth item gives a badly unrepresentative sample, and say what property of the list caused it.
Hint: Make the list itself repeat with a period.
Answer:
Take a list of flats in a block where every fourth flat is a corner unit — flat 1, 5, 9 and so on. A systematic sample with an interval of four will consist entirely of corner units, or entirely of non-corner units, depending only on where it started.
Corner flats are larger and more expensive, so a survey of rents would come out badly wrong in one direction or the other, and it would do so with a perfectly random starting point.
The property that caused it is periodicity: the list has a repeating pattern whose period divides the sampling interval. Systematic sampling is safe when the list order is unrelated to the variable being measured, and dangerous when it is not — which is why an alphabetical list is usually fine and a list ordered by date, floor or shift is worth checking.
Fill the middle
Complete the two definitions so that the contrast is visible.
Fill in the blanks
\textall \textsome \; ___ \; \text___ \qquad \text___ \text___ \; ___ \; \text___
Why: The two definitions are word-for-word mirror images, and writing them one above the other is the fastest way to keep them apart. Stratified guarantees coverage of every group at the cost of effort; cluster buys cheapness by accepting that most groups will be missed entirely.
Section
Section 4
Concept
A type of sampling that is non-random is convenience sampling, which uses results that are readily available. Separately, sampling with replacement returns each chosen member to the population so it may be chosen again; sampling without replacement does not. Surveys are typically done without replacement.
convenience sampling — Sampling that uses results which are readily available, with no chance process involved. A software shop interviewing customers who happen to be browsing is the book's example. The results may be very good in some cases and highly biased in others.
\[ \text{with replacement: } P(\text{chosen}) \text{ constant} \qquad \text{without: } P \text{ changes as members are removed} \]
The replacement question is a mathematical issue only when the population is small. Most samples are taken from large populations and are small in comparison, so the chance of picking the same individual twice is very low and sampling without replacement is approximately the same as sampling with replacement. This is why chapter 4 can use the binomial distribution, which assumes replacement, for survey data that was collected without it.
Figure (svg): Two columns separating sampling error, which is unavoidable, from sampling bias, which is a fault in how the sample was chosen
OpenStax Introductory Statistics 2e, §1.2 Data, Sampling, and Variation in Data and Sampling §1.2, pp. 19-20 — convenience sampling, and sampling with and without replacement
Picture it
Two things that both make a sample wrong, and only one of them is a mistake.
Figure (svg): Two columns separating sampling error, which is unavoidable, from sampling bias, which is a fault in how the sample was chosen
The row to memorise is the third: sampling error shrinks as the sample grows, and sampling bias does not. That single asymmetry explains why the call-in radio poll with 4,000 responses is worse evidence than a properly drawn sample of 400, and why 'we surveyed a lot of people' is never an answer to the objection that the wrong people were surveyed.
Worked example
The book is careful here, and so should we be.
\[ \text{A software shop interviews customers browsing in the shop} \]
Identify what the sample can support
Why: These are people who buy software in shops.
Identify what it cannot support
Why: It excludes everyone who buys online or does not buy at all.
Recall the book's own verdict
Why: Results may be very good in some cases and highly biased in others.
State the deciding question
Why: Does the convenient group coincide with the population of interest?
Figure (svg): The solution to Worked example is convenience sampling always useless shown as a ladder of expressions, one row per legal move
\[ \text{convenient group} = \text{population?} \;\Longrightarrow\; \text{usable} \]
Verify: confirm with a case where convenience is the right choice
Why: If the shop wants to know how its own browsing customers rate the store layout, the browsing customers are the population, and interviewing them is not a convenience sample masquerading as something better — it is a census attempt on the right group. The method becomes biased only when the results are read as describing software buyers generally. So the fault is rarely in the interviewing; it is in the population silently attached to the answer afterwards.
OpenStax Introductory Statistics 2e, §1.2 Data, Sampling, and Variation in Data and Sampling §1.2, p. 19
Discrimination
Ask whether a larger sample, chosen the same way, would help.
Sort into buckets
Sort each problem.
Worked example
Two draws, two population sizes, and the difference between them.
\[ \text{Drawing 2 members at random from a group of 10, then from a group of 10,000} \]
Small population, with replacement
Why: The second draw sees all 10 again.
Small population, without replacement
Why: The second draw sees 9.
Large population, with replacement
Why: The second draw sees all 10,000.
Large population, without replacement
Why: The second draw sees 9,999.
Figure (svg): The solution to Worked example when replacement actually matters shown as a ladder of expressions, one row per legal move
\[ \frac{1}{10} \text{ against } \frac{1}{10000} \]
Verify: confirm why this licenses a later approximation
Why: The chance of drawing the same person twice from a large population is so small that the two schemes give nearly identical answers, which is exactly the book's point. It is not merely a curiosity: chapter 4's binomial distribution assumes each trial has the same probability, which is true with replacement and false without it. Survey data are collected without replacement and analysed with the binomial anyway, and this paragraph is the justification. Chapter 4's hypergeometric distribution is the version that does the arithmetic exactly, for the small populations where the approximation fails.
OpenStax Introductory Statistics 2e, §1.2 Data, Sampling, and Variation in Data and Sampling §1.2, p. 20
Trap
\[ \text{a call-in poll of } 400 \text{ gives } 78\% \text{ opposed} \]
Worry that 400 is too few and gather 4,000 instead
Why: The student treats every problem with a sample as a problem of size.
\[ n: 400 \to 4000 \quad \text{(the bias is unchanged)} \]
All 4,000 arrived through the same filter. The estimate is now more precisely centred on the wrong value.
\[ \text{redesign the sampling, then } n = 400 \text{ suffices} \]
Ask first whether every member had a chance of selection
Why: Only if the answer is yes does sample size become the relevant question.
Sampling error and sampling bias respond to completely different remedies. Error is reduced by taking more people; bias is reduced by taking DIFFERENT people, chosen by a chance process. Increasing the size of a biased sample narrows the confidence interval around a wrong number, which makes the result worse rather than better, because it is now wrong and confident.
Sorting
The book lists common problems to watch for. Match each scenario to one.
Sort into buckets
Sort each study.
The book's own instruction on the self-funded case is worth quoting: do not automatically assume the study is bad, but do not automatically assume it is good either. Evaluate it on its merits and the work done. That is harder than a rule and it is the honest position.
Two truths and a lie
All three concern samples going wrong.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and it is the most expensive misconception in the section. Bias comes from some members being less likely to be chosen than others, and enlarging the sample while keeping the method leaves that inequality exactly as it was. The larger sample simply estimates the biased quantity more precisely. Only changing how members are selected can remove bias.
Socratic
The book calls sampling with replacement truly random.
Discussion prompt
If sampling with replacement is the truly random version, why do real surveys almost always sample without it?
Hint: Imagine the interviewer arriving at a door for the second time.
Answer:
Because asking the same person the same question twice adds no information and costs as much as asking a new person. A survey of 1,000 draws that happened to include somebody three times would really be a survey of 998 people at the price of 1,000.
The mathematics prefers replacement because it keeps every draw identical and independent, which makes the formulas simpler. Practice prefers no replacement because it maximises information per unit of cost.
The resolution is the approximation in this section: for a large population the two schemes give nearly the same answers, so surveys take the practical option and the theory is applied as though the other had been used. When the population is small enough for that to fail, chapter 4's hypergeometric distribution does the exact calculation instead.
Section
Section 5
Concept
In reality a sample will never be exactly representative of the population, so there will always be some sampling error. As a rule, the larger the sample, the smaller the sampling error. Variation is present in data itself, in the samples drawn from a population, and in the results those samples produce.
variation — Differences among data values, among samples drawn from the same population, and among the results those samples give. Variation in data is a property of the world; variation among samples is a consequence of sampling, and it is what sampling error measures.
\[ \text{larger } n \;\Longrightarrow\; \text{smaller sampling error, but never zero} \]
Two soft drink cans from the same production line do not contain identical amounts, and two samples of 500 doctors do not give identical proportions. The first is variation in the data and it is a fact about the manufacturing process; the second is variation among samples and it is a fact about sampling. Distinguishing them matters, because the first cannot be reduced by studying harder and the second can be reduced by sampling more.
Figure (svg): Two columns separating sampling error, which is unavoidable, from sampling bias, which is a fault in how the sample was chosen
OpenStax Introductory Statistics 2e, §1.2 Data, Sampling, and Variation in Data and Sampling §1.2, pp. 20-25 — sampling error, variation in data and in samples
Picture it
One column shrinks with sample size. The other does not move.
Figure (svg): Two columns separating sampling error, which is unavoidable, from sampling bias, which is a fault in how the sample was chosen
Everything in chapters 7 and 8 is machinery for the left column: given a properly drawn sample, how much does the statistic move, and how wide an interval must be reported to account for it? There is no chapter for the right column, because there is no calculation that repairs a sample which never gave some members a chance.
Worked example
The book's collaborative exercise. For each, ask who could not have been chosen.
\[ \text{To find the mean grade point average of all students at a university, sample only honours students.} \]
Ask who is in the population
Why: All students at the university.
Ask who could be selected
Why: Only honours students, who are selected FOR high grades.
Predict the direction of the distortion
Why: Honours students have higher averages by definition.
Name the fault
Why: Some members had no chance of selection.
Figure (svg): The solution to Worked example judging four proposed samples shown as a ladder of expressions, one row per legal move
\[ \bar{x}_{\text{honours}} > \mu_{\text{all students}} \]
Verify: confirm the direction can be predicted in advance
Why: This bias is not merely present but signed: because honours status is awarded for high grades, the sample's mean must exceed the population's. Being able to predict the direction is diagnostic — when you can say in advance which way an estimate will be wrong, the sample is certainly biased rather than merely unlucky, since genuine sampling error is as likely to fall one way as the other.
OpenStax Introductory Statistics 2e, §1.2 Data, Sampling, and Variation in Data and Sampling §1.2, pp. 20-21
Prediction
Commit before reasoning.
Predict first
A call-in poll gets 4,000 responses; a properly drawn random sample gets 400. Which is better evidence about public opinion?
Correct: The random sample of 400.
Why: Size reduces sampling error and does nothing about bias, and the call-in poll's problem is entirely bias: only people who listened and felt strongly enough to call are in it. The random sample's 400 responses carry a larger sampling error, which is measurable and can be reported as an interval, and that is a far better position than an unmeasurable distortion of unknown size. The last option is too pessimistic — a properly drawn sample of 400 supports a national estimate to within about five points.
Worked example
The book's second scenario, which is subtler than the first.
\[ \text{To find the most popular cereal among under-tens, speak to every twentieth child entering a supermarket for three hours.} \]
Note what is done well
Why: Every twentieth child is a genuine systematic rule, not a hand-picked choice.
Ask what the list of children really is
Why: Children entering ONE supermarket in ONE three-hour window.
Identify who cannot appear
Why: Children who shop elsewhere, at other times, or not at all.
Name the fault
Why: The rule is fair; the group it is applied to is not the population.
Figure (svg): The solution to Worked example a sample that looks systematic and is not fine shown as a ladder of expressions, one row per legal move
\[ \text{fair rule} + \text{wrong frame} = \text{biased sample} \]
Verify: confirm why this one fools people
Why: The systematic rule is genuinely good and does the job it can do: within that supermarket in that window, no child was favoured. What fails is one step earlier, in the list the rule was applied to. A three-hour window on a weekday afternoon catches children of parents who shop then, and one supermarket serves one neighbourhood at one price level. This is worth internalising, because a defensible selection method applied to an indefensible frame is the commonest way a study looks rigorous and is not.
OpenStax Introductory Statistics 2e, §1.2 Data, Sampling, and Variation in Data and Sampling §1.2, p. 21
Trap
\[ \text{sample 1: } 0.43, \quad \text{sample 2: } 0.47 \]
Conclude that one of the two studies was done wrongly
Why: The student expects two correct procedures to give the same answer.
\[ \text{one of these must be an error} \quad \text{(wrong)} \]
Both may be flawless. Two different sets of people were asked, and different people answer differently.
\[ \text{both correct; the gap is sampling variability} \]
Expect disagreement, and ask instead how large it should be
Why: The question is not whether the samples differ but whether they differ by more than chance explains.
This reframing is the doorway to the rest of the book. Once disagreement is expected rather than alarming, the useful question becomes quantitative: how far apart can two honest samples fall? Chapter 7 answers it with the Central Limit Theorem, chapter 8 turns the answer into an interval, and chapter 9 turns it into a test for whether an observed gap is larger than chance can account for.
Matching
Each problem, and the thing that actually addresses it.
Match the pairs
Why: Only the first is cured by a larger sample. The book's own advice covers the second: it is better for the person conducting the survey to select the respondents than to let them select themselves. The third is a nonsampling error, caused by a factor unrelated to the sampling process, and no amount of statistical care compensates for a broken instrument.
Estimation
Two independent random samples of 500 doctors, from the same directory, estimating the same proportion.
Predict first
What should you expect of their two proportions?
Correct: Close, but almost certainly not equal.
Why: Proper sampling does not produce agreement; it produces disagreement of a predictable size. Two samples of 500 estimating the same proportion typically land within a few percentage points of one another, and exact equality would be a mild surprise. The third option gives up too much — the disagreement is bounded in a way chapter 7 makes precise — and the fourth is meaningless here, since the samples are the same size and neither is authoritative.
Explain it
A classmate says a survey of a million people must be better than one of a thousand.
Discussion prompt
In four sentences or fewer, give them a case where the survey of a million is worse, and name the property that decides it.
Hint: Let the million be collected in a way that excludes people.
Answer:
Suppose the million responses come from a pop-up on one website and the thousand come from a random draw of the electoral roll. The million tells you about visitors to that site who like pop-up surveys, and no larger number of them will ever tell you about anyone else.
The property that decides it is whether every member of the population had a chance of being selected. If they did, size improves the estimate; if they did not, size only sharpens a distortion.
The historical example worth knowing: a 1936 magazine poll of over two million people predicted the wrong winner of a United States presidential election, while a properly drawn sample of a few thousand got it right. The two million were drawn from car and telephone owners, who were not typical voters in 1936.
Comparison
Fill the blanks. The last column is the one that decides whether a study is worth reading.
Comparison matrix
| Method | How members are chosen | Random? |
|---|---|---|
| Simple random | n drawn from the whole population at random | yes: every group of size n equally likely |
| Stratified | a proportionate random draw from EVERY stratum | yes |
| Cluster | clusters drawn at random, then ALL members of those | yes, though whole groups are missed |
| Systematic | a random start, then every nth on a list | yes, unless the list is periodic |
| Convenience | whoever is readily available | no: no chance process at all |
Only the last row is non-random, and it is the only one where no later chapter can quantify what went wrong. The four above it all give every member an initial chance of selection, which is precisely the property that makes inference legitimate.
Pattern
Six questions, asked before the numbers are looked at. The book's critical-evaluation list, put in the order that finds problems fastest.
If the answer to the first question reveals a group with no chance of selection, stop. Nothing later in the analysis can repair it, and reading the rest of the study will only tell you how precisely the wrong quantity was estimated.
OpenStax Introductory Business Statistics 2e, §1.2 Data, Sampling, and Variation in Data and Sampling §1.2 Data, Sampling, and Variation in Data and Sampling
Check
Classify before you display.
Check your understanding
A researcher records the exact time in seconds that each of 50 users spends on a page. What type of data is this?
Answer: A
Why: Time is measured rather than counted, and a finer instrument would give more decimal places, so the data are continuous.
Check
Name the method by what happens after the grouping.
Check your understanding
A city has 400 blocks. Twelve blocks are chosen at random and every household in those twelve is surveyed. Which method is this?
Answer: A
Why: Groups were drawn at random and then every member of the chosen groups was taken, which is exactly cluster sampling. The 388 blocks not chosen contribute nobody.
Check
Error or bias.
Check your understanding
A survey is mailed to 5,000 households and 600 are returned, mostly by people with strong views. What should the researcher do?
Answer: A
Why: The 600 responses are a self-selected sample, which is bias rather than error. The remedy is to change how respondents are chosen, and the book advises that the person conducting the survey should select them.
Real world
A restaurant chain wants to know how satisfied its customers are. Three designs are proposed: a card on each table inviting diners to rate the meal online; a tablet handed to every twentieth departing customer by a staff member; and a random draw from the loyalty-card database, contacted by telephone.
Discussion prompt
Rank the three by the quality of evidence they would produce, name the method each uses, and say for each exactly which group is excluded.
Hint: For each design, ask who has no chance of being counted.
Answer:
The table card is worst, and it is self-selected. Only diners motivated enough to go online afterwards respond, and motivation is concentrated among the delighted and the furious. The excluded group is everyone with an ordinary opinion, which is most people, so the results will look polarised whatever the truth is.
The loyalty-card draw is a simple random sample, but of the wrong population. The chance process is sound and every member of the database can be reached. What is excluded is every customer who never signed up — often occasional or first-time diners, who may be exactly the ones with the least favourable view. The frame is the problem, not the method.
The tablet handed to every twentieth departing customer is best. It is systematic sampling with a staff member doing the selecting rather than the customer, so self-selection is removed, and the frame is everyone who actually ate there. Its weakness is periodicity: if the interval happens to align with parties leaving together, it could over-sample large groups.
\[ \text{self-selected} \;<\; \text{random draw, wrong frame} \;<\; \text{systematic, right frame} \]
The ranking turns on the first question of the evaluation procedure every time — who could not have been counted — and the middle case is the instructive one, because it uses the strongest chance process of the three and still comes second. A good method applied to the wrong list is not rescued by the method.
Commit first
Answer, then rate your confidence honestly.
Predict first
Which of these does a larger sample NOT improve?
Correct: The bias from a self-selected sample.
\[ n \uparrow \;\Longrightarrow\; \text{error} \downarrow, \quad \text{bias unchanged} \]
Why: The first, second and fourth are three descriptions of the same thing, and all improve as the sample grows: more data means less sampling error, which means a tighter interval and a more precise estimate. Bias is different in kind. It arises because some members of the population were less likely to be selected than others, and drawing more members by the same faulty method reproduces the same inequality on a larger scale, giving a more precise estimate of the wrong quantity.
Explain it
They cannot keep stratified and cluster sampling apart, and have re-read the definitions four times.
Discussion prompt
In three sentences or fewer, give them a test that settles it every time, and an example of each.
Hint: Count the groups that end up with somebody in the sample.
Answer:
Tell them to count how many groups have at least one person in the final sample, and compare that with how many groups the population was divided into. If the two numbers are equal it is stratified; if the sample's number is smaller it is cluster.
Stratified: divide a college into its twenty departments and take a few faculty from each, so all twenty are represented. Cluster: pick four of those twenty departments and survey everyone in them, so sixteen departments contribute nobody.
\[ \text{stratified: some from ALL groups} \qquad \text{cluster: ALL from SOME groups} \]
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For data in digits, ask whether adding two values and halving would mean anything. For the pie chart, add the percentages and check they total exactly 100 with no category omitted and no overlap. For stratified against cluster, count the groups represented in the sample and compare with the number of groups in the population. For error against bias, ask whether a larger sample chosen the same way would help — if yes it is error, if no it is bias. Do five examples of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top, draw the data classification as a tree: one node splitting into qualitative and quantitative, and the quantitative branch splitting again into discrete and continuous. Write the question asked at each split, and put two examples in each of the three leaves, including one qualitative example written entirely in digits. Below that, draw a four-by-ten grid of dots four times over, and shade the selected members to show simple random, stratified, cluster and systematic sampling; beside each write one sentence naming what that method cannot detect. To the right, write the stratified and cluster definitions one directly above the other so the words ALL and SOME line up in opposite positions. At the bottom, make a two-column table headed sampling error and sampling bias, with four rows: what causes it, whether every sample has it, whether a larger sample helps, and whether any later chapter can measure it. Finally, write the six evaluation questions as a numbered list and apply the first one to this study: a website asks visitors to rate a new feature and reports that 81 percent approve.
If your two shaded grids for stratified and cluster look similar, redraw them: the cluster grid should have entire rows completely empty, and the stratified grid should have at least one shaded dot in every row. That visual difference is the definition.
Recap
Five things, and the last one is the reason this section is the longest in the chapter.
| If you see | Then |
|---|---|
| Values that name a group, digits or not | Qualitative: pie chart or bar graph, never a mean |
| Values from counting | Quantitative discrete |
| Values from measuring, with decimals possible | Quantitative continuous |
| Percentages totalling more than 100 | Categories overlap: bar graph only |
| Percentages totalling less than 100 | A category is missing: restore it or use a bar graph |
| Some members taken from every group | Stratified |
| All members taken from some groups | Cluster |
| Respondents who chose to reply | Self-selected: bias, and no sample size fixes it |
Section 1.3 takes the data this section classified and organises them into a frequency table, which is the first genuine summary in the course and the direct ancestor of every histogram in chapter 2.
OpenStax Introductory Statistics 2e, §1.2 Data, Sampling, and Variation in Data and Sampling §1.2, pp. 10-25 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.