The chapter's last section changes the design rather than the parameter. The three previous tests all required the two samples to be independent, and here they are not: two measurements are drawn from the same pair of individuals or objects, so a subject who scores high on the first will tend to score high on the second. The remedy is to calculate a difference for each pair, and those differences become the data. The population mean of the differences is then tested using a Student t test for a single population mean with n minus one degrees of freedom, where n is the number of differences — which is chapter 9's one-sample test applied to a single column of numbers. The design has a real advantage beyond correctness: because each subject serves as their own control, the variation between subjects is cancelled and only each subject's own change remains to be analysed.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 10 — Hypothesis Testing with Two Samples
Matched or Paired Samples
Objectives
Five outcomes, and the second is what makes the whole method work.
OpenStax Introductory Statistics 2e, §10.4 Matched or Paired Samples §10.4, pp. 527-531 — the section these objectives are drawn from
Warm-up
Sections 10.1 to 10.3 all began by requiring the two samples to be independent.
Discussion prompt
Eight subjects are measured for pain before hypnotism and again after. Are the two sets of eight measurements independent samples?
Hint: If you knew a subject's before score, would that tell you anything about their after score?
Answer:
No. A subject who felt a great deal of pain beforehand will probably still feel comparatively more afterwards, so knowing one measurement tells you a good deal about the other. The two columns are strongly linked, subject by subject.
That breaks the independence every test in this chapter has assumed so far, and the formulas built on it do not apply. Using them would treat sixteen linked numbers as sixteen independent ones and understate the uncertainty.
The fix is simpler than a new formula. Take each subject's after minus before, and analyse that single column — the pairing is used up in forming the differences, and what remains is one ordinary sample of eight numbers.
Concept
In a hypothesis test for matched or paired samples, subjects are matched in pairs and differences are calculated. The differences are the data. The population mean for the differences is then tested using a Student t test for a single population mean with n minus one degrees of freedom, where n is the number of differences.
the differences are the data — The book's own phrase. Once each pair's difference is computed, the two original columns play no further part and the test is chapter 9's one-sample t on a single column.
\[ t = \frac{\bar{x}_d - \mu_d}{s_d/\sqrt{n}}, \qquad \text{df} = n - 1 \]
The collapse is worth appreciating for what it avoids. There is no new distribution, no new standard error, and no new degrees-of-freedom formula — the whole of section 10.1's Aspin-Welch apparatus is unnecessary here because there is only one sample once the differences are taken. A design problem has been solved by an arithmetic step rather than by new theory.
Figure (svg): Two columns contrasting independent samples with matched pairs
OpenStax Introductory Statistics 2e, §10.4 Matched or Paired Samples §10.4, pp. 527-528
Section
Section 1
Concept
The characteristics are that simple random sampling is used, that sample sizes are often small, and that two measurements are drawn from the same pair of individuals or objects. Differences are then calculated, and those differences form the sample used for the test.
matched or paired — Either the same subject measured twice — before and after — or two subjects deliberately matched on characteristics that matter, with one assigned to each condition.
\[ \text{pair } i: \; (x_{1i}, x_{2i}) \;\to\; d_i = x_{2i} - x_{1i} \]
The design comes in two forms and both are covered by the same test. The commonest is a before-and-after measurement on one subject, as in Example 10.11. The other is genuine matching — twins, or patients paired on age and severity, with one of each pair receiving each treatment — where the pairing is created by the experimenter rather than by repeated measurement.
Figure (svg): Two columns contrasting independent samples with matched pairs
OpenStax Introductory Statistics 2e, §10.4 Matched or Paired Samples §10.4, pp. 527-528 — the six characteristics
Picture it
Six rows, and the fourth is the one that decides.
Figure (svg): Two columns contrasting independent samples with matched pairs
The third row is a useful practical tell: independent samples may differ in size and paired samples never can, since every pair contributes exactly one difference. Two groups of unequal size cannot be paired, whatever else is true of them.
Worked example
Example 10.11's study of hypnotism.
\[ \text{eight subjects, each measured before and after} \]
Count the groups of subjects
Why: One group of eight.
Count measurements per subject
Why: Before and after.
Ask about independence
Why: Same person twice.
Choose the method
Why: Take differences.
Figure (svg): The solution to Worked example identifying the design shown as a ladder of expressions, one row per legal move
\[ n = 8 \text{ differences}, \quad \text{df} = 7 \]
Verify: confirm the sample size is eight rather than sixteen
Why: There are sixteen measurements but only eight independent pieces of information about the effect of hypnotism, because the two measurements on each subject are linked. Counting sixteen would give fifteen degrees of freedom instead of seven and would understate the uncertainty substantially — the error that treating paired data as independent always makes.
OpenStax Introductory Statistics 2e, §10.4 Matched or Paired Samples §10.4, p. 528
Sorting
Ask whether the two observations in a pair are linked.
Sort into buckets
Sort each design.
Item (c) is worth noticing: two eyes on one patient are as linked as two measurements on one person, and eye studies are a standard example of paired data. Anything measured twice on the same unit belongs on the left.
Worked example
Distinguishing repeated measures from deliberate matching.
\[ \text{(a) one subject measured twice}; \quad \text{(b) twins split between treatments} \]
Case (a)
Why: The same person.
Case (b)
Why: Two people, matched.
What they share
Why: Linked observations.
The method
Why: Differences within each pair.
Figure (svg): The solution to Worked example the other kind of pairing shown as a ladder of expressions, one row per legal move
\[ \text{one sample of } n \text{ differences either way} \]
Verify: confirm why deliberate matching counts as pairing
Why: Twins share genetics and upbringing, so their responses are linked in the same way a person's two measurements are — and that link is precisely what the pairing is designed to exploit. What makes a design paired is not that the same body was measured twice but that the two observations in a pair are more alike than two randomly chosen ones would be.
OpenStax Introductory Statistics 2e, §10.4 Matched or Paired Samples §10.4, pp. 527-528
Trap
\[ \text{eight before-scores against eight after-scores, independently} \]
Treat the two columns as two samples
Why: There are eight numbers in each.
\[ \text{but the columns are linked subject by subject} \]
The between-subject spread — from 2.0 to 11.6 — would swamp the effect, and the test would find nothing.
\[ \text{take the eight differences and test their mean} \]
Use the pairing, which is the whole point of collecting the data that way
Why: Each subject is their own control.
An independent-samples t on these data gives a p-value around 0.10 rather than 0.0095 — a real effect goes undetected because the between-subject variation is left in. The pairing was designed into the study to remove exactly that variation, and ignoring it discards the design's main advantage.
Fill the middle
The book's fifth characteristic.
Fill in the blanks
\textdifferences ___ \text___
Why: The differences. Once they are computed the two original columns are set aside, and what remains is a single sample analysed by chapter 9's one-sample t test.
Two truths and a lie
All three concern the design.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. Eight pairs give eight differences and seven degrees of freedom, because the two measurements on a subject are not two independent pieces of information. Counting all sixteen would substantially understate the uncertainty.
Prediction
Commit before reasoning.
Predict first
What happens if paired data are analysed with an independent-samples test?
Correct: A real effect may go undetected.
Why: The independent-samples standard error includes all the variation between subjects, which pairing was designed to remove. For Example 10.11 that turns a p-value of 0.0095 into roughly 0.10 — the same data, and an opposite conclusion, because the design's advantage was discarded.
Section
Section 2
Concept
For each matched pair, subtract one measurement from the other. The book's instruction for Example 10.11 is to calculate after minus before, and the direction must be the same for every pair and must match the direction in the hypotheses.
the direction of subtraction — Fixed once for the whole problem. After minus before makes an improvement negative when a lower score is better, which is what the hypotheses must then reflect.
\[ d_i = \text{after}_i - \text{before}_i \]
The choice of direction is free but its consequences are not. Subtracting before minus after would flip every difference's sign, turn the mean from minus 3.125 to plus 3.125, and require a right-tailed alternative instead of a left-tailed one. Both routes give the same p-value; mixing them gives the wrong tail.
Figure (svg): Three columns of eight numbers, with the third boxed and an arrow showing it becoming the data for a one-sample test
OpenStax Introductory Statistics 2e, §10.4 Matched or Paired Samples §10.4, p. 528 — the difference table and the instruction to compute after minus before
Picture it
Example 10.11's table, with the column the test uses.
Figure (svg): Three columns of eight numbers, with the third boxed and an arrow showing it becoming the data for a one-sample test
The boxed column is the entire input to the test. Everything downstream — the mean, the standard deviation, the statistic, the degrees of freedom — is computed from those eight numbers and from nothing else.
Worked example
Example 10.11's eight subjects.
\[ \text{after minus before, for each of the eight} \]
Subject A
Why: 6.8 minus 6.6.
\[ 0.2 \]
Subject B
Why: 2.4 minus 6.5.
\[ -4.1 \]
Continue through H
Why: Each pair in turn.
Read the pattern
Why: Seven negative.
Figure (svg): The solution to Worked example computing the differences shown as a ladder of expressions, one row per legal move
\[ \{0.2, -4.1, -1.6, -1.8, -3.2, -2, -2.9, -9.6\} \]
Verify: confirm the signs mean what the problem requires
Why: A lower score indicates less pain, so a negative difference is an improvement — and seven of the eight are negative. Subject A's positive 0.2 is the one who got slightly worse. Checking that the signs point the way the context expects, before computing anything, is what catches a subtraction done in the wrong direction.
OpenStax Introductory Statistics 2e, §10.4 Matched or Paired Samples §10.4, p. 528
Faded example
Subject H: before 11.6, after 2.0, using after minus before.
Fill in the blanks
d = 2.0 - 11.6 = -9.6, \textimprovement ___
Why: Minus 9.6, the largest improvement in the study — a lower score means less pain, so a strongly negative difference is a strongly improved subject.
Worked example
The book asks the reader to verify these.
\[ \text{the eight differences} \]
Add them
Why: The total.
\[ -25 \]
Divide by eight
Why: The mean difference.
\[ -3.125 \]
Compute the sample sd
Why: From the eight values.
\[ 2.9114 \]
Note the sample size
Why: Eight differences.
\[ n = 8 \]
Figure (svg): The solution to Worked example the summary statistics shown as a ladder of expressions, one row per legal move
\[ \bar{x}_d = -3.125, \quad s_d = 2.9114 \]
Verify: confirm the mean is a plausible summary of the eight values
Why: The differences run from 0.2 down to minus 9.6, and a mean of minus 3.125 sits sensibly among them — closer to the middle of the cluster than to the extreme, as the single value of minus 9.6 pulls it down somewhat. The standard deviation of 2.91 is also plausible for values with that spread, and a much smaller one would suggest the outlying minus 9.6 had been dropped.
OpenStax Introductory Statistics 2e, §10.4 Matched or Paired Samples §10.4, p. 528
Error analysis
Before is 6.5 and after is 2.4; the book's convention is after minus before.
Annotate
On: \( \begin{aligned} &(1)\; 6.5 - 2.4 = +4.1 \\ &(2)\; |2.4 - 6.5| = 4.1 \\ &(3)\; \tfrac{2.4}{6.5} \approx 0.37 \\ &(4)\; 2.4 - 6.5 = -4.1 \end{aligned} \)
Error (2) is the most damaging because it looks tidy. Absolute differences are all positive, so their mean is positive whatever happened, and a test on them could never detect a direction of change.
Two truths and a lie
All three concern the differences.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and destroys the test. Absolute differences are all positive, so their mean is positive whether subjects improved or worsened — and the question of whether pain decreased could no longer be asked at all.
Prediction
Commit before reasoning.
Predict first
If before minus after were used instead, what changes?
Correct: The mean flips sign and the tail flips with it.
Why: Every difference changes sign, so the mean becomes plus 3.125 and improvement now shows as a positive difference — requiring a right-tailed alternative. The p-value is identical either way, provided the hypotheses are reversed to match. Reversing one without the other is what produces a wrong answer.
Faded example
The eight differences total minus 25.
Fill in the blanks
\bar8_d = \frac-3.125___} = ___
Why: Minus 3.125, the book's mean difference. The divisor is the number of PAIRS, not the number of measurements — eight rather than sixteen.
Section
Section 3
Concept
The hypotheses concern mu-d, the population mean of the differences. For Example 10.11 the null is that mu-d is zero or positive, meaning the same or more pain after hypnotism, and the alternative is that it is negative, meaning less pain.
mu sub d — The population mean of the differences. The subscript d denotes differences, and it is a single parameter — which is why chapter 9's one-sample machinery applies unchanged.
\[ H_0: \mu_d \ge 0, \qquad H_a: \mu_d < 0 \]
The book explains both hypotheses in words as well as symbols, and the explanation is worth following: the null being zero or positive means the subject shows no improvement, while the alternative being negative means the score should be lower after hypnotism, so the difference ought to be negative to indicate improvement. Getting that sign right is where the direction of subtraction and the direction of the claim have to agree.
Figure (svg): A t distribution with seven degrees of freedom and its left tail beyond minus three shaded
OpenStax Introductory Statistics 2e, §10.4 Matched or Paired Samples §10.4, p. 529 — the hypotheses and their explanations
Picture it
Example 10.11: are the measurements lower after hypnotism?
Figure (svg): A t distribution with seven degrees of freedom and its left tail beyond minus three shaded
The tail is on the left because improvement means a negative difference under the after-minus-before convention. Had the differences been formed the other way, the identical evidence would sit in the right tail — the picture would be mirrored and the p-value unchanged.
Worked example
Example 10.11, following the book's reasoning.
\[ \text{are the measurements, on average, LOWER after hypnotism?} \]
Fix the subtraction
Why: After minus before.
Read the claim
Why: Lower after.
Write the alternative
Why: The claim.
\[ H a: \mu _{d} < 0 \]
Write the null
Why: Its complement.
\[ H 0: \mu _{d} \ge 0 \]
Figure (svg): The solution to Worked example writing the hypotheses shown as a ladder of expressions, one row per legal move
\[ H_0: \mu_d \ge 0, \qquad H_a: \mu_d < 0 \]
Verify: confirm the null covers no improvement and worsening together
Why: The book's gloss is that the null being zero or positive means there is the same or more pain felt after hypnotism — so the null covers both no change and a worsening, and the alternative covers improvement alone. That partition is what section 9.1 required, and reading it in words rather than symbols is the check that the signs were assigned correctly.
OpenStax Introductory Statistics 2e, §10.4 Matched or Paired Samples §10.4, p. 529
Faded example
Differences are formed as after minus before, and a lower score is better.
Fill in the blanks
\textnegative < \text___ H_a: \mu_d ___ 0
Why: A lower after-score makes the difference negative, so the alternative claiming improvement is that mu-d is below zero — a left-tailed test.
Worked example
The same design with a different question.
\[ \text{does hypnotism CHANGE the measurements?} \]
Read the claim
Why: Change, either way.
Write the alternative
Why: Not equal to zero.
\[ H a: \mu _{d} \ne 0 \]
Write the null
Why: Equal to zero.
\[ H 0: \mu _{d} = 0 \]
Note the tail
Why: Both.
Figure (svg): The solution to Worked example a two-tailed paired test shown as a ladder of expressions, one row per legal move
\[ H_0: \mu_d = 0, \qquad H_a: \mu_d \ne 0 \]
Verify: confirm what the two-tailed version would give here
Why: The one-tailed p-value was 0.0095, so the two-tailed version is about 0.019 — still significant at 5 percent but not at 1 percent, where the one-tailed test would have been. Asking a directionless question costs statistical power, which is why a direction should be claimed when the science genuinely supplies one, and only then.
OpenStax Introductory Statistics 2e, §10.4 Matched or Paired Samples §10.4, p. 529
Trap
\[ d = \text{after} - \text{before}, \text{ but } H_a: \mu_d > 0 \text{ for improvement} \]
Write a greater-than because improvement sounds positive
Why: Improving is a good outcome.
\[ \text{but a LOWER score after gives a NEGATIVE difference} \]
The test would look in the right tail while all the evidence sits in the left, returning a p-value near 0.99.
\[ d = \text{after} - \text{before}, \; H_a: \mu_d < 0 \]
Let the subtraction's direction decide the sign of an improvement
Why: Then write the alternative to match.
The reliable order is to fix the subtraction first, work out what an improvement looks like in those terms, and only then write the alternative. Writing the hypotheses first and the subtraction afterwards is how the two come to disagree — and a p-value near 1 when the data plainly support the claim is the symptom.
Discrimination
Differences are always after minus before.
Sort into buckets
Sort each claim by its tail.
Two truths and a lie
All three concern the hypotheses.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. Once the differences are taken, the test is a one-sample test on a single parameter — the mean difference. The two original means never appear in the hypotheses, and there is no two-sample apparatus anywhere in this method.
Prediction
Commit before reasoning.
Predict first
The one-tailed p-value is 0.0095. What is the two-tailed p-value?
Correct: About 0.019.
Why: Doubling the one-sided area gives about 0.019, which still rejects at 5 percent but no longer at 1 percent. Asking a directionless question costs power, which is the reason to claim a direction when the science supplies one — and the reason a direction must never be chosen after seeing the data.
Section
Section 4
Concept
With the differences computed and the hypotheses written, the test is exactly chapter 9's one-sample t test. The statistic is the mean difference minus the hypothesised value, divided by the sample standard deviation of the differences over the square root of their number.
the paired t statistic — Identical in form to section 9.5's one-sample t. The only thing that identifies it as paired is where the data came from, which the arithmetic no longer knows about.
\[ t = \frac{\bar{x}_d - 0}{s_d/\sqrt{n}}, \qquad \text{df} = n - 1 \]
It is worth noticing how little remains to learn. There is no new distribution, no new standard error and no new degrees-of-freedom rule — a student who worked section 9.5 can run this test immediately once the differences exist. The whole content of this section is the recognition and the subtraction.
Figure (svg): A t distribution with seven degrees of freedom and its left tail beyond minus three shaded
OpenStax Introductory Statistics 2e, §10.4 Matched or Paired Samples §10.4, pp. 528-530 — the test statistic and Example 10.11
Picture it
A left-tailed t on seven degrees of freedom.
Figure (svg): A t distribution with seven degrees of freedom and its left tail beyond minus three shaded
The shaded left tail beyond minus 3.04 holds about 0.0095, comfortably below a 5 percent threshold. The mean improvement of 3.125 points is roughly three standard errors from zero, which is why the evidence is strong despite only eight subjects.
Worked example
Example 10.11, at the 5 percent level.
\[ \bar{x}_d = -3.125, \; s_d = 2.9114, \; n = 8 \]
Find the standard error
Why: 2.9114 over root 8.
\[ 1.0294 \]
Compute the statistic
Why: Minus 3.125 over that.
\[ t = -3.04 \]
Set the degrees of freedom
Why: Eight minus one.
\[ 7 \]
Find the p-value and decide
Why: Left tail; alpha exceeds it.
\[ 0.0095;\text{ reject} \]
Figure (svg): The solution to Worked example the hypnotism test in full shown as a ladder of expressions, one row per legal move
\[ t = \frac{-3.125}{2.9114/\sqrt{8}} \approx -3.04, \quad p \approx 0.0095 \]
Verify: confirm the result against the raw differences
Why: Seven of the eight subjects improved and the eighth worsened by only 0.2, so a strong result is unsurprising — and the mean improvement of over three points is large relative to the differences' own spread of 2.91. When a test's conclusion matches what the raw data plainly show, the arithmetic has probably been done correctly; a p-value near 1 on data like these would signal a wrong tail.
OpenStax Introductory Statistics 2e, §10.4 Matched or Paired Samples §10.4, pp. 528-530
Faded example
Mean difference minus 3.125, standard deviation 2.9114, eight pairs.
Fill in the blanks
t = \frac1.029-3.04} = \frac______} \approx ___
Why: The standard error is 2.9114 over about 2.828, which is 1.029, so the statistic is about minus 3.04 on seven degrees of freedom.
Worked example
Comparing with an independent-samples analysis of the same numbers.
\[ \text{the same sixteen measurements, analysed both ways} \]
Paired analysis
Why: On eight differences.
\[ p\text{ about } 0.0095 \]
Independent analysis
Why: Eight against eight.
\[ p\text{ about } 0.10 \]
Say why they differ
Why: Between-subject spread.
Note the consequence
Why: One rejects, one does not.
Figure (svg): The solution to Worked example what pairing bought shown as a ladder of expressions, one row per legal move
\[ 0.0095 \quad\text{against}\quad \approx 0.10 \]
Verify: confirm the paired analysis is the correct one here
Why: The data were collected as pairs, so the paired analysis is the valid one and the independent test would be simply wrong — not merely less powerful. But the comparison shows what the design achieved: subjects varied from 2.0 to 11.6 in their raw scores, and removing that variation is what made an effect of three points detectable with only eight people.
OpenStax Introductory Statistics 2e, §10.4 Matched or Paired Samples §10.4, pp. 528-530
Trap
\[ n = 16 \text{ measurements} \;\Rightarrow\; \text{df} = 15 \]
Count every number in the table
Why: There are sixteen measurements.
\[ \text{but there are only eight differences} \]
The t statistic would be divided by the root of 16 instead of 8, understating the standard error by about 40 percent.
\[ n = 8 \text{ differences} \;\Rightarrow\; \text{df} = 7 \]
Count the PAIRS, since the differences are the data
Why: Each pair contributes one number.
The book's phrase — where n is the number of differences — exists to forestall exactly this. Once the difference column is written down, the temptation disappears: it plainly has eight entries, which is why forming the column explicitly before doing anything else is worth the moment it takes.
Two truths and a lie
All three concern running the test.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The standard error divides by the square root of the number of DIFFERENCES — eight, not sixteen. Using sixteen would shrink the standard error by about 40 percent and overstate the evidence correspondingly.
Estimation
A mean difference of minus 3.125 with a standard error of about 1.03.
Predict first
Roughly how many standard errors from zero is that?
Correct: About 3.
Why: 3.125 divided by 1.03 is about 3.04. Three standard errors is a substantial departure on any distribution, so a small p-value was predictable before computing — which is the check worth making on every test in this chapter.
Explain it
A classmate used sixteen for n because the table has sixteen measurements.
Discussion prompt
In two sentences or fewer, correct them.
Hint: Ask how many numbers are in their difference column.
Answer:
Their difference column has eight entries, because each pair of measurements produces exactly one difference — and the differences are the data.
So n is 8 and the degrees of freedom are 7; using 16 would shrink the standard error by about 40 percent and make the evidence look far stronger than it is.
Section
Section 5
Concept
Because both measurements come from the same subject, the differences remove everything about that subject that does not change — their baseline sensitivity, their circumstances, their measurement idiosyncrasies. What remains is the change itself, which is what the question asks about.
between-subject variation — The spread among subjects' baseline levels, which in Example 10.11 runs from about 2 to nearly 12. Pairing cancels it entirely, leaving only the within-subject change.
\[ \text{spread of raw scores} \;\gg\; \text{spread of differences about their mean} \]
This is why paired designs are used wherever they are practical. Eight subjects sufficed to establish an effect here that an independent-samples design would have needed far more people to detect — and the book's second characteristic, that sample sizes are often small, is a consequence of that efficiency rather than a limitation of the method.
Figure (svg): Two dot plots, the left showing widely scattered raw measurements and the right showing the differences clustered near one value
OpenStax Introductory Statistics 2e, §10.4 Matched or Paired Samples §10.4, pp. 527-530 — the design's characteristics and Example 10.11
Picture it
The raw measurements against the differences.
Figure (svg): Two dot plots, the left showing widely scattered raw measurements and the right showing the differences clustered near one value
The left cloud is wide because people differ; the right one is tight because the effect of hypnotism is fairly consistent across them. Analysing the left as two independent groups would ask the effect to stand out against all that between-person spread, and it would not.
Worked example
Example 10.11's raw scores against its differences.
\[ \text{sixteen raw measurements against eight differences} \]
Range of raw scores
Why: 2.0 to 11.6.
\[ \text{about } 9.6 \]
Spread of differences
Why: Their standard deviation.
\[ 2.91 \]
Compare
Why: Much tighter.
Say what was removed
Why: Baseline differences.
Figure (svg): The solution to Worked example comparing the two spreads shown as a ladder of expressions, one row per legal move
\[ s_d = 2.91 \;\ll\; \text{the raw range of } 9.6 \]
Verify: confirm this is what makes eight subjects enough
Why: The test needed the mean difference to be large relative to the differences' spread, and 3.125 against 2.91 is comfortably so. Had the analysis used the raw scores' spread instead, the same effect would have been buried — which is precisely the calculation an independent-samples test performs, and why it fails to reject on these data.
OpenStax Introductory Statistics 2e, §10.4 Matched or Paired Samples §10.4, pp. 528-530
Sorting
Ask whether each unit can appear in both conditions.
Sort into buckets
Sort each situation.
Item (e) is a standard trick of experimental design: putting both treatments on the same car removes everything about that car's driver and roads. Wherever a unit has two comparable parts, pairing is usually available and usually worth taking.
Worked example
Recognising the design's limits.
\[ \text{comparing two medications given to different patients} \]
Ask what would be paired
Why: Each patient with themselves.
Note the reason
Why: A patient gets one medication.
Consider matching instead
Why: Pair on age and severity.
Otherwise
Why: Independent samples.
\[ \text{section } 10.3' s\text{ test} \]
Figure (svg): The solution to Worked example when pairing is not available shown as a ladder of expressions, one row per legal move
\[ \text{no pairing} \;\Rightarrow\; \text{an independent-samples test} \]
Verify: confirm why some studies cannot pair even in principle
Why: A treatment that permanently changes a patient — surgery, or a cure — cannot be undone to try the alternative, so a before-and-after design is unavailable and matching is the only route to pairing. Where subjects cannot be matched well either, an independent-samples design with a larger sample is the honest answer, and the extra subjects are the price of the lost efficiency.
OpenStax Introductory Statistics 2e, §10.4 Matched or Paired Samples §10.4, pp. 527-528
Trap
\[ \text{sort both groups and subtract the } i\text{th values} \]
Create pairs by matching positions after the fact
Why: Both groups have the same number of subjects.
\[ \text{the pairing is arbitrary and the differences meaningless} \]
Sorting manufactures an apparent link between unrelated subjects, and the resulting standard error understates the real uncertainty.
\[ \text{pair only what the DESIGN paired} \]
Use the pairing built into the data collection
Why: It has to exist before the data are analysed.
Pairing is a property of how a study was run, not an operation that can be applied to any two equal-sized samples. Inventing it afterwards — especially by sorting, which forces the differences to be small — produces a spuriously significant result from data that support nothing.
Two truths and a lie
All three concern the design's value.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. Pairing must be built into how the data were collected — the same subject measured twice, or subjects deliberately matched. Manufacturing pairs afterwards, by sorting or by arbitrary position, produces differences that mean nothing.
Estimation
The paired p-value is 0.0095; the independent-samples one on the same numbers is about 0.10.
Predict first
What does that comparison show?
Correct: Removing between-subject variation revealed the effect.
Why: Both analyses use identical numbers; only the pairing differs. The independent test asks the effect to stand out against the whole spread of people's baseline pain levels, and it cannot — while the paired test asks only whether each person improved, which they overwhelmingly did.
Prediction
Commit before reasoning.
Predict first
Pairing helps most when what is true?
Correct: When subjects differ a great deal.
Why: The more between-subject variation there is, the more pairing removes — and the bigger the advantage over an independent design. When subjects are nearly identical there is little to cancel, and the two designs perform similarly. Example 10.11's subjects ranged from 2.0 to 11.6, which is exactly the case where pairing pays.
Comparison
Fill the blanks. The last one is a design rather than a new parameter.
Comparison matrix
| Section | What is compared | Method |
|---|---|---|
| 10.1 | two means, sigmas unknown | t with Aspin-Welch df; do not pool |
| 10.2 | two means, sigmas known | a normal; no degrees of freedom |
| 10.3 | two proportions | a normal, with a pooled proportion |
| 10.4 | one group measured twice | a one-sample t on the differences |
The first three all require independence and the fourth exists because that requirement fails. Recognising which row a problem belongs to is the chapter's main skill, and it is settled by the design and the data type rather than by any calculation.
Pattern
Six steps, and the first two are where the method is decided.
Check that the sign of the mean difference matches what the raw data plainly show. A p-value near 1 on data that visibly support the claim means the tail and the subtraction disagree.
OpenStax Introductory Business Statistics 2e, §10.6 Matched or Paired Samples §10.6 Matched or Paired Samples
Check
The sample size.
Check your understanding
Eight subjects are each measured twice. How many degrees of freedom does the paired test have?
Answer: A
Why: There are eight differences, and the degrees of freedom are the number of differences minus one.
Check
The design.
Check your understanding
Which of these calls for a matched-pairs test?
Answer: A
Why: The same subjects are measured under both conditions, so the two sets of measurements are linked.
Check
The tail.
Check your understanding
Differences are after minus before, and lower scores are better. Which alternative claims improvement?
Answer: A
Why: A lower after-score makes the difference negative, so improvement corresponds to a negative mean difference.
Real world
A company trials a new keyboard layout. Twenty typists are timed on the old layout, then trained and timed again on the new one. Mean speed rises from 62 to 66 words per minute. An analyst runs an independent-samples t test on the two columns of twenty, gets a p-value of 0.21, and reports no evidence that the new layout is faster.
Discussion prompt
Assess the analysis, and say what should have been done.
Hint: Ask whether the two columns are independent.
Answer:
The wrong test was used. The same twenty typists appear in both columns, so the two sets of measurements are strongly linked — a fast typist is fast on both layouts. The independence that an independent-samples t test requires simply does not hold.
The consequence is predictable and one-directional. Typists vary enormously in baseline speed, perhaps from 40 to 90 words per minute, and the independent-samples standard error includes all of that variation. The four-word improvement has to stand out against it, and it cannot — which is exactly why the p-value came out at 0.21.
\[ d_i = \text{new}_i - \text{old}_i, \qquad t = \frac{\bar{x}_d}{s_d/\sqrt{20}}, \qquad \text{df} = 19 \]
The right analysis takes each typist's own improvement. If the twenty differences average 4 words per minute with a standard deviation of 6, the standard error is 6 over the square root of 20, about 1.34, giving a statistic near 3 and a one-tailed p-value under 0.005 — a clear result from the same data, because the between-typist variation has been removed.
Two cautions belong in a full answer, and both concern the design rather than the arithmetic. Every typist used the old layout first and the new one second, so improvement could reflect practice rather than the layout — a proper trial would randomise the order, half the typists starting with each. And a paired design cannot distinguish the layout's effect from anything else that changed between the two sessions, which is why the order effect matters enough to design around.
Commit first
Answer, then rate your confidence honestly.
Predict first
Why does a matched-pairs test use n minus one degrees of freedom rather than a two-sample formula?
Correct: Because the differences are a single sample.
\[ t = \frac{\bar{x}_d}{s_d/\sqrt{n}} \sim t_{n-1}, \quad n = \text{the number of PAIRS} \]
Why: Taking one difference per pair collapses two linked columns into one column of numbers, and the test is then chapter 9's one-sample t on that column — with n the number of differences. No two-sample machinery is involved at all, which is why there is no Aspin-Welch formula and no pooling question. Equal sample sizes are a consequence of pairing rather than its justification.
Explain it
They analysed before-and-after data with an independent-samples t test and found nothing.
Discussion prompt
In two sentences or fewer, explain what went wrong.
Hint: Ask whether the same people appear in both columns.
Answer:
The same subjects are in both columns, so the two sets of measurements are linked — and an independent-samples test asks the effect to stand out against all the variation between subjects.
Taking each subject's own difference removes that variation, and testing the mean of those differences will usually find an effect the other test buried.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, ask whether the same unit appears in both columns. For the second, fix the subtraction first, then work out what an improvement looks like, then write the alternative. For the third, count the difference column. For the fourth, pairing cancels each subject's baseline, and it helps most when subjects differ a lot. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Twenty-five minutes. This one closes the chapter, so make it a decision map.
Draw it
At the top, draw a four-way branch for the chapter's four tests, with the questions that select each: are the samples independent, are the data means or proportions, and are the standard deviations known. Under the fourth limb, write that the differences become the data. In the middle of the page, reproduce Example 10.11's table with its three columns — after, before, and difference — and box the difference column. Compute its mean of minus 3.125 and standard deviation of 2.9114, then the standard error of 1.029 and the statistic of minus 3.04, and draw a t curve with seven degrees of freedom shading the LEFT tail. Write the hypotheses in symbols and in words, noting that a negative difference means improvement. Beside that, draw two dot plots: the sixteen raw measurements scattered from 2 to 11.6, and the eight differences clustered near minus 3 — and write one sentence on what pairing removed. At the bottom, note the independent-samples p-value of about 0.10 for the same numbers, and write one sentence on why the two disagree.
Check the difference column by confirming it has eight entries, not sixteen — that single count settles the degrees of freedom and is where the commonest error in the section occurs. Check the tail by confirming the shading is on the side where the observed mean difference of minus 3.125 actually falls.
Recap
Five things, and the whole method rests on the second.
| If you see | Then |
|---|---|
| Before-and-after measurements | A matched-pairs test |
| Two measurements on the same unit | Paired, even if the unit is a car or a patient |
| Twins or deliberately matched subjects | Also paired |
| Two separate groups, one measurement each | Sections 10.1 to 10.3 |
| A difference column of n entries | df is n minus one |
| A p-value near 1 on convincing data | The tail and the subtraction disagree |
| Equal-sized groups with no real pairing | Do not manufacture pairs |
That closes chapter 10 and the tests built on means and proportions. Chapter 11 introduces a different distribution altogether: the chi-square, which compares an entire set of observed counts against what a model predicts, and so tests claims that no single mean or proportion could express.
OpenStax Introductory Statistics 2e, §10.4 Matched or Paired Samples §10.4, pp. 527-531 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.