What a study must do before its results can support a claim about cause: the anatomy of a randomized experiment with its explanatory and response variables, treatments and experimental units; lurking variables and why random assignment rather than random selection is what neutralises them; the control group and the placebo that make the power of suggestion measurable; blinding and double-blinding and what each one protects against; and the studies that cannot be randomized at all, where an observational design shows association and cannot establish cause. Closes with the ethical obligations a researcher carries toward the people being studied.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 1 — Sampling and Data
Experimental Design and Ethics
Objectives
Five outcomes. The third and fourth are the two devices that separate a trial from an anecdote.
OpenStax Introductory Statistics 2e, §1.4 Experimental Design and Ethics §1.4, pp. 34-37 — the section these objectives are drawn from
Warm-up
Section 1.2 listed causality among the things to evaluate critically: a relationship between two variables does not mean one causes the other.
Discussion prompt
Children with larger vocabularies have larger shoe sizes. Nobody thinks vocabulary grows feet. What is going on, and what would a study have to do to establish a genuine cause?
Hint: Ask what else differs between the children with small feet and the children with large feet.
Answer:
Age is driving both. Older children have bigger feet and know more words, so the two move together without either affecting the other. A third variable is producing the association.
This section gives that third variable a name — a lurking variable — and gives the remedy. To show that one thing causes another, a researcher must arrange matters so that the groups being compared differ in exactly one respect: the treatment imposed.
Section 1.2 could only warn you that correlation is not causation. This section supplies the design that earns the causal claim, and the whole of it turns on one step: assigning subjects to groups at random.
Concept
The purpose of an experiment is to investigate the relationship between two variables. When one causes change in another, the first is the explanatory variable and the affected one is the response variable. In a randomized experiment the researcher manipulates the explanatory variable and measures the resulting change, and the different values of the explanatory variable are the treatments.
randomized experiment — A study in which the researcher assigns experimental units to treatment groups at random and manipulates the explanatory variable, rather than merely observing what subjects do. The random assignment spreads all potential lurking variables equally among the groups, so the only remaining difference is the treatment.
\[ \text{explanatory variable} \;\longrightarrow\; \text{response variable} \]
The logic is worth stating exactly, because it is the argument the rest of the course serves. If the groups are alike in every respect except the treatment, and their outcomes differ, then the treatment is the only available explanation. Random assignment is what makes the first clause true — not by removing lurking variables, which is impossible, but by distributing them so evenly that none of them can account for a difference between the groups.
Figure (svg): A flow diagram of a randomized experiment: subjects recruited, randomly assigned to a treatment group and a control group, then the response measured in each
OpenStax Introductory Statistics 2e, §1.4 Experimental Design and Ethics §1.4, p. 34
Section
Section 1
Concept
An experimental unit is a single object or individual to be measured. The explanatory variable is the one the researcher manipulates; the response variable is the one measured for change. The different values of the explanatory variable are the treatments.
explanatory and response variables — The explanatory variable is manipulated by the researcher and is suspected of causing change. The response variable is measured to detect that change. The values the explanatory variable takes are called treatments, and an experimental unit is a single object or individual measured.
\[ \text{treatments} = \text{the values of the explanatory variable} \]
In the book's aspirin study, 400 people aged 50 to 84 are recruited and split at random into two groups; one takes aspirin daily for three years and the other takes a placebo, and the researchers count heart attacks in each. The population is people aged 50 to 84, the sample is the 400 participants, the experimental units are the individual people, the explanatory variable is oral medication, the treatments are aspirin and placebo, and the response variable is whether a subject had a heart attack.
Figure (svg): A flow diagram of a randomized experiment: subjects recruited, randomly assigned to a treatment group and a control group, then the response measured in each
OpenStax Introductory Statistics 2e, §1.4 Experimental Design and Ethics §1.4, pp. 34-35 — the parts of an experiment, with the aspirin study
Picture it
Left to right, and the order of the steps is part of the design.
Figure (svg): A flow diagram of a randomized experiment: subjects recruited, randomly assigned to a treatment group and a control group, then the response measured in each
Note where the randomization sits: after recruitment and before treatment. A study that recruits at random and then lets subjects choose their group has randomized the wrong step and has bought nothing. The placement of that single box is what separates an experiment from an observational study dressed up as one.
Worked example
Example 1.19. Work through the six in a fixed order every time.
\[ \text{400 people aged 50-84, split at random; one group takes aspirin daily, the other a placebo, for 3 years} \]
Population and sample
Why: The question is about people aged 50 to 84; 400 took part.
\[ \text{population } 50 - 84;\text{ sample the } 400 \]
Experimental units
Why: A single individual to be measured.
Explanatory variable
Why: What the researcher manipulates.
Treatments and response
Why: Its two values, and what gets measured.
Figure (svg): The solution to Worked example naming the six parts shown as a ladder of expressions, one row per legal move
\[ \text{explanatory} = \text{oral medication} \;\longrightarrow\; \text{response} = \text{heart attack} \]
Verify: confirm the placebo really is a treatment and not the absence of one
Why: It is easy to write the treatments as 'aspirin' alone and call the other group untreated, but the control group takes a pill every day for three years just as the treatment group does. The explanatory variable is oral medication, and it takes two values; placebo is one of them. Getting this right matters because the comparison being made is aspirin against placebo, not aspirin against nothing — and those are different questions, as the next idea shows.
OpenStax Introductory Statistics 2e, §1.4 Experimental Design and Ethics §1.4, p. 35
Sorting
Ask which variable the researcher could set in advance.
Sort into buckets
Sort each variable from a described study.
The pairs are one study each, and in every case the explanatory variable is the boring-sounding one. What a study is 'about' is usually its response variable, which is why naming the explanatory one requires the deliberate question of what could be assigned.
Worked example
Example 1.20. The scented mask study, which randomizes the ORDER rather than the group.
\[ \text{Subjects run mazes three times in floral-scented masks and three times in unscented masks.} \]
Name the explanatory and response variables
Why: Scent is manipulated; completion time is measured.
Name the treatments
Why: The two values of the explanatory variable.
Find what was randomized
Why: Which subjects wore the floral mask first.
Say why that suffices
Why: All subjects experienced both treatments, so the groups cannot differ in composition.
Figure (svg): The solution to Worked example an experiment where subjects get both treatments shown as a ladder of expressions, one row per legal move
\[ \text{scent} \;\longrightarrow\; \text{maze completion time} \]
Verify: confirm why randomizing the order was necessary at all
Why: Without it, every subject would wear the floral mask in the same position — say the last three trials — and practice would be perfectly confounded with scent. Subjects get faster at mazes with repetition, so the floral trials would be quicker whatever the scent did. Randomizing which half of the trials carries the scent spreads the practice effect evenly across both treatments, and it is the same logic as random assignment applied to a within-subject design.
OpenStax Introductory Statistics 2e, §1.4 Experimental Design and Ethics §1.4, pp. 35-36
Trap
\[ \text{does aspirin reduce heart attacks?} \]
Write the explanatory variable as 'heart attack' because that is what the study is about
Why: The student names the interesting variable rather than the manipulated one.
\[ \text{explanatory} = \text{heart attack} \quad \text{(wrong)} \]
Nobody assigned subjects to have heart attacks. The variable the researcher controls is which pill each person takes.
\[ \text{explanatory} = \text{medication}, \quad \text{response} = \text{heart attack} \]
Ask which variable the researcher sets, and which one is merely watched
Why: The one you can assign is explanatory; the one you wait for is the response.
A one-question test that never fails: could the researcher decide this variable's value for each subject before the study began? If yes it is explanatory, and if no it is the response. It also settles the studies where randomization is impossible — nobody can assign a birth order, which is exactly why that study cannot be an experiment.
Fill the middle
The vocabulary for the values an explanatory variable takes.
Fill in the blanks
\texttreatments ___
Why: Treatments are the values, not the groups. In the aspirin study the explanatory variable is oral medication and its two treatments are aspirin and placebo. Recognising the placebo as a treatment rather than as an absence is what keeps the comparison honest.
Matching
Example 1.19, part by part.
Match the pairs
Why: The population and the experimental units are easy to conflate, and the difference is scope: the population is the whole group the conclusion is meant to serve, while an experimental unit is one individual measured. The sample, not listed here, is the 400 who took part — sitting between the two.
Prediction
Commit before reasoning.
Predict first
In the aspirin study, what comparison do the results actually support?
Correct: Aspirin against a placebo taken under identical conditions.
Why: Both groups took a daily pill for three years and were watched equally closely, so everything about being in the study is common to them and the only difference is the pill's content. That is a narrower and much more defensible claim than aspirin against nothing, which would confound the drug with the routine and the attention. The last option describes an observational study, which is exactly the design this one was built to avoid.
Section
Section 2
Concept
Additional variables that can cloud a study are called lurking variables. To prove that the explanatory variable is causing a change in the response, the explanatory variable must be isolated: the experiment must be designed so there is only one difference between the groups being compared, the planned treatments. This is accomplished by random assignment of experimental units to treatment groups.
lurking variable — An additional variable, not among those being studied, that can cloud a study by offering an alternative explanation for a difference between groups. Random assignment spreads potential lurking variables equally among the groups, so none of them can account for a difference in the response.
\[ \text{groups alike} + \text{outcomes differ} \;\Longrightarrow\; \text{the treatment did it} \]
The book's example is exact. Recruit subjects, ask whether they regularly take vitamin E, and observe that the takers are healthier. This proves nothing, because people who take vitamin E regularly often take other steps to improve their health — exercise, diet, other supplements, not smoking — and any one of those could be producing the difference. The two groups differ in many ways at once, and the study cannot say which difference matters.
Figure (svg): A diagram showing vitamin E use and better health connected by several possible paths through exercise, diet and not smoking
OpenStax Introductory Statistics 2e, §1.4 Experimental Design and Ethics §1.4, p. 34 — lurking variables and the vitamin E study
Picture it
The claimed arrow is only one of several routes from left to right.
Figure (svg): A diagram showing vitamin E use and better health connected by several possible paths through exercise, diet and not smoking
Each lurking variable supplies a complete alternative explanation, and the data cannot distinguish between them because they all predict the same observation. Adding more subjects makes the association more precisely measured and leaves every alternative explanation exactly as available as before — the same asymmetry section 1.2 drew between sampling error and bias.
Worked example
The same lurking variables, and why they stop mattering.
\[ \text{Assign the 400 subjects to aspirin or placebo at random.} \]
Note that lurking variables still exist
Why: Some subjects smoke, some exercise, some eat well.
Ask where they end up
Why: Assignment ignores every one of those characteristics.
State the consequence for the two groups
Why: Both contain about the same share of smokers and exercisers.
Identify the only remaining difference
Why: The one the researcher imposed.
Figure (svg): The solution to Worked example what random assignment does to the picture shown as a ladder of expressions, one row per legal move
\[ \text{random assignment} \;\Longrightarrow\; \text{groups differ only by treatment} \]
Verify: confirm the argument does not claim the groups are identical
Why: Random assignment does not produce two identical groups, and cannot: one group will have slightly more smokers than the other by chance. What it produces is groups that differ only by chance rather than systematically, and the size of that chance difference is exactly what chapters 7 through 10 learn to measure. So the conclusion is not 'the groups were the same' but 'any difference between them beyond what chance explains must come from the treatment' — which is precisely the form a hypothesis test will take.
OpenStax Introductory Statistics 2e, §1.4 Experimental Design and Ethics §1.4, p. 34
Discrimination
Ask what the randomness decided: who is in the study, or what they received.
Sort into buckets
Sort each description.
Worked example
Example 1.21. Some explanatory variables cannot be assigned.
\[ \text{A researcher wants to study the effect of birth order on personality.} \]
Identify the explanatory variable
Why: Birth order is what the study is about.
Ask whether it can be assigned
Why: Nobody can be allocated to be a firstborn.
State what is therefore lost
Why: Random assignment is what eliminates lurking variables.
Name what remains possible
Why: Differences can still be observed and reported.
Figure (svg): The solution to Worked example a study that cannot be randomized shown as a ladder of expressions, one row per legal move
\[ \text{cannot assign} \;\Longrightarrow\; \text{observational study} \;\Longrightarrow\; \text{association only} \]
Verify: confirm what the lurking variables would be here
Why: Firstborn children differ from later-born children in more than birth order: family size, parental age, and the economic position of the household when they were raised all vary systematically with it. A personality difference between the groups could come from any of those. The book's own statement of the problem is the general one — when subjects cannot be assigned to treatment groups at random, there will be differences between the groups other than the explanatory variable — and it applies to sex, nationality, and long-term smoking just as much as to birth order.
OpenStax Introductory Statistics 2e, §1.4 Experimental Design and Ethics §1.4, p. 36
Trap
\[ \text{the 400 subjects were selected at random from the population} \]
Conclude the study can establish cause
Why: The word random appears, so the student credits it with everything randomness can do.
\[ \text{random selection} \;\Longrightarrow\; \text{cause} \quad \text{(wrong)} \]
If those randomly selected subjects then chose for themselves whether to take vitamin E, the two groups are self-selected and every lurking variable is back.
\[ \text{random SELECTION: the sample represents the population} \]
Keep the two randomizations separate: one for who is studied, one for who gets what
Why: They happen at different steps and buy different things.
Random selection is section 1.2's subject and it earns generalisation: what is true of the sample is probably true of the population. Random assignment is this section's subject and it earns causation: a difference between the groups must come from the treatment. A study can have either, both, or neither, and only a study with both can say that this treatment causes this effect in that population.
Two truths and a lie
All three concern lurking variables.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and it repeats section 1.2's lesson in a new setting. If people choose their own group, a bigger study simply measures the confounded association more precisely: ten million vitamin E takers still exercise more than non-takers. Only assignment breaks the link between the treatment and the lurking variable, and no sample size substitutes for it.
Socratic
A study reports that people who eat breakfast weigh less than people who skip it, and concludes that eating breakfast prevents weight gain.
Discussion prompt
Name two lurking variables that could produce this association without breakfast causing anything, and describe the experiment that would settle it.
Hint: Ask what else is true of people who eat breakfast regularly.
Answer:
People who eat breakfast regularly tend to keep regular schedules generally, and to exercise more; both are associated with lower weight. And people already trying to lose weight sometimes skip meals, so the causation may run backwards — weight concern causing breakfast-skipping rather than the reverse.
The experiment: recruit subjects, assign them at random to eat breakfast or to skip it for a fixed period, and measure weight change in both groups. Because assignment ignores schedule, exercise habit and prior weight concern, all three are spread evenly and none can explain a difference.
Trials of roughly this design have been run, and they find much weaker effects than the observational studies did — which is the usual outcome when a widely reported association is finally randomized, and a good reason to ask which kind of study a headline is based on.
Fill the middle
The condition an experiment must satisfy to establish cause.
Fill in the blanks
\textone ___ \text___
Why: The whole design exists to produce exactly one systematic difference. If the groups differ in two respects, an outcome difference cannot be attributed to either one, and no amount of statistical analysis afterwards can separate them — which is why confounding appears on section 1.2's critical-evaluation list as a problem of design rather than of arithmetic.
Section
Section 3
Concept
The power of suggestion can have an important influence on the outcome of an experiment; the expectation of a participant can be as important as the actual medication. To counter it, researchers set aside one treatment group as a control group, which is given a placebo — an active treatment that cannot directly influence the response variable.
control group and placebo — A control group is a treatment group set aside to receive a placebo: an active treatment that cannot directly influence the response variable. The control group balances the effects of being in an experiment against the effects of the active treatments, so the comparison isolates the treatment's own effect.
\[ \text{observed effect} = \text{treatment effect} + \text{effect of being in a study} \]
The book cites a study of performance-enhancing drugs in which believing one had taken the substance produced times almost as fast as actually consuming it, while taking the drug without knowledge yielded no significant improvement. The expectation was doing nearly all the work. A control group does not remove that effect; it reproduces it, so that subtracting one group's outcome from the other's leaves the treatment's own contribution.
Figure (svg): A flow diagram of a randomized experiment: subjects recruited, randomly assigned to a treatment group and a control group, then the response measured in each
OpenStax Introductory Statistics 2e, §1.4 Experimental Design and Ethics §1.4, pp. 34-35 — the power of suggestion, the control group and the placebo
Picture it
The control group does everything the treatment group does except receive the active ingredient.
Figure (svg): A flow diagram of a randomized experiment: subjects recruited, randomly assigned to a treatment group and a control group, then the response measured in each
Notice the phrase the book uses for a placebo: an ACTIVE treatment that cannot directly influence the response. It is active in the sense that the subject takes a pill daily, attends the same appointments and is watched the same way. Everything about the experience is reproduced except the one thing under test, which is what makes the subtraction legitimate.
Worked example
The difference between a placebo group and an untreated group is the whole design.
\[ \text{Compare: group B takes a placebo daily, against group C takes nothing.} \]
List what group A, on aspirin, experiences
Why: A daily pill, three years of appointments, being observed.
List what group B, on placebo, experiences
Why: A daily pill, the same appointments, the same observation.
List what group C, untreated, experiences
Why: None of it.
Ask which comparison isolates the drug
Why: A minus B leaves only the drug; A minus C leaves the drug plus everything else.
Figure (svg): The solution to Worked example why the control group is not 'no treatment' shown as a ladder of expressions, one row per legal move
\[ A - B = \text{drug effect} \qquad A - C = \text{drug} + \text{study experience} \]
Verify: confirm with the book's own performance study
Why: In the drug study the book cites, believing one had taken the substance produced almost the full performance gain. If that trial had used an untreated comparison group, the entire suggestion effect would have been added to the drug's own effect and reported as the drug's. Comparing against a placebo group reproduces the belief in both arms so it cancels, and the residual difference is what the chemistry did. The comparison chosen determines what quantity the study estimates, which is why the control group's design is a scientific decision rather than an administrative one.
OpenStax Introductory Statistics 2e, §1.4 Experimental Design and Ethics §1.4, p. 35
Elimination
A trial tests whether a new painkiller reduces headache severity.
Eliminate the wrong options
Eliminate the three unsuitable controls and keep the one that works.
Survives elimination: B
Why: The survivor reproduces every part of the experience except the active ingredient, and the assignment is random so the groups are comparable. Choice D is worth noting because it is not a mistake, merely a different design: dose-comparison trials are common and they answer whether more is better, not whether the drug does anything at all.
Worked example
The placebo must reproduce the experience without the active ingredient, whatever the treatment is.
\[ \text{Testing whether a new exercise programme reduces back pain} \]
Identify what the treatment group experiences
Why: Attention from an instructor, time spent, an expectation of improvement.
Reject the obvious control
Why: A group told to carry on as usual gets none of that.
Design a sham treatment
Why: The same contact time doing gentle stretching with no therapeutic claim.
Check what the comparison now isolates
Why: Both groups get attention and time; only one gets the programme.
Figure (svg): The solution to Worked example a placebo that is not a pill shown as a ladder of expressions, one row per legal move
\[ \text{placebo} = \text{everything except the active ingredient} \]
Verify: confirm why this is genuinely hard outside drug trials
Why: A pill placebo is easy because an inert tablet is indistinguishable from an active one. An exercise placebo is not: a participant may suspect that gentle stretching is the sham, and the instructor certainly knows. This is why surgical and behavioural trials are harder to run than drug trials and why their evidence is often weaker — the design constraint is real rather than a matter of effort, and the honest response is to report which blinding was achievable rather than to claim more.
OpenStax Introductory Statistics 2e, §1.4 Experimental Design and Ethics §1.4, p. 35
Trap
\[ \text{placebo group: } 30\% \text{ improved} \quad \text{treatment group: } 35\% \text{ improved} \]
Report that the treatment produced a 35 percent improvement rate
Why: The treatment group's number is quoted on its own.
\[ \text{effect} = 35\% \quad \text{(wrong)} \]
Thirty percent improved on an inert pill, so most of that 35 is not the treatment doing anything.
\[ \text{effect} = 35\% - 30\% = 5 \text{ percentage points} \]
Report the DIFFERENCE between the groups, never one group's level
Why: The control group's level is the baseline the treatment must beat.
The control group exists precisely so that this subtraction can be made, and quoting the treatment arm alone throws away the information it was collected to provide. Whether five percentage points is a real effect or chance variation is the question chapter 10 answers, but the quantity to be tested is always the difference — which is why that chapter is about comparing two groups rather than describing one.
Prediction
Commit before reasoning.
Predict first
A placebo group shows 30 percent improvement. What has been measured?
Correct: Everything apart from the active ingredient.
Why: The 30 percent bundles the power of suggestion, the effect of being observed and cared for, and the natural course of the condition — headaches resolve on their own. All of those would also be present in the treatment group, which is exactly why the baseline is needed. The last option is the misconception the control group exists to defeat: the inert arm is the most informative group in the trial, because it says how much improvement happens for reasons the treatment can claim no credit for.
Fill the middle
The book's careful phrasing.
Fill in the blanks
\textresponse ___ \text___
Why: Calling it active is deliberate: the subject genuinely takes something and genuinely participates. What it cannot do is act directly on the outcome being measured. That combination — a full experience with no active ingredient — is what makes it the right comparison.
Estimation
A trial reports 62 percent improvement on the drug and 48 percent on placebo.
Predict first
What is the drug's estimated effect?
Correct: 14 percentage points.
Why: The effect is the difference between the arms: 62 minus 48. The first option quotes the treatment arm alone and credits the drug with the placebo response as well. The last option divides the two figures, which produces a relative risk of a sort but is not what 'the effect' usually means and is easy to report misleadingly — a jump from 1 percent to 2 percent is also a doubling. Whether 14 points is more than chance explains is chapter 10's question.
Section
Section 4
Concept
If you are participating in a study and you know you are receiving a pill containing no medication, the power of suggestion is no longer a factor. Blinding, or masking, preserves it: a blinded person does not know who is receiving the active treatment and who is receiving the placebo. A double-blind experiment is one in which both the subjects and the researchers involved with the subjects are blinded.
blinding and double-blinding — Blinding means a person involved in the study does not know who receives the active treatment and who receives the placebo. In a single-blind design the subjects are blinded; in a double-blind design the researchers working with the subjects are blinded as well.
\[ \text{single-blind: subjects} \qquad \text{double-blind: subjects and researchers} \]
The logic runs backwards from the placebo. A placebo works because the subject believes it might be the real treatment, so telling them it is inert destroys the very effect the control group was constructed to reproduce. Blinding the researchers as well guards against a different leak: an investigator who knows which subject is on the drug may probe harder for improvement, record borderline outcomes more generously, or convey optimism without intending to.
Figure (svg): Two columns describing a single-blind experiment where only subjects are blinded and a double-blind experiment where researchers are blinded too
OpenStax Introductory Statistics 2e, §1.4 Experimental Design and Ethics §1.4, p. 35 — blinding, masking and double-blind designs
Picture it
The same trial, with one more person kept in the dark.
Figure (svg): Two columns describing a single-blind experiment where only subjects are blinded and a double-blind experiment where researchers are blinded too
Double-blinding is the standard for clinical trials because the two leaks are independent: blinding subjects stops their expectations affecting their outcomes, and blinding researchers stops the researchers' expectations affecting what gets recorded. Someone must still hold the key, of course — the allocation is kept by a third party and opened after the data are collected.
Worked example
Example 1.20 again. Ask separately for each person involved.
\[ \text{Subjects run mazes wearing floral-scented or unscented masks.} \]
Ask whether subjects can be blinded
Why: A subject will plainly know whether they can smell flowers.
Ask whether researchers can be blinded
Why: The person timing the maze need not know which mask is worn.
Classify the achievable design
Why: Only one of the two parties is blinded, and it is not the subjects.
Say what remains uncontrolled
Why: Subject expectation about scent still affects performance.
Figure (svg): The solution to Worked example can this study be blinded shown as a ladder of expressions, one row per legal move
\[ \text{subjects: no} \qquad \text{researchers: yes} \]
Verify: confirm the study is still worth running
Why: The limitation is real and it does not make the design useless. Blinding the timers removes one whole channel of bias, and randomizing which trials carry the scent removes the practice effect. What remains is that a subject who expects flowers to help may try harder, so the study cannot separate the scent's direct effect from the subject's belief about it. Stating that limitation is the honest report, and it is more useful than either abandoning the study or pretending the problem is absent.
OpenStax Introductory Statistics 2e, §1.4 Experimental Design and Ethics §1.4, p. 36
Sorting
Ask separately whether the subject and the observer can be kept unaware.
Sort into buckets
Sort each study.
Every study in the right-hand bucket can still blind its assessors, and doing so is worth real effort: it removes an entire channel of bias at almost no cost. A trial that could have blinded its assessors and did not has given something away for nothing.
Worked example
A blind can fail during a trial without anyone breaking a rule.
\[ \text{A drug causes a distinctive dry mouth; the placebo does not.} \]
Note what subjects can infer
Why: Those with a dry mouth conclude they are on the drug.
Note what follows for the placebo group
Why: Those without it conclude they are on the placebo.
Identify the consequence
Why: The suggestion effect is no longer equal in both arms.
Name the standard remedy
Why: Use a placebo producing a similar harmless sensation.
Figure (svg): The solution to Worked example what breaks a blind shown as a ladder of expressions, one row per legal move
\[ \text{distinctive side effect} \;\Longrightarrow\; \text{subjects unblind themselves} \]
Verify: confirm this is a design problem rather than a subject failure
Why: Nobody has cheated: the subjects have simply noticed something true about their own bodies. That is why the fix is in the design rather than in the instructions, and why trials sometimes ask subjects at the end which arm they think they were in — if their guesses are much better than chance, the blind did not hold, and the results must be read with that in mind. An honest trial reports the answer to that question.
OpenStax Introductory Statistics 2e, §1.4 Experimental Design and Ethics §1.4, p. 35
Trap
\[ \text{double-blind: neither subjects nor researchers know the assignment} \]
Conclude that the assignment is unknown to everyone until the end
Why: Nobody in the room knows, so the student assumes nobody anywhere does.
\[ \text{nobody knows who got what} \quad \text{(then how is it analysed?)} \]
If the allocation were genuinely lost, the data could never be split into two groups and the trial would produce nothing.
\[ \text{a third party holds the key, sealed until the data are collected} \]
Blind the people who could influence or record the outcome, not the record itself
Why: The book's phrase is researchers INVOLVED WITH THE SUBJECTS.
The allocation is recorded and held by someone with no contact with the subjects — a pharmacist or an independent data centre — and opened once the outcomes are locked. There is also a safety reason for the key to exist: if a subject is harmed, the treating doctor must be able to find out immediately what they were given.
Two truths and a lie
All three concern blinding.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The assignment is always recorded and held by a party with no contact with the subjects, then revealed after outcomes are locked. It has to be, both so the data can be analysed and so that a subject who is harmed can be treated appropriately.
Matching
Each blind, and the specific leak it closes.
Match the pairs
Why: Four devices, four distinct threats, and none of them substitutes for another. This is why a well-designed trial is described with a string of adjectives — randomized, double-blind, placebo-controlled — each word naming a specific problem that has been dealt with rather than a general claim of quality.
Explain it
A classmate says giving sick people a fake pill is pointless and unkind.
Discussion prompt
In four sentences or fewer, explain what the placebo group is measuring and why the trial would be uninterpretable without it.
Hint: Ask what fraction of patients improve on the fake pill.
Answer:
Tell them that in many trials 30 or 40 percent of the placebo group improves — from expectation, from the attention of being in a study, and because many conditions get better on their own. Without a placebo arm, all of that improvement gets credited to the drug.
So the placebo group is not a group that receives nothing; it is the measurement of how much improvement happens for reasons the drug can claim no credit for, and the drug's actual effect is the difference between the two arms.
On the kindness question: participants are told in advance that they may receive a placebo and consent to that, nobody is denied an existing effective treatment to be in a trial, and trials are stopped early when one arm is clearly better.
Section
Section 5
Concept
When subjects cannot be assigned to treatment groups at random, there will be differences between the groups other than the explanatory variable. Such a study can establish association but not cause. Separately, the researcher carries obligations to the people studied: informed consent, honest reporting, and the accurate presentation of data.
observational study — A study in which the researcher records what happens without assigning treatments, because the explanatory variable cannot be manipulated. It can establish that two variables are associated, but not that one causes the other, since the groups differ in ways beyond the explanatory variable.
\[ \text{association} \;\ne\; \text{causation, without random assignment} \]
Birth order is the book's example and the point generalises: sex, nationality, having smoked for twenty years, and having survived a disease are all unassignable. For these questions an observational study is the only available design, and its results are worth having — the link between smoking and lung cancer was established this way, through many studies with converging evidence rather than one trial. What is never available is the single-study causal claim a randomized experiment can support.
Figure (svg): Two columns contrasting an observational study, where the researcher only records what happens, with an experiment, where the researcher assigns treatments
OpenStax Introductory Statistics 2e, §1.4 Experimental Design and Ethics §1.4, pp. 36-37 — studies that cannot be randomized, and research ethics
Picture it
The two designs, and what each is entitled to claim.
Figure (svg): Two columns contrasting an observational study, where the researcher only records what happens, with an experiment, where the researcher assigns treatments
The right column is stronger and the left is often the only one available, so the useful skill is not preferring experiments but reading a reported study to see which it was. A newspaper reporting that a habit 'reduces risk' has almost always described an observational study, and the appropriate reading is that the habit is associated with lower risk among people who already have it.
Worked example
The smoking and lung cancer case, which was never randomized and is nonetheless settled.
\[ \text{Smokers develop lung cancer far more often than non-smokers.} \]
Note why no experiment was possible
Why: Assigning people to smoke for twenty years is not permissible.
Note what a single such study establishes
Why: An association, with lurking variables uncontrolled.
List what strengthened the case
Why: A large effect, a dose-response relation, consistency across many populations, and a biological mechanism.
State the standard reached
Why: Not one study's proof, but a body of evidence no alternative explains.
Figure (svg): The solution to Worked example what an observational study can still support shown as a ladder of expressions, one row per legal move
\[ \text{size} + \text{dose-response} + \text{consistency} + \text{mechanism} \]
Verify: confirm why the dose-response relation carries so much weight
Why: Heavier smokers develop lung cancer more often than lighter smokers, and those who quit reduce their risk over time. A lurking variable would have to be present in proportion to the amount smoked and to decline after quitting, which is a far more demanding thing to postulate than a variable that merely accompanies smoking. Each additional pattern of this kind narrows the space of alternative explanations, which is how observational evidence accumulates into a causal conclusion without any single study being able to deliver one.
OpenStax Introductory Statistics 2e, §1.4 Experimental Design and Ethics §1.4, p. 36
Discrimination
Ask whether the researcher assigned the explanatory variable.
Sort into buckets
Sort each study.
Worked example
Design is not only a technical matter.
\[ \text{A researcher plans a trial of a new treatment on human subjects.} \]
Before recruiting
Why: Subjects must understand what participation involves, including that they may receive a placebo, and agree freely.
During the study
Why: Data must be recorded as collected, without discarding inconvenient observations.
When reporting
Why: The design, the sample size and the failures must be reported alongside the result.
Throughout
Why: Someone independent reviews the protocol and can stop the study.
Figure (svg): The solution to Worked example the obligations a study carries shown as a ladder of expressions, one row per legal move
\[ \text{consent} \to \text{honest collection} \to \text{complete reporting} \]
Verify: confirm why selective reporting is a statistical problem and not only a moral one
Why: If a researcher runs twenty analyses and publishes the one that came out favourable, the reported result carries none of the reliability its p-value claims — chapter 9 shows that testing enough hypotheses guarantees some will look significant by chance alone. So suppressing the other nineteen does not merely mislead the reader about what was done; it invalidates the statistic that was reported. This is why the obligation to report completely sits in a statistics textbook rather than only in an ethics one.
OpenStax Introductory Statistics 2e, §1.4 Experimental Design and Ethics §1.4, pp. 36-37
Trap
\[ \text{people who drink coffee live longer} \]
Conclude that drinking more coffee will extend your life
Why: The headline is phrased as advice, so the student reads a causal claim.
\[ \text{coffee} \;\longrightarrow\; \text{longer life} \quad \text{(not established)} \]
Nobody was assigned to drink coffee. Coffee drinkers differ from non-drinkers in income, occupation and health at the outset, and one common reason to have given up coffee is already being ill.
\[ \text{coffee drinking is ASSOCIATED with longer life in these data} \]
Ask first whether the explanatory variable was assigned
Why: If subjects chose their own group, the study reports association.
The reverse-causation possibility in this example is worth naming because it is so easy to miss: people who become ill often stop drinking coffee, which would produce exactly this association with the causal arrow pointing backwards. Good observational research works hard to exclude such explanations, and the strongest of them can support a causal conclusion, but a single study and a headline never do.
Prediction
Commit before reasoning.
Predict first
A large, careful observational study finds a strong association between a habit and an outcome. What has it established?
Correct: That the two are associated, with the cause still open.
Why: Association is a real finding and often an important one; what remains open is whether the habit causes the outcome, the outcome causes the habit, or a third variable causes both. The third option is the overcorrection worth avoiding — observational studies established the link between smoking and lung cancer, and dismissing them wholesale would discard most of what is known in epidemiology, where randomization is frequently impossible.
Constraint
You want to know whether working night shifts damages long-term health.
Discussion prompt
Explain why this cannot be a randomized experiment, then describe the best observational design you can, saying which lurking variables you would try to account for.
Hint: Ask what you would have to assign, and for how long.
Answer:
Randomizing would mean assigning people to work nights for a decade, which no researcher can impose and no ethics committee would approve. So the explanatory variable cannot be manipulated and the design must be observational.
The best available design follows a large group forward in time, recording shift patterns as they occur and health outcomes as they arise, rather than asking people to recall past work. Following forward avoids the recall problem and establishes that the exposure came before the outcome.
The lurking variables to account for are the ones that travel with night work: income and occupation, smoking, sleep duration, and the fact that people in poorer health may move off nights — which would make night workers look healthier for a reason that has nothing to do with the shifts.
Even done well this supports association and a carefully argued causal case, never the single-study proof a randomized trial would give. Saying so plainly is part of reporting it honestly.
Two truths and a lie
All three concern the limits of study designs.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and it is the section's central point restated. Size does not create comparability between groups that formed themselves. What can build a causal case from observational data is converging evidence of several kinds — a large effect, a dose-response relationship, consistency across populations, and a plausible mechanism — and that is a different thing from one large study.
Comparison
Fill the blanks. Each device answers a different threat, and none substitutes for another.
Comparison matrix
| Device | The threat it answers | What it does NOT do |
|---|---|---|
| Random selection | the sample not representing the population | it does not make the groups comparable |
| Random assignment | lurking variables piling into one group | it does not make the sample represent the population |
| Control group with placebo | the power of suggestion and natural recovery | it does not stop researchers scoring outcomes generously |
| Double-blinding | expectations of subjects AND researchers shaping outcomes | it does not remove lurking variables |
This is why a well-designed trial is described with a string of adjectives — randomized, double-blind, placebo-controlled — rather than one word. Each adjective names a specific threat that has been dealt with, and a missing adjective names one that has not.
Pattern
Six questions, and the first two settle what kind of claim the study can possibly support.
If the answer to the first question is that subjects chose their own groups, no later answer can rescue a causal claim. Read the rest for what the association is worth, and read the word 'causes' in the headline as an error.
OpenStax Introductory Business Statistics 2e, §1.4 Experimental Design and Ethics §1.4 Experimental Design and Ethics
Check
Name the manipulated variable.
Check your understanding
Ninety adults are divided at random into three groups, each given a different medicine for six months, and the average change in height is measured. What is the explanatory variable?
Answer: A
Why: The researcher decides which medicine each subject receives, so that is the manipulated variable, and its three values A, B and C are the treatments.
Check
Which randomization does the work here?
Check your understanding
A study finds that people who take a supplement are healthier than those who do not. The 5,000 participants were selected at random from the population. Can the study establish that the supplement improves health?
Answer: A
Why: Random selection makes the sample represent the population; it does nothing to make the two supplement groups comparable, because subjects sorted themselves into those.
Check
Report the effect.
Check your understanding
In a trial, 55 percent of the treatment group improved and 42 percent of the placebo group improved. What is the estimated treatment effect?
Answer: A
Why: The effect is the difference between the arms, 55 minus 42, because the placebo arm measures all the improvement that happens for reasons other than the active treatment.
Real world
Two headlines appear the same week. The first: 'People who take daily walks have 30 percent lower rates of heart disease' — from a study following 40,000 adults for ten years, recording their own exercise habits. The second: 'New drug cuts heart attacks by 4 percent' — from a trial of 20,000 people randomly assigned the drug or a placebo, double-blind.
Discussion prompt
Which finding is larger, which is better evidence of cause, and what should a reader do with each?
Hint: Ask for each study whether the researcher assigned the explanatory variable.
Answer:
The walking finding is much larger and much weaker. Nobody was assigned to walk. People who walk daily differ from those who do not in income, existing health, occupation and much else, and people who are already unwell walk less — so the causal arrow may point backwards. A 30 percent difference is real as an association and its cause is entirely open.
The drug finding is small and much stronger. Assignment was random, so the two groups were comparable at the outset; the placebo control means the 4 percent is a difference between arms rather than a raw rate; and double-blinding stops both subjects' and researchers' expectations shaping the outcome. Four percentage points attributable to the drug is a genuine causal claim.
What to do with each. The drug result supports a decision about taking the drug. The walking result supports the weaker statement that walking is associated with lower heart disease — which is worth knowing, is consistent with a real benefit, and is not evidence that walking more would lower a given person's risk by 30 percent.
\[ \text{effect size} \;\ne\; \text{strength of evidence} \]
The general lesson is that the size of a number and the quality of the design behind it are independent, and headlines report the first while burying the second. The reliable move is to find, in any report, the sentence saying whether people were assigned to their groups — and if that sentence is absent, the study was almost certainly observational.
Commit first
Answer, then rate your confidence honestly.
Predict first
Which single feature of a study is what licenses a claim that the treatment CAUSED the difference?
Correct: Random assignment of units to treatment groups.
\[ \text{random assignment} \;\Longrightarrow\; \text{groups differ only by treatment} \;\Longrightarrow\; \text{cause} \]
Why: Random assignment is what makes the groups comparable, so that any systematic difference in outcome has only one available explanation. Random selection earns generalisation to the population instead; a double-blind design protects the measurement from expectation but cannot make self-selected groups comparable; and size improves precision without touching comparability. All four are worth having and only one of them buys cause.
Explain it
They cannot see why a study needs to give half the subjects a fake pill.
Discussion prompt
In four sentences or fewer, explain what would go wrong without the placebo group, using a number.
Hint: Tell them how many people improve on an inert pill.
Answer:
Say that in a typical trial around a third of the people on the fake pill get better — from expecting to, from the attention of being in a study, and because many conditions improve on their own.
Without a placebo group you would see 40 percent improving on the drug and have no way of knowing whether that was the drug or the 33 percent who would have improved regardless.
With the placebo group you can subtract: 40 minus 33 leaves 7 percentage points that the drug can claim, and that subtraction is the only thing that turns a rate into an effect.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the variables, ask which one the researcher could set for each subject before the study began. For the two randomizations, ask whether the randomness decided who is studied or what they received. For the control group, remember it measures everything that improves outcomes apart from the active ingredient. For causation, ask the single question of whether the explanatory variable was assigned. Do five described studies of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top, draw the anatomy of a randomized experiment left to right: subjects, a box labelled random assignment, a treatment group and a control group, and the response measured in each. Label all six parts using the aspirin study, and mark which box is the one that buys a causal claim. Below, draw the vitamin E problem: two boxes for taking vitamin E and better health with a direct arrow between them, and three lurking variables underneath with arrows into both, then write one sentence saying why the direct arrow cannot be isolated. To the right, make a two-column table headed random selection and random assignment with three rows: what the randomness decides, what it buys, and what it does not buy. In the middle, write the four design devices — randomization, control group, placebo, blinding — and beside each the single threat it answers. At the bottom, take this study and decide what it can claim: 40,000 adults are followed for ten years, their exercise habits recorded, and daily walkers are found to have 30 percent less heart disease. Write which design it is, one lurking variable, and the sentence you would use to report it honestly.
Check the table in the middle of your page: if the row for random selection and the row for random assignment say the same thing in the 'what it buys' column, look again. One makes a sample stand for a population; the other makes two groups comparable with each other, and no study gets both from one randomization.
Recap
Five things, and the second one is the whole reason experiments exist.
| If you see | Then |
|---|---|
| Subjects chose their own group | Observational: association only, whatever the sample size |
| The researcher assigned treatments at random | Experiment: a causal claim is available |
| Random selection from a population | The sample generalises; the groups may still be incomparable |
| No control group | The treatment effect is confounded with being in a study |
| A control group taking nothing | Confounded with the pill, the routine and the attention |
| Only the treatment arm's rate reported | Ask for the placebo arm and subtract |
| An unassignable explanatory variable | No experiment is possible; look for converging evidence |
Chapter 1 is finished: you can name what a study is about, how its data were gathered, how to organise them, and whether the design supports the claim being made. Chapter 2 turns to the data themselves, and asks how a set of numbers is described — by a picture, by a centre, and by a spread.
OpenStax Introductory Statistics 2e, §1.4 Experimental Design and Ethics §1.4, pp. 34-37 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.