The section that supplies the logic behind a hypothesis test. The reasoning is the rare-event argument: make an assumption about the population, gather data, and if the sample has properties that would be very unlikely to occur when the assumption is true, conclude that the assumption is probably incorrect. The p-value measures that unlikeliness — it is the probability that, if the null hypothesis is true, another randomly selected sample would give results as extreme or more extreme than the ones obtained. A large p-value indicates the null should not be rejected, and the smaller the p-value the stronger the evidence against it. Comparing the p-value against a significance level chosen before the data were collected turns the measurement into a decision, and the conclusion must then be written in the words of the problem. The p-value is conditional on the null being true, which is what makes almost every common misreading of it wrong.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 9 — Hypothesis Testing with One Sample
Rare Events, the Sample, Decision and Conclusion
Objectives
Five outcomes, and two of them are about what the p-value does not say.
OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, pp. 467-469 — the section these objectives are drawn from
Warm-up
Section 9.3 chose a distribution; sections 9.1 and 9.2 set up the hypotheses and named the errors.
Discussion prompt
Two friends are told a basket of 200 bubbles contains exactly one hundred-dollar bill. The first to reach in draws it. What should the second friend conclude?
Hint: How likely was that, if what they were told is true?
Answer:
If the claim is true, the chance of drawing the one bill first is one in two hundred — 0.005. That is a rare event, and it just happened on the first attempt.
So the second friend has grounds to doubt the claim. Either something very unlikely occurred, or the assumption that there is only one bill is wrong — and the second explanation is the more comfortable one.
That is the whole logic of a hypothesis test, and the book uses exactly this example to introduce it. What remains is to make unlikely precise, which is what a p-value does, and to fix how unlikely is unlikely enough, which is what alpha does.
Concept
Suppose you make an assumption about a property of the population — this assumption is the null hypothesis — and then gather sample data randomly. If the sample has properties that would be very unlikely to occur if the assumption is true, then you would conclude that your assumption about the population is probably incorrect.
the rare-event argument — The reasoning behind every hypothesis test. The book's parenthesis states the asymmetry: the assumption is just an assumption and may or may not be true, but the sample data are real and show a fact that seems to contradict it.
\[ P(\text{data this extreme} \mid H_0) \text{ small} \;\Longrightarrow\; \text{doubt } H_0 \]
The argument's shape is worth noticing because it explains the asymmetry section 9.1 established. It can only run one way: unlikely data cast doubt on the assumption, but ordinary data are consistent with the assumption and with a great many alternatives, so they cast doubt on nothing. That is precisely why a test can reject a null and never confirm one.
Figure (svg): Three linked boxes showing the rare-event argument: assume the null, ask how likely the data would be, and doubt the assumption if the answer is small
OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, p. 467
Section
Section 1
Concept
The book's parenthesis carries the argument: remember that your assumption is just an assumption — it is not a fact and it may or may not be true. But your sample data are real, and the data are showing you a fact that seems to contradict your assumption.
a rare event — An outcome whose probability would be small if the assumption held. Its occurrence is grounds for doubting the assumption rather than for accepting that something unlikely happened.
\[ \text{unlikely data} \;+\; \text{data are real} \;\Longrightarrow\; \text{the assumption gives way} \]
Nothing here is unique to statistics. It is the ordinary reasoning of anyone who checks a claim against evidence: told that a coin is fair and seeing twenty heads in a row, most people conclude the coin is not fair rather than that they witnessed something with probability one in a million. The test simply makes the threshold explicit and the probability computable.
Figure (svg): Three linked boxes showing the rare-event argument: assume the null, ask how likely the data would be, and doubt the assumption if the answer is small
OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, p. 467 — the rare-events passage and Didi and Ali
Picture it
Assume, compute, and judge.
Figure (svg): Three linked boxes showing the rare-event argument: assume the null, ask how likely the data would be, and doubt the assumption if the answer is small
The third box is where the asymmetry lives. A small probability licenses doubting the assumption; a large one licenses nothing at all, because ordinary data are exactly what many different assumptions would predict. Section 9.1's rule against accepting the null is this observation restated.
Worked example
The book's introduction to the argument.
\[ 200 \text{ bubbles, one holding a } \$100 \text{ bill; Didi draws it first} \]
State the assumption
Why: What they were told.
Compute the probability
Why: One in two hundred.
\[ 0.005 \]
Judge it
Why: Very unlikely.
Draw the conclusion
Why: Doubt the assumption.
Figure (svg): The solution to Worked example the hundred-dollar bill shown as a ladder of expressions, one row per legal move
\[ P = \frac{1}{200} = 0.005 \]
Verify: confirm what the argument does NOT establish
Why: It does not prove there is more than one bill — Didi may simply have been lucky, and one time in two hundred she would be. What the reasoning gives is grounds for doubt proportional to how unlikely the event was, which is exactly what a p-value quantifies. Treating it as proof would be the same error as accepting a null on a large p-value, in the opposite direction.
OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, p. 467
Prediction
Commit before reasoning.
Predict first
In the rare-event argument, why does the assumption give way rather than the data?
Correct: Because the data are real and the assumption is only an assumption.
Why: The book states it in exactly those terms. One of the two must give when they conflict, and the observed data are a fact while the null is a supposition entertained for the sake of argument. This does not mean data are never wrong — a biased sample is bad data — but within the test's own frame the assumption is the movable part.
Worked example
Example 9.9, the same argument with a computed probability.
\[ H_0: \mu \le 15; \; n = 10, \; \bar{x} = 17, \; \sigma = 0.5 \]
Assume the null
Why: The mean is 15.
Find the sampling distribution
Why: By the central limit theorem.
\[ N(15, 0.158) \]
Locate the sample mean
Why: 17 cm.
\[ \text{about } 12.6\text{ SEs out} \]
Judge the probability
Why: The area beyond.
Figure (svg): The solution to Worked example the baker's loaves shown as a ladder of expressions, one row per legal move
\[ p\text{-value} = P(\bar{x} > 17) \approx 0 \]
Verify: confirm the standard error is what makes 17 so extreme
Why: Two centimetres above 15 sounds modest, and for a single loaf it would be — four standard deviations, which is unusual but not extraordinary. For a MEAN of ten loaves the standard error is only 0.158 cm, so the same two centimetres is 12.6 standard errors. Chapter 7's concentration of sample means is what turns a moderate-looking gap into an impossible one.
OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, pp. 467-468
Trap
\[ P = 0.005 \;\Rightarrow\; \text{it cannot have happened by chance} \]
Read a small probability as ruling chance out
Why: One in two hundred is very unlikely.
\[ \text{but one time in two hundred, it does happen} \]
The argument gives grounds for doubt, not a proof, and the strength of those grounds is what the probability measures.
\[ P = 0.005 \;\Rightarrow\; \text{strong grounds to doubt the assumption} \]
Read the probability as a measure of how strong the doubt is
Why: Smaller means stronger, never certain.
This is section 9.2's Type I error seen from the other side: a test that rejects at a 5 percent level will reject a true null one time in twenty, and that is not a malfunction but the stated error rate. Any language of proof in reporting a test is overstating what the argument can deliver.
Two truths and a lie
All three concern the argument.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. A rare event is grounds for doubt, and rare events do occur — one draw in two hundred finds the bill even when the claim is true. Reading the argument as proof is what produces overconfident claims from single studies.
Faded example
One bill among 200 bubbles, drawn first.
Fill in the blanks
P = \frac2000.005} = ___
Why: One in two hundred is 0.005 — small enough that Ali reasonably doubts the claim, though not small enough to rule out simple luck.
Explain it
A classmate asks why unlikely data can undermine a claim but ordinary data cannot support it.
Discussion prompt
In two sentences or fewer, explain.
Hint: Ask what other claims would predict the same ordinary data.
Answer:
Unlikely data are hard to explain under the null, so the null becomes hard to hold — that is a real constraint.
Ordinary data are exactly what the null predicts, but also what a dozen nearby alternatives predict, so they do not pick the null out from among them.
Section
Section 2
Concept
The p-value is the probability that, if the null hypothesis is true, the results from another randomly selected sample will be as extreme or more extreme than the results obtained from the given sample.
p-value — A conditional probability: it assumes the null is true and asks how often data at least this extreme would arise. Smaller means the observed data are harder to explain under the null.
\[ p\text{-value} = P(\text{a result at least this extreme} \mid H_0 \text{ true}) \]
Two parts of the definition do all the work and are both routinely dropped. If the null hypothesis is true is the condition that makes every common misreading wrong. As extreme or more extreme is why the p-value is a tail area rather than the probability of the exact result observed — which for a continuous variable would be zero, as section 5.1 established.
Figure (svg): A narrow normal curve centred at fifteen with the sample mean of seventeen far out in the right tail
OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, pp. 467-468 — the definition and Example 9.9
Picture it
Example 9.9's sampling distribution if the baker's null held.
Figure (svg): A narrow normal curve centred at fifteen with the sample mean of seventeen far out in the right tail
The curve is centred at 15 because the null is being assumed, and it is narrow because it describes means of ten loaves rather than single loaves. The observed 17 lies so far right that it is off the picture entirely, which is what a p-value of essentially zero looks like.
Worked example
Example 9.9, computed.
\[ \bar{X} \sim N(15, 0.158); \text{ find } P(\bar{x} > 17) \]
Find the standard error
Why: 0.5 over root 10.
\[ 0.158 \]
Standardise the observation
Why: 17 minus 15, over 0.158.
\[ z = 12.65 \]
Take the right tail
Why: Beyond that z.
Interpret
Why: Almost no sample would.
Figure (svg): The solution to Worked example the baker's p-value shown as a ladder of expressions, one row per legal move
\[ p\text{-value} = P(\bar{x} > 17) \approx 0 \]
Verify: confirm the interpretation names the condition
Why: The book's own reading is that almost zero percent of all loaves of bread would be at least as high as 17 cm purely by chance HAD the population mean height really been 15. The conditional clause is not decoration — without it the sentence would be a claim about how often the mean is 15, which is a different and unavailable quantity.
OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, p. 468
Fill the middle
The clause that makes the definition precise.
Fill in the blanks
\texttrue ___
Why: True. Every p-value is computed on the null's own distribution, which is why it cannot be a probability about whether the null holds.
Worked example
Try It 9.9, where the evidence is strong but finite.
\[ \sigma = 1, \; n = 36, \; \bar{x} = 12.5; \; H_0: \mu \le 12 \]
Find the standard error
Why: One over six.
\[ 0.1667 \]
Standardise
Why: 0.5 over 0.1667.
\[ z = 3 \]
Take the right tail
Why: Beyond z of 3.
\[ 0.0013 \]
Interpret
Why: About one in 750.
Figure (svg): The solution to Worked example a p-value that is small but not tiny shown as a ladder of expressions, one row per legal move
\[ p\text{-value} = P(\bar{x} > 12.5) \approx 0.0013 \]
Verify: confirm the value against the Empirical Rule
Why: A z of exactly 3 puts the observation three standard errors above the mean, and chapter 6's rule says about 0.3 percent of a normal distribution lies more than three standard deviations from the centre — half of that, about 0.15 percent, in one tail. That matches 0.0013 closely, and it is a way of sanity-checking any p-value whose z-score is a whole number.
OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, p. 468
Trap
\[ p\text{-value} = P(\bar{x} = 17) \]
Ask how likely the observed result was
Why: The p-value measures how surprising the data are.
\[ \text{but for a continuous variable that is exactly zero} \]
Section 5.1 established that a single value has no width and therefore no area, whatever the data turn out to be.
\[ p\text{-value} = P(\bar{x} > 17), \text{ the tail} \]
Include everything at least as extreme, which makes it an area
Why: The definition says as extreme or more extreme.
The as-or-more-extreme clause is what makes the quantity meaningful and computable. It also explains why the direction of the alternative matters: it decides which tail counts as extreme, and section 9.5 will make that the difference between a left, right and two-tailed test.
Faded example
A right-tailed test with a standard error of 0.1667 and a sample mean 0.5 above the null.
Fill in the blanks
z = \frac30.0013 = ___, \quad p\text___ \approx ___
Why: A z of 3 leaves about 0.0013 in the right tail, matching the Empirical Rule's estimate of roughly 0.15 percent beyond three standard deviations on one side.
Two truths and a lie
All three concern the definition.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and for a continuous variable it would always be zero. The p-value is a tail area covering the observed result and everything further from the null, which is what makes it a measure of how extreme the data are rather than of how likely one particular value was.
Estimation
Two tests report p-values of 0.001 and 0.20.
Predict first
What do these say about the evidence against the null?
Correct: The first gives much stronger evidence.
Why: The book's rule is that the smaller the p-value, the more unlikely the outcome and the stronger the evidence against the null. Data with a p-value of 0.001 are very hard to explain under the null; data with 0.20 are entirely ordinary under it.
Section
Section 3
Concept
A systematic way to decide is to compare the p-value with a preset significance level alpha. If alpha exceeds the p-value, reject the null and call the results significant. If alpha is less than or equal to the p-value, do not reject, and the results are not significant.
significance level — The preset alpha, which is the probability of a Type I error. It may or may not be given at the start of a problem, and if none is given a common standard is 0.05.
\[ \alpha > p \;\Rightarrow\; \text{reject}; \qquad \alpha \le p \;\Rightarrow\; \text{do not reject} \]
That alpha is preset is the part that matters and the part most easily lost. It is chosen before the data are seen, which is what makes the Type I error rate of section 9.2 a genuine guarantee. Choosing alpha after seeing the p-value — settling on 0.10 because the result came out at 0.08 — destroys that guarantee entirely, since the threshold has then been fitted to the data.
Figure (svg): Two boxes giving the decision rule, with the book's rhyming memory device beneath them
OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, pp. 468-469 — the decision rule and Example 9.10
Picture it
The two cases, with the book's memory device.
Figure (svg): Two boxes giving the decision rule, with the book's rhyming memory device beneath them
The book's third bullet is the one that carries section 9.1's warning forward: when you do not reject, it does not mean you should believe the null is true — it simply means the sample data have failed to provide sufficient evidence to cast serious doubt on it.
Worked example
The baker's test at a 5 percent level.
\[ p\text{-value} \approx 0, \quad \alpha = 0.05 \]
State alpha
Why: The preset level.
\[ 0.05 \]
Compare
Why: Alpha against the p-value.
\[ 0.05 > 0 \]
Apply the rule
Why: Alpha exceeds it.
Describe the result
Why: In the book's terms.
Figure (svg): The solution to Worked example applying the rule shown as a ladder of expressions, one row per legal move
\[ \alpha > p \;\Rightarrow\; \text{reject } H_0 \]
Verify: confirm the conclusion is about evidence rather than truth
Why: Rejecting says the data are hard to reconcile with a mean of 15, not that the mean is definitely more. The book's own conclusion is phrased accordingly: there is sufficient evidence that the true mean height is greater than 15 cm. Sufficient evidence is a claim about the data; is greater would be a claim about the world.
OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, p. 468
Sorting
Compare each p-value against its stated alpha.
Sort into buckets
Sort each case.
All five are from the book's own worked examples in section 9.5. Notice that case (c) at 0.0396 rejects at 0.05 but would not at 0.025 — the same data, two thresholds, opposite decisions, which is why the level belongs in every reported conclusion.
Worked example
The opposite case at the same level.
\[ p\text{-value} = 0.1331, \quad \alpha = 0.025 \]
Compare
Why: Alpha against the p-value.
\[ 0.025 < 0.1331 \]
Apply the rule
Why: Alpha does not exceed it.
Describe the result
Why: Not significant.
State what is NOT concluded
Why: The null is true.
Figure (svg): The solution to Worked example a result that does not reject shown as a ladder of expressions, one row per legal move
\[ \alpha \le p \;\Rightarrow\; \text{do not reject } H_0 \]
Verify: confirm the same data could reject at a different level
Why: A p-value of 0.1331 would still not reject at 0.05 or even 0.10, but it would at 0.15. That the decision depends on a threshold chosen in advance is why the threshold has to be stated in the conclusion — a result described simply as not significant is incomplete without the level it was tested at.
OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, pp. 468-469
Error analysis
Which correctly state the decision and its meaning?
Annotate
On: \( \begin{aligned} &(1)\; \text{do not reject, since } 0.03 < 0.05 \\ &(2)\; \text{reject; there is a 3 percent chance } H_0 \text{ is true} \\ &(3)\; \text{reject; the effect is large} \\ &(4)\; \text{reject; the results are significant at the 5 percent level} \end{aligned} \)
Errors (2) and (3) both reach the correct decision, which is what makes them durable — nothing about the conclusion reveals the faulty reasoning, and both appear regularly in published summaries of research.
Fill the middle
The book's memory device.
Fill in the blanks
\textgo ___
Why: Go — a low p-value rejects. The companion line is that if the p-value is high, the null must fly, meaning it survives.
Two truths and a lie
All three concern the rule.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. Fitting the threshold to the result destroys the error-rate guarantee that made the threshold meaningful — a test where alpha is chosen afterwards has no controlled Type I error rate at all. The book calls alpha preset for exactly this reason.
Prediction
Commit before reasoning.
Predict first
A problem states no significance level. What should be used?
Correct: 0.05, stated explicitly.
Why: The book gives 0.05 as the common standard when no level is specified. What matters as much as the number is saying which level was used, since the same data can reject at one level and not at another — a conclusion without its level is not interpretable.
Section
Section 4
Concept
The p-value assumes the null is true and measures how extreme the data are under that assumption. It is therefore not the probability that the null is true, not the probability that a decision is wrong, and not a measure of how large any effect is.
the conditioning error — Reading P(data given the null) as P(the null given the data). They are different quantities, and the second depends on how plausible the null was to begin with — information the test never has.
\[ P(\text{data} \mid H_0) \;\ne\; P(H_0 \mid \text{data}) \]
The confusion is the same one section 8.1 identified for confidence intervals, and it has the same root: the randomness lives in the data and the procedure, not in the hypothesis. A null hypothesis is either true or false; it has no probability from the test's point of view, so no output of the test can be one.
Figure (svg): Two columns separating what a p-value is from what it is commonly mistaken for
OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, pp. 467-469 — the definition, and the note on judgement
Picture it
Every entry on the right appears regularly in print.
Figure (svg): Two columns separating what a p-value is from what it is commonly mistaken for
The fourth row is the one that misleads most often in practice. A very large study will produce a tiny p-value for an effect far too small to matter, and a small study will produce a large p-value for an effect that matters a great deal — so significance and importance are close to independent.
Worked example
Why a small p-value need not mean a large effect.
\[ \text{a difference of } 0.1 \text{ points, } n = 100\,000 \]
Note the effect size
Why: A tenth of a point.
Note the sample
Why: A hundred thousand.
Ask about the standard error
Why: Sigma over root n.
Conclude
Why: The z-score is large.
Figure (svg): The solution to Worked example significance against size shown as a ladder of expressions, one row per legal move
\[ \text{significant} \;\ne\; \text{important} \]
Verify: confirm the reverse case is equally possible
Why: A study of eight patients could show a difference large enough to change clinical practice and still return a p-value of 0.3, because the standard error is enormous. So the two failures run in both directions: significance without importance in large samples, and importance without significance in small ones. Reporting an effect size alongside a p-value is what separates them.
OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, p. 469
Sorting
For a p-value of 0.02.
Sort into buckets
Sort each statement.
Items (b) and (d) are the two conditioning errors and (e) is the size error. All three appear routinely in press coverage of research, and (d) is the most insidious because it sounds like a modest, careful claim.
Worked example
Why 0.03 is not a three percent chance the null holds.
\[ p\text{-value} = 0.03 \]
Write what was computed
Why: Data given the null.
Write the misreading
Why: Null given the data.
Ask what the second needs
Why: Prior plausibility.
Conclude
Why: They are different.
Figure (svg): The solution to Worked example the conditioning error shown as a ladder of expressions, one row per legal move
\[ P(\text{data} \mid H_0) = 0.03 \]
Verify: confirm the two can differ enormously
Why: Consider testing a thousand hypotheses of which only ten are false. At a 5 percent level about fifty of the 990 true nulls will be rejected, alongside perhaps eight of the ten false ones — so most rejections are of true nulls, despite every p-value being below 0.05. How often a rejection is mistaken depends on how many nulls were false to begin with, which the test never sees.
OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, pp. 467-468
Trap
\[ p = 0.03 \;\Rightarrow\; \text{a 3 percent chance the null is true} \]
Attach the probability to the hypothesis
Why: The p-value is the only probability in sight.
\[ p = P(\text{data} \mid H_0), \text{ not } P(H_0 \mid \text{data}) \]
The null was assumed true in order to compute the number, so it cannot also be a probability about whether it is true.
\[ \text{if the null holds, data this extreme arise 3 percent of the time} \]
Keep the conditional clause in the sentence
Why: The p-value is computed under the null.
The corrected sentence is longer and says less, which is precisely why the wrong one is popular. But the difference is not pedantic: the quantity people want — how likely the null is given the data — depends on how plausible the null was beforehand, and no hypothesis test in this book computes it.
Two truths and a lie
All three concern misreadings.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and combines both errors: it inverts the conditioning and then converts to a probability about a hypothesis. Nothing in a hypothesis test produces a probability that either hypothesis is true.
Estimation
Two studies find the same 0.1-point difference, one with n = 20 and one with n = 100,000.
Predict first
What is likely to differ between them?
Correct: The p-values, dramatically.
Why: The effect is the same 0.1 points in both, but the standard error is about seventy times smaller in the large study, so its p-value will be minute while the small study's is large. That divergence is the clearest demonstration that a p-value measures evidence about a difference existing, not how big the difference is.
Explain it
A classmate writes: our p-value was 0.04, so there is only a 4 percent chance we are wrong.
Discussion prompt
In two sentences or fewer, correct it.
Hint: Ask what was assumed in order to compute the 0.04.
Answer:
The 0.04 was computed by assuming the null is true, so it cannot also measure the chance that assumption is wrong.
What it says is that if the null holds, data at least this extreme would arise about four times in a hundred — which is grounds for doubt, but not a probability about their conclusion.
Section
Section 5
Concept
After making the decision, write a thoughtful conclusion about the hypotheses in terms of the given problem. The conclusion names the significance level, says whether the evidence was sufficient, and describes the claim in the language of the situation rather than in symbols.
a thoughtful conclusion — A sentence a reader who saw no arithmetic could act on. It states the level, whether the evidence was sufficient, and what the claim was — and it never says the null was accepted or proved.
\[ \text{“At the 5 percent level, there is sufficient evidence that ...”} \]
The book adds a note about judgement that is easy to skip and worth keeping: a data analyst should have more confidence in a rejection with a p-value of 0.001 than one at 0.04, even using the same 0.05 threshold, and more confidence in a non-rejection at 0.4 than at 0.056. This makes the analyst use judgement rather than mindlessly applying rules — so the threshold produces the decision, and the p-value still carries information the decision throws away.
Figure (svg): The procedure from p-value to written conclusion
OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, p. 469 — the conclusion, and the note on judgement
Picture it
Four p-values against a 5 percent line.
Figure (svg): A horizontal scale of p-values with a dashed threshold at nought point nought five and four values marked, two of them straddling the line very closely
The pair at 0.04 and 0.056 land on opposite sides of the line and are almost indistinguishable as evidence; the pair at 0.001 and 0.04 land on the same side and are not remotely comparable. A decision rule has to draw a line somewhere, and reporting the p-value alongside the decision is what preserves what the line discards.
Worked example
Example 9.9's result, stated properly.
\[ \text{reject } H_0 \text{ at } \alpha = 0.05 \]
Name the level
Why: Five percent.
\[ \text{at the } 5 \%\text{ level} \]
Say what the evidence did
Why: It sufficed.
State the claim in context
Why: Bread height.
\[ \text{the true mean exceeds } 15 \text{cm} \]
Avoid forbidden wording
Why: No proof, no acceptance.
Figure (svg): The solution to Worked example writing the baker's conclusion shown as a ladder of expressions, one row per legal move
\[ \text{sufficient evidence that } \mu > 15 \]
Verify: confirm the sentence would still make sense to a reader who saw no arithmetic
Why: It names the standard applied, what was found, and what it concerns — a customer could act on it without knowing what a z-score is. A conclusion reading merely reject H-nought communicates nothing to such a reader, which is why the book asks for the claim to be stated in terms of the given problem.
OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, pp. 468-469
Discrimination
For a test that rejected at the 5 percent level.
Sort into buckets
Sort each sentence.
Worked example
The opposite case, worded correctly.
\[ \text{do not reject at } \alpha = 0.025, \; p = 0.1331 \]
Name the level
Why: Two and a half percent.
\[ \text{at the } 2.5 \%\text{ level} \]
Say what the evidence did not do
Why: It did not suffice.
State the claim
Why: In context.
\[ \text{that the mean exceeds } 275 \]
Stop there
Why: No claim about the null.
Figure (svg): The solution to Worked example a conclusion that does not reject shown as a ladder of expressions, one row per legal move
\[ \text{not sufficient evidence that } \mu > 275 \]
Verify: confirm the sentence makes no claim about the null being true
Why: It says the evidence did not suffice for the alternative, and stops — it never asserts that the mean IS 275. That restraint is what section 9.1's rule requires, and writing the sentence in this shape makes the rule automatic rather than something to remember separately.
OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, p. 469
Trap
\[ \text{the result was not significant} \]
Report the decision alone
Why: That is what the test produced.
\[ \text{significant against WHAT threshold?} \]
A p-value of 0.06 is not significant at 0.05 and is significant at 0.10, so the bare word conveys nothing.
\[ \text{not significant at the 5 percent level; } p = 0.06 \]
Give the level, and the p-value itself where possible
Why: The level fixes the decision; the p-value preserves the evidence.
Reporting both is better practice than either alone, and it is what the book's note about judgement is pointing at. The decision tells a reader what the analyst concluded; the p-value lets them see how close the call was and reach their own view if their threshold differs.
Two truths and a lie
All three concern the conclusion.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. A rejection at 5 percent is wrong one time in twenty when the null is true — that is the stated error rate rather than a malfunction, and it is exactly why the language of proof is unavailable.
Estimation
Four p-values: 0.001, 0.04, 0.056 and 0.40, tested at alpha = 0.05.
Predict first
Which pair represents the most similar strength of evidence?
Correct: 0.04 and 0.056.
Why: They differ by less than two hundredths and yet fall on opposite sides of the threshold, while 0.001 and 0.04 share a decision and differ by a factor of forty. This is the book's point about judgement: the threshold produces a decision, and the p-value carries the evidence.
Faded example
A test at the 5 percent level rejects the null that mean bread height is at most 15 cm.
Fill in the blanks
\text5 sufficient \text___ ___ \text___
Why: The sentence names the level, says the evidence was sufficient, and states the claim in the problem's own terms — the three things the book asks a conclusion to do.
Comparison
Fill the blanks. Each row needs the level attached to mean anything.
Comparison matrix
| p-value below alpha | p-value at or above alpha | |
|---|---|---|
| Decision | reject the null | do not reject the null |
| The results are | significant | not significant |
| Conclusion wording | sufficient evidence that ... | not sufficient evidence that ... |
| What is NOT concluded | that the alternative is proved | that the null is true |
The last row is where both columns fail in the same way if the wording slips. Neither decision settles a hypothesis; one finds the evidence sufficient and the other does not, both at a threshold chosen in advance.
Pattern
Six steps, and the first happens before the data arrive.
Report the p-value as well as the decision. The threshold produces a verdict; the p-value preserves how close the call was, which the verdict discards.
OpenStax Introductory Business Statistics 2e, §9.4 Full Hypothesis Test Examples §9.4 Full Hypothesis Test Examples
Check
The definition.
Check your understanding
What does a p-value measure?
Answer: A
Why: It is conditional on the null and covers everything at least as extreme as what was seen.
Check
The decision rule.
Check your understanding
A p-value of 0.0396 is obtained with a preset alpha of 0.05. What is the decision?
Answer: A
Why: Alpha exceeds the p-value, so the null is rejected and the results are significant at that level.
Check
Interpretation.
Check your understanding
Which statement correctly reports a p-value of 0.02?
Answer: A
Why: The p-value is conditional on the null and describes how extreme the data are under it.
Real world
A technology company runs an A/B test on a checkout page, comparing a new design against the current one across 400,000 sessions. The new design increases the conversion rate from 4.00 percent to 4.06 percent, with a p-value of 0.008. The team reports a highly significant improvement and prepares to roll it out.
Discussion prompt
Assess the report, and say what the team should establish before rolling out.
Hint: Separate the two questions a p-value cannot both answer.
Answer:
The p-value is correct and the description highly significant is defensible as a statistical statement. With 400,000 sessions the standard error on a proportion near 0.04 is about 0.0004, so a difference of 0.0006 is around 1.5 standard errors on each arm and comfortably detectable. The data really are hard to explain by chance alone.
But significance here says almost nothing about importance. The improvement is six hundredths of a percentage point — one and a half extra conversions per ten thousand sessions. A sample this large would detect a difference far too small to matter, which is exactly the failure mode the lesson's fourth idea describes.
\[ \text{effect} = 0.0006, \qquad \text{SE} \approx 0.0004, \qquad p = 0.008 \]
What the team should establish is whether the effect is worth having, which is a business question rather than a statistical one. A confidence interval for the difference would be the right tool: it would show the plausible range of improvement in the units that matter, and the team could then ask whether even the optimistic end of that range justifies the engineering and risk of a rollout.
Two further points belong in a careful review. Very large A/B tests routinely produce significant results for changes of no practical consequence, so a pre-specified minimum effect worth detecting should be set before the test rather than after — otherwise every test eventually succeeds. And running many such tests raises the number of false positives regardless of any single p-value: at a 5 percent level, twenty tests of ineffective changes will on average produce one significant result, which is section 9.2's Type I error rate operating exactly as advertised.
Commit first
Answer, then rate your confidence honestly.
Predict first
A p-value of 0.03 is obtained. What exactly does the 0.03 measure?
Correct: The probability of data at least this extreme, assuming the null.
\[ p = P(\text{result at least this extreme} \mid H_0 \text{ true}) \]
Why: The null is assumed in order to compute the number, so the number cannot also be a probability about whether the null holds. How often a rejection is mistaken depends on how plausible the null was beforehand, which the test never sees — and the effect's size is a separate quantity that a p-value does not report at all.
Explain it
They wrote: our p-value was 0.001, so the effect must be large.
Discussion prompt
In two sentences or fewer, correct them.
Hint: Ask how big their sample was.
Answer:
A p-value measures how hard the data are to explain under the null, and a large sample makes even a trivial difference hard to explain that way.
The size of the effect is a separate quantity that has to be reported separately — best as a confidence interval, which gives the plausible range in the units that matter.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, unlikely data under an assumption cast doubt on the assumption. For the second, always include if the null were true. For the third, reject when alpha exceeds the p-value, and never change alpha afterwards. For the fourth, name the level and say sufficient evidence rather than proof. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top, write the rare-event argument as three linked boxes — assume, compute, judge — and beneath it the Didi and Ali example with its probability of 0.005 and what Ali concludes. In the middle, work Example 9.9 completely: write the hypotheses, compute the standard error of 0.158, draw the sampling distribution centred at 15, mark where 17 falls in standard errors, shade the tail, and write the p-value. Then write the full definition of a p-value underneath with its conditional clause underlined. To the right, draw a horizontal p-value scale with a dashed line at 0.05 and mark four values: 0.001, 0.04, 0.056 and 0.40 — then write one sentence on which pair is closest as evidence and one on which pair shares a decision. At the bottom, make two columns headed what a p-value IS and what it is NOT, with at least four entries each, and finish by writing out two full conclusions in the words of a problem: one that rejects and one that does not, each naming its significance level.
Check the middle section by confirming your shaded tail is invisible at any sensible scale — a z of 12.6 is what an essentially zero p-value looks like, and if your shading is visible the standard error has probably been left as sigma. Check the two conclusions by reading them aloud and confirming neither uses the words prove or accept.
Recap
Five things, and two of them are about avoiding misreadings.
| If you see | Then |
|---|---|
| Data very unlikely under the null | Grounds to doubt the null, not proof |
| A p-value below alpha | Reject: the results are significant at that level |
| A p-value at or above alpha | Do not reject; claim nothing about the null |
| No alpha stated | Use 0.05 and say so in the conclusion |
| A p-value read as P(null is true) | The conditioning has been inverted |
| A tiny p-value from a huge sample | Check the effect size before calling it important |
| A conclusion without a level | Incomplete: the same data may reject at another |
Section 9.5 puts all five sections together. With the hypotheses written, the distribution chosen and the decision rule in hand, the remaining question is which tail the p-value occupies — and the alternative hypothesis answers it, making every test left-tailed, right-tailed or two-tailed.
OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, pp. 467-469 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.