9.4 Rare Events, the Sample, Decision and Conclusion

The section that supplies the logic behind a hypothesis test. The reasoning is the rare-event argument: make an assumption about the population, gather data, and if the sample has properties that would be very unlikely to occur when the assumption is true, conclude that the assumption is probably incorrect. The p-value measures that unlikeliness — it is the probability that, if the null hypothesis is true, another randomly selected sample would give results as extreme or more extreme than the ones obtained. A large p-value indicates the null should not be rejected, and the smaller the p-value the stronger the evidence against it. Comparing the p-value against a significance level chosen before the data were collected turns the measurement into a decision, and the conclusion must then be written in the words of the problem. The p-value is conditional on the null being true, which is what makes almost every common misreading of it wrong.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 9.4 Rare Events, the Sample, Decision and Conclusion

Title

Statistics · Chapter 9 — Hypothesis Testing with One Sample

Rare Events, the Sample, Decision and Conclusion

2. By the end of this lesson you can

Objectives

Five outcomes, and two of them are about what the p-value does not say.

OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, pp. 467-469 — the section these objectives are drawn from

3. What you already have

Warm-up

Section 9.3 chose a distribution; sections 9.1 and 9.2 set up the hypotheses and named the errors.

Discussion prompt

Two friends are told a basket of 200 bubbles contains exactly one hundred-dollar bill. The first to reach in draws it. What should the second friend conclude?

Hint: How likely was that, if what they were told is true?

Answer:

If the claim is true, the chance of drawing the one bill first is one in two hundred — 0.005. That is a rare event, and it just happened on the first attempt.

So the second friend has grounds to doubt the claim. Either something very unlikely occurred, or the assumption that there is only one bill is wrong — and the second explanation is the more comfortable one.

That is the whole logic of a hypothesis test, and the book uses exactly this example to introduce it. What remains is to make unlikely precise, which is what a p-value does, and to fix how unlikely is unlikely enough, which is what alpha does.

4. If the data would be very unlikely under the assumption, doubt the assumption

Concept

Suppose you make an assumption about a property of the population — this assumption is the null hypothesis — and then gather sample data randomly. If the sample has properties that would be very unlikely to occur if the assumption is true, then you would conclude that your assumption about the population is probably incorrect.

the rare-event argument — The reasoning behind every hypothesis test. The book's parenthesis states the asymmetry: the assumption is just an assumption and may or may not be true, but the sample data are real and show a fact that seems to contradict it.

\[ P(\text{data this extreme} \mid H_0) \text{ small} \;\Longrightarrow\; \text{doubt } H_0 \]

The argument's shape is worth noticing because it explains the asymmetry section 9.1 established. It can only run one way: unlikely data cast doubt on the assumption, but ordinary data are consistent with the assumption and with a great many alternatives, so they cast doubt on nothing. That is precisely why a test can reject a null and never confirm one.

Figure (svg): Three linked boxes showing the rare-event argument: assume the null, ask how likely the data would be, and doubt the assumption if the answer is small

The book's own example: a rare event happened, so Ali doubts the assumption that made it rare.

OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, p. 467

5. The rare-event argument

Section

Section 1

6. The data are real; the assumption is not

Concept

The book's parenthesis carries the argument: remember that your assumption is just an assumption — it is not a fact and it may or may not be true. But your sample data are real, and the data are showing you a fact that seems to contradict your assumption.

a rare event — An outcome whose probability would be small if the assumption held. Its occurrence is grounds for doubting the assumption rather than for accepting that something unlikely happened.

\[ \text{unlikely data} \;+\; \text{data are real} \;\Longrightarrow\; \text{the assumption gives way} \]

Nothing here is unique to statistics. It is the ordinary reasoning of anyone who checks a claim against evidence: told that a coin is fair and seeing twenty heads in a row, most people conclude the coin is not fair rather than that they witnessed something with probability one in a million. The test simply makes the threshold explicit and the probability computable.

Figure (svg): Three linked boxes showing the rare-event argument: assume the null, ask how likely the data would be, and doubt the assumption if the answer is small

The book's own example: a rare event happened, so Ali doubts the assumption that made it rare.

OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, p. 467 — the rare-events passage and Didi and Ali

7. Three steps

Picture it

Assume, compute, and judge.

Figure (svg): Three linked boxes showing the rare-event argument: assume the null, ask how likely the data would be, and doubt the assumption if the answer is small

The book's own example: a rare event happened, so Ali doubts the assumption that made it rare.

The third box is where the asymmetry lives. A small probability licenses doubting the assumption; a large one licenses nothing at all, because ordinary data are exactly what many different assumptions would predict. Section 9.1's rule against accepting the null is this observation restated.

8. Worked example: the hundred-dollar bill

Worked example

The book's introduction to the argument.

\[ 200 \text{ bubbles, one holding a } \$100 \text{ bill; Didi draws it first} \]

State the assumption

Why: What they were told.

Compute the probability

Why: One in two hundred.

\[ 0.005 \]

Judge it

Why: Very unlikely.

Draw the conclusion

Why: Doubt the assumption.

Figure (svg): The solution to Worked example the hundred-dollar bill shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P = \frac{1}{200} = 0.005 \]

Verify: confirm what the argument does NOT establish

Why: It does not prove there is more than one bill — Didi may simply have been lucky, and one time in two hundred she would be. What the reasoning gives is grounds for doubt proportional to how unlikely the event was, which is exactly what a p-value quantifies. Treating it as proof would be the same error as accepting a null on a large p-value, in the opposite direction.

OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, p. 467

9. What licenses the doubt?

Prediction

Commit before reasoning.

Predict first

In the rare-event argument, why does the assumption give way rather than the data?

  • Because the data are real and the assumption is only an assumption
  • Because data are always correct
  • Because assumptions are usually false
  • Because the probability was computed

Correct: Because the data are real and the assumption is only an assumption.

Why: The book states it in exactly those terms. One of the two must give when they conflict, and the observed data are a fact while the null is a supposition entertained for the sake of argument. This does not mean data are never wrong — a biased sample is bad data — but within the test's own frame the assumption is the movable part.

10. Worked example: the baker's loaves

Worked example

Example 9.9, the same argument with a computed probability.

\[ H_0: \mu \le 15; \; n = 10, \; \bar{x} = 17, \; \sigma = 0.5 \]

Assume the null

Why: The mean is 15.

Find the sampling distribution

Why: By the central limit theorem.

\[ N(15, 0.158) \]

Locate the sample mean

Why: 17 cm.

\[ \text{about } 12.6\text{ SEs out} \]

Judge the probability

Why: The area beyond.

Figure (svg): The solution to Worked example the baker's loaves shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ p\text{-value} = P(\bar{x} > 17) \approx 0 \]

Verify: confirm the standard error is what makes 17 so extreme

Why: Two centimetres above 15 sounds modest, and for a single loaf it would be — four standard deviations, which is unusual but not extraordinary. For a MEAN of ten loaves the standard error is only 0.158 cm, so the same two centimetres is 12.6 standard errors. Chapter 7's concentration of sample means is what turns a moderate-looking gap into an impossible one.

OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, pp. 467-468

11. Trap: treating an unlikely outcome as impossible

Trap

The trap

\[ P = 0.005 \;\Rightarrow\; \text{it cannot have happened by chance} \]

Read a small probability as ruling chance out

Why: One in two hundred is very unlikely.

\[ \text{but one time in two hundred, it does happen} \]

The argument gives grounds for doubt, not a proof, and the strength of those grounds is what the probability measures.

The fix

\[ P = 0.005 \;\Rightarrow\; \text{strong grounds to doubt the assumption} \]

Read the probability as a measure of how strong the doubt is

Why: Smaller means stronger, never certain.

This is section 9.2's Type I error seen from the other side: a test that rejects at a 5 percent level will reject a true null one time in twenty, and that is not a malfunction but the stated error rate. Any language of proof in reporting a test is overstating what the argument can deliver.

12. One of these is false

Two truths and a lie

All three concern the argument.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The probability is computed assuming the null is true
  • C. Ordinary data cast doubt on nothing
  • B. A rare event proves the assumption false

Survives elimination: B

Why: The survivor is false. A rare event is grounds for doubt, and rare events do occur — one draw in two hundred finds the bill even when the claim is true. Reading the argument as proof is what produces overconfident claims from single studies.

13. The bubble probability

Faded example

One bill among 200 bubbles, drawn first.

Fill in the blanks

P = \frac2000.005} = ___

Why: One in two hundred is 0.005 — small enough that Ali reasonably doubts the claim, though not small enough to rule out simple luck.

14. Explain the asymmetry

Explain it

A classmate asks why unlikely data can undermine a claim but ordinary data cannot support it.

Discussion prompt

In two sentences or fewer, explain.

Hint: Ask what other claims would predict the same ordinary data.

Answer:

Unlikely data are hard to explain under the null, so the null becomes hard to hold — that is a real constraint.

Ordinary data are exactly what the null predicts, but also what a dozen nearby alternatives predict, so they do not pick the null out from among them.

15. The p-value

Section

Section 2

16. How extreme the data are, assuming the null

Concept

The p-value is the probability that, if the null hypothesis is true, the results from another randomly selected sample will be as extreme or more extreme than the results obtained from the given sample.

p-value — A conditional probability: it assumes the null is true and asks how often data at least this extreme would arise. Smaller means the observed data are harder to explain under the null.

\[ p\text{-value} = P(\text{a result at least this extreme} \mid H_0 \text{ true}) \]

Two parts of the definition do all the work and are both routinely dropped. If the null hypothesis is true is the condition that makes every common misreading wrong. As extreme or more extreme is why the p-value is a tail area rather than the probability of the exact result observed — which for a continuous variable would be zero, as section 5.1 established.

Figure (svg): A narrow normal curve centred at fifteen with the sample mean of seventeen far out in the right tail

The p-value is the area beyond 17, which is essentially zero: almost no sample of ten loaves would rise that high if the true mean were 15.

OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, pp. 467-468 — the definition and Example 9.9

17. The tail area, on the null's own curve

Picture it

Example 9.9's sampling distribution if the baker's null held.

Figure (svg): A narrow normal curve centred at fifteen with the sample mean of seventeen far out in the right tail

The p-value is the area beyond 17, which is essentially zero: almost no sample of ten loaves would rise that high if the true mean were 15.

The curve is centred at 15 because the null is being assumed, and it is narrow because it describes means of ten loaves rather than single loaves. The observed 17 lies so far right that it is off the picture entirely, which is what a p-value of essentially zero looks like.

18. Worked example: the baker's p-value

Worked example

Example 9.9, computed.

\[ \bar{X} \sim N(15, 0.158); \text{ find } P(\bar{x} > 17) \]

Find the standard error

Why: 0.5 over root 10.

\[ 0.158 \]

Standardise the observation

Why: 17 minus 15, over 0.158.

\[ z = 12.65 \]

Take the right tail

Why: Beyond that z.

Interpret

Why: Almost no sample would.

Figure (svg): The solution to Worked example the baker's p-value shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ p\text{-value} = P(\bar{x} > 17) \approx 0 \]

Verify: confirm the interpretation names the condition

Why: The book's own reading is that almost zero percent of all loaves of bread would be at least as high as 17 cm purely by chance HAD the population mean height really been 15. The conditional clause is not decoration — without it the sentence would be a claim about how often the mean is 15, which is a different and unavailable quantity.

OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, p. 468

19. The condition

Fill the middle

The clause that makes the definition precise.

Fill in the blanks

\texttrue ___

Why: True. Every p-value is computed on the null's own distribution, which is why it cannot be a probability about whether the null holds.

20. Worked example: a p-value that is small but not tiny

Worked example

Try It 9.9, where the evidence is strong but finite.

\[ \sigma = 1, \; n = 36, \; \bar{x} = 12.5; \; H_0: \mu \le 12 \]

Find the standard error

Why: One over six.

\[ 0.1667 \]

Standardise

Why: 0.5 over 0.1667.

\[ z = 3 \]

Take the right tail

Why: Beyond z of 3.

\[ 0.0013 \]

Interpret

Why: About one in 750.

Figure (svg): The solution to Worked example a p-value that is small but not tiny shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ p\text{-value} = P(\bar{x} > 12.5) \approx 0.0013 \]

Verify: confirm the value against the Empirical Rule

Why: A z of exactly 3 puts the observation three standard errors above the mean, and chapter 6's rule says about 0.3 percent of a normal distribution lies more than three standard deviations from the centre — half of that, about 0.15 percent, in one tail. That matches 0.0013 closely, and it is a way of sanity-checking any p-value whose z-score is a whole number.

OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, p. 468

21. Trap: computing the probability of the exact result

Trap

The trap

\[ p\text{-value} = P(\bar{x} = 17) \]

Ask how likely the observed result was

Why: The p-value measures how surprising the data are.

\[ \text{but for a continuous variable that is exactly zero} \]

Section 5.1 established that a single value has no width and therefore no area, whatever the data turn out to be.

The fix

\[ p\text{-value} = P(\bar{x} > 17), \text{ the tail} \]

Include everything at least as extreme, which makes it an area

Why: The definition says as extreme or more extreme.

The as-or-more-extreme clause is what makes the quantity meaningful and computable. It also explains why the direction of the alternative matters: it decides which tail counts as extreme, and section 9.5 will make that the difference between a left, right and two-tailed test.

22. Compute a p-value

Faded example

A right-tailed test with a standard error of 0.1667 and a sample mean 0.5 above the null.

Fill in the blanks

z = \frac30.0013 = ___, \quad p\text___ \approx ___

Why: A z of 3 leaves about 0.0013 in the right tail, matching the Empirical Rule's estimate of roughly 0.15 percent beyond three standard deviations on one side.

23. One of these is false

Two truths and a lie

All three concern the definition.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. It includes results more extreme than the one observed
  • C. A smaller p-value is stronger evidence against the null
  • B. It is the probability of the exact result observed

Survives elimination: B

Why: The survivor is false, and for a continuous variable it would always be zero. The p-value is a tail area covering the observed result and everything further from the null, which is what makes it a measure of how extreme the data are rather than of how likely one particular value was.

24. Read the strength

Estimation

Two tests report p-values of 0.001 and 0.20.

Predict first

What do these say about the evidence against the null?

  • The first gives much stronger evidence against it
  • The first gives much weaker evidence
  • They give the same evidence
  • Neither says anything about evidence

Correct: The first gives much stronger evidence.

Why: The book's rule is that the smaller the p-value, the more unlikely the outcome and the stronger the evidence against the null. Data with a p-value of 0.001 are very hard to explain under the null; data with 0.20 are entirely ordinary under it.

25. The decision rule

Section

Section 3

26. Compare the p-value with a preset alpha

Concept

A systematic way to decide is to compare the p-value with a preset significance level alpha. If alpha exceeds the p-value, reject the null and call the results significant. If alpha is less than or equal to the p-value, do not reject, and the results are not significant.

significance level — The preset alpha, which is the probability of a Type I error. It may or may not be given at the start of a problem, and if none is given a common standard is 0.05.

\[ \alpha > p \;\Rightarrow\; \text{reject}; \qquad \alpha \le p \;\Rightarrow\; \text{do not reject} \]

That alpha is preset is the part that matters and the part most easily lost. It is chosen before the data are seen, which is what makes the Type I error rate of section 9.2 a genuine guarantee. Choosing alpha after seeing the p-value — settling on 0.10 because the result came out at 0.08 — destroys that guarantee entirely, since the threshold has then been fitted to the data.

Figure (svg): Two boxes giving the decision rule, with the book's rhyming memory device beneath them

Example 9.10's memory aid: a p-value below alpha rejects, and one above it does not.

OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, pp. 468-469 — the decision rule and Example 9.10

27. The rule, and the rhyme

Picture it

The two cases, with the book's memory device.

Figure (svg): Two boxes giving the decision rule, with the book's rhyming memory device beneath them

Example 9.10's memory aid: a p-value below alpha rejects, and one above it does not.

The book's third bullet is the one that carries section 9.1's warning forward: when you do not reject, it does not mean you should believe the null is true — it simply means the sample data have failed to provide sufficient evidence to cast serious doubt on it.

28. Worked example: applying the rule

Worked example

The baker's test at a 5 percent level.

\[ p\text{-value} \approx 0, \quad \alpha = 0.05 \]

State alpha

Why: The preset level.

\[ 0.05 \]

Compare

Why: Alpha against the p-value.

\[ 0.05 > 0 \]

Apply the rule

Why: Alpha exceeds it.

Describe the result

Why: In the book's terms.

Figure (svg): The solution to Worked example applying the rule shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \alpha > p \;\Rightarrow\; \text{reject } H_0 \]

Verify: confirm the conclusion is about evidence rather than truth

Why: Rejecting says the data are hard to reconcile with a mean of 15, not that the mean is definitely more. The book's own conclusion is phrased accordingly: there is sufficient evidence that the true mean height is greater than 15 cm. Sufficient evidence is a claim about the data; is greater would be a claim about the world.

OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, p. 468

29. Reject, or not?

Sorting

Compare each p-value against its stated alpha.

Sort into buckets

Sort each case.

Reject the null
p = 0.0187, alpha = 0.05; p = 0.0396, alpha = 0.05; p = 0.0013, alpha = 0.01
Do not reject
p = 0.1331, alpha = 0.025; p = 0.5485, alpha = 0.01
rej
The p-value falls below the significance level, so the data are more surprising than the threshold allows.
no
The p-value exceeds the level, so the data are not surprising enough to cast doubt.

All five are from the book's own worked examples in section 9.5. Notice that case (c) at 0.0396 rejects at 0.05 but would not at 0.025 — the same data, two thresholds, opposite decisions, which is why the level belongs in every reported conclusion.

30. Worked example: a result that does not reject

Worked example

The opposite case at the same level.

\[ p\text{-value} = 0.1331, \quad \alpha = 0.025 \]

Compare

Why: Alpha against the p-value.

\[ 0.025 < 0.1331 \]

Apply the rule

Why: Alpha does not exceed it.

Describe the result

Why: Not significant.

State what is NOT concluded

Why: The null is true.

Figure (svg): The solution to Worked example a result that does not reject shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \alpha \le p \;\Rightarrow\; \text{do not reject } H_0 \]

Verify: confirm the same data could reject at a different level

Why: A p-value of 0.1331 would still not reject at 0.05 or even 0.10, but it would at 0.15. That the decision depends on a threshold chosen in advance is why the threshold has to be stated in the conclusion — a result described simply as not significant is incomplete without the level it was tested at.

OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, pp. 468-469

31. Error analysis: four readings of a p-value of 0.03 at alpha = 0.05

Error analysis

Which correctly state the decision and its meaning?

Annotate

On: \( \begin{aligned} &(1)\; \text{do not reject, since } 0.03 < 0.05 \\ &(2)\; \text{reject; there is a 3 percent chance } H_0 \text{ is true} \\ &(3)\; \text{reject; the effect is large} \\ &(4)\; \text{reject; the results are significant at the 5 percent level} \end{aligned} \)

  • (1) has the comparison backwards. A p-value BELOW alpha rejects, since the data are more surprising than the threshold allows.
  • (2) reaches the right decision by wrong reasoning: the p-value is not a probability that the null is true.
  • (3) confuses significance with size. A tiny effect gives a small p-value in a large enough sample.
  • (4) is correct: the decision, with the level it was made at.

Errors (2) and (3) both reach the correct decision, which is what makes them durable — nothing about the conclusion reveals the faulty reasoning, and both appear regularly in published summaries of research.

32. The rule

Fill the middle

The book's memory device.

Fill in the blanks

\textgo ___

Why: Go — a low p-value rejects. The companion line is that if the p-value is high, the null must fly, meaning it survives.

33. One of these is false

Two truths and a lie

All three concern the rule.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Alpha is chosen before the data are collected
  • C. Not rejecting does not mean the null is true
  • B. Alpha may be adjusted once the p-value is known

Survives elimination: B

Why: The survivor is false. Fitting the threshold to the result destroys the error-rate guarantee that made the threshold meaningful — a test where alpha is chosen afterwards has no controlled Type I error rate at all. The book calls alpha preset for exactly this reason.

34. What if no alpha is given?

Prediction

Commit before reasoning.

Predict first

A problem states no significance level. What should be used?

  • A common standard of 0.05, stated explicitly in the conclusion
  • Whatever makes the result significant
  • The p-value itself
  • No decision can be made

Correct: 0.05, stated explicitly.

Why: The book gives 0.05 as the common standard when no level is specified. What matters as much as the number is saying which level was used, since the same data can reject at one level and not at another — a conclusion without its level is not interpretable.

35. What a p-value is not

Section

Section 4

36. Conditional on the null, and silent about everything else

Concept

The p-value assumes the null is true and measures how extreme the data are under that assumption. It is therefore not the probability that the null is true, not the probability that a decision is wrong, and not a measure of how large any effect is.

the conditioning error — Reading P(data given the null) as P(the null given the data). They are different quantities, and the second depends on how plausible the null was to begin with — information the test never has.

\[ P(\text{data} \mid H_0) \;\ne\; P(H_0 \mid \text{data}) \]

The confusion is the same one section 8.1 identified for confidence intervals, and it has the same root: the randomness lives in the data and the procedure, not in the hypothesis. A null hypothesis is either true or false; it has no probability from the test's point of view, so no output of the test can be one.

Figure (svg): Two columns separating what a p-value is from what it is commonly mistaken for

Every entry on the right is a common misreading in print. The first is the most frequent: a p-value is conditional ON the null, not a probability OF it.

OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, pp. 467-469 — the definition, and the note on judgement

37. Six things it is, and six it is not

Picture it

Every entry on the right appears regularly in print.

Figure (svg): Two columns separating what a p-value is from what it is commonly mistaken for

Every entry on the right is a common misreading in print. The first is the most frequent: a p-value is conditional ON the null, not a probability OF it.

The fourth row is the one that misleads most often in practice. A very large study will produce a tiny p-value for an effect far too small to matter, and a small study will produce a large p-value for an effect that matters a great deal — so significance and importance are close to independent.

38. Worked example: significance against size

Worked example

Why a small p-value need not mean a large effect.

\[ \text{a difference of } 0.1 \text{ points, } n = 100\,000 \]

Note the effect size

Why: A tenth of a point.

Note the sample

Why: A hundred thousand.

Ask about the standard error

Why: Sigma over root n.

Conclude

Why: The z-score is large.

Figure (svg): The solution to Worked example significance against size shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{significant} \;\ne\; \text{important} \]

Verify: confirm the reverse case is equally possible

Why: A study of eight patients could show a difference large enough to change clinical practice and still return a p-value of 0.3, because the standard error is enormous. So the two failures run in both directions: significance without importance in large samples, and importance without significance in small ones. Reporting an effect size alongside a p-value is what separates them.

OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, p. 469

39. Correct reading?

Sorting

For a p-value of 0.02.

Sort into buckets

Sort each statement.

A correct reading
if the null were true, data this extreme would arise 2 percent of the time; the data are surprising under the null
A misreading
there is a 2 percent chance the null is true; there is a 2 percent chance the decision is wrong; the effect is large
ok
The statement is conditional on the null and describes the data.
no
It attaches a probability to the hypothesis or the decision, or confuses significance with effect size.

Items (b) and (d) are the two conditioning errors and (e) is the size error. All three appear routinely in press coverage of research, and (d) is the most insidious because it sounds like a modest, careful claim.

40. Worked example: the conditioning error

Worked example

Why 0.03 is not a three percent chance the null holds.

\[ p\text{-value} = 0.03 \]

Write what was computed

Why: Data given the null.

Write the misreading

Why: Null given the data.

Ask what the second needs

Why: Prior plausibility.

Conclude

Why: They are different.

Figure (svg): The solution to Worked example the conditioning error shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(\text{data} \mid H_0) = 0.03 \]

Verify: confirm the two can differ enormously

Why: Consider testing a thousand hypotheses of which only ten are false. At a 5 percent level about fifty of the 990 true nulls will be rejected, alongside perhaps eight of the ten false ones — so most rejections are of true nulls, despite every p-value being below 0.05. How often a rejection is mistaken depends on how many nulls were false to begin with, which the test never sees.

OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, pp. 467-468

41. Trap: reading the p-value as the chance the null is true

Trap

The trap

\[ p = 0.03 \;\Rightarrow\; \text{a 3 percent chance the null is true} \]

Attach the probability to the hypothesis

Why: The p-value is the only probability in sight.

\[ p = P(\text{data} \mid H_0), \text{ not } P(H_0 \mid \text{data}) \]

The null was assumed true in order to compute the number, so it cannot also be a probability about whether it is true.

The fix

\[ \text{if the null holds, data this extreme arise 3 percent of the time} \]

Keep the conditional clause in the sentence

Why: The p-value is computed under the null.

The corrected sentence is longer and says less, which is precisely why the wrong one is popular. But the difference is not pedantic: the quantity people want — how likely the null is given the data — depends on how plausible the null was beforehand, and no hypothesis test in this book computes it.

42. One of these is false

Two truths and a lie

All three concern misreadings.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. A tiny p-value can accompany a trivial effect
  • C. The p-value is conditional on the null being true
  • B. A p-value of 0.01 means the alternative is 99 percent likely

Survives elimination: B

Why: The survivor is false and combines both errors: it inverts the conditioning and then converts to a probability about a hypothesis. Nothing in a hypothesis test produces a probability that either hypothesis is true.

43. Significance and size

Estimation

Two studies find the same 0.1-point difference, one with n = 20 and one with n = 100,000.

Predict first

What is likely to differ between them?

  • The p-values, dramatically — while the effect is identical
  • The effect sizes
  • Nothing
  • Only the confidence levels

Correct: The p-values, dramatically.

Why: The effect is the same 0.1 points in both, but the standard error is about seventy times smaller in the large study, so its p-value will be minute while the small study's is large. That divergence is the clearest demonstration that a p-value measures evidence about a difference existing, not how big the difference is.

44. Correct the sentence

Explain it

A classmate writes: our p-value was 0.04, so there is only a 4 percent chance we are wrong.

Discussion prompt

In two sentences or fewer, correct it.

Hint: Ask what was assumed in order to compute the 0.04.

Answer:

The 0.04 was computed by assuming the null is true, so it cannot also measure the chance that assumption is wrong.

What it says is that if the null holds, data at least this extreme would arise about four times in a hundred — which is grounds for doubt, but not a probability about their conclusion.

45. Writing the conclusion

Section

Section 5

46. In the words of the problem, at a stated level

Concept

After making the decision, write a thoughtful conclusion about the hypotheses in terms of the given problem. The conclusion names the significance level, says whether the evidence was sufficient, and describes the claim in the language of the situation rather than in symbols.

a thoughtful conclusion — A sentence a reader who saw no arithmetic could act on. It states the level, whether the evidence was sufficient, and what the claim was — and it never says the null was accepted or proved.

\[ \text{“At the 5 percent level, there is sufficient evidence that ...”} \]

The book adds a note about judgement that is easy to skip and worth keeping: a data analyst should have more confidence in a rejection with a p-value of 0.001 than one at 0.04, even using the same 0.05 threshold, and more confidence in a non-rejection at 0.4 than at 0.056. This makes the analyst use judgement rather than mindlessly applying rules — so the threshold produces the decision, and the p-value still carries information the decision throws away.

Figure (svg): The procedure from p-value to written conclusion

Step two's instruction to draw the graph is the book's: the test is easier to perform because you see the problem more clearly.

OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, p. 469 — the conclusion, and the note on judgement

47. How much the threshold discards

Picture it

Four p-values against a 5 percent line.

Figure (svg): A horizontal scale of p-values with a dashed threshold at nought point nought five and four values marked, two of them straddling the line very closely

The threshold is a decision rule, not a description of the evidence — which is why the book asks for judgement alongside it.

The pair at 0.04 and 0.056 land on opposite sides of the line and are almost indistinguishable as evidence; the pair at 0.001 and 0.04 land on the same side and are not remotely comparable. A decision rule has to draw a line somewhere, and reporting the p-value alongside the decision is what preserves what the line discards.

48. Worked example: writing the baker's conclusion

Worked example

Example 9.9's result, stated properly.

\[ \text{reject } H_0 \text{ at } \alpha = 0.05 \]

Name the level

Why: Five percent.

\[ \text{at the } 5 \%\text{ level} \]

Say what the evidence did

Why: It sufficed.

State the claim in context

Why: Bread height.

\[ \text{the true mean exceeds } 15 \text{cm} \]

Avoid forbidden wording

Why: No proof, no acceptance.

Figure (svg): The solution to Worked example writing the baker's conclusion shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{sufficient evidence that } \mu > 15 \]

Verify: confirm the sentence would still make sense to a reader who saw no arithmetic

Why: It names the standard applied, what was found, and what it concerns — a customer could act on it without knowing what a z-score is. A conclusion reading merely reject H-nought communicates nothing to such a reader, which is why the book asks for the claim to be stated in terms of the given problem.

OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, pp. 468-469

49. Is this conclusion acceptable?

Discrimination

For a test that rejected at the 5 percent level.

Sort into buckets

Sort each sentence.

Acceptable
at the 5 percent level there is sufficient evidence that the mean exceeds 15; the data support the claim, at the 5 percent significance level; we reject the null hypothesis
Overstates the result
we have proved the mean is greater than 15; the null hypothesis is false
ok
It reports evidence and a decision at a stated level, without claiming certainty.
no
It claims the hypothesis has been settled, which no test can do.

50. Worked example: a conclusion that does not reject

Worked example

The opposite case, worded correctly.

\[ \text{do not reject at } \alpha = 0.025, \; p = 0.1331 \]

Name the level

Why: Two and a half percent.

\[ \text{at the } 2.5 \%\text{ level} \]

Say what the evidence did not do

Why: It did not suffice.

State the claim

Why: In context.

\[ \text{that the mean exceeds } 275 \]

Stop there

Why: No claim about the null.

Figure (svg): The solution to Worked example a conclusion that does not reject shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{not sufficient evidence that } \mu > 275 \]

Verify: confirm the sentence makes no claim about the null being true

Why: It says the evidence did not suffice for the alternative, and stops — it never asserts that the mean IS 275. That restraint is what section 9.1's rule requires, and writing the sentence in this shape makes the rule automatic rather than something to remember separately.

OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, p. 469

51. Trap: reporting the decision without the level

Trap

The trap

\[ \text{the result was not significant} \]

Report the decision alone

Why: That is what the test produced.

\[ \text{significant against WHAT threshold?} \]

A p-value of 0.06 is not significant at 0.05 and is significant at 0.10, so the bare word conveys nothing.

The fix

\[ \text{not significant at the 5 percent level; } p = 0.06 \]

Give the level, and the p-value itself where possible

Why: The level fixes the decision; the p-value preserves the evidence.

Reporting both is better practice than either alone, and it is what the book's note about judgement is pointing at. The decision tells a reader what the analyst concluded; the p-value lets them see how close the call was and reach their own view if their threshold differs.

52. One of these is false

Two truths and a lie

All three concern the conclusion.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The conclusion should name the significance level
  • C. The p-value carries information the decision discards
  • B. A rejection may be reported as proof

Survives elimination: B

Why: The survivor is false. A rejection at 5 percent is wrong one time in twenty when the null is true — that is the stated error rate rather than a malfunction, and it is exactly why the language of proof is unavailable.

53. Which two are closest as evidence?

Estimation

Four p-values: 0.001, 0.04, 0.056 and 0.40, tested at alpha = 0.05.

Predict first

Which pair represents the most similar strength of evidence?

  • 0.04 and 0.056, despite opposite decisions
  • 0.001 and 0.04, since both reject
  • 0.056 and 0.40, since neither rejects
  • 0.001 and 0.40

Correct: 0.04 and 0.056.

Why: They differ by less than two hundredths and yet fall on opposite sides of the threshold, while 0.001 and 0.04 share a decision and differ by a factor of forty. This is the book's point about judgement: the threshold produces a decision, and the p-value carries the evidence.

54. Write the conclusion

Faded example

A test at the 5 percent level rejects the null that mean bread height is at most 15 cm.

Fill in the blanks

\text5 sufficient \text___ ___ \text___

Why: The sentence names the level, says the evidence was sufficient, and states the claim in the problem's own terms — the three things the book asks a conclusion to do.

55. The two decisions, fully stated

Comparison

Fill the blanks. Each row needs the level attached to mean anything.

Comparison matrix

p-value below alphap-value at or above alpha
Decisionreject the nulldo not reject the null
The results aresignificantnot significant
Conclusion wordingsufficient evidence that ...not sufficient evidence that ...
What is NOT concludedthat the alternative is provedthat the null is true

The last row is where both columns fail in the same way if the wording slips. Neither decision settles a hypothesis; one finds the evidence sufficient and the other does not, both at a threshold chosen in advance.

56. From data to conclusion, in order

Pattern

Six steps, and the first happens before the data arrive.

  1. Fix alpha before collecting or examining the data; if none is stated, use 0.05 and say so.
  2. Assume the null is true and identify the sampling distribution it implies.
  3. Compute the p-value as the probability of a result at least as extreme as the one observed, and draw the graph.
  4. Compare: reject if alpha exceeds the p-value, and do not reject otherwise.
  5. State the decision in the language of rejection, never of acceptance or proof.
  6. Write the conclusion in the words of the problem, naming the significance level and reporting the p-value.

Report the p-value as well as the decision. The threshold produces a verdict; the p-value preserves how close the call was, which the verdict discards.

OpenStax Introductory Business Statistics 2e, §9.4 Full Hypothesis Test Examples §9.4 Full Hypothesis Test Examples

57. Check yourself 1 of 3

Check

The definition.

Check your understanding

What does a p-value measure?

  • A. The chance of a result at least this extreme, if the null is true (correct)
  • B. The chance that the null hypothesis is true
  • C. The chance of the exact result observed
  • D. The size of the effect

Answer: A

Why: It is conditional on the null and covers everything at least as extreme as what was seen.

Why B tempts people
That inverts the conditioning; the null was assumed in order to compute the p-value.
Why C tempts people
For a continuous variable that probability is exactly zero.
Why D tempts people
A tiny effect gives a small p-value in a large enough sample; the two are close to independent.

58. Check yourself 2 of 3

Check

The decision rule.

Check your understanding

A p-value of 0.0396 is obtained with a preset alpha of 0.05. What is the decision?

  • A. Reject the null: the results are significant (correct)
  • B. Do not reject the null
  • C. Accept the null hypothesis
  • D. Raise alpha to 0.10 and retest

Answer: A

Why: Alpha exceeds the p-value, so the null is rejected and the results are significant at that level.

Why B tempts people
That would follow only if the p-value were at or above alpha.
Why C tempts people
No test accepts a null, whatever the p-value.
Why D tempts people
Alpha is preset; adjusting it after seeing the result destroys the error-rate guarantee.

59. Check yourself 3 of 3

Check

Interpretation.

Check your understanding

Which statement correctly reports a p-value of 0.02?

  • A. If the null were true, data this extreme would arise about 2 percent of the time (correct)
  • B. There is a 2 percent chance the null is true
  • C. There is a 2 percent chance of an error
  • D. The effect is 2 percent in size

Answer: A

Why: The p-value is conditional on the null and describes how extreme the data are under it.

Why B tempts people
This inverts the conditioning and asks for a quantity the test does not compute.
Why C tempts people
The chance a decision is wrong depends on how plausible the null was beforehand.
Why D tempts people
A p-value is not an effect size and carries no information about magnitude.

60. Where this shows up outside the textbook

Real world

A technology company runs an A/B test on a checkout page, comparing a new design against the current one across 400,000 sessions. The new design increases the conversion rate from 4.00 percent to 4.06 percent, with a p-value of 0.008. The team reports a highly significant improvement and prepares to roll it out.

Discussion prompt

Assess the report, and say what the team should establish before rolling out.

Hint: Separate the two questions a p-value cannot both answer.

Answer:

The p-value is correct and the description highly significant is defensible as a statistical statement. With 400,000 sessions the standard error on a proportion near 0.04 is about 0.0004, so a difference of 0.0006 is around 1.5 standard errors on each arm and comfortably detectable. The data really are hard to explain by chance alone.

But significance here says almost nothing about importance. The improvement is six hundredths of a percentage point — one and a half extra conversions per ten thousand sessions. A sample this large would detect a difference far too small to matter, which is exactly the failure mode the lesson's fourth idea describes.

\[ \text{effect} = 0.0006, \qquad \text{SE} \approx 0.0004, \qquad p = 0.008 \]

What the team should establish is whether the effect is worth having, which is a business question rather than a statistical one. A confidence interval for the difference would be the right tool: it would show the plausible range of improvement in the units that matter, and the team could then ask whether even the optimistic end of that range justifies the engineering and risk of a rollout.

Two further points belong in a careful review. Very large A/B tests routinely produce significant results for changes of no practical consequence, so a pre-specified minimum effect worth detecting should be set before the test rather than after — otherwise every test eventually succeeds. And running many such tests raises the number of false positives regardless of any single p-value: at a 5 percent level, twenty tests of ineffective changes will on average produce one significant result, which is section 9.2's Type I error rate operating exactly as advertised.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

A p-value of 0.03 is obtained. What exactly does the 0.03 measure?

  • The probability that the null hypothesis is true
  • The probability of data at least this extreme, computed assuming the null is true
  • The probability that the decision to reject is wrong
  • The proportion of the effect explained

Correct: The probability of data at least this extreme, assuming the null.

\[ p = P(\text{result at least this extreme} \mid H_0 \text{ true}) \]

Why: The null is assumed in order to compute the number, so the number cannot also be a probability about whether the null holds. How often a rejection is mistaken depends on how plausible the null was beforehand, which the test never sees — and the effect's size is a separate quantity that a p-value does not report at all.

62. Explain it to someone a year behind you

Explain it

They wrote: our p-value was 0.001, so the effect must be large.

Discussion prompt

In two sentences or fewer, correct them.

Hint: Ask how big their sample was.

Answer:

A p-value measures how hard the data are to explain under the null, and a large sample makes even a trivial difference hard to explain that way.

The size of the effect is a separate quantity that has to be reported separately — best as a confidence interval, which gives the plausible range in the units that matter.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Stating the rare-event argument and what licenses the doubt
  • Defining a p-value with its conditional clause intact
  • Applying the decision rule against a preset alpha
  • Writing a conclusion that overstates nothing

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the first, unlikely data under an assumption cast doubt on the assumption. For the second, always include if the null were true. For the third, reject when alpha exceeds the p-value, and never change alpha afterwards. For the fourth, name the level and say sufficient evidence rather than proof. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Fifteen minutes.

Draw it

At the top, write the rare-event argument as three linked boxes — assume, compute, judge — and beneath it the Didi and Ali example with its probability of 0.005 and what Ali concludes. In the middle, work Example 9.9 completely: write the hypotheses, compute the standard error of 0.158, draw the sampling distribution centred at 15, mark where 17 falls in standard errors, shade the tail, and write the p-value. Then write the full definition of a p-value underneath with its conditional clause underlined. To the right, draw a horizontal p-value scale with a dashed line at 0.05 and mark four values: 0.001, 0.04, 0.056 and 0.40 — then write one sentence on which pair is closest as evidence and one on which pair shares a decision. At the bottom, make two columns headed what a p-value IS and what it is NOT, with at least four entries each, and finish by writing out two full conclusions in the words of a problem: one that rejects and one that does not, each naming its significance level.

Check the middle section by confirming your shaded tail is invisible at any sensible scale — a z of 12.6 is what an essentially zero p-value looks like, and if your shading is visible the standard error has probably been left as sigma. Check the two conclusions by reading them aloud and confirming neither uses the words prove or accept.

65. What you can do now

Recap

Five things, and two of them are about avoiding misreadings.

If you seeThen
Data very unlikely under the nullGrounds to doubt the null, not proof
A p-value below alphaReject: the results are significant at that level
A p-value at or above alphaDo not reject; claim nothing about the null
No alpha statedUse 0.05 and say so in the conclusion
A p-value read as P(null is true)The conditioning has been inverted
A tiny p-value from a huge sampleCheck the effect size before calling it important
A conclusion without a levelIncomplete: the same data may reject at another

Section 9.5 puts all five sections together. With the hypotheses written, the distribution chosen and the decision rule in hand, the remaining question is which tail the p-value occupies — and the alternative hypothesis answers it, making every test left-tailed, right-tailed or two-tailed.

OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion §9.4, pp. 467-469 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §9.4 Rare Events, the Sample, Decision and Conclusion — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 467-469
  2. OpenStax Introductory Business Statistics 2e, §9.4 Full Hypothesis Test Examples — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108