Chapter 8 asked which values of a parameter the data supports. Chapter 9 starts from a specific claim about the parameter and asks whether the data are consistent with it, and the whole apparatus rests on stating two contradictory hypotheses correctly before any arithmetic begins. The null hypothesis is a statement of no difference, often the status quo, and always contains a symbol with an equal in it. The alternative hypothesis contradicts it, never contains an equal sign, and is usually what the researcher is trying to show. Because the two are contradictory, evidence in the form of sample data can support one over the other, and the decision is always either to reject the null hypothesis or to decline to reject it. It is never to accept the null, because failing to find evidence against a claim is not evidence for it.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 9 — Hypothesis Testing with One Sample
Null and Alternative Hypotheses
Objectives
Five outcomes, and the last is the one most often stated wrongly.
OpenStax Introductory Statistics 2e, §9.1 Null and Alternative Hypotheses §9.1, pp. 462-464 — the section these objectives are drawn from
Warm-up
Chapter 8 built an interval of plausible values for an unknown parameter from a sample.
Discussion prompt
A 95 percent confidence interval for a mean comes out as (67.02, 68.98). Someone claims the true mean is 65. What can be said about that claim, and what would it take to say it more formally?
Hint: Is 65 among the values the interval calls plausible?
Answer:
The value 65 lies outside the interval, so it is not among the values the data supports at that level of confidence. That is already an informal test of the claim, and it points at what this chapter formalises.
What is missing is a procedure. The interval was built without reference to any particular claim, and reading a verdict off it afterwards leaves several things unstated: which claim was being tested, how strong the evidence against it is, and what conclusion is warranted if it had fallen inside.
Chapter 9 supplies that procedure, and it begins where every test begins — with two carefully written sentences. Get those wrong and no amount of correct arithmetic afterwards will produce a meaningful answer.
Concept
The test begins by considering two hypotheses containing opposing viewpoints. The null hypothesis is a statement of no difference between the variables, which can often be considered the status quo. The alternative hypothesis is a claim about the population contradictory to the null, and is what we conclude when we reject the null.
null and alternative hypotheses — A contradictory pair, written H-nought and H-a. Since they contradict, evidence for one is evidence against the other, and sample data decide whether there is enough to reject the null.
\[ H_0 \text{ and } H_a \text{ contradict: exactly one can hold} \]
The asymmetry between them is deliberate and is the point of the whole construction. The null is assumed while the test runs, and the question asked is whether the data would be surprising if it were true. That means the test can find a claim implausible but can never confirm one — which is why the two available decisions are reject and do not reject, and why the phrase accept the null does not appear.
Figure (svg): Two columns contrasting the null hypothesis with the alternative hypothesis
OpenStax Introductory Statistics 2e, §9.1 Null and Alternative Hypotheses §9.1, p. 462
Section
Section 1
Concept
The null hypothesis states no difference — the variables are not related, nothing has changed, the parameter equals its assumed value. The alternative is what the researcher is usually trying to prove, and it is what we conclude when we reject the null.
the status quo reading — The book notes that the null can often be considered the status quo, so that if you cannot accept the null it requires some action. That framing explains why the burden of proof sits where it does.
\[ H_0: \text{no difference}; \qquad H_a: \text{a difference, in a stated direction} \]
Reading the null as the status quo makes the asymmetry feel less arbitrary. A drug is assumed ineffective until evidence says otherwise; a process is assumed in control until evidence says otherwise. In each case the default costs nothing to maintain, and the test asks whether the data justify departing from it — which is why strong evidence is demanded before the null is abandoned.
Figure (svg): Two columns contrasting the null hypothesis with the alternative hypothesis
OpenStax Introductory Statistics 2e, §9.1 Null and Alternative Hypotheses §9.1, p. 462 — the definitions of both hypotheses
Picture it
What each hypothesis is for.
Figure (svg): Two columns contrasting the null hypothesis with the alternative hypothesis
The third row of each column is the operational rule and the next idea's subject. Every other property follows from the pair being contradictory, but the equal-sign rule is a convention worth learning as a rule because it settles which of the two a given statement belongs in.
Worked example
Example 9.4's setting, read for which claim is the default.
\[ 6.6 \text{ percent of U.S. students take advanced placement exams} \]
Find the established figure
Why: The reported rate.
\[ 6.6 \% \]
Find what is being tested
Why: Whether it is more.
\[ \text{more than } 6.6 \% \]
Assign the alternative
Why: The claim under test.
\[ p > 0.066 \]
Assign the null
Why: Its contradiction.
\[ p \le 0.066 \]
Figure (svg): The solution to Worked example identifying the status quo shown as a ladder of expressions, one row per legal move
\[ H_0: p \le 0.066, \qquad H_a: p > 0.066 \]
Verify: confirm the pair covers every possibility with no overlap
Why: Every proportion is either at most 0.066 or above it, and none is both — so the two statements partition the possibilities exactly. A pair that left a gap, such as p below 0.066 against p above it, would leave the value 0.066 itself unaccounted for, and a pair that overlapped would allow both to be true at once. Checking the partition is the fastest test of a correctly written pair.
OpenStax Introductory Statistics 2e, §9.1 Null and Alternative Hypotheses §9.1, p. 463
Sorting
Ask which statement is the default and which is being argued for.
Sort into buckets
Sort each statement by where it belongs.
Items (b) and (d) are both things somebody wants to show, which is the tell for an alternative. Item (c) reads as a claim too, but it is the assumption in force until evidence overturns it — and its natural symbolic form contains an equality.
Worked example
Example 9.2, where no direction is specified.
\[ \text{is the mean GPA different from } 2.0? \]
Find the parameter
Why: A mean GPA.
Find the comparison value
Why: Two point zero.
\[ 2.0 \]
Read the direction
Why: Different, either way.
Write the pair
Why: Equality against inequality.
\[ \mu = 2.0\text{ against } \mu \ne 2.0 \]
Figure (svg): The solution to Worked example a two-sided claim shown as a ladder of expressions, one row per legal move
\[ H_0: \mu = 2.0, \qquad H_a: \mu \ne 2.0 \]
Verify: confirm why this pair uses a plain equals rather than an inequality
Why: The word different admits departures in both directions, so the alternative has to be not-equal — and the only statement contradicting not-equal is plain equality. The one-sided examples use at most or at least in the null because their alternatives point one way. Reading the direction word before choosing the symbols is what selects the right row of the book's table.
OpenStax Introductory Statistics 2e, §9.1 Null and Alternative Hypotheses §9.1, p. 463
Trap
\[ H_0: p > 0.066, \qquad H_a: p \le 0.066 \]
Put the claim being investigated in the null
Why: It is the statement the study is about.
\[ \text{but } H_0 \text{ must contain an equal sign} \]
The pair is also backwards in purpose: the test would now demand evidence against the very thing being proposed.
\[ H_0: p \le 0.066, \qquad H_a: p > 0.066 \]
Put the claim being tested for in the ALTERNATIVE
Why: The null is the default it is being tested against.
The equal-sign rule catches this instantly, which is why it is worth applying mechanically before any thought about the science. It also encodes something real: the burden of proof falls on the new claim, so the new claim goes where evidence has to be found for it.
Fill the middle
The book's definition.
Fill in the blanks
\textdifference ___ \text___
Why: No difference — the variables are not related, or the parameter equals its assumed value. That default is what the sample evidence is weighed against.
Two truths and a lie
All three concern the pair.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The book's rule is absolute: the null always has a symbol with an equal in it, and the alternative never does. That convention is what makes it possible to compute a p-value at all, since a test needs a single specific value to assume.
Prediction
Commit before reasoning.
Predict first
Why must the two hypotheses be contradictory?
Correct: So that evidence against one is evidence for the other.
Why: If the two could both be true, rejecting the null would say nothing about the alternative, and the test would have no conclusion to draw. The partition is what turns a statement about how surprising the data are into a statement about which claim the data supports.
Section
Section 2
Concept
The null hypothesis always has a symbol with an equal in it: equals, greater than or equal to, or less than or equal to. The alternative never does: not equal, greater than, or less than. The choice among them depends on the wording of the hypothesis test.
the equal-sign rule — H-nought always contains an equality and H-a never does. It settles which of a contradictory pair is which, without any reference to the science of the problem.
\[ H_0 \in \{=, \ge, \le\}, \qquad H_a \in \{\ne, >, <\} \]
The book adds a practical note worth knowing: many researchers use a plain equals in the null even when the alternative carries a greater-than or less-than, and it calls this acceptable because we only make the decision to reject or not reject the null. Both conventions give the same test, since the p-value is computed at the boundary value either way.
Figure (svg): A two-column table pairing each null hypothesis symbol with its matching alternative
OpenStax Introductory Statistics 2e, §9.1 Null and Alternative Hypotheses §9.1, p. 462 — the symbol table and the note on conventions
Picture it
The book's own table of matching symbols.
Figure (svg): A two-column table pairing each null hypothesis symbol with its matching alternative
The pairings are forced once you know the direction: an alternative of greater-than must face a null of at-most, and so on. Getting the direction from the wording is the only judgement required, and the symbols follow mechanically from it.
Worked example
Example 9.3, a less-than claim.
\[ \text{do college students take less than five years to graduate?} \]
Read the direction
Why: Less than.
\[ \text{H-a uses } < \]
Find the matching null
Why: From the book's table.
\[ H - 0\text{ uses } \ge \]
Write the alternative
Why: Under five years.
\[ \mu < 5 \]
Write the null
Why: At least five.
\[ \mu \ge 5 \]
Figure (svg): The solution to Worked example matching the symbols shown as a ladder of expressions, one row per legal move
\[ H_0: \mu \ge 5, \qquad H_a: \mu < 5 \]
Verify: confirm the equal sign sits in the right hypothesis
Why: The greater-than-or-equal-to is in the null and the plain less-than is in the alternative, which is the book's second row. Swapping them would put the equality in H-a, which the rule forbids and which would leave the test with no single value to assume when computing a p-value.
OpenStax Introductory Statistics 2e, §9.1 Null and Alternative Hypotheses §9.1, p. 463
Faded example
A test of whether more than 40 percent pass a driving test on the first try.
Fill in the blanks
H_0: p <= 0.40, \qquad H_a: p > 0.40
Why: The claim is that more than 40 percent pass, so that goes in the alternative with a plain greater-than, and the null takes the matching at-most.
Worked example
The book's note about researchers who write a plain equals.
\[ H_a: \mu < 5 \]
The book's table form
Why: At least five.
\[ \mu \ge 5 \]
The common research form
Why: Exactly five.
\[ \mu = 5 \]
Ask what changes
Why: The p-value calculation.
Say why
Why: Both compute at the boundary.
\[ \mu = 5 \]
Figure (svg): The solution to Worked example the two acceptable conventions shown as a ladder of expressions, one row per legal move
\[ H_0: \mu \ge 5 \quad\text{or}\quad H_0: \mu = 5 \]
Verify: confirm the arithmetic really is unaffected
Why: A p-value asks how likely the observed data would be if the null were true, and it is computed at mu equal to 5 in both versions — because among all values at least 5, the one nearest the alternative is 5 itself, and that is the hardest case for the alternative to beat. The book's justification is that we only make the decision to reject or not reject, so the distinction never surfaces in a conclusion.
OpenStax Introductory Statistics 2e, §9.1 Null and Alternative Hypotheses §9.1, p. 462
Error analysis
The book's answer is p at most 0.30 against p greater than 0.30.
Annotate
On: \( \begin{aligned} &(1)\; H_0: p > 0.30, \; H_a: p \le 0.30 \\ &(2)\; H_0: p < 0.30, \; H_a: p > 0.30 \\ &(3)\; H_0: p \le 0.30, \; H_a: p \ge 0.30 \\ &(4)\; H_0: p \le 0.30, \; H_a: p > 0.30 \end{aligned} \)
Errors (2) and (3) are the interesting ones because both symbols look plausible in isolation. Checking that the pair partitions the possible values — every value in exactly one — catches both, and it is the same check that confirms (4) is right.
Discrimination
Apply the equal-sign rule alone.
Sort into buckets
Sort each statement by which hypothesis it could belong to.
Two truths and a lie
All three concern symbols.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The alternative never contains an equality of any kind — only strict inequalities or not-equal. An at-least in the alternative would overlap the null at the boundary value, so the two would no longer contradict.
Prediction
Commit before reasoning.
Predict first
Why does the null hypothesis need a symbol with an equal in it?
Correct: Because the test must assume one specific value.
Why: A p-value asks how likely the data would be IF the null were true, and that requires a single parameter value to build a distribution around. A statement like mu greater than 5 names no particular value, so no distribution could be constructed and no p-value computed — which is exactly why the equality lives in the hypothesis that gets assumed.
Section
Section 3
Concept
The choice of symbol depends on the wording of the hypothesis test. The reliable order is to find the parameter, find the comparison value, write the alternative from the direction the wording gives, and then write the null as its exact contradiction.
direction words — More than and greater than give a greater-than alternative; less than and fewer than give a less-than one; different from, changed and not equal give a two-sided alternative.
\[ \text{“more than”} \to H_a: > ; \quad \text{“less than”} \to H_a: < ; \quad \text{“different”} \to H_a: \ne \]
Writing the alternative first is worth insisting on because the wording describes it directly, while the null is only defined by contradiction. Attempting the null first means working out the complement of a sentence before it has been symbolised, which is where most translation errors originate.
Figure (svg): A procedure for turning a worded claim into a pair of hypotheses
OpenStax Introductory Statistics 2e, §9.1 Null and Alternative Hypotheses §9.1, pp. 462-463 — the wording rule and Examples 9.1 through 9.4
Picture it
The procedure, with the alternative written first.
Figure (svg): A procedure for turning a worded claim into a pair of hypotheses
The final step is the check the previous idea introduced: the pair must partition the possible values. Applying it every time costs a few seconds and catches the two errors that survive the symbol rule — a gap at the boundary, and an overlap at it.
Worked example
The book's Examples 9.1 through 9.4, side by side.
\[ \text{four worded claims} \]
Voters, more than 30 percent
Why: A proportion, one-sided up.
\[ p \le 0.30\text{ against } p > 0.30 \]
GPA, different from 2.0
Why: A mean, two-sided.
\[ \mu = 2.0\text{ against } \mu \ne 2.0 \]
Graduation, less than 5 years
Why: A mean, one-sided down.
\[ \mu \ge 5\text{ against } \mu < 5 \]
AP exams, more than 6.6 percent
Why: A proportion, one-sided up.
\[ p \le 0.066\text{ against } p > 0.066 \]
Figure (svg): The solution to Worked example four claims translated shown as a ladder of expressions, one row per legal move
\[ H_a \text{ is the claim; } H_0 \text{ is what is left} \]
Verify: confirm the direction word alone determined each pair
Why: More than gave a greater-than, less than gave a less-than, and different from gave a not-equal — three different direction words producing three different shapes, with the parameter and the comparison value simply read off. No knowledge of voting, education or graduation rates was needed at any point, which is what makes the translation mechanical once the order is fixed.
OpenStax Introductory Statistics 2e, §9.1 Null and Alternative Hypotheses §9.1, pp. 462-463
Matching
Match each phrase to the symbol it produces in H-a.
Match the pairs
Why: The fourth row is the one to watch: a claim of sameness is an equality, so it goes in the null and the alternative becomes not-equal. Rows three and four produce the same alternative from opposite-sounding wordings.
Worked example
A wording that puts the researcher's interest in the null.
\[ \text{a trial tests whether a medicine reduces cholesterol by } 25 \text{ percent} \]
Find the parameter
Why: A percentage reduction.
Read the claim
Why: Reduces by 25 percent.
Place the equality
Why: It must be the null.
\[ H - 0:\text{ reduction } = 0.25 \]
Write the alternative
Why: Its contradiction.
\[ H - a:\text{ reduction } \ne 0.25 \]
Figure (svg): The solution to Worked example a claim of no change shown as a ladder of expressions, one row per legal move
\[ H_0: p = 0.25, \qquad H_a: p \ne 0.25 \]
Verify: confirm this does not contradict the rule about researcher's claims
Why: Usually the researcher's claim goes in the alternative, but here the claim itself is an equality, and equalities can only live in the null. The test then asks whether the data cast doubt on the 25 percent figure — which is a legitimate question, though it is worth noticing that such a test can never confirm the figure, only fail to refute it. That limitation is the subject of the last idea.
OpenStax Introductory Statistics 2e, §9.1 Null and Alternative Hypotheses §9.1, p. 462
Trap
\[ \text{“less than five years”} \;\to\; H_0: \mu < 5 \]
Symbolise the sentence and call it the null
Why: It is the first hypothesis on the page.
\[ \text{but } H_0 \text{ cannot contain a strict inequality} \]
The sentence describes the claim being tested, which is the alternative, and the null is its complement.
\[ H_a: \mu < 5 \;\Rightarrow\; H_0: \mu \ge 5 \]
Symbolise the sentence as the ALTERNATIVE, then negate it
Why: The wording describes what is being tested for.
Working in this order removes the need to negate a sentence in English, which is where the errors come from — the complement of takes less than five years is not takes more than five years but takes at least five years, and that boundary is easy to lose. Negating a symbol is mechanical; negating a sentence is not.
Faded example
A test of whether it takes fewer than 45 minutes to teach a lesson plan.
Fill in the blanks
H_a: \mu < 45, \qquad H_0: \mu >= 45
Why: Fewer than gives a less-than alternative, and the null takes the matching at-least. Writing H-a first and negating the symbol avoids having to negate the English sentence.
Two truths and a lie
All three concern translation.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false as an absolute. It holds when the claim asserts a difference, but a claim of no change is an equality and equalities can only be nulls. The rule about equal signs takes precedence, since it is what makes a p-value computable.
Prediction
Commit before reasoning.
Predict first
What is the exact complement of the claim that a mean is less than five?
Correct: The mean is at least five.
Why: The value five itself is not less than five, so it belongs in the complement — which makes the complement at-least rather than more-than. Losing the boundary value is the commonest translation error, and it produces a pair that fails the partition check.
Section
Section 4
Concept
The hypotheses are written about a parameter, and which parameter it is determines the distribution used, the point estimate and the standard error. A claim about an average concerns a mean; a claim about a percentage or a fraction concerns a proportion.
the parameter — The population quantity the hypotheses are about, written mu for a mean and p for a proportion. Sample statistics never appear in a hypothesis, since the hypotheses are claims about the population.
\[ H_0: \mu = \mu_0 \quad\text{or}\quad H_0: p = p_0 \]
The distinction is the same one section 8.3 drew for confidence intervals, and it works the same way here. The book's Examples 9.1 and 9.4 are about proportions because they concern percentages of people; 9.2 and 9.3 are about means because they concern averages. Section 9.3 will attach a distribution to each case.
Figure (svg): A three-column table listing four worded claims with their null and alternative hypotheses in symbols
OpenStax Introductory Statistics 2e, §9.1 Null and Alternative Hypotheses §9.1, pp. 462-463 — the four examples, two of each kind
Picture it
The book's examples with their hypotheses in symbols.
Figure (svg): A three-column table listing four worded claims with their null and alternative hypotheses in symbols
Notice that all four hypotheses are about population parameters and none mentions a sample. That is a rule with no exceptions: a hypothesis such as x-bar equals 16 would be a statement about data already collected, which is either true or false by inspection and needs no test.
Worked example
Two of the book's examples, contrasted.
\[ \text{(a) mean GPA differs from } 2.0; \quad \text{(b) more than } 6.6\text{ percent take AP exams} \]
Read (a)
Why: An average GPA.
Symbolise (a)
Why: With mu.
\[ \mu \ne 2.0 \]
Read (b)
Why: A percentage of students.
Symbolise (b)
Why: With p.
\[ p > 0.066 \]
Figure (svg): The solution to Worked example spotting the parameter shown as a ladder of expressions, one row per legal move
\[ H_a: \mu \ne 2.0 \qquad\text{against}\qquad H_a: p > 0.066 \]
Verify: confirm the tell is the same as in section 8.3
Why: A proportion problem counts how many are in a category and reports a fraction or percent; a mean problem measures a quantity for each subject and averages it. The book's earlier test applies unchanged: if there is no mention of a mean or average and the data are yes-or-no, the parameter is a proportion. Getting this wrong changes the distribution, the standard error and the conclusion.
OpenStax Introductory Statistics 2e, §9.1 Null and Alternative Hypotheses §9.1, pp. 462-463
Sorting
Ask what kind of quantity is claimed.
Sort into buckets
Sort each claim.
Item (c) is the one worth a second look: it names no average explicitly, but taking less than five years to graduate is a claim about how long students take on average, which is a mean. The word average is often implied rather than written.
Worked example
Distinguishing a parameter from a statistic.
\[ H_0: \bar{x} = 16.43 \]
Identify the symbol
Why: x-bar is a sample mean.
Ask whether it is unknown
Why: It is computed from data.
Say what a test would mean
Why: Testing a known value.
Correct it
Why: Use the population mean.
\[ \mu = 16.43 \]
Figure (svg): The solution to Worked example what a hypothesis may not contain shown as a ladder of expressions, one row per legal move
\[ H_0: \mu = 16.43 \]
Verify: confirm why the distinction matters operationally
Why: The test computes how likely the observed x-bar would be if the hypothesised mu were true — so x-bar is the evidence and mu is what is on trial. Putting the sample mean in the hypothesis collapses the two roles and leaves nothing to test. Every hypothesis in this chapter and the next uses a population symbol: mu, p, sigma, or a difference between two of them.
OpenStax Introductory Statistics 2e, §9.1 Null and Alternative Hypotheses §9.1, pp. 462-463
Trap
\[ H_0: \bar{x} = 65, \qquad H_a: \bar{x} > 65 \]
Use the quantity the data will supply
Why: The sample mean is what gets compared against 65.
\[ \text{but } \bar{x} \text{ is known once the data are in} \]
A statement about the sample is settled by looking at it, so there is nothing for a test to decide.
\[ H_0: \mu = 65, \qquad H_a: \mu > 65 \]
Write the hypotheses about the POPULATION parameter
Why: The sample mean is the evidence, not the claim.
The division of roles is worth stating explicitly: mu is unknown and on trial, x-bar is observed and is the evidence. Every quantity in a hypothesis test falls on one side or the other of that line, and mixing them is the error that makes a test meaningless rather than merely wrong.
Two truths and a lie
All three concern the parameter.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The sample mean is the evidence the test weighs, not the claim it weighs. Once the data are collected x-bar is known exactly, so a hypothesis about it could be settled by inspection and there would be nothing to test.
Fill the middle
A claim about the percentage of voters.
Fill in the blanks
H_a: p > 0.30
Why: The parameter p, since the claim is about a fraction of a group. Using mu here would signal a mean and would send the rest of the test to the wrong distribution.
Prediction
Commit before reasoning.
Predict first
Beyond the symbol used, what does the choice of parameter determine?
Correct: The distribution, the point estimate and the standard error.
Why: Section 9.3 attaches a distribution to each case: a normal for a mean with sigma known, a t for a mean with sigma unknown, and a normal for a proportion with its own standard error. The tail direction comes from the alternative rather than from the parameter, and the confidence level is chosen by the analyst.
Section
Section 5
Concept
After determining which hypothesis the sample supports, there are two options for a decision: reject the null if the sample information favours the alternative, or do not reject it — also written decline to reject — if the sample information is insufficient.
decline to reject — The correct phrase for the second decision. It records that the evidence was insufficient, which is not the same as evidence that the null is true.
\[ \text{reject } H_0 \quad\text{or}\quad \text{do not reject } H_0 \]
The asymmetry is the most consequential idea in the chapter. A test assumes the null and asks whether the data would be surprising under it, so surprising data casts doubt on the null while unsurprising data merely fails to. A study that finds no effect has not shown there is none — it may simply have been too small to detect one, which is exactly what section 9.2's Type II error describes.
Figure (svg): Two boxes giving the two permitted decisions, above a red box giving the phrase that must never be used
OpenStax Introductory Statistics 2e, §9.1 Null and Alternative Hypotheses §9.1, p. 462 — the two decisions
Picture it
And the phrase that never appears.
Figure (svg): Two boxes giving the two permitted decisions, above a red box giving the phrase that must never be used
The wording matters in practice as well as in exams. Reporting that a study proved no difference overstates what any test can deliver, and it is a frequent misreading of published results — particularly of small studies, where failing to detect an effect is the expected outcome even when one exists.
Worked example
Two ways of reporting the same result.
\[ \text{a test does not reject } H_0: \mu = 65 \]
The wrong wording
Why: Accepting the null.
\[ \text{the mean } IS 65 \]
Why it fails
Why: No evidence was found either way.
The right wording
Why: Insufficient evidence.
Note what remains open
Why: The mean may still differ.
Figure (svg): The solution to Worked example stating a conclusion correctly shown as a ladder of expressions, one row per legal move
\[ \text{do not reject } H_0 \;\ne\; H_0 \text{ is true} \]
Verify: confirm the asymmetry has a concrete cause
Why: A test computes how likely the data are under the null, so it can only tell you when they are unlikely. Data that are perfectly ordinary under the null are also perfectly ordinary under many nearby alternatives, so they distinguish nothing. That is why the evidence can point away from a claim but never toward it, and it is a property of the logic rather than a limitation of any particular test.
OpenStax Introductory Statistics 2e, §9.1 Null and Alternative Hypotheses §9.1, p. 462
Sorting
Apply the two-decision rule.
Sort into buckets
Sort each statement.
Items (c) and (e) say the same thing in different registers — one in plain language and one in the formal phrase — and both are correct. The two rejected wordings differ only in how strongly they overstate, and both are common in reports of published research.
Worked example
The practical consequence of the asymmetry.
\[ \text{a trial of } 8 \text{ patients finds no significant effect} \]
State the decision
Why: Not rejected.
\[ \text{do not reject } H - 0 \]
Ask what could cause it
Why: No effect, or too little data.
Note which is excluded
Why: Neither.
State the conclusion
Why: Insufficient evidence.
Figure (svg): The solution to Worked example why a small study proves nothing shown as a ladder of expressions, one row per legal move
\[ \text{no evidence of an effect} \;\ne\; \text{evidence of no effect} \]
Verify: confirm the distinction has a name
Why: The second explanation is section 9.2's Type II error: failing to reject a null that is actually false. Its probability depends on the sample size, which is why a study that finds nothing is uninformative unless it was large enough to have found something. Reporting a non-significant result without saying how large an effect the study could have detected leaves the reader unable to tell the two cases apart.
OpenStax Introductory Statistics 2e, §9.1 Null and Alternative Hypotheses §9.1, pp. 462-464
Trap
\[ p\text{-value} = 0.42 \;\Rightarrow\; \text{the mean IS } 65 \]
Read a failure to reject as confirmation
Why: The data were consistent with the null.
\[ \text{they were consistent with many other values too} \]
A mean of 64.8 or 65.3 would have produced equally unsurprising data, so nothing singles out 65.
\[ \text{there is insufficient evidence that the mean differs from } 65 \]
Report the absence of evidence, not the presence of sameness
Why: The book: the data have failed to cast serious doubt.
A useful discipline is to ask what OTHER null values the same data would also have failed to reject — usually a whole interval of them, which is precisely the confidence interval of chapter 8. Seeing that interval makes it obvious that no single value was confirmed, and it is why reporting an interval alongside a test is better practice than reporting the test alone.
Two truths and a lie
All three concern the decision.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and is the chapter's most consequential misreading. Unsurprising data are consistent with the null and with many nearby alternatives at once, so they single out nothing. Absence of evidence is not evidence of absence.
Prediction
Commit before reasoning.
Predict first
A study fails to reject the null. What extra information makes that result meaningful?
Correct: How large an effect it could have detected.
Why: A large study that could have detected a small effect and did not gives real evidence that any effect is small; a study of eight gives almost none. That capability is the power of the test, which section 9.2 defines as one minus the probability of a Type II error, and reporting it is what separates an informative null result from a vacuous one.
Explain it
A classmate writes: the p-value was 0.42, so we accept that the mean is 65.
Discussion prompt
In two sentences or fewer, fix it.
Hint: Ask what other values the data would also be consistent with.
Answer:
The same data would have been equally unsurprising if the mean were 64.8 or 65.4, so nothing about it picks out 65 in particular.
The conclusion should be that there is insufficient evidence to conclude the mean differs from 65 — a failure to reject, not an acceptance.
Comparison
Fill the blanks. Every property follows from the pair being contradictory.
Comparison matrix
| H-nought | H-a | |
|---|---|---|
| What it asserts | no difference; the status quo | a difference, in a stated direction |
| Symbols allowed | equals, at least, at most | not equal, greater than, less than |
| Role in the test | assumed true while testing | concluded only if the null is rejected |
| Possible verdicts | rejected, or not rejected | supported, or not supported |
The last row is where the wording has to be exact. A null is never accepted and an alternative is never proved — the verdicts available are rejection and its absence, which is a narrower thing than truth and falsity.
Pattern
Five steps, and the third is where the direction is decided.
If the claim itself is an equality, it belongs in the null and the alternative becomes not-equal. The equal-sign rule takes precedence over the usual placement of the researcher's claim.
OpenStax Introductory Business Statistics 2e, §9.1 Null and Alternative Hypotheses §9.1 Null and Alternative Hypotheses
Check
The symbol rule.
Check your understanding
Which statement could be a null hypothesis?
Answer: A
Why: The null always contains a symbol with an equal in it, and at most is less-than-or-equal-to.
Check
Translating a claim.
Check your understanding
A researcher tests whether more than 40 percent pass on the first try. What is H-a?
Answer: A
Why: The claim being tested is that the proportion exceeds 0.40, and the claim goes in the alternative with no equality.
Check
The decision.
Check your understanding
A test does not reject the null. What may be concluded?
Answer: A
Why: Failing to reject records that the evidence was insufficient, not that the null is true.
Real world
A pharmaceutical company runs a trial of a new painkiller against the existing standard treatment. The trial finds no statistically significant difference, and the marketing department proposes advertising the new drug as proven equally effective as the leading brand.
Discussion prompt
Assess the proposed claim, and say what the company would have to do to support something like it.
Hint: What were the two hypotheses, and what does failing to reject one establish?
Answer:
The claim is not supported by the trial. The hypotheses would have been that there is no difference between the drugs against that there is one, and the trial failed to reject the first. That records an absence of evidence for a difference, not evidence that the drugs are equivalent.
\[ H_0: \mu_1 = \mu_2, \quad H_a: \mu_1 \ne \mu_2; \qquad \text{not rejected} \;\ne\; \text{equal} \]
The result is equally consistent with a real difference the trial was too small to find. If the trial enrolled few patients, or the outcome was highly variable, a clinically meaningful difference could easily have gone undetected — which is a Type II error, and its probability is exactly what a small trial makes large.
Supporting an equivalence claim requires a different design, not a different reading of this one. Equivalence and non-inferiority trials invert the logic: the null becomes that the drugs DIFFER by at least some pre-specified clinically meaningful margin, and rejecting that null supports equivalence within the margin. That is a real test with a real conclusion, and it typically needs a larger sample than a difference trial.
Two points make this more than a technicality. Regulators require the equivalence design precisely because the misreading is so tempting and so common, and a company advertising on the strength of a non-significant difference test would be making a claim its own data cannot bear. And the margin has to be chosen in advance and justified clinically — an equivalence claim is only as meaningful as the margin it is equivalent within, which is why the honest version of the marketing claim would have to name a number.
Commit first
Answer, then rate your confidence honestly.
Predict first
Why can a hypothesis test never confirm the null hypothesis?
Correct: Because unsurprising data are consistent with many values at once.
\[ \text{reject } H_0 \quad\text{or}\quad \text{do not reject } H_0; \quad \text{never accept} \]
Why: A test assumes the null and asks how likely the observed data would be under it. Surprising data cast doubt on that assumption, but ordinary data would have been just as ordinary under a range of nearby values — so they single nothing out. The asymmetry is a feature of the logic rather than a limitation of sample size, though a small sample makes failing to reject far more likely.
Explain it
They wrote H-nought as mu less than 5 and H-a as mu greater than or equal to 5.
Discussion prompt
In two sentences or fewer, locate the error.
Hint: Ask which hypothesis carries the equal sign.
Answer:
The equality has ended up in the alternative, and the rule is the other way round: the null always contains a symbol with an equal in it and the alternative never does.
Swapping them gives H-nought as mu at least 5 and H-a as mu less than 5, which also puts the claim being tested in the alternative where it belongs.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, the claim being tested for is the alternative. For the second, the null carries the equality and the pair must cover every value exactly once. For the third, write H-a first and negate the symbol rather than the sentence. For the fourth, say insufficient evidence rather than accept. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top, draw two columns headed H-nought and H-a, and fill each with what it asserts, which symbols it may contain, its role in the test, and what verdicts are available for it. Beneath, write the book's three-row symbol table pairing each null symbol with its alternative. In the middle of the page, work all four of the book's examples: the voters, the GPA, the graduation times and the AP exams — writing for each the parameter, the comparison value, the alternative first and then the null, and drawing a small number line showing that the pair partitions the possible values with no gap and no overlap. Beside them, write one wrong version of each and say which rule it breaks. At the bottom, write the two permitted decisions in a box and the two forbidden phrases in a box beneath it, and then write out one full sentence reporting a failure to reject — correctly. Finish with one sentence on why a small study that finds nothing has not shown there is nothing.
Check the four examples by confirming each pair covers the boundary value exactly once: 0.30, 2.0, 5 and 0.066 must each appear in one hypothesis and not the other. Check the bottom by reading your reporting sentence aloud and confirming it never says the null is true.
Recap
Five things, and the last is the one that gets misreported in public.
| If you see | Then |
|---|---|
| A claim of more than, less than, different from | That claim is the alternative |
| A claim of sameness or no change | That claim is the null, since it is an equality |
| An equality in a hypothesis | It can only be the null |
| A claim about a percentage | The parameter is p |
| A claim about an average | The parameter is mu |
| x-bar or p-prime in a hypothesis | An error: hypotheses are about parameters |
| A non-significant result | Insufficient evidence, not proof of no effect |
Section 9.2 takes up what can go wrong. Because a test decides on incomplete evidence, two kinds of error are possible — rejecting a true null and failing to reject a false one — and they have probabilities named alpha and beta. Which of the two matters more depends entirely on the consequences, which is a judgement the statistics cannot make.
OpenStax Introductory Statistics 2e, §9.1 Null and Alternative Hypotheses §9.1, pp. 462-464 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.