The chapter assembled. Sections 9.1 through 9.4 supplied the hypotheses, the two error types, the distributions and the decision rule, and this section adds the one remaining piece and then works complete tests. The remaining piece is the tail: when the p-value is drawn it occupies the left tail, the right tail, or is split evenly between both, and the alternative hypothesis is what decides which. A less-than alternative gives a left-tailed test, a greater-than gives a right-tailed one, and a not-equal gives a two-tailed test whose p-value is twice the area beyond the observed statistic. The significance level is chosen before the data are collected, and a common standard when none is given is five percent. The section then works four full tests, one for each of the three distributions, ending each with a conclusion in the words of the problem and an identification of the two error types in context.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 9 — Hypothesis Testing with One Sample
Additional Information and Full Hypothesis Test Examples
Objectives
Five outcomes, and the first is what the previous four sections left open.
OpenStax Introductory Statistics 2e, §9.5 Additional Information and Full Hypothesis Test Examples §9.5, pp. 469-482 — the section these objectives are drawn from
Warm-up
The previous four sections gave the hypotheses, the errors, the distribution and the decision rule.
Discussion prompt
A test has H-nought that the mean is 16.43 and H-a that it is less. The observed sample mean is 16. Which values count as at least as extreme as what was seen?
Hint: Extreme means far from the null in the direction the alternative points.
Answer:
Values of 16 and below. The alternative claims the mean is LESS than 16.43, so a sample mean of 15.8 would be even stronger evidence for it than 16 was, while 16.9 would be evidence in the opposite direction entirely.
So the p-value is the area to the left of 16 — a left tail. Had the alternative pointed the other way, the same observation of 16 would have counted for almost nothing and the tail would have been on the right.
That is the piece section 9.4 left open. The p-value's definition says as extreme or more extreme, and the alternative hypothesis is what fixes which direction extreme runs in.
Concept
When you calculate the p-value and draw the picture, the p-value is the area in the left tail, the right tail, or split evenly between the two tails. For this reason we call the hypothesis test left, right, or two tailed. The alternative hypothesis tells you which, and it is the key to conducting the appropriate test.
left, right and two-tailed tests — A less-than alternative puts the p-value in the left tail; a greater-than alternative puts it in the right; a not-equal alternative splits it between both, so the p-value is twice the one-sided area.
\[ H_a: < \;\to\; \text{left}; \quad H_a: > \;\to\; \text{right}; \quad H_a: \ne \;\to\; \text{two} \]
The book states several other conventions in the same place. A phrase such as the level of significance is 1 percent names the preset alpha; the statistician selects alpha before collecting the data; and if no level of significance is given, a common standard to use is 0.05. It also repeats that H-a never has a symbol containing an equal sign, which is what guarantees the alternative always points somewhere.
Figure (svg): Three bell curves side by side with the left tail shaded on the first, the right tail on the second, and both tails on the third
OpenStax Introductory Statistics 2e, §9.5 Additional Information and Full Hypothesis Test Examples §9.5, p. 469
Section
Section 1
Concept
The direction of the alternative fixes which values count as extreme. A less-than alternative makes small values extreme, a greater-than alternative makes large ones extreme, and a not-equal alternative makes departures in either direction extreme.
the tail of a test — The region whose area is the p-value. It lies on the side the alternative points toward, or on both sides when the alternative is two-sided.
\[ \text{extreme} = \text{far from } H_0 \text{ in the direction } H_a \text{ points} \]
The book's three short examples make the pattern explicit: mu less than 5 is left-tailed, p greater than 0.2 is right-tailed, and p not equal to 50 is two-tailed. Because the alternative never contains an equality, it always points somewhere or points both ways — which is why the equal-sign rule of section 9.1 is what makes this step possible at all.
Figure (svg): Three bell curves side by side with the left tail shaded on the first, the right tail on the second, and both tails on the third
OpenStax Introductory Statistics 2e, §9.5 Additional Information and Full Hypothesis Test Examples §9.5, pp. 469-471 — Examples 9.11, 9.12 and 9.13
Picture it
The book's three illustrations of a p-value's picture.
Figure (svg): Three bell curves side by side with the left tail shaded on the first, the right tail on the second, and both tails on the third
Drawing the picture before computing anything is worth the few seconds. It fixes which tail is wanted, it gives a rough expectation of the answer's size, and it catches the commonest error in the section — computing the area on the wrong side.
Worked example
The book's Examples 9.11 through 9.13.
\[ \text{(a) } \mu < 5; \; \text{(b) } p > 0.2; \; \text{(c) } p \ne 50 \]
Read (a)
Why: Less than.
Read (b)
Why: Greater than.
Read (c)
Why: Not equal.
Note what decided each
Why: The alternative alone.
Figure (svg): The solution to Worked example three alternatives, three pictures shown as a ladder of expressions, one row per legal move
\[ < \;\to\; \text{left}; \quad > \;\to\; \text{right}; \quad \ne \;\to\; \text{two} \]
Verify: confirm the data play no part in choosing the tail
Why: The tail is fixed before any sample is examined, because it comes from the alternative and the alternative was written from the claim. Choosing the tail after seeing which way the data went would be fitting the test to the result — the same error as adjusting alpha afterwards, and with the same effect of destroying the stated error rate.
OpenStax Introductory Statistics 2e, §9.5 Additional Information and Full Hypothesis Test Examples §9.5, pp. 470-471
Sorting
Read the symbol in the alternative.
Sort into buckets
Sort each test by its tail.
Only the not-equal alternative gives a two-tailed test, and it is the one whose p-value has to be doubled. Everything else in the sort is one-tailed, with the side read straight off the inequality.
Worked example
Why the tail changes everything.
\[ \bar{x} = 16 \text{ against } \mu_0 = 16.43 \]
Under H-a: mu < 16.43
Why: 16 is below the null.
Under H-a: mu > 16.43
Why: 16 is below the null.
Compute the second
Why: The right tail beyond 16.
\[ 0.9813 \]
Compare
Why: Opposite conclusions.
Figure (svg): The solution to Worked example the same observation, two directions shown as a ladder of expressions, one row per legal move
\[ p_{\text{left}} = 0.0187, \qquad p_{\text{right}} = 0.9813 \]
Verify: confirm the two p-values sum to one, and why
Why: For a continuous distribution the two one-sided tails at the same point are complementary, so they must sum to 1 — and 0.0187 plus 0.9813 does. That gives a quick check: if a one-tailed p-value comes out above 0.5 when the data point in the direction the alternative claims, the wrong tail has been used.
OpenStax Introductory Statistics 2e, §9.5 Additional Information and Full Hypothesis Test Examples §9.5, pp. 470-472
Trap
\[ \bar{x} = 16 \text{ is below } 16.43, \text{ so use the left tail} \]
Pick the tail the data fell into
Why: The evidence points that way.
\[ \text{but the tail comes from } H_a, \text{ written beforehand} \]
If the alternative had claimed the mean was higher, this observation would be evidence against the claim, not for it.
\[ H_a: \mu < 16.43 \;\Rightarrow\; \text{left-tailed, before any data} \]
Read the tail from the alternative, which was written from the claim
Why: The direction is part of the question, not of the answer.
The two happen to agree here, which is what makes the error easy to make and hard to notice. It shows itself when the data go the other way: a right-tailed test on data that fell left returns a p-value near 1, and the honest conclusion is that the evidence points against the claim rather than that a different test should be run.
Fill the middle
What fixes the tail.
Fill in the blanks
\textalternative ___ \text___
Why: The alternative. The book calls it the key to conducting the appropriate test, and since it never contains an equality it always points somewhere.
Two truths and a lie
All three concern tails.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and is a form of fitting the test to the result. The direction belongs to the question being asked; choosing it afterwards guarantees a small p-value whichever way the data fall, and destroys the Type I error rate the test claims.
Prediction
Commit before reasoning.
Predict first
A right-tailed test observes a sample mean well BELOW the null value. What is the p-value?
Correct: Large, above 0.5.
Why: The right tail beyond a value below the mean covers more than half the distribution, so the p-value exceeds 0.5 and the null is not rejected. The data are far from the null but in the direction the alternative did not claim, which is evidence against the alternative rather than for it.
Section
Section 2
Concept
When the population standard deviation is given, the sampling distribution of the mean is normal with standard error sigma over the square root of n, and the value assumed for the mean comes from the null hypothesis rather than from the data.
the value from H-nought — The book stresses that mu comes from H-nought and not the data. The sample supplies the observation being judged; the null supplies the assumption it is judged against.
\[ \bar{X} \sim N\left(\mu_0, \frac{\sigma}{\sqrt{n}}\right) \]
The distinction between what comes from the null and what comes from the data runs through every test in the chapter. The centre of the sampling distribution is the null's claimed value; the observation plotted on it is the sample's. Mixing them — centring the distribution on the sample mean — would produce a p-value of about a half every time, since the observation would sit at the centre by construction.
Figure (svg): A normal curve centred at sixteen point four three with the left tail below sixteen shaded
OpenStax Introductory Statistics 2e, §9.5 Additional Information and Full Hypothesis Test Examples §9.5, pp. 471-474 — Examples 9.14 and 9.15
Picture it
Example 9.14: did the goggles make Jeffrey faster?
Figure (svg): A normal curve centred at sixteen point four three with the left tail below sixteen shaded
The book's own interpretation of the p-value is worth copying: if the null is true, there is a 1.87 percent probability that Jeffrey's mean time is 16 seconds or less — and because 1.87 percent is small, that outcome is unlikely to have happened randomly. It is a rare event.
Worked example
Example 9.14, a left-tailed z test at 5 percent.
\[ \mu_0 = 16.43, \; \sigma = 0.8, \; n = 15, \; \bar{x} = 16, \; \alpha = 0.05 \]
Set up the hypotheses
Why: Faster means less time.
\[ H 0: \mu = 16.43, H a: \mu < 16.43 \]
Find the distribution
Why: Sigma is known.
\[ N(16.43, 0.2066) \]
Compute the p-value
Why: The area to the left of 16.
\[ 0.0187 \]
Decide
Why: Alpha exceeds it.
\[ \text{reject } H 0 \]
Figure (svg): The solution to Worked example Jeffrey's swim times shown as a ladder of expressions, one row per legal move
\[ p = P(\bar{x} < 16) = 0.0187 < 0.05 \]
Verify: confirm both error types in this context, as the book does
Why: The Type I error is concluding Jeffrey swims in less than 16.43 seconds on average when he actually does not — rejecting a true null. The Type II error is failing to find evidence that he swims faster when in fact he does. Naming both in the problem's own words is the last step of the book's worked solutions, and it forces the conclusion to be read as a decision under uncertainty rather than as a fact.
OpenStax Introductory Statistics 2e, §9.5 Additional Information and Full Hypothesis Test Examples §9.5, pp. 471-472
Faded example
H0: mu = 16.43, sigma = 0.8, n = 15, sample mean 16.
Fill in the blanks
\text0.2066 = \frac0.0187___} \approx ___, \quad p = P(\bar___ < 16) = ___
Why: The standard error is about 0.2066, so 16 sits about 2.08 standard errors below 16.43, leaving 0.0187 in the left tail.
Worked example
Example 9.15, a right-tailed z test at 2.5 percent.
\[ \mu_0 = 275, \; \sigma = 55, \; n = 30, \; \alpha = 0.025 \]
Compute the sample mean
Why: From the frequency table.
\[ 286.17 \]
Set up the hypotheses
Why: More than 275.
\[ H 0: \mu = 275, H a: \mu > 275 \]
Compute the p-value
Why: The right tail.
\[ 0.1331 \]
Decide
Why: Alpha is below it.
Figure (svg): The solution to Worked example the bench press shown as a ladder of expressions, one row per legal move
\[ p = P(\bar{x} > 286.17) = 0.1331 > 0.025 \]
Verify: resolve the book's two printed p-values
Why: The book prints 0.1331 in its calculator output and 0.1323 in its comparison line. The exact sample mean of these thirty values is 286.1667, which gives 0.1331; rounding the mean to 286.2 before computing gives 0.1323. So 0.1331 is the accurate figure and 0.1323 is an intermediate-rounding artefact. The decision is the same either way, but carrying full precision until the final step is the better habit.
OpenStax Introductory Statistics 2e, §9.5 Additional Information and Full Hypothesis Test Examples §9.5, pp. 473-474
Trap
\[ \bar{X} \sim N(16, 0.2066), \text{ using the observed mean} \]
Build the distribution around the data
Why: The sample mean is the number computed from the sample.
\[ p\text{-value} = P(\bar{x} < 16) = 0.5 \]
The observation now sits exactly at the centre, so the p-value is a half whatever the data were.
\[ \bar{X} \sim N(16.43, 0.2066), \text{ centred on } H_0 \]
Centre the distribution on the null's value
Why: Mu comes from H-nought, not the data.
The book flags this explicitly — mu equals 16.43 comes from H-nought and not the data — because the error produces a p-value that looks ordinary while being uninformative. A p-value of almost exactly 0.5 on data that visibly differ from the null is the symptom.
Two truths and a lie
All three concern the z test.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and would make every p-value about a half, since the observation would sit at the centre by construction. The null supplies the assumption; the sample supplies the evidence judged against it.
Estimation
A sample mean sits about two standard errors below a null value, in a left-tailed test.
Predict first
Roughly what p-value should be expected?
Correct: About 0.02.
Why: Chapter 6's Empirical Rule puts about 2.5 percent of a normal distribution beyond two standard deviations on one side, so a p-value near 0.02 is right — and Example 9.14's 0.0187 at 2.08 standard errors matches. Making this estimate before computing catches a wrong tail or a wrong standard error immediately.
Explain it
A classmate centred the sampling distribution on their sample mean and got a p-value of 0.5.
Discussion prompt
In two sentences or fewer, explain what went wrong.
Hint: Ask where their observation sits on their own curve.
Answer:
Their observation is sitting exactly at the centre of their curve, so half the area lies beyond it no matter what the data were — which is why the answer came out at 0.5.
The curve has to be centred on the null's value, since the whole question is how surprising the data would be IF the null were true.
Section
Section 3
Concept
When no population standard deviation is given and only sample data are available, the distribution for the test is a Student t with n minus one degrees of freedom, and the sample standard deviation is computed from the data alongside the sample mean.
recognising a t test — The book's own tell: reading the problem carefully shows there is no population standard deviation given, only n sample data values — so the distribution for the test is a Student t.
\[ t = \frac{\bar{x} - \mu_0}{s/\sqrt{n}} \sim t_{n-1} \]
The recognition step deserves the emphasis the book gives it, because nothing else in the problem announces it. A z test and a t test look identical on the page until you notice whether a sigma was supplied, and section 9.3 established that using a z where a t belongs always understates the p-value — making the evidence look stronger than it is.
Figure (svg): A t distribution with nine degrees of freedom, its right tail beyond a test statistic of about two shaded
OpenStax Introductory Statistics 2e, §9.5 Additional Information and Full Hypothesis Test Examples §9.5, pp. 475-476 — Example 9.16
Picture it
Example 9.16: is the mean first-test score above 65?
Figure (svg): A t distribution with nine degrees of freedom, its right tail beyond a test statistic of about two shaded
The p-value of 0.0396 rejects at 5 percent but would not at 2.5 percent, which is exactly the marginal case section 9.4's note about judgement had in mind. Reporting the p-value alongside the decision lets a reader see how close the call was.
Worked example
Example 9.16, a right-tailed t test at 5 percent.
\[ \text{scores } 65, 65, 70, 67, 66, 63, 63, 68, 72, 71; \; \mu_0 = 65 \]
Notice no sigma is given
Why: Only ten data values.
Compute the statistics
Why: Mean and sample sd.
\[ 67\text{ and } 3.1972 \]
Set the degrees of freedom
Why: Ten minus one.
\[ 9 \]
Compute the p-value
Why: The right tail.
\[ 0.0396 \]
Figure (svg): The solution to Worked example the statistics test scores shown as a ladder of expressions, one row per legal move
\[ p = P(\bar{x} > 67) = 0.0396 < 0.05 \]
Verify: confirm how much a z would have understated it
Why: Using a normal instead of a t with 9 degrees of freedom would give about 0.0239 rather than 0.0396 — a p-value nearly 40 percent smaller, on the same data. The decision at 5 percent would be unchanged, but at 2.5 percent the two disagree: the z rejects and the correct t does not. That is the concrete cost of the wrong distribution.
OpenStax Introductory Statistics 2e, §9.5 Additional Information and Full Hypothesis Test Examples §9.5, pp. 475-476
Faded example
Ten scores, mean 67, sample standard deviation 3.1972, null mean 65.
Fill in the blanks
t = \frac1.9789} \approx ___, \quad \text___ = ___
Why: The standard error is about 1.011, so the t statistic is about 1.978 on nine degrees of freedom, giving a right-tail p-value of 0.0396.
Worked example
Example 9.16's interpretation, in the book's words.
\[ p\text{-value} = 0.0396 \]
State the condition
Why: If the null is true.
\[ \text{assuming } \mu = 65 \]
State the probability
Why: The p-value.
\[ 0.0396 \]
State the event
Why: As extreme or more.
\[ a\text{ sample mean of } 67\text{ or more} \]
Combine
Why: The book's sentence.
\[ 3.96 \% \]
Figure (svg): The solution to Worked example interpreting the p-value shown as a ladder of expressions, one row per legal move
\[ P(\bar{x} \ge 67 \mid \mu = 65) = 0.0396 \]
Verify: confirm the sentence contains both required clauses
Why: It names the condition — if the null is true — and it covers 67 or more rather than exactly 67. Section 9.4 established that dropping either clause turns a correct statement into one of the standard misreadings, and writing the interpretation in this fixed shape makes both automatic.
OpenStax Introductory Statistics 2e, §9.5 Additional Information and Full Hypothesis Test Examples §9.5, p. 475
Error analysis
Ten scores with mean 67 and sample sd 3.1972, testing a null of 65.
Annotate
On: \( \begin{aligned} &(1)\; \text{normalcdf using } z = 1.978 \;\to\; 0.0239 \\ &(2)\; \text{tcdf with df} = 10 \;\to\; 0.0381 \\ &(3)\; \text{left tail of the } t \;\to\; 0.9604 \\ &(4)\; \text{right tail of } t_9 \;\to\; 0.0396 \end{aligned} \)
Errors (1) and (2) both understate the p-value and would go unnoticed; error (3) is caught instantly by the check that a p-value above 0.5 contradicts data pointing the way the alternative claims.
Two truths and a lie
All three concern the t test.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and has the direction backwards. The t has heavier tails, so it leaves MORE area beyond any statistic — the z gives the smaller p-value, which is why the error always makes evidence look stronger.
Discrimination
Read what the problem supplies.
Sort into buckets
Sort each test.
Estimation
The test rejects at 5 percent.
Predict first
What would happen at a 2.5 percent significance level?
Correct: It would not reject.
Why: The p-value is a property of the data and does not change with alpha, but the decision does — 0.0396 is below 0.05 and above 0.025. That the same evidence gives opposite verdicts at two conventional thresholds is why the level must be stated and the p-value reported.
Section
Section 4
Concept
When the alternative says the parameter is different from a value rather than above or below it, departures in either direction count as extreme. The p-value is then the area in both tails, which by symmetry is twice the area beyond the observed statistic.
a two-tailed p-value — Twice the one-sided area. For Example 9.17 the observed proportion of 0.53 sits 0.03 above the null's 0.50, so the mirror boundary is 0.47 and the p-value covers everything below 0.47 or above 0.53.
\[ p = P(p' < 0.47 \text{ or } p' > 0.53) = 2P(p' > 0.53) \]
The book spells out the mirroring: since the curve is symmetrical and the test is two-tailed, the boundary for the left tail is 0.50 minus 0.03, which is 0.47, where 0.03 is the difference between 0.53 and 0.50. That construction is worth following once, because it explains why doubling the one-sided area is correct rather than merely conventional.
Figure (svg): A bell curve with both tails shaded, cut at sample proportions of nought point four seven and nought point five three
OpenStax Introductory Statistics 2e, §9.5 Additional Information and Full Hypothesis Test Examples §9.5, pp. 476-477 — Example 9.17
Picture it
Example 9.17: are half of first-time brides younger than their grooms?
Figure (svg): A bell curve with both tails shaded, cut at sample proportions of nought point four seven and nought point five three
The p-value of 0.5485 is enormous, and the picture says why: the observed 0.53 is only 0.6 standard errors from 0.50, so most of the distribution is at least that far out. The data are entirely ordinary under the null.
Worked example
Example 9.17, a two-tailed proportion test at 1 percent.
\[ p_0 = 0.50, \; n = 100, \; x = 53, \; \alpha = 0.01 \]
Set up the hypotheses
Why: Same or different.
\[ H 0: p = 0.50, H a: p \ne 0.50 \]
Form the point estimate
Why: 53 over 100.
\[ p' = 0.53 \]
Find the two boundaries
Why: Mirrored about 0.50.
\[ 0.47\text{ and } 0.53 \]
Compute the p-value
Why: Both tails.
\[ 0.5485 \]
Figure (svg): The solution to Worked example the first-time brides shown as a ladder of expressions, one row per legal move
\[ p = 0.5485 \;>\; 0.01 \]
Verify: confirm the doubling by computing one tail
Why: The standard error is the square root of 0.25 over 100, which is 0.05, so 0.53 sits 0.6 standard errors above 0.50. The area beyond that is about 0.2743, and twice it is 0.5485 — matching the book exactly. Checking a two-tailed p-value by halving it and confirming the one-sided area is a quick guard against forgetting to double.
OpenStax Introductory Statistics 2e, §9.5 Additional Information and Full Hypothesis Test Examples §9.5, pp. 476-477
Faded example
A two-tailed test gives a one-sided area of 0.2743.
Fill in the blanks
p\text2 = 0.5485 \times 0.2743 = ___
Why: Two-tailed p-values double the one-sided area, giving 0.5485 — far above any conventional alpha, so the null is not rejected.
Worked example
Confirming section 9.3's conditions before proceeding.
\[ n = 100, \; p_0 = 0.50 \]
Note the parameter
Why: A proportion.
Choose the distribution
Why: The third row.
Check np
Why: One hundred times a half.
\[ 50 \]
Check nq
Why: The same.
\[ 50 \]
Figure (svg): The solution to Worked example why this test uses a normal shown as a ladder of expressions, one row per legal move
\[ np = nq = 50 > 5 \]
Verify: confirm the value of p used in the check
Why: The check uses 0.50 from the null hypothesis rather than the observed 0.53, because the whole test is computed under the assumption that the null is true. Here the two give nearly identical answers, but for a proportion near zero or one they can differ enough to change whether the condition passes.
OpenStax Introductory Statistics 2e, §9.5 Additional Information and Full Hypothesis Test Examples §9.5, pp. 476-477
Trap
\[ p\text{-value} = P(p' > 0.53) = 0.2743 \]
Compute the tail beyond the observation and stop
Why: That is where the data fell.
\[ \text{but } H_a \text{ says DIFFERENT, so both directions count} \]
A proportion of 0.47 would have been equally extreme evidence against a null of 0.50, and it has been left out.
\[ p\text{-value} = 2(0.2743) = 0.5485 \]
Double the one-sided area for a two-tailed test
Why: The two tails are mirror images about the null.
Halving a p-value always makes evidence look stronger, so this error runs in the dangerous direction. It matters most near a threshold: a one-sided area of 0.03 would reject at 5 percent while its correct two-tailed value of 0.06 would not, from identical data.
Two truths and a lie
All three concern two-tailed tests.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. A two-tailed test counts both directions because the alternative claims a difference without specifying which way — using only the observed side halves the p-value and overstates the evidence.
Prediction
Commit before reasoning.
Predict first
The same data are tested one-tailed and two-tailed in the direction they fell. Which gives the smaller p-value?
Correct: The one-tailed test, by a factor of two.
Why: The two-tailed p-value doubles the one-sided area, so it is always the larger and always the harder to reject with. This is exactly why the direction must be fixed from the claim before the data are seen: choosing one-tailed afterwards, in the direction the data went, halves the p-value and doubles the true Type I error rate.
Estimation
A two-tailed test gives a p-value of 0.5485.
Predict first
What does that say about the data?
Correct: They are entirely ordinary under the null.
Why: More than half of all samples would be at least this far from 0.50 if the null were true, so the observation is unremarkable. The third option is section 9.4's conditioning error: the p-value says nothing directly about how likely the null is.
Section
Section 5
Concept
Every full example in the section follows the same order: set up the hypotheses, determine the distribution, calculate the p-value, draw the graph, compare with alpha, decide, and write the conclusion in context.
the standard sequence — Hypotheses, tail, distribution, p-value with a graph, comparison, decision, conclusion. The book's four full examples differ only in which distribution step three selects.
\[ H_0, H_a \to \text{tail} \to \text{distribution} \to p \to \text{decide} \to \text{conclude} \]
Following a fixed order matters more here than in most topics, because the errors this chapter invites are errors of sequence rather than of arithmetic. Choosing a tail after seeing the data, adjusting alpha after seeing the p-value, or centring a distribution on the sample mean are all failures to keep the assumption-setting steps ahead of the data-reading ones.
Figure (svg): The full sequence of a hypothesis test
OpenStax Introductory Statistics 2e, §9.5 Additional Information and Full Hypothesis Test Examples §9.5, pp. 469-477 — the structure common to Examples 9.14 through 9.17
Picture it
The order every worked example follows.
Figure (svg): The full sequence of a hypothesis test
Steps one and two happen before the data are examined and step five happens after — and alpha belongs with the first group even though it is used in the fifth. Keeping that boundary is what makes the stated error rate a real guarantee rather than a formality.
Worked example
What differs and what does not across the section's examples.
\[ \text{Examples } 9.14, 9.15, 9.16, 9.17 \]
9.14
Why: Mean, sigma known, left.
9.15
Why: Mean, sigma known, right.
9.16
Why: Mean, sigma unknown, right.
9.17
Why: Proportion, two-tailed.
Figure (svg): The solution to Worked example the four tests compared shown as a ladder of expressions, one row per legal move
\[ \text{same sequence; different step three and step two} \]
Verify: confirm two reject and two do not, and what that shows
Why: The four split evenly, which is deliberate on the book's part: a chapter whose examples all rejected would suggest that a properly conducted test usually finds something. Two of these four end in insufficient evidence, and writing those conclusions correctly — without accepting the null — is as much of the skill as computing the p-values.
OpenStax Introductory Statistics 2e, §9.5 Additional Information and Full Hypothesis Test Examples §9.5, pp. 471-477
Sorting
Which steps must be settled in advance?
Sort into buckets
Sort each step.
The tail belongs in the left column even though it is drawn when the p-value is computed, because it follows from the alternative — which was written before any data existed. Moving it to the right is the section's most consequential sequencing error.
Worked example
The book's closing step for Example 9.14.
\[ H_0: \mu = 16.43, \; H_a: \mu < 16.43 \]
Type I
Why: Reject a true null.
Type II
Why: Fail to reject a false one.
Attach probabilities
Why: Alpha and beta.
\[ 0.05\text{ and unknown} \]
Say which matters
Why: Depends on stakes.
Figure (svg): The solution to Worked example naming the errors in context shown as a ladder of expressions, one row per legal move
\[ \alpha = 0.05; \quad \beta \text{ depends on the true improvement} \]
Verify: confirm why beta cannot be stated as a number here
Why: Section 9.2 established that beta depends on how false the null actually is, and nothing in the problem says how much faster Jeffrey really swims. Alpha is 0.05 because it was chosen; beta would have to be computed for a specified improvement, which is what a power calculation does at the design stage.
OpenStax Introductory Statistics 2e, §9.5 Additional Information and Full Hypothesis Test Examples §9.5, p. 472
Trap
\[ \text{compute the statistic, then choose the tail and alpha to suit} \]
Let the data guide the setup
Why: It is more efficient to look first.
\[ \text{the stated error rate no longer holds} \]
Both choices were meant to be made in ignorance of the result, and fitting them to it guarantees significance.
\[ \text{hypotheses and alpha first; data second; decision third} \]
Keep the assumption-setting steps ahead of the data-reading ones
Why: That ordering is what the error rate depends on.
A test run in the wrong order can produce every number correctly and still mean nothing, because the guarantee it claims — that a true null is rejected only alpha of the time — was built on choices made in advance. This is the deepest reason the chapter insists on a fixed sequence.
Two truths and a lie
All three concern the sequence.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The tail comes from the alternative, which was written from the claim before any data were collected. Choosing it afterwards, in the direction the data fell, halves the p-value and doubles the real Type I error rate.
Faded example
The six steps of a full test.
Fill in the blanks
\textp-value alpha \text___ ___ \text___
Why: The p-value is computed on the chosen distribution and compared with the preset alpha, which produces the decision and then the written conclusion.
Explain it
A classmate asks why it matters whether alpha is chosen before or after the p-value, since the comparison is the same either way.
Discussion prompt
In two sentences or fewer, explain.
Hint: Ask what alpha is supposed to guarantee.
Answer:
Alpha is meant to guarantee that a true null is rejected only that fraction of the time, and that guarantee depends on the threshold being fixed without knowledge of the result.
Choosing it afterwards means picking whichever side of the p-value suits, so the effective rejection rate is no longer alpha and the test's central claim is void.
Comparison
Fill the blanks. Same six steps, different step two and step three.
Comparison matrix
| Example | Distribution | Tail and outcome |
|---|---|---|
| 9.14 Jeffrey's swim | normal, sigma known | left-tailed; reject |
| 9.15 bench press | normal, sigma known | right-tailed; do not reject |
| 9.16 test scores | Student t, 9 df | right-tailed; reject |
| 9.17 first-time brides | normal, a proportion | two-tailed; do not reject |
Two reject and two do not, which is worth noticing: writing an insufficient-evidence conclusion without accepting the null is half the skill, and a chapter of successful tests would never exercise it.
Pattern
Six steps, and the first three happen before the data are read.
Check the p-value against the picture before accepting it. A one-tailed p-value above 0.5 when the data point the way H-a claims means the wrong tail was used.
OpenStax Introductory Business Statistics 2e, §9.4 Full Hypothesis Test Examples §9.4 Full Hypothesis Test Examples
Check
The tail.
Check your understanding
A test has H-a: mu less than 5. Where does the p-value lie?
Answer: A
Why: A less-than alternative makes small values extreme, so the p-value is the area to the left of the observed statistic.
Check
A two-tailed p-value.
Check your understanding
A two-tailed test gives a one-sided area of 0.03. What is the p-value?
Answer: A
Why: Both tails count, and by symmetry the p-value is twice the one-sided area.
Check
Centring the distribution.
Check your understanding
In a test of H-nought that mu equals 16.43, where is the sampling distribution centred?
Answer: A
Why: The test asks how surprising the data would be if the null were true, so the distribution is built around the null's claimed value.
Real world
A supplement manufacturer tests whether its product raises average daily energy scores above the population norm of 60. The trial of 45 users gives a mean of 62.1 with a sample standard deviation of 8.4. The analyst planned a two-tailed test at 5 percent, obtained a p-value of 0.098, then switched to a one-tailed test — since the data went the expected way — and reported p = 0.049 as significant.
Discussion prompt
Assess what was done, and say what the honest report would be.
Hint: The arithmetic is right at both stages. The problem is the order.
Answer:
Both p-values are computed correctly. With s of 8.4 and n of 45 the standard error is about 1.252, so the t statistic is about 1.677 on 44 degrees of freedom, giving a one-sided area of about 0.049 and a two-sided p-value of about 0.098. Nothing in the arithmetic is wrong.
\[ t = \frac{62.1 - 60}{8.4/\sqrt{45}} \approx 1.677, \qquad p_{\text{two}} \approx 0.098, \; p_{\text{one}} \approx 0.049 \]
The switch is the problem, and it invalidates the stated error rate. The tail was chosen after the data were seen, in the direction they happened to fall. A test conducted that way rejects a true null about 10 percent of the time rather than the 5 percent it claims — because a departure in either direction would have been converted into a one-tailed test pointing that way.
The honest report is the pre-specified two-tailed result: p = 0.098, not significant at the 5 percent level. A one-tailed test would have been perfectly legitimate had it been chosen in advance on the grounds that only an increase was of interest — but that decision has to be made and recorded before the data arrive, precisely because it cannot be verified afterwards.
Two further points belong in a careful review. The effect of 2.1 points with a standard error of 1.25 is imprecisely estimated, and a confidence interval — roughly 62.1 give or take 2.5 — would show the plausible range including values close to no effect at all, which is more informative than either p-value. And the episode is a small instance of a general problem: analytic choices made after seeing data inflate false-positive rates across a literature, which is why pre-registration of the hypothesis, the tail and the significance level has become standard practice in fields that depend on these tests.
Commit first
Answer, then rate your confidence honestly.
Predict first
What determines whether a test is left-tailed, right-tailed or two-tailed?
Correct: The alternative hypothesis.
\[ H_a: < \;\to\; \text{left}; \quad > \;\to\; \text{right}; \quad \ne \;\to\; \text{two, and double the area} \]
Why: The book calls H-a the key to conducting the appropriate test: a less-than gives a left tail, a greater-than a right, and a not-equal both. Choosing from the data instead halves the p-value in whichever direction the sample fell, and doubles the real Type I error rate. Whether sigma is known selects the distribution, not the tail, and alpha is the threshold rather than the shape.
Explain it
For a two-tailed test they computed the area beyond their statistic and reported it as the p-value.
Discussion prompt
In two sentences or fewer, correct them.
Hint: Ask what a result equally far on the other side would have meant.
Answer:
Their alternative says the parameter is DIFFERENT from the null value, so a result equally far out on the other side would have been just as strong evidence and has to be counted too.
The p-value is twice the area they computed, and reporting only one tail halves it — which always makes the evidence look stronger than it is.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, less-than is left, greater-than is right, not-equal is both. For the second, a two-tailed p-value is twice the one-sided area. For the third, a t when only data are given, and always centre on the null's value. For the fourth, name the level, say sufficient or not sufficient evidence, and describe both errors in the problem's own words. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Twenty-five minutes. This one closes the chapter, so make it the whole procedure.
Draw it
Down the left, write the six steps of a complete test in order, and draw a line separating the steps that happen before the data are seen from those that happen after. Across the top, draw three small bell curves shaded left, right and both, and label each with the alternative-hypothesis symbol that produces it. In the middle of the page, work Example 9.14 completely: hypotheses, distribution centred at 16.43 with a standard error of 0.2066, the shaded left tail, the p-value of 0.0187, the comparison with 0.05, the decision, a full conclusion in words, and both error types described in Jeffrey's own terms. Beside it, work Example 9.16 the same way — noticing first that no sigma is given, using a t with nine degrees of freedom, and reaching 0.0396 — then write beneath it what a z would have given and which direction that error runs. At the bottom, work Example 9.17: compute the standard error of 0.05, mark 0.53 and its mirror at 0.47, shade both tails, compute one tail and double it to 0.5485, and write the conclusion. Finish with a two-line note on why the tail and alpha must be fixed before the data arrive.
Check the middle two by confirming the shaded regions are on the sides their alternatives point to — left for less than, right for greater than. Check the bottom by halving your p-value and confirming the result matches the single-tail area you computed, which is the guard against forgetting to double.
Recap
Five things, and the first is what the earlier sections left open.
| If you see | Then |
|---|---|
| H-a with a less-than | Left-tailed: the p-value is the left area |
| H-a with a greater-than | Right-tailed: the p-value is the right area |
| H-a with a not-equal | Two-tailed: double the one-sided area |
| A sigma stated | A z test, centred on the null's value |
| Only data values given | A t test with n minus one degrees of freedom |
| A one-tailed p-value above 0.5 | Probably the wrong tail |
| A tail chosen after seeing the data | The stated error rate no longer holds |
That closes chapter 9 and the one-sample methods. Chapter 10 keeps the same six steps and changes what is being compared: two populations rather than one, whose means or proportions are tested against each other rather than against a fixed value. The hypotheses, the errors, the p-value and the decision rule all carry over unchanged.
OpenStax Introductory Statistics 2e, §9.5 Additional Information and Full Hypothesis Test Examples §9.5, pp. 469-482 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.