A hypothesis test decides on incomplete evidence, so it can be wrong in two distinct ways. Crossing the two possible decisions with the two possible truths gives four outcomes: two of them correct, and two of them errors. A Type I error is rejecting the null hypothesis when it is true, and its probability is called alpha. A Type II error is failing to reject the null when it is false, and its probability is called beta. Both should be as small as possible and neither is ever zero, and they trade against each other — moving the decision boundary to reduce one increases the other. Only a larger sample reduces both at once. The probability of correctly rejecting a false null is one minus beta, and is called the power of the test. Which of the two errors matters more is a judgement about consequences that the statistics cannot supply.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 9 — Hypothesis Testing with One Sample
Outcomes and the Type I and Type II Errors
Objectives
Five outcomes, and the fourth is the one that is not a mathematical question at all.
OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, pp. 464-466 — the section these objectives are drawn from
Warm-up
Section 9.1 established that a test ends in one of two decisions: reject the null, or decline to reject it.
Discussion prompt
A smoke alarm is a hypothesis test. Its null hypothesis is that there is no fire. In what two ways can it be wrong, and which failure would you rather it made?
Hint: The alarm can sound, or stay silent, and the house can be on fire, or not.
Answer:
It can sound when there is no fire — a false alarm, which is rejecting a true null. Or it can stay silent when there is a fire — a missed detection, which is failing to reject a false null.
Almost everyone would rather have the false alarm, because burnt toast setting off a siren is an annoyance while an undetected fire is a catastrophe. That preference is why smoke alarms are deliberately made sensitive.
This section names those two failures Type I and Type II, gives each a probability, and shows that they trade against each other. What it cannot do is decide which one to prefer — that came from the consequences, not from any statistics.
Concept
When you perform a hypothesis test there are four possible outcomes, depending on the actual truth or falseness of the null hypothesis and the decision to reject or not. Rejecting a true null is a Type I error with probability alpha; failing to reject a false null is a Type II error with probability beta. Both should be as small as possible, and they are rarely zero.
Type I and Type II errors — The two off-diagonal outcomes. A Type I error rejects a null that is true; a Type II error fails to reject one that is false. Their probabilities are alpha and beta.
\[ \alpha = P(\text{Type I}), \qquad \beta = P(\text{Type II}) \]
The fourth outcome — correctly rejecting a false null — has a name of its own, the power of the test, and equals one minus beta. It is the outcome a study is usually run to achieve, so a test with low power is one that is unlikely to find what it is looking for even when it is there. Increasing the sample size increases the power, which is the practical reason sample size matters.
Figure (svg): A two-by-two table crossing the decision with the truth, giving two correct outcomes and the two error types
OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464
Section
Section 1
Concept
The decision is either to reject or not to reject the null, and the null is either true or false. Crossing them gives four outcomes: not rejecting a true null and rejecting a false one are correct, while rejecting a true null and failing to reject a false one are the two errors.
the outcome table — A two-by-two grid whose rows are the decision and whose columns are the truth. The two diagonal cells are correct decisions and the two off-diagonal cells are the errors.
\[ \text{decision} \times \text{truth} \;\Longrightarrow\; 4 \text{ outcomes} \]
It is worth being clear about what is and is not known in practice. The decision is known — you made it — but the truth never is, so you can never tell which of the four cells you are in. What the probabilities alpha and beta describe is how often each error would occur across many repetitions, which is the same kind of statement a confidence level made in chapter 8.
Figure (svg): A two-by-two table crossing the decision with the truth, giving two correct outcomes and the two error types
OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464 — the outcome table and the four cases
Picture it
Both errors and both correct outcomes in one picture.
Figure (svg): A two-by-two table crossing the decision with the truth, giving two correct outcomes and the two error types
Notice that the two errors sit in different columns, which means they cannot both occur on the same occasion — if the null is true only a Type I error is possible, and if it is false only a Type II. That is why alpha and beta are not probabilities of the same event and do not sum to anything meaningful.
Worked example
For a null hypothesis that a drug has no effect.
\[ H_0: \text{the drug has no effect} \]
Do not reject, null true
Why: Correctly finding nothing.
Reject, null true
Why: Claiming an effect that is not there.
Do not reject, null false
Why: Missing a real effect.
Reject, null false
Why: Correctly finding a real effect.
Figure (svg): The solution to Worked example naming the four cells shown as a ladder of expressions, one row per legal move
\[ \text{Type I: false alarm}; \quad \text{Type II: missed detection} \]
Verify: confirm which cell a real study hopes to land in
Why: A trial is usually run because the researcher believes the drug works, so the hoped-for cell is rejecting a false null — the power cell. That is why power matters at the design stage: a trial with power of 0.4 will miss a real effect more often than it finds one, and running it is close to a waste of resources whatever the result turns out to be.
OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464
Sorting
Cross the decision with the truth.
Sort into buckets
Sort each outcome.
Item (e) is item (b) restated in context, and it is the outcome most studies are run to achieve — which is why it has its own name, the power of the test.
Worked example
Distinguishing the decision from the truth.
\[ \text{a test rejects } H_0 \]
Note what is known
Why: The decision.
Note what is unknown
Why: Whether the null was true.
List the possible cells
Why: Two remain.
Say what alpha describes
Why: The long-run rate.
Figure (svg): The solution to Worked example what is knowable shown as a ladder of expressions, one row per legal move
\[ \text{decision known}; \quad \text{truth unknown} \]
Verify: confirm alpha is not the probability that this rejection is wrong
Why: Alpha is the probability of rejecting GIVEN that the null is true — a statement about the procedure, not about the case in hand. Reading a 5 percent alpha as a 5 percent chance that this particular rejection is mistaken is the same error section 8.1 warned about for confidence levels, and it has the same cause: the randomness lives in the procedure and not in the fixed truth.
OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464
Trap
\[ \alpha = 0.05 \;\Rightarrow\; \text{a 5 percent chance my rejection is mistaken} \]
Attach the probability to the conclusion just reached
Why: Alpha is the error rate, so it should describe errors.
\[ \alpha = P(\text{reject} \mid H_0 \text{ true}) \]
It is conditional on the null being true, which is exactly what is not known.
\[ \text{across many tests where } H_0 \text{ holds, 5 percent would reject it} \]
Read alpha as a long-run rate of the procedure
Why: It describes the method, not this instance.
The distinction matters because the chance that a particular rejection is wrong depends on how often the null is true among the questions being asked — something the test never sees. A field where most hypotheses tested are false will have far fewer mistaken rejections than one where most are true, at the same alpha.
Fill the middle
Rejecting a null hypothesis that is in fact true.
Fill in the blanks
\textI ___ \text___
Why: Type I — rejecting a true null, the false alarm. Its probability is alpha, and it is the error a researcher controls directly by choosing the significance level.
Two truths and a lie
All three concern the grid.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. Alpha and beta are conditional on different things — alpha assumes the null is true and beta assumes it is false — so they are not complementary and their sum has no meaning. What does sum to one is beta and the power.
Prediction
Commit before reasoning.
Predict first
A researcher believes a new treatment works. Which outcome are they hoping for?
Correct: Rejecting a false null.
Why: If the treatment really works then the null is false, and the desired outcome is to reject it — which is the power cell. That is why power is the quantity a study is designed around: a trial with low power will probably fail to detect the very effect it was run to find.
Section
Section 2
Concept
A Type I error is the decision to reject the null hypothesis when it is in fact true. Alpha is its probability: the probability of rejecting the null hypothesis when the null hypothesis is true.
alpha — The probability of a Type I error, and also the significance level chosen for the test. It is the one error probability the researcher sets directly, before any data are collected.
\[ \alpha = P(\text{reject } H_0 \mid H_0 \text{ true}) \]
Alpha's double role is worth noticing early. It is the probability of a Type I error, and it is also the significance level a researcher picks — usually 0.05 or 0.01. Those are the same number because choosing a significance level IS choosing how often you are willing to raise a false alarm, and section 9.4 will use it as the threshold a p-value is compared against.
Figure (svg): Two overlapping bell curves with a vertical cut-off between them, the right tail of the left curve shaded red for alpha and the left region of the right curve shaded orange for beta
OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464 — the definition of alpha
Picture it
The right tail of the null's distribution, beyond the cut-off.
Figure (svg): Two overlapping bell curves with a vertical cut-off between them, the right tail of the left curve shaded red for alpha and the left region of the right curve shaded orange for beta
The red region sits entirely on the left curve, which is the distribution the data would have IF the null were true. That is what the conditional in alpha's definition means, drawn: it is an area under the null's own curve and says nothing about how the data would behave if the alternative held.
Worked example
Example 9.5, the rock climbing equipment.
\[ H_0: \text{Navah's rock climbing equipment is safe} \]
Recall the definition
Why: Reject a true null.
Say what is concluded
Why: The rejection.
Combine them
Why: The error in words.
Name the probability
Why: Its chance.
Figure (svg): The solution to Worked example a Type I error in context shown as a ladder of expressions, one row per legal move
\[ \alpha = P(\text{thinks unsafe} \mid \text{is safe}) \]
Verify: confirm the phrasing keeps the two halves in the right order
Why: The sentence has the shape thinks X when in fact Y, where Y is the null being true and X is the rejection. Writing it in that order every time prevents the two errors from being swapped, which is the commonest mistake in this section. Reversed, the same words would describe the Type II error instead.
OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464
Faded example
H-nought: the blood cultures contain no traces of pathogen X.
Fill in the blanks
\textcontain do not \text___ ___
Why: Rejecting a true null: concluding the pathogen is present when it is not. Both halves of the sentence are needed, since an error is a mismatch between conclusion and truth.
Worked example
Example 9.6, the accident victim.
\[ H_0: \text{the victim is alive on arrival} \]
Describe the error
Why: Reject a true null.
Ask what follows
Why: The action taken.
Assess the cost
Why: A treatable patient dies.
Compare with Type II
Why: Uncertainty about a dead victim.
Figure (svg): The solution to Worked example the consequence of a Type I error shown as a ladder of expressions, one row per legal move
\[ \text{Type I here is the worse error} \]
Verify: confirm the assessment came from consequences rather than probabilities
Why: Nothing about alpha or beta entered this reasoning at all — the judgement rested entirely on what each mistake would cause someone to do. That is the section's central point, and the reason the book's four examples do not all give the same answer. The statistics supply the probabilities; the situation supplies the costs.
OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 465
Trap
\[ \text{Type I: she thinks the equipment is unsafe} \]
State what was concluded and stop
Why: That is what the decision was.
\[ \text{but that could be correct} \]
Thinking the equipment is unsafe is only an error if the equipment is in fact safe, and the sentence never says so.
\[ \text{Type I: she thinks it may not be safe when in fact it IS safe} \]
Always state both halves: the conclusion AND the truth
Why: An error is a mismatch, so both sides must appear.
Every description of a Type I or Type II error needs the phrase when in fact, or something equivalent. Without it the sentence describes a decision rather than an error, and there is no way to tell which of the two types is meant — which is precisely what the question is asking.
Two truths and a lie
All three concern alpha.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The book says both error probabilities are rarely zero, and alpha in particular is fixed by the researcher's choice rather than by the data — a larger sample reduces beta but leaves alpha exactly where it was set. Making alpha zero would mean never rejecting, and the test would be useless.
Prediction
Commit before reasoning.
Predict first
What determines the probability of a Type I error?
Correct: The significance level chosen.
Why: Alpha is set directly by the choice of significance level, before any data are seen — it is the same number. Sample size affects beta and the power but not alpha, and the p-value is computed from the data afterwards and compared against alpha rather than determining it.
Estimation
A hundred researchers each test a null hypothesis that happens to be true, at a 5 percent level.
Predict first
Roughly how many would reject it?
Correct: About 5.
Why: Alpha is the rejection rate among true nulls, so about five of the hundred would raise a false alarm despite doing everything correctly. This is why a single significant result in a field testing many hypotheses is weak evidence, and why replication matters more than any one p-value.
Section
Section 3
Concept
A Type II error is the decision not to reject the null hypothesis when in fact it is false. Beta is its probability: the probability of not rejecting the null hypothesis when the null hypothesis is false.
beta — The probability of a Type II error. Unlike alpha it is not chosen directly, and it depends on the sample size, the significance level, and how false the null actually is.
\[ \beta = P(\text{do not reject } H_0 \mid H_0 \text{ false}) \]
Beta is harder to pin down than alpha because it depends on something unknown: how far from the null the truth actually lies. A null that is badly wrong is easy to reject, so beta is small; a null that is wrong by a hair is nearly impossible to reject, so beta is close to one minus alpha. That is why beta is always quoted for a specified effect size rather than in the abstract.
Figure (svg): Two overlapping bell curves with a vertical cut-off between them, the right tail of the left curve shaded red for alpha and the left region of the right curve shaded orange for beta
OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464 — the definition of beta
Picture it
The region below the cut-off, on the alternative's distribution.
Figure (svg): Two overlapping bell curves with a vertical cut-off between them, the right tail of the left curve shaded red for alpha and the left region of the right curve shaded orange for beta
The orange region sits on the right-hand curve, which is where the data would fall if the alternative were true. Because the two shaded areas live on different curves, shifting the cut-off shrinks one and grows the other — and only narrowing both curves, which means collecting more data, shrinks both at once.
Worked example
Example 9.5 again, the other error.
\[ H_0: \text{Navah's rock climbing equipment is safe} \]
Recall the definition
Why: Fail to reject a false null.
Say what is concluded
Why: No rejection.
Combine them
Why: The error in words.
Ask what follows
Why: The action.
Figure (svg): The solution to Worked example a Type II error in context shown as a ladder of expressions, one row per legal move
\[ \beta = P(\text{thinks safe} \mid \text{is unsafe}) \]
Verify: confirm this is the more serious error here, and why
Why: The book says so explicitly, and gives the reason in one clause: if Navah thinks her rock climbing equipment is safe, she will go ahead and use it. The consequence of the Type I error is that she replaces gear that did not need replacing — an expense. The consequence of the Type II error is a fall. Comparing consequences rather than probabilities is what settles it.
OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464
Faded example
H-nought: the experimental drug has a cure rate of at least 75 percent.
Fill in the blanks
\textat least less \text___ ___
Why: Failing to reject a false null: believing the drug is as effective as claimed when it is not. The book calls this the more serious error here, since the belief influences the choice of treatment.
Worked example
Why beta cannot be quoted without an effect size.
\[ H_0: \mu = 100; \text{ the truth is } \mu = 100.01 \]
Note the null is false
Why: By a hundredth.
Ask how the data look
Why: Almost identical to the null.
Say what the test does
Why: Almost never rejects.
\[ \beta\text{ near } 1 - \alpha \]
Contrast a large error
Why: If mu were 130.
Figure (svg): The solution to Worked example when the truth is barely false shown as a ladder of expressions, one row per legal move
\[ \beta \text{ depends on how false } H_0 \text{ is} \]
Verify: confirm this is a feature rather than a defect
Why: A test failing to detect a departure of 0.01 is behaving sensibly: such a difference is almost certainly of no practical importance, and a method that flagged it would flag everything. What the dependence means in practice is that a power calculation must begin by naming the smallest effect worth detecting — which forces a judgement about what matters, before any data are collected.
OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464
Error analysis
H-nought is that the equipment is safe. Which describes the Type II error?
Annotate
On: \( \begin{aligned} &(1)\; \text{she thinks it is unsafe} \\ &(2)\; \text{she thinks it is unsafe when it is safe} \\ &(3)\; \text{she thinks it is safe when it is unsafe} \\ &(4)\; \text{the equipment fails} \end{aligned} \)
Items (2) and (3) use the same six words in different orders, which is why the when in fact clause has to be written out in full every time. Item (4) is a different kind of confusion: a Type II error is a mistaken conclusion, not the accident that may follow from it.
Two truths and a lie
All three concern beta.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. Alpha is set directly; beta follows from the sample size, the significance level and the true effect size, none of which the researcher fully controls. Beta is targeted at the design stage by choosing a sample size, but it is never simply set.
Discrimination
For a null hypothesis that a patient is not sick.
Sort into buckets
Sort each mistaken conclusion.
Prediction
Commit before reasoning.
Predict first
For a fixed sample and significance level, when is a Type II error most likely?
Correct: When the null is false by a small amount.
Why: A tiny departure produces data almost indistinguishable from what the null predicts, so the test almost never rejects and beta is close to its maximum. A large departure is easy to detect and beta is near zero. The third option is a category error: if the null is true, a Type II error is impossible by definition.
Section
Section 4
Concept
The book presents four examples and asks in each which error has the greater consequence. The answers alternate, and in every case the reasoning turns on what each mistake would cause someone to do rather than on any probability.
consequence, not probability — Alpha and beta say how often each error occurs; they say nothing about what each costs. Deciding which to guard against harder requires knowing the situation, which is why the significance level is chosen rather than derived.
\[ \text{cost} = P(\text{error}) \times \text{consequence} \]
The practical consequence is that the conventional 0.05 is a convention rather than a law. A test whose Type I error would shut down a production line unnecessarily might reasonably use a smaller alpha; a screening test where a missed case is fatal might use a larger one, accepting more false alarms to reduce misses. Neither choice is more statistical than the other.
Figure (svg): A three-column table listing four null hypotheses with which error type is worse in each case and why
OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, pp. 464-466 — Examples 9.5 through 9.8
Picture it
The book's cases and which error each identifies as worse.
Figure (svg): A three-column table listing four null hypotheses with which error type is worse in each case and why
Two say Type I and two say Type II, which is as clear a demonstration as could be arranged that there is no general answer. Anyone who learns a rule such as Type I is always worse will get half of these wrong.
Worked example
Example 9.7, the Genetic Labs claim.
\[ H_0: \text{Genetic Labs has no effect on sex outcome} \]
Describe Type I
Why: Reject a true null.
Describe Type II
Why: Fail to reject a false one.
Ask what follows from each
Why: The actions taken.
Compare
Why: Money spent on nothing.
Figure (svg): The solution to Worked example comparing the two consequences shown as a ladder of expressions, one row per legal move
\[ \text{Type I: a useless product is bought} \]
Verify: confirm the comparison would flip in a different setting
Why: If the product were free and harmless, the Type I error would cost almost nothing and the Type II — missing a genuine effect — might matter more to the science. The consequence depends on what follows from the belief, not on the belief itself, which is why the same null hypothesis can have different answers in different circumstances.
OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 465
Sorting
Ask what each mistake causes someone to do.
Sort into buckets
Sort each situation by which error carries the greater consequence.
Item (e) is not in the book but follows the same reasoning as the climbing equipment: a null of safety makes the missed detection the dangerous error. Nulls asserting that something is safe almost always put the weight on the Type II error.
Worked example
Example 9.8, the experimental cancer drug.
\[ H_0: \text{the cure rate is at least } 75 \text{ percent} \]
Describe Type I
Why: Believe the rate is under 75.
\[ \text{when it is at least } 75 \]
Describe Type II
Why: Believe the rate is at least 75.
Ask what follows from Type II
Why: The treatment choice.
Compare
Why: A missed better option.
Figure (svg): The solution to Worked example the opposite verdict shown as a ladder of expressions, one row per legal move
\[ \text{Type II: a weak drug is chosen} \]
Verify: confirm why this reverses Example 9.7's answer
Why: In the Genetic Labs case the error that leads to action is the Type I; here the null is a claim of effectiveness, so it is the failure to reject it that leads someone to act. Which error triggers the costly action depends on how the null was written — and since the null could have been written the other way round, the labelling is partly a matter of framing. What does not change is which real-world mistake is worse.
OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 466
Trap
\[ \text{Type I errors are always the more serious} \]
Generalise from the conventional emphasis on alpha
Why: Alpha is the one always reported, so it must matter more.
\[ \text{but two of the book's four examples say otherwise} \]
The rock climber and the cancer patient both face a worse Type II error, for reasons specific to their situations.
\[ \text{ask what each error would cause someone to DO} \]
Compare consequences case by case
Why: The statistics supply probabilities, not costs.
Alpha gets more attention partly for a practical reason: it is the error a researcher can control directly and report honestly, while beta depends on an unknown effect size. That convenience has nothing to do with which error matters more, and mistaking one for the other is how tests get designed to guard against the wrong failure.
Two truths and a lie
All three concern consequences.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and confuses how often with how bad. A rare error with catastrophic consequences may matter far more than a common one with trivial ones — which is precisely the case for the climbing equipment, where beta may be small and still dominate.
Prediction
Commit before reasoning.
Predict first
A team decides the Type I error is far more costly. What should they do?
Correct: Choose a smaller alpha.
Why: Alpha is the one error probability set directly, so guarding harder against a Type I error means lowering it — and since the two trade against each other at a fixed sample size, beta rises as a result. The only way to reduce both is a larger sample, which is the next idea.
Explain it
A classmate asks why the book does not just say which error type is worse.
Discussion prompt
In two sentences or fewer, explain.
Hint: Ask what each error causes in the climbing case and in the ambulance case.
Answer:
In the climbing case the dangerous mistake is believing unsafe gear is safe, and in the ambulance case it is believing a living victim is dead — one is a Type II and the other a Type I.
Which error leads to the harmful action depends entirely on the situation and on how the null was written, so no general answer exists and the question has to be asked afresh each time.
Section
Section 5
Concept
The power of the test is one minus beta: the probability of correctly rejecting a false null. Ideally we want a high power, as close to one as possible, and increasing the sample size can increase it.
the power of the test — The probability of the fourth outcome — rejecting the null when it is false. It is what a study is usually run to achieve, and it is bought with sample size.
\[ \text{power} = 1 - \beta \]
The trade-off between the two errors is worth seeing clearly. At a fixed sample size, moving the decision boundary reduces one error probability and raises the other, because the two areas sit on different curves. Only a larger sample narrows both curves and reduces both at once — which is why sample size is the lever that matters and why a low-powered study cannot be rescued by adjusting alpha.
Figure (svg): A rising curve showing the power of a test climbing from about a fifth at a sample of five toward one at a sample of a hundred
OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464 — the power of the test
Picture it
For a true effect of half a standard deviation at a five percent level.
Figure (svg): A rising curve showing the power of a test climbing from about a fifth at a sample of five toward one at a sample of a hundred
The curve is steep in the middle and flat at both ends, which has a practical reading. Below about twenty observations the study will usually miss this effect; above about seventy it will usually find it; and the range in between is where an extra handful of observations changes the answer most.
Worked example
Converting between the two.
\[ \beta = 0.30 \]
Recall the relationship
Why: Power is one minus beta.
\[ 1 - 0.30 \]
Compute
Why: The power.
\[ 0.70 \]
State what it means
Why: The chance of detecting.
\[ 70 \% \]
Assess it
Why: Below the usual target.
Figure (svg): The solution to Worked example reading power from beta shown as a ladder of expressions, one row per legal move
\[ \text{power} = 1 - 0.30 = 0.70 \]
Verify: confirm against the conventional target
Why: Studies are commonly designed for a power of 0.80, so 0.70 falls short — the test would miss the effect three times in ten. Whether that is acceptable depends on the cost of missing it, which is the previous idea's question again. A power of 0.70 also means that a non-significant result from this study is weak evidence of absence, since failure was fairly likely even if the effect is real.
OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464
Faded example
A test has a Type II error probability of 0.20.
Fill in the blanks
\text0.20 = 1 - 0.80 = ___
Why: Power of 0.80, which is the level studies are conventionally designed to reach. It means the test would detect the specified effect four times in five.
Worked example
What moving the cut-off does.
\[ \text{a fixed } n; \text{ alpha lowered from } 0.05 \text{ to } 0.01 \]
Note the cut-off moves
Why: Further into the tail.
Effect on Type I
Why: Fewer false alarms.
Effect on Type II
Why: More missed detections.
Effect on power
Why: One minus a larger beta.
Figure (svg): The solution to Worked example the trade-off at a fixed sample shown as a ladder of expressions, one row per legal move
\[ \alpha \downarrow \;\Rightarrow\; \beta \uparrow \;\Rightarrow\; \text{power} \downarrow \]
Verify: confirm why only more data escapes the trade-off
Why: Moving the boundary shifts area between two fixed curves, so what one gains the other loses. A larger sample narrows both curves, pulling each away from the boundary, so both tails shrink together. That is the whole reason sample size is worth paying for, and it is why the answer to a study that finds nothing is usually a bigger study rather than a looser threshold.
OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464
Trap
\[ \text{use } \alpha = 0.001 \text{ so the conclusion is more trustworthy} \]
Reduce the error probability to reduce errors
Why: A smaller alpha means fewer mistakes.
\[ \beta \text{ rises, so more real effects are missed} \]
It reduces one kind of mistake by increasing the other, and at a fixed sample size that is all it can do.
\[ \text{choose } \alpha \text{ from consequences; buy power with } n \]
Set alpha by what a false alarm costs, and set n by what a miss costs
Why: The two levers do different jobs.
An extremely small alpha on a small sample produces a test that almost never rejects anything, true or false — which looks rigorous and is actually uninformative. The honest way to make a study trustworthy is to make it large enough that both error rates are acceptable, and to say what effect size it was powered to detect.
Two truths and a lie
All three concern power and the trade-off.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. At a fixed sample size the two trade against each other — moving the cut-off to reduce false alarms necessarily increases missed detections. Only more data reduces both, which is why sample size rather than significance level is the lever that improves a study.
Estimation
A trial is calculated to have power 0.30 for the smallest effect worth detecting.
Predict first
What does that mean for the study?
Correct: It will miss a real effect about seven times in ten.
Why: Power of 0.30 means beta is 0.70, so the effect goes undetected in seventy percent of such studies even when it is genuinely there. Running it risks producing a non-significant result that tells nobody anything, which is why power calculations belong before data collection rather than after.
Prediction
Commit before reasoning.
Predict first
Which change increases the power of a test without increasing the Type I error rate?
Correct: A larger sample.
Why: More data narrows the sampling distribution under both hypotheses, pulling both tails away from the boundary — so beta falls and power rises while alpha stays exactly where it was set. Every other option either trades one error for the other or makes the effect harder to detect.
Comparison
Fill the blanks. They are conditional on opposite things, which is why they never sum to anything.
Comparison matrix
| Type I | Type II | |
|---|---|---|
| The mistake | rejecting a true null | failing to reject a false null |
| Probability | alpha | beta |
| Set or derived? | chosen directly by the researcher | follows from n, alpha and the effect size |
| Everyday name | a false positive, or false alarm | a false negative, or missed detection |
The third row explains why alpha is always reported and beta rarely is: one is a decision and the other depends on an unknown. That asymmetry in reporting has nothing to do with which error matters more.
Pattern
Five steps, and the fifth is not a calculation.
Every error description needs a when in fact clause. A sentence stating only a conclusion describes a decision, not an error, and cannot be assigned a type.
OpenStax Introductory Business Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2 Outcomes and the Type I and Type II Errors
Check
Naming the error.
Check your understanding
A test rejects a null hypothesis that is actually true. What is this?
Answer: A
Why: Rejecting a true null is the Type I error, and its probability is alpha — the significance level chosen for the test.
Check
The trade-off.
Check your understanding
At a fixed sample size, what happens to beta if alpha is reduced?
Answer: A
Why: The two areas sit on different curves, so moving the cut-off to shrink one enlarges the other. Only a larger sample reduces both.
Check
Which error is worse.
Check your understanding
For the null hypothesis that a water supply is uncontaminated, which error is more serious?
Answer: A
Why: The Type II error leads people to drink contaminated water, while the Type I error leads to unnecessary treatment or testing.
Real world
An airport installs a scanner that flags bags for manual search. Its null hypothesis is that a bag is harmless. The vendor advertises a false alarm rate of only 0.5 percent, and the security manager proposes tightening the threshold further to reduce the queues that searches cause.
Discussion prompt
Analyse the proposal in terms of the two errors, and say what the manager should consider instead.
Hint: The advertised figure describes only one of the two errors.
Answer:
The advertised 0.5 percent is alpha alone. It is the probability that a harmless bag is flagged — the false alarm rate, and the one that generates queues. The vendor's figure says nothing whatever about the other error: the probability that a dangerous bag is passed, which is beta.
Tightening the threshold reduces alpha and increases beta. Fewer harmless bags would be searched, shortening queues, and more dangerous bags would go through undetected. At a fixed scanner and a fixed inspection process, that trade is unavoidable — the two error rates sit on different distributions and only move in opposite directions.
\[ \text{threshold} \uparrow \;\Rightarrow\; \alpha \downarrow, \; \beta \uparrow, \; \text{power} \downarrow \]
The consequences here are wildly asymmetric, which settles the direction. A Type I error costs a passenger a few minutes; a Type II error is a weapon on an aircraft. Security systems are therefore deliberately set to a high false alarm rate — the opposite of what the manager proposes — and the queue is the accepted price of a small beta.
What the manager should pursue instead is anything that improves both errors at once, which means better information rather than a different threshold: a more sensitive scanner, a second independent screening stage, or more staff so that a given alarm rate produces shorter queues. This is the same conclusion as the lesson's trade-off figure — moving the boundary redistributes the errors, and only more or better data reduces both. It is also worth noting that the false alarm rate alone is an incomplete specification for any detection system, and a vendor quoting it without a detection rate has described half the product.
Commit first
Answer, then rate your confidence honestly.
Predict first
At a fixed sample size, why can alpha and beta not both be made small?
Correct: Because the two areas sit on different distributions.
\[ \alpha = P(\text{reject} \mid H_0), \qquad \beta = P(\text{not reject} \mid H_a) \]
Why: Alpha is a tail area under the null's distribution and beta is an area under the alternative's, on opposite sides of the same boundary. Moving the boundary therefore shrinks one and grows the other. They do not sum to one, since they are conditional on different things — and the only way to reduce both is to narrow both distributions, which means more data.
Explain it
They wrote: the Type II error is that Navah thinks her equipment is unsafe.
Discussion prompt
In two sentences or fewer, correct it.
Hint: Ask what the equipment is actually like in their sentence.
Answer:
Their sentence gives a conclusion but never says what the truth is, so it does not describe an error of either type — and thinking the gear is unsafe is a rejection, which points at Type I anyway.
The Type II error is thinking the equipment may be safe when in fact it is not, which is the failure to reject a false null and the one that gets her hurt.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, cross the decision with the truth. For the second, always include a when in fact clause. For the third, ask what each mistake makes someone do. For the fourth, moving the cut-off trades one error for the other and only more data shrinks both. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top, draw the two-by-two grid with the decision down the side and the truth across the top, and fill all four cells with their names and probabilities, marking which cell is the power. Beneath it, draw two overlapping bell curves with a vertical cut-off between them, shade the right tail of the left curve for alpha and the left region of the right curve for beta, and write one sentence on what happens to each area when the cut-off moves right. Then write a second sentence on what happens to both when the sample grows. In the middle of the page, take the book's four examples in turn — the climbing equipment, the accident victim, Genetic Labs and the cancer drug — and for each write the null as a sentence, then the Type I error and the Type II error each with a when in fact clause, then which is worse and what action makes it so. At the bottom, write power equals one minus beta, and sketch a rising curve of power against sample size, marking roughly where a study would go from usually missing an effect to usually finding it.
Check your four examples by confirming that two say Type I is worse and two say Type II — if all four agree, one of the descriptions has probably been swapped. Check the two shaded areas by confirming they sit on different curves, since that is what makes the trade-off unavoidable.
Recap
Five things, and the third is a judgement rather than a calculation.
| If you see | Then |
|---|---|
| Rejecting a true null | A Type I error, probability alpha |
| Failing to reject a false null | A Type II error, probability beta |
| A description with no when-in-fact clause | A decision, not an error |
| A null asserting something is safe | The Type II error is usually the dangerous one |
| A request to lower alpha at fixed n | Beta will rise; power will fall |
| A study that found nothing | Ask what its power was |
| A wish to reduce both errors | Only a larger sample does that |
Section 9.3 turns from what can go wrong to what machinery to use. Three tests have appeared so far — a mean with sigma known, a mean with sigma unknown, and a proportion — and each carries its own distribution and its own assumptions, which the next section sets out in a single table.
OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, pp. 464-466 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.