9.2 Outcomes and the Type I and Type II Errors

A hypothesis test decides on incomplete evidence, so it can be wrong in two distinct ways. Crossing the two possible decisions with the two possible truths gives four outcomes: two of them correct, and two of them errors. A Type I error is rejecting the null hypothesis when it is true, and its probability is called alpha. A Type II error is failing to reject the null when it is false, and its probability is called beta. Both should be as small as possible and neither is ever zero, and they trade against each other — moving the decision boundary to reduce one increases the other. Only a larger sample reduces both at once. The probability of correctly rejecting a false null is one minus beta, and is called the power of the test. Which of the two errors matters more is a judgement about consequences that the statistics cannot supply.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 9.2 Outcomes and the Type I and Type II Errors

Title

Statistics · Chapter 9 — Hypothesis Testing with One Sample

Outcomes and the Type I and Type II Errors

2. By the end of this lesson you can

Objectives

Five outcomes, and the fourth is the one that is not a mathematical question at all.

OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, pp. 464-466 — the section these objectives are drawn from

3. What you already have

Warm-up

Section 9.1 established that a test ends in one of two decisions: reject the null, or decline to reject it.

Discussion prompt

A smoke alarm is a hypothesis test. Its null hypothesis is that there is no fire. In what two ways can it be wrong, and which failure would you rather it made?

Hint: The alarm can sound, or stay silent, and the house can be on fire, or not.

Answer:

It can sound when there is no fire — a false alarm, which is rejecting a true null. Or it can stay silent when there is a fire — a missed detection, which is failing to reject a false null.

Almost everyone would rather have the false alarm, because burnt toast setting off a siren is an annoyance while an undetected fire is a catastrophe. That preference is why smoke alarms are deliberately made sensitive.

This section names those two failures Type I and Type II, gives each a probability, and shows that they trade against each other. What it cannot do is decide which one to prefer — that came from the consequences, not from any statistics.

4. Two ways to be wrong, with different probabilities

Concept

When you perform a hypothesis test there are four possible outcomes, depending on the actual truth or falseness of the null hypothesis and the decision to reject or not. Rejecting a true null is a Type I error with probability alpha; failing to reject a false null is a Type II error with probability beta. Both should be as small as possible, and they are rarely zero.

Type I and Type II errors — The two off-diagonal outcomes. A Type I error rejects a null that is true; a Type II error fails to reject one that is false. Their probabilities are alpha and beta.

\[ \alpha = P(\text{Type I}), \qquad \beta = P(\text{Type II}) \]

The fourth outcome — correctly rejecting a false null — has a name of its own, the power of the test, and equals one minus beta. It is the outcome a study is usually run to achieve, so a test with low power is one that is unlikely to find what it is looking for even when it is there. Increasing the sample size increases the power, which is the practical reason sample size matters.

Figure (svg): A two-by-two table crossing the decision with the truth, giving two correct outcomes and the two error types

Alpha is the probability of rejecting a true null; beta the probability of failing to reject a false one. Neither is ever zero.

OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464

5. The four outcomes

Section

Section 1

6. Two decisions crossed with two truths

Concept

The decision is either to reject or not to reject the null, and the null is either true or false. Crossing them gives four outcomes: not rejecting a true null and rejecting a false one are correct, while rejecting a true null and failing to reject a false one are the two errors.

the outcome table — A two-by-two grid whose rows are the decision and whose columns are the truth. The two diagonal cells are correct decisions and the two off-diagonal cells are the errors.

\[ \text{decision} \times \text{truth} \;\Longrightarrow\; 4 \text{ outcomes} \]

It is worth being clear about what is and is not known in practice. The decision is known — you made it — but the truth never is, so you can never tell which of the four cells you are in. What the probabilities alpha and beta describe is how often each error would occur across many repetitions, which is the same kind of statement a confidence level made in chapter 8.

Figure (svg): A two-by-two table crossing the decision with the truth, giving two correct outcomes and the two error types

Alpha is the probability of rejecting a true null; beta the probability of failing to reject a false one. Neither is ever zero.

OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464 — the outcome table and the four cases

7. The grid

Picture it

Both errors and both correct outcomes in one picture.

Figure (svg): A two-by-two table crossing the decision with the truth, giving two correct outcomes and the two error types

Alpha is the probability of rejecting a true null; beta the probability of failing to reject a false one. Neither is ever zero.

Notice that the two errors sit in different columns, which means they cannot both occur on the same occasion — if the null is true only a Type I error is possible, and if it is false only a Type II. That is why alpha and beta are not probabilities of the same event and do not sum to anything meaningful.

8. Worked example: naming the four cells

Worked example

For a null hypothesis that a drug has no effect.

\[ H_0: \text{the drug has no effect} \]

Do not reject, null true

Why: Correctly finding nothing.

Reject, null true

Why: Claiming an effect that is not there.

Do not reject, null false

Why: Missing a real effect.

Reject, null false

Why: Correctly finding a real effect.

Figure (svg): The solution to Worked example naming the four cells shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{Type I: false alarm}; \quad \text{Type II: missed detection} \]

Verify: confirm which cell a real study hopes to land in

Why: A trial is usually run because the researcher believes the drug works, so the hoped-for cell is rejecting a false null — the power cell. That is why power matters at the design stage: a trial with power of 0.4 will miss a real effect more often than it finds one, and running it is close to a waste of resources whatever the result turns out to be.

OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464

9. Correct or an error?

Sorting

Cross the decision with the truth.

Sort into buckets

Sort each outcome.

An error
reject a true null; do not reject a false null
A correct outcome
reject a false null; do not reject a true null; reject when the drug really works
err
The decision disagrees with the truth: either a false alarm or a missed detection.
ok
The decision matches the truth.

Item (e) is item (b) restated in context, and it is the outcome most studies are run to achieve — which is why it has its own name, the power of the test.

10. Worked example: what is knowable

Worked example

Distinguishing the decision from the truth.

\[ \text{a test rejects } H_0 \]

Note what is known

Why: The decision.

Note what is unknown

Why: Whether the null was true.

List the possible cells

Why: Two remain.

Say what alpha describes

Why: The long-run rate.

Figure (svg): The solution to Worked example what is knowable shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{decision known}; \quad \text{truth unknown} \]

Verify: confirm alpha is not the probability that this rejection is wrong

Why: Alpha is the probability of rejecting GIVEN that the null is true — a statement about the procedure, not about the case in hand. Reading a 5 percent alpha as a 5 percent chance that this particular rejection is mistaken is the same error section 8.1 warned about for confidence levels, and it has the same cause: the randomness lives in the procedure and not in the fixed truth.

OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464

11. Trap: reading alpha as the chance this conclusion is wrong

Trap

The trap

\[ \alpha = 0.05 \;\Rightarrow\; \text{a 5 percent chance my rejection is mistaken} \]

Attach the probability to the conclusion just reached

Why: Alpha is the error rate, so it should describe errors.

\[ \alpha = P(\text{reject} \mid H_0 \text{ true}) \]

It is conditional on the null being true, which is exactly what is not known.

The fix

\[ \text{across many tests where } H_0 \text{ holds, 5 percent would reject it} \]

Read alpha as a long-run rate of the procedure

Why: It describes the method, not this instance.

The distinction matters because the chance that a particular rejection is wrong depends on how often the null is true among the questions being asked — something the test never sees. A field where most hypotheses tested are false will have far fewer mistaken rejections than one where most are true, at the same alpha.

12. Name the error

Fill the middle

Rejecting a null hypothesis that is in fact true.

Fill in the blanks

\textI ___ \text___

Why: Type I — rejecting a true null, the false alarm. Its probability is alpha, and it is the error a researcher controls directly by choosing the significance level.

13. One of these is false

Two truths and a lie

All three concern the grid.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Only one of the two errors is possible on any given occasion
  • C. You can never tell which cell you are in
  • B. Alpha and beta sum to one

Survives elimination: B

Why: The survivor is false. Alpha and beta are conditional on different things — alpha assumes the null is true and beta assumes it is false — so they are not complementary and their sum has no meaning. What does sum to one is beta and the power.

14. Which cell does a study want?

Prediction

Commit before reasoning.

Predict first

A researcher believes a new treatment works. Which outcome are they hoping for?

  • Rejecting a false null: the power cell
  • Not rejecting a true null
  • A Type I error
  • A Type II error

Correct: Rejecting a false null.

Why: If the treatment really works then the null is false, and the desired outcome is to reject it — which is the power cell. That is why power is the quantity a study is designed around: a trial with low power will probably fail to detect the very effect it was run to find.

15. The Type I error

Section

Section 2

16. A false alarm, with probability alpha

Concept

A Type I error is the decision to reject the null hypothesis when it is in fact true. Alpha is its probability: the probability of rejecting the null hypothesis when the null hypothesis is true.

alpha — The probability of a Type I error, and also the significance level chosen for the test. It is the one error probability the researcher sets directly, before any data are collected.

\[ \alpha = P(\text{reject } H_0 \mid H_0 \text{ true}) \]

Alpha's double role is worth noticing early. It is the probability of a Type I error, and it is also the significance level a researcher picks — usually 0.05 or 0.01. Those are the same number because choosing a significance level IS choosing how often you are willing to raise a false alarm, and section 9.4 will use it as the threshold a p-value is compared against.

Figure (svg): Two overlapping bell curves with a vertical cut-off between them, the right tail of the left curve shaded red for alpha and the left region of the right curve shaded orange for beta

The two shaded areas sit on different curves, which is why they cannot both be made small by moving the boundary.

OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464 — the definition of alpha

17. Alpha as an area

Picture it

The right tail of the null's distribution, beyond the cut-off.

Figure (svg): Two overlapping bell curves with a vertical cut-off between them, the right tail of the left curve shaded red for alpha and the left region of the right curve shaded orange for beta

The two shaded areas sit on different curves, which is why they cannot both be made small by moving the boundary.

The red region sits entirely on the left curve, which is the distribution the data would have IF the null were true. That is what the conditional in alpha's definition means, drawn: it is an area under the null's own curve and says nothing about how the data would behave if the alternative held.

18. Worked example: a Type I error in context

Worked example

Example 9.5, the rock climbing equipment.

\[ H_0: \text{Navah's rock climbing equipment is safe} \]

Recall the definition

Why: Reject a true null.

Say what is concluded

Why: The rejection.

Combine them

Why: The error in words.

Name the probability

Why: Its chance.

Figure (svg): The solution to Worked example a Type I error in context shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \alpha = P(\text{thinks unsafe} \mid \text{is safe}) \]

Verify: confirm the phrasing keeps the two halves in the right order

Why: The sentence has the shape thinks X when in fact Y, where Y is the null being true and X is the rejection. Writing it in that order every time prevents the two errors from being swapped, which is the commonest mistake in this section. Reversed, the same words would describe the Type II error instead.

OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464

19. Describe a Type I error

Faded example

H-nought: the blood cultures contain no traces of pathogen X.

Fill in the blanks

\textcontain do not \text___ ___

Why: Rejecting a true null: concluding the pathogen is present when it is not. Both halves of the sentence are needed, since an error is a mismatch between conclusion and truth.

20. Worked example: the consequence of a Type I error

Worked example

Example 9.6, the accident victim.

\[ H_0: \text{the victim is alive on arrival} \]

Describe the error

Why: Reject a true null.

Ask what follows

Why: The action taken.

Assess the cost

Why: A treatable patient dies.

Compare with Type II

Why: Uncertainty about a dead victim.

Figure (svg): The solution to Worked example the consequence of a Type I error shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{Type I here is the worse error} \]

Verify: confirm the assessment came from consequences rather than probabilities

Why: Nothing about alpha or beta entered this reasoning at all — the judgement rested entirely on what each mistake would cause someone to do. That is the section's central point, and the reason the book's four examples do not all give the same answer. The statistics supply the probabilities; the situation supplies the costs.

OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 465

21. Trap: describing the error without naming the truth

Trap

The trap

\[ \text{Type I: she thinks the equipment is unsafe} \]

State what was concluded and stop

Why: That is what the decision was.

\[ \text{but that could be correct} \]

Thinking the equipment is unsafe is only an error if the equipment is in fact safe, and the sentence never says so.

The fix

\[ \text{Type I: she thinks it may not be safe when in fact it IS safe} \]

Always state both halves: the conclusion AND the truth

Why: An error is a mismatch, so both sides must appear.

Every description of a Type I or Type II error needs the phrase when in fact, or something equivalent. Without it the sentence describes a decision rather than an error, and there is no way to tell which of the two types is meant — which is precisely what the question is asking.

22. One of these is false

Two truths and a lie

All three concern alpha.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Alpha is chosen by the researcher before the data
  • C. Alpha is conditional on the null being true
  • B. Alpha can be made zero with enough data

Survives elimination: B

Why: The survivor is false. The book says both error probabilities are rarely zero, and alpha in particular is fixed by the researcher's choice rather than by the data — a larger sample reduces beta but leaves alpha exactly where it was set. Making alpha zero would mean never rejecting, and the test would be useless.

23. What sets alpha?

Prediction

Commit before reasoning.

Predict first

What determines the probability of a Type I error?

  • The significance level the researcher chooses
  • The sample size
  • The true value of the parameter
  • The p-value

Correct: The significance level chosen.

Why: Alpha is set directly by the choice of significance level, before any data are seen — it is the same number. Sample size affects beta and the power but not alpha, and the p-value is computed from the data afterwards and compared against alpha rather than determining it.

24. How often, at 5 percent?

Estimation

A hundred researchers each test a null hypothesis that happens to be true, at a 5 percent level.

Predict first

Roughly how many would reject it?

  • About 5
  • None
  • About 50
  • About 95

Correct: About 5.

Why: Alpha is the rejection rate among true nulls, so about five of the hundred would raise a false alarm despite doing everything correctly. This is why a single significant result in a field testing many hypotheses is weak evidence, and why replication matters more than any one p-value.

25. The Type II error

Section

Section 3

26. A missed detection, with probability beta

Concept

A Type II error is the decision not to reject the null hypothesis when in fact it is false. Beta is its probability: the probability of not rejecting the null hypothesis when the null hypothesis is false.

beta — The probability of a Type II error. Unlike alpha it is not chosen directly, and it depends on the sample size, the significance level, and how false the null actually is.

\[ \beta = P(\text{do not reject } H_0 \mid H_0 \text{ false}) \]

Beta is harder to pin down than alpha because it depends on something unknown: how far from the null the truth actually lies. A null that is badly wrong is easy to reject, so beta is small; a null that is wrong by a hair is nearly impossible to reject, so beta is close to one minus alpha. That is why beta is always quoted for a specified effect size rather than in the abstract.

Figure (svg): Two overlapping bell curves with a vertical cut-off between them, the right tail of the left curve shaded red for alpha and the left region of the right curve shaded orange for beta

The two shaded areas sit on different curves, which is why they cannot both be made small by moving the boundary.

OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464 — the definition of beta

27. Beta as an area on the other curve

Picture it

The region below the cut-off, on the alternative's distribution.

Figure (svg): Two overlapping bell curves with a vertical cut-off between them, the right tail of the left curve shaded red for alpha and the left region of the right curve shaded orange for beta

The two shaded areas sit on different curves, which is why they cannot both be made small by moving the boundary.

The orange region sits on the right-hand curve, which is where the data would fall if the alternative were true. Because the two shaded areas live on different curves, shifting the cut-off shrinks one and grows the other — and only narrowing both curves, which means collecting more data, shrinks both at once.

28. Worked example: a Type II error in context

Worked example

Example 9.5 again, the other error.

\[ H_0: \text{Navah's rock climbing equipment is safe} \]

Recall the definition

Why: Fail to reject a false null.

Say what is concluded

Why: No rejection.

Combine them

Why: The error in words.

Ask what follows

Why: The action.

Figure (svg): The solution to Worked example a Type II error in context shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \beta = P(\text{thinks safe} \mid \text{is unsafe}) \]

Verify: confirm this is the more serious error here, and why

Why: The book says so explicitly, and gives the reason in one clause: if Navah thinks her rock climbing equipment is safe, she will go ahead and use it. The consequence of the Type I error is that she replaces gear that did not need replacing — an expense. The consequence of the Type II error is a fall. Comparing consequences rather than probabilities is what settles it.

OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464

29. Describe a Type II error

Faded example

H-nought: the experimental drug has a cure rate of at least 75 percent.

Fill in the blanks

\textat least less \text___ ___

Why: Failing to reject a false null: believing the drug is as effective as claimed when it is not. The book calls this the more serious error here, since the belief influences the choice of treatment.

30. Worked example: when the truth is barely false

Worked example

Why beta cannot be quoted without an effect size.

\[ H_0: \mu = 100; \text{ the truth is } \mu = 100.01 \]

Note the null is false

Why: By a hundredth.

Ask how the data look

Why: Almost identical to the null.

Say what the test does

Why: Almost never rejects.

\[ \beta\text{ near } 1 - \alpha \]

Contrast a large error

Why: If mu were 130.

Figure (svg): The solution to Worked example when the truth is barely false shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \beta \text{ depends on how false } H_0 \text{ is} \]

Verify: confirm this is a feature rather than a defect

Why: A test failing to detect a departure of 0.01 is behaving sensibly: such a difference is almost certainly of no practical importance, and a method that flagged it would flag everything. What the dependence means in practice is that a power calculation must begin by naming the smallest effect worth detecting — which forces a judgement about what matters, before any data are collected.

OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464

31. Error analysis: four descriptions of an error for a safety null

Error analysis

H-nought is that the equipment is safe. Which describes the Type II error?

Annotate

On: \( \begin{aligned} &(1)\; \text{she thinks it is unsafe} \\ &(2)\; \text{she thinks it is unsafe when it is safe} \\ &(3)\; \text{she thinks it is safe when it is unsafe} \\ &(4)\; \text{the equipment fails} \end{aligned} \)

  • (1) states a conclusion without a truth, so it describes a decision rather than an error of either type.
  • (2) is the Type I error: rejecting a true null.
  • (3) is the Type II error: failing to reject a false null.
  • (4) describes an event in the world rather than a decision, and is not an error of inference at all.

Items (2) and (3) use the same six words in different orders, which is why the when in fact clause has to be written out in full every time. Item (4) is a different kind of confusion: a Type II error is a mistaken conclusion, not the accident that may follow from it.

32. One of these is false

Two truths and a lie

All three concern beta.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Beta depends on how false the null actually is
  • C. Beta is reduced by a larger sample
  • B. Beta is chosen by the researcher, like alpha

Survives elimination: B

Why: The survivor is false. Alpha is set directly; beta follows from the sample size, the significance level and the true effect size, none of which the researcher fully controls. Beta is targeted at the design stage by choosing a sample size, but it is never simply set.

33. Which error is this?

Discrimination

For a null hypothesis that a patient is not sick.

Sort into buckets

Sort each mistaken conclusion.

Type I: rejecting a true null
diagnosing illness in a healthy patient; treating a healthy patient unnecessarily; a false positive on a screening test
Type II: failing to reject a false null
clearing a patient who is in fact sick; missing a disease that is present
one
The null of not sick is true and has been rejected: a false positive.
two
The null of not sick is false and has not been rejected: a false negative.

34. When is beta largest?

Prediction

Commit before reasoning.

Predict first

For a fixed sample and significance level, when is a Type II error most likely?

  • When the null is false by only a very small amount
  • When the null is badly false
  • When the null is true
  • Beta does not vary

Correct: When the null is false by a small amount.

Why: A tiny departure produces data almost indistinguishable from what the null predicts, so the test almost never rejects and beta is close to its maximum. A large departure is easy to detect and beta is near zero. The third option is a category error: if the null is true, a Type II error is impossible by definition.

35. Which error costs more

Section

Section 4

36. A judgement the statistics cannot make

Concept

The book presents four examples and asks in each which error has the greater consequence. The answers alternate, and in every case the reasoning turns on what each mistake would cause someone to do rather than on any probability.

consequence, not probability — Alpha and beta say how often each error occurs; they say nothing about what each costs. Deciding which to guard against harder requires knowing the situation, which is why the significance level is chosen rather than derived.

\[ \text{cost} = P(\text{error}) \times \text{consequence} \]

The practical consequence is that the conventional 0.05 is a convention rather than a law. A test whose Type I error would shut down a production line unnecessarily might reasonably use a smaller alpha; a screening test where a missed case is fatal might use a larger one, accepting more false alarms to reduce misses. Neither choice is more statistical than the other.

Figure (svg): A three-column table listing four null hypotheses with which error type is worse in each case and why

Two say Type I and two say Type II, which is exactly why the question has to be asked afresh every time.

OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, pp. 464-466 — Examples 9.5 through 9.8

37. Four examples, two answers

Picture it

The book's cases and which error each identifies as worse.

Figure (svg): A three-column table listing four null hypotheses with which error type is worse in each case and why

Two say Type I and two say Type II, which is exactly why the question has to be asked afresh every time.

Two say Type I and two say Type II, which is as clear a demonstration as could be arranged that there is no general answer. Anyone who learns a rule such as Type I is always worse will get half of these wrong.

38. Worked example: comparing the two consequences

Worked example

Example 9.7, the Genetic Labs claim.

\[ H_0: \text{Genetic Labs has no effect on sex outcome} \]

Describe Type I

Why: Reject a true null.

Describe Type II

Why: Fail to reject a false one.

Ask what follows from each

Why: The actions taken.

Compare

Why: Money spent on nothing.

Figure (svg): The solution to Worked example comparing the two consequences shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{Type I: a useless product is bought} \]

Verify: confirm the comparison would flip in a different setting

Why: If the product were free and harmless, the Type I error would cost almost nothing and the Type II — missing a genuine effect — might matter more to the science. The consequence depends on what follows from the belief, not on the belief itself, which is why the same null hypothesis can have different answers in different circumstances.

OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 465

39. Which error is worse?

Sorting

Ask what each mistake causes someone to do.

Sort into buckets

Sort each situation by which error carries the greater consequence.

Type I is worse
H0: the accident victim is alive; H0: Genetic Labs has no effect
Type II is worse
H0: the climbing equipment is safe; H0: the drug cures at least 75 percent; H0: the water supply is uncontaminated
one
The false alarm leads to the damaging action — withholding treatment, or buying a useless product.
two
The missed detection leads to the damaging action — using unsafe gear, taking a weak drug, or drinking contaminated water.

Item (e) is not in the book but follows the same reasoning as the climbing equipment: a null of safety makes the missed detection the dangerous error. Nulls asserting that something is safe almost always put the weight on the Type II error.

40. Worked example: the opposite verdict

Worked example

Example 9.8, the experimental cancer drug.

\[ H_0: \text{the cure rate is at least } 75 \text{ percent} \]

Describe Type I

Why: Believe the rate is under 75.

\[ \text{when it is at least } 75 \]

Describe Type II

Why: Believe the rate is at least 75.

Ask what follows from Type II

Why: The treatment choice.

Compare

Why: A missed better option.

Figure (svg): The solution to Worked example the opposite verdict shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{Type II: a weak drug is chosen} \]

Verify: confirm why this reverses Example 9.7's answer

Why: In the Genetic Labs case the error that leads to action is the Type I; here the null is a claim of effectiveness, so it is the failure to reject it that leads someone to act. Which error triggers the costly action depends on how the null was written — and since the null could have been written the other way round, the labelling is partly a matter of framing. What does not change is which real-world mistake is worse.

OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 466

41. Trap: looking for a general rule

Trap

The trap

\[ \text{Type I errors are always the more serious} \]

Generalise from the conventional emphasis on alpha

Why: Alpha is the one always reported, so it must matter more.

\[ \text{but two of the book's four examples say otherwise} \]

The rock climber and the cancer patient both face a worse Type II error, for reasons specific to their situations.

The fix

\[ \text{ask what each error would cause someone to DO} \]

Compare consequences case by case

Why: The statistics supply probabilities, not costs.

Alpha gets more attention partly for a practical reason: it is the error a researcher can control directly and report honestly, while beta depends on an unknown effect size. That convenience has nothing to do with which error matters more, and mistaking one for the other is how tests get designed to guard against the wrong failure.

42. One of these is false

Two truths and a lie

All three concern consequences.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Which error is worse depends on the situation
  • C. The judgement rests on consequences rather than probabilities
  • B. The error with the larger probability is the more serious

Survives elimination: B

Why: The survivor is false and confuses how often with how bad. A rare error with catastrophic consequences may matter far more than a common one with trivial ones — which is precisely the case for the climbing equipment, where beta may be small and still dominate.

43. What follows from the judgement?

Prediction

Commit before reasoning.

Predict first

A team decides the Type I error is far more costly. What should they do?

  • Choose a smaller alpha, accepting a larger beta
  • Choose a larger alpha
  • Collect no more data
  • Swap the hypotheses

Correct: Choose a smaller alpha.

Why: Alpha is the one error probability set directly, so guarding harder against a Type I error means lowering it — and since the two trade against each other at a fixed sample size, beta rises as a result. The only way to reduce both is a larger sample, which is the next idea.

44. Explain the alternation

Explain it

A classmate asks why the book does not just say which error type is worse.

Discussion prompt

In two sentences or fewer, explain.

Hint: Ask what each error causes in the climbing case and in the ambulance case.

Answer:

In the climbing case the dangerous mistake is believing unsafe gear is safe, and in the ambulance case it is believing a living victim is dead — one is a Type II and the other a Type I.

Which error leads to the harmful action depends entirely on the situation and on how the null was written, so no general answer exists and the question has to be asked afresh each time.

45. Power, and the trade-off

Section

Section 5

46. One minus beta, and how to buy it

Concept

The power of the test is one minus beta: the probability of correctly rejecting a false null. Ideally we want a high power, as close to one as possible, and increasing the sample size can increase it.

the power of the test — The probability of the fourth outcome — rejecting the null when it is false. It is what a study is usually run to achieve, and it is bought with sample size.

\[ \text{power} = 1 - \beta \]

The trade-off between the two errors is worth seeing clearly. At a fixed sample size, moving the decision boundary reduces one error probability and raises the other, because the two areas sit on different curves. Only a larger sample narrows both curves and reduces both at once — which is why sample size is the lever that matters and why a low-powered study cannot be rescued by adjusting alpha.

Figure (svg): A rising curve showing the power of a test climbing from about a fifth at a sample of five toward one at a sample of a hundred

Drawn for a true effect of half a standard deviation at a five percent significance level: a small study would miss it more often than not.

OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464 — the power of the test

47. Power against sample size

Picture it

For a true effect of half a standard deviation at a five percent level.

Figure (svg): A rising curve showing the power of a test climbing from about a fifth at a sample of five toward one at a sample of a hundred

Drawn for a true effect of half a standard deviation at a five percent significance level: a small study would miss it more often than not.

The curve is steep in the middle and flat at both ends, which has a practical reading. Below about twenty observations the study will usually miss this effect; above about seventy it will usually find it; and the range in between is where an extra handful of observations changes the answer most.

48. Worked example: reading power from beta

Worked example

Converting between the two.

\[ \beta = 0.30 \]

Recall the relationship

Why: Power is one minus beta.

\[ 1 - 0.30 \]

Compute

Why: The power.

\[ 0.70 \]

State what it means

Why: The chance of detecting.

\[ 70 \% \]

Assess it

Why: Below the usual target.

Figure (svg): The solution to Worked example reading power from beta shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{power} = 1 - 0.30 = 0.70 \]

Verify: confirm against the conventional target

Why: Studies are commonly designed for a power of 0.80, so 0.70 falls short — the test would miss the effect three times in ten. Whether that is acceptable depends on the cost of missing it, which is the previous idea's question again. A power of 0.70 also means that a non-significant result from this study is weak evidence of absence, since failure was fairly likely even if the effect is real.

OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464

49. Power from beta

Faded example

A test has a Type II error probability of 0.20.

Fill in the blanks

\text0.20 = 1 - 0.80 = ___

Why: Power of 0.80, which is the level studies are conventionally designed to reach. It means the test would detect the specified effect four times in five.

50. Worked example: the trade-off at a fixed sample

Worked example

What moving the cut-off does.

\[ \text{a fixed } n; \text{ alpha lowered from } 0.05 \text{ to } 0.01 \]

Note the cut-off moves

Why: Further into the tail.

Effect on Type I

Why: Fewer false alarms.

Effect on Type II

Why: More missed detections.

Effect on power

Why: One minus a larger beta.

Figure (svg): The solution to Worked example the trade-off at a fixed sample shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \alpha \downarrow \;\Rightarrow\; \beta \uparrow \;\Rightarrow\; \text{power} \downarrow \]

Verify: confirm why only more data escapes the trade-off

Why: Moving the boundary shifts area between two fixed curves, so what one gains the other loses. A larger sample narrows both curves, pulling each away from the boundary, so both tails shrink together. That is the whole reason sample size is worth paying for, and it is why the answer to a study that finds nothing is usually a bigger study rather than a looser threshold.

OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, p. 464

51. Trap: lowering alpha to make a test more reliable

Trap

The trap

\[ \text{use } \alpha = 0.001 \text{ so the conclusion is more trustworthy} \]

Reduce the error probability to reduce errors

Why: A smaller alpha means fewer mistakes.

\[ \beta \text{ rises, so more real effects are missed} \]

It reduces one kind of mistake by increasing the other, and at a fixed sample size that is all it can do.

The fix

\[ \text{choose } \alpha \text{ from consequences; buy power with } n \]

Set alpha by what a false alarm costs, and set n by what a miss costs

Why: The two levers do different jobs.

An extremely small alpha on a small sample produces a test that almost never rejects anything, true or false — which looks rigorous and is actually uninformative. The honest way to make a study trustworthy is to make it large enough that both error rates are acceptable, and to say what effect size it was powered to detect.

52. One of these is false

Two truths and a lie

All three concern power and the trade-off.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Power is one minus beta
  • C. A larger sample reduces both error probabilities
  • B. Lowering alpha at a fixed sample also lowers beta

Survives elimination: B

Why: The survivor is false. At a fixed sample size the two trade against each other — moving the cut-off to reduce false alarms necessarily increases missed detections. Only more data reduces both, which is why sample size rather than significance level is the lever that improves a study.

53. Is this study worth running?

Estimation

A trial is calculated to have power 0.30 for the smallest effect worth detecting.

Predict first

What does that mean for the study?

  • It will miss a real effect about seven times in ten
  • It will find a real effect about seven times in ten
  • It has a 30 percent false alarm rate
  • It is adequately powered

Correct: It will miss a real effect about seven times in ten.

Why: Power of 0.30 means beta is 0.70, so the effect goes undetected in seventy percent of such studies even when it is genuinely there. Running it risks producing a non-significant result that tells nobody anything, which is why power calculations belong before data collection rather than after.

54. What buys power?

Prediction

Commit before reasoning.

Predict first

Which change increases the power of a test without increasing the Type I error rate?

  • A larger sample
  • A smaller alpha
  • A smaller effect size
  • A more extreme cut-off

Correct: A larger sample.

Why: More data narrows the sampling distribution under both hypotheses, pulling both tails away from the boundary — so beta falls and power rises while alpha stays exactly where it was set. Every other option either trades one error for the other or makes the effect harder to detect.

55. The two errors side by side

Comparison

Fill the blanks. They are conditional on opposite things, which is why they never sum to anything.

Comparison matrix

Type IType II
The mistakerejecting a true nullfailing to reject a false null
Probabilityalphabeta
Set or derived?chosen directly by the researcherfollows from n, alpha and the effect size
Everyday namea false positive, or false alarma false negative, or missed detection

The third row explains why alpha is always reported and beta rarely is: one is a decision and the other depends on an unknown. That asymmetry in reporting has nothing to do with which error matters more.

56. Describing the errors for a given null, in order

Pattern

Five steps, and the fifth is not a calculation.

  1. Write the null hypothesis as a plain sentence.
  2. For the Type I error, write: concluding the null is false when in fact it is true, in the words of the problem.
  3. For the Type II error, write: concluding the null cannot be rejected when in fact it is false, in the same words.
  4. Attach alpha to the first and beta to the second.
  5. Ask what action each mistake would lead to, and compare those consequences to say which is worse.

Every error description needs a when in fact clause. A sentence stating only a conclusion describes a decision, not an error, and cannot be assigned a type.

OpenStax Introductory Business Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2 Outcomes and the Type I and Type II Errors

57. Check yourself 1 of 3

Check

Naming the error.

Check your understanding

A test rejects a null hypothesis that is actually true. What is this?

  • A. A Type I error, with probability alpha (correct)
  • B. A Type II error, with probability beta
  • C. A correct decision
  • D. The power of the test

Answer: A

Why: Rejecting a true null is the Type I error, and its probability is alpha — the significance level chosen for the test.

Why B tempts people
A Type II error is failing to reject a null that is false, which is the opposite mismatch.
Why C tempts people
The decision disagrees with the truth, so it is an error.
Why D tempts people
The power is correctly rejecting a FALSE null, which is the other cell in that row.

58. Check yourself 2 of 3

Check

The trade-off.

Check your understanding

At a fixed sample size, what happens to beta if alpha is reduced?

  • A. Beta increases (correct)
  • B. Beta decreases
  • C. Beta is unchanged
  • D. Beta becomes zero

Answer: A

Why: The two areas sit on different curves, so moving the cut-off to shrink one enlarges the other. Only a larger sample reduces both.

Why B tempts people
That would require both tails to shrink at once, which needs more data rather than a moved boundary.
Why C tempts people
The cut-off's position affects both error probabilities, so beta must move.
Why D tempts people
Neither error probability is ever zero for a test that can reject at all.

59. Check yourself 3 of 3

Check

Which error is worse.

Check your understanding

For the null hypothesis that a water supply is uncontaminated, which error is more serious?

  • A. Type II: concluding it is clean when it is contaminated (correct)
  • B. Type I: concluding it is contaminated when it is clean
  • C. They are equally serious
  • D. It depends on alpha

Answer: A

Why: The Type II error leads people to drink contaminated water, while the Type I error leads to unnecessary treatment or testing.

Why B tempts people
That is costly but far less so: it wastes resources rather than causing illness.
Why C tempts people
The consequences are plainly asymmetric here, as they are in most safety settings.
Why D tempts people
Alpha is a probability and says nothing about consequences.

60. Where this shows up outside the textbook

Real world

An airport installs a scanner that flags bags for manual search. Its null hypothesis is that a bag is harmless. The vendor advertises a false alarm rate of only 0.5 percent, and the security manager proposes tightening the threshold further to reduce the queues that searches cause.

Discussion prompt

Analyse the proposal in terms of the two errors, and say what the manager should consider instead.

Hint: The advertised figure describes only one of the two errors.

Answer:

The advertised 0.5 percent is alpha alone. It is the probability that a harmless bag is flagged — the false alarm rate, and the one that generates queues. The vendor's figure says nothing whatever about the other error: the probability that a dangerous bag is passed, which is beta.

Tightening the threshold reduces alpha and increases beta. Fewer harmless bags would be searched, shortening queues, and more dangerous bags would go through undetected. At a fixed scanner and a fixed inspection process, that trade is unavoidable — the two error rates sit on different distributions and only move in opposite directions.

\[ \text{threshold} \uparrow \;\Rightarrow\; \alpha \downarrow, \; \beta \uparrow, \; \text{power} \downarrow \]

The consequences here are wildly asymmetric, which settles the direction. A Type I error costs a passenger a few minutes; a Type II error is a weapon on an aircraft. Security systems are therefore deliberately set to a high false alarm rate — the opposite of what the manager proposes — and the queue is the accepted price of a small beta.

What the manager should pursue instead is anything that improves both errors at once, which means better information rather than a different threshold: a more sensitive scanner, a second independent screening stage, or more staff so that a given alarm rate produces shorter queues. This is the same conclusion as the lesson's trade-off figure — moving the boundary redistributes the errors, and only more or better data reduces both. It is also worth noting that the false alarm rate alone is an incomplete specification for any detection system, and a vendor quoting it without a detection rate has described half the product.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

At a fixed sample size, why can alpha and beta not both be made small?

  • Because they sum to one
  • Because the two areas sit on different distributions, so moving the cut-off shifts area from one to the other
  • Because sample sizes are always too small
  • Because alpha is chosen and beta is not

Correct: Because the two areas sit on different distributions.

\[ \alpha = P(\text{reject} \mid H_0), \qquad \beta = P(\text{not reject} \mid H_a) \]

Why: Alpha is a tail area under the null's distribution and beta is an area under the alternative's, on opposite sides of the same boundary. Moving the boundary therefore shrinks one and grows the other. They do not sum to one, since they are conditional on different things — and the only way to reduce both is to narrow both distributions, which means more data.

62. Explain it to someone a year behind you

Explain it

They wrote: the Type II error is that Navah thinks her equipment is unsafe.

Discussion prompt

In two sentences or fewer, correct it.

Hint: Ask what the equipment is actually like in their sentence.

Answer:

Their sentence gives a conclusion but never says what the truth is, so it does not describe an error of either type — and thinking the gear is unsafe is a rejection, which points at Type I anyway.

The Type II error is thinking the equipment may be safe when in fact it is not, which is the failure to reject a false null and the one that gets her hurt.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Placing an outcome correctly in the four-cell grid
  • Writing an error description with both halves of the sentence
  • Judging which error carries the greater consequence
  • Explaining the trade-off, and what power costs

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the first, cross the decision with the truth. For the second, always include a when in fact clause. For the third, ask what each mistake makes someone do. For the fourth, moving the cut-off trades one error for the other and only more data shrinks both. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Fifteen minutes.

Draw it

At the top, draw the two-by-two grid with the decision down the side and the truth across the top, and fill all four cells with their names and probabilities, marking which cell is the power. Beneath it, draw two overlapping bell curves with a vertical cut-off between them, shade the right tail of the left curve for alpha and the left region of the right curve for beta, and write one sentence on what happens to each area when the cut-off moves right. Then write a second sentence on what happens to both when the sample grows. In the middle of the page, take the book's four examples in turn — the climbing equipment, the accident victim, Genetic Labs and the cancer drug — and for each write the null as a sentence, then the Type I error and the Type II error each with a when in fact clause, then which is worse and what action makes it so. At the bottom, write power equals one minus beta, and sketch a rising curve of power against sample size, marking roughly where a study would go from usually missing an effect to usually finding it.

Check your four examples by confirming that two say Type I is worse and two say Type II — if all four agree, one of the descriptions has probably been swapped. Check the two shaded areas by confirming they sit on different curves, since that is what makes the trade-off unavoidable.

65. What you can do now

Recap

Five things, and the third is a judgement rather than a calculation.

If you seeThen
Rejecting a true nullA Type I error, probability alpha
Failing to reject a false nullA Type II error, probability beta
A description with no when-in-fact clauseA decision, not an error
A null asserting something is safeThe Type II error is usually the dangerous one
A request to lower alpha at fixed nBeta will rise; power will fall
A study that found nothingAsk what its power was
A wish to reduce both errorsOnly a larger sample does that

Section 9.3 turns from what can go wrong to what machinery to use. Three tests have appeared so far — a mean with sigma known, a mean with sigma unknown, and a proportion — and each carries its own distribution and its own assumptions, which the next section sets out in a single table.

OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors §9.2, pp. 464-466 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 464-466
  2. OpenStax Introductory Business Statistics 2e, §9.2 Outcomes and the Type I and Type II Errors — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108