9.3 Probability Distribution Needed for Hypothesis Testing

A short section that does one job: it says which probability distribution each of the three one-sample tests uses, and what has to be true for that choice to be legitimate. A test of a mean when the population standard deviation is known uses a normal distribution; a test of a mean when it is unknown and estimated by the sample standard deviation uses a Student t; and a test of a proportion uses a normal distribution built on the binomial. The three rows are chapter 8's three confidence intervals in the same order and for the same reasons. What the section adds is an explicit statement of the assumptions behind each: a simple random sample in every case, approximate normality of the population for the two mean tests, and for a proportion the binomial conditions together with the requirement that n times p and n times q both exceed five, so that the binomial is close enough in shape to the normal for the approximation to hold.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 9.3 Probability Distribution Needed for Hypothesis Testing

Title

Statistics · Chapter 9 — Hypothesis Testing with One Sample

Probability Distribution Needed for Hypothesis Testing

2. By the end of this lesson you can

Objectives

Five outcomes, and the last is the condition the earlier chapters left implicit.

OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, pp. 466-467 — the section these objectives are drawn from

3. What you already have

Warm-up

Chapter 8 built three confidence intervals: a z for a mean with sigma known, a t for a mean without it, and a z for a proportion.

Discussion prompt

A hypothesis test asks a different question from a confidence interval. Does it need a different set of distributions?

Hint: What does each procedure actually compute an area under?

Answer:

No. A confidence interval starts from a point estimate and works outward to find the values consistent with it; a test starts from a claimed value and works outward to see whether the point estimate is consistent with THAT. Both operate on the same sampling distribution.

So the three rows of this section's table are exactly chapter 8's three cases, and the deciding questions are the same: is the parameter a mean or a proportion, and if a mean, is sigma known?

What this section adds is the assumptions, stated in one place for the first time — including one condition for proportions that chapter 8 never spelled out, and which turns out to matter.

4. Three tests, three distributions, one deciding question each

Concept

Particular distributions are associated with various types of hypothesis testing. A test for a mean with the population standard deviation known uses the normal distribution; a test for a mean with it unknown uses the Student t; and a test for proportions uses the normal distribution.

the three one-sample tests — A mean with sigma known, a mean with sigma unknown, and a proportion. The parameter and the point estimate differ across them, and the distribution follows from the pair.

\[ \bar{x} \to N \text{ or } t; \qquad p' \to N \]

It is worth noticing that two of the three rows use a normal distribution, and for quite different reasons. The first uses it because sigma is known, so nothing beyond the mean is being estimated; the third uses it because a proportion's spread follows from its own point estimate, so again nothing extra is estimated. Only the middle row estimates a second quantity, and only it needs a t.

Figure (svg): A four-column table listing the three one-sample tests with their parameters, point estimates and distributions

Two rows use a normal distribution and one uses a t, and the deciding question is the same one section 8.2 asked.

OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 466

5. The three rows

Section

Section 1

6. Parameter, point estimate, distribution

Concept

Each row of the book's table names three things: the population parameter being tested, the sample statistic that estimates it, and the probability distribution used to compute the test. The first two determine the third.

the three rows — Mean with sigma known, estimated by the sample mean, using a normal. Mean with sigma unknown, estimated by the sample mean, using a t. Proportion, estimated by the sample proportion, using a normal.

\[ (\mu, \bar{x}, N); \quad (\mu, \bar{x}, t); \quad (p, p', N) \]

The table is worth learning as a whole rather than as three separate facts, because the questions that select a row are asked in a fixed order. First, is the parameter a mean or a proportion — that settles rows one and two against row three. Then, for a mean, is sigma known — that settles row one against row two. Two questions, three answers, and no other information is needed.

Figure (svg): A four-column table listing the three one-sample tests with their parameters, point estimates and distributions

Two rows use a normal distribution and one uses a t, and the deciding question is the same one section 8.2 asked.

OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 466 — the table of tests and distributions

7. The table

Picture it

Three tests, with what each uses.

Figure (svg): A four-column table listing the three one-sample tests with their parameters, point estimates and distributions

Two rows use a normal distribution and one uses a t, and the deciding question is the same one section 8.2 asked.

The parameter column contains only mu and p, and the point estimate column only x-bar and p prime — which is a reminder from section 9.1 that hypotheses concern parameters while the data supply statistics. Every test in this chapter and the next has that same two-column structure.

8. Worked example: selecting a row

Worked example

Three problems, asked in the fixed order.

\[ \text{(a) } \sigma = 0.8 \text{ given}; \; \text{(b) only data}; \; \text{(c) } 53 \text{ of } 100 \]

Case (a)

Why: A mean, sigma known.

Case (b)

Why: A mean, sigma unknown.

Case (c)

Why: A proportion.

Note the order of questions

Why: Parameter, then sigma.

Figure (svg): The solution to Worked example selecting a row shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ N, \quad t, \quad N \]

Verify: confirm the two questions are enough

Why: Nothing else was needed: not the sample size, not the confidence level, not the direction of the alternative. Sample size affects the degrees of freedom in case (b) and the validity check in case (c), but it never changes which row applies. Keeping the selection to two questions makes it hard to get wrong.

OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 466

9. Which distribution?

Sorting

Ask about the parameter first, then about sigma.

Sort into buckets

Sort each test.

Normal
a mean, with sigma = 0.8 given; a proportion, 53 of 100; a proportion, 421 of 500
Student t
a mean, from ten data values only; a mean, with s computed from the sample
n
Either sigma is known, or the parameter is a proportion — in both cases only one quantity is estimated.
t
The parameter is a mean and its spread is estimated from the sample.

Both proportion problems use a normal regardless of their sample sizes, which is the distinction the previous trap turns on. Sample size matters for proportions through the np and nq check rather than through the choice of curve.

10. Worked example: why two rows share a distribution

Worked example

Rows one and three both use a normal, for different reasons.

\[ \text{row 1: } \sigma \text{ known}; \quad \text{row 3: a proportion} \]

Row one

Why: Sigma is given.

Row three

Why: The spread comes from p prime.

Row two

Why: s estimates sigma.

Conclude

Why: Only row two needs a t.

Figure (svg): The solution to Worked example why two rows share a distribution shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{one estimate} \to N; \quad \text{two estimates} \to t \]

Verify: confirm this matches section 8.2's account

Why: Section 8.2 introduced the t precisely because s varies from sample to sample, adding a second source of uncertainty that a z cannot account for. A proportion has no such second source, since the binomial's spread is determined by its probability — which is why no t distribution appears anywhere in the proportion row, however small the sample.

OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, pp. 466-467

11. Trap: using a t because the proportion sample is small

Trap

The trap

\[ n = 25 \text{ for a proportion} \;\Rightarrow\; \text{use } t_{24} \]

Apply section 8.2's small-sample rule

Why: A t handles small samples.

\[ \text{but nothing beyond } p' \text{ is being estimated} \]

The t corrects for an estimated standard deviation, and a proportion has none to estimate.

The fix

\[ \text{use a normal, and check } np > 5 \text{ and } nq > 5 \]

Use a normal, with the binomial condition as the small-sample safeguard

Why: That is what the third row says.

Small samples are a real problem for proportions, but the remedy is different: the np and nq check, and if it fails, the plus-four adjustment of section 8.3 or an exact binomial method. Reaching for a t is a reasonable instinct applied to the wrong row.

12. The point estimate

Fill the middle

For a hypothesis test about a population proportion.

Fill in the blanks

\textproportion ___

Why: The sample proportion, p prime — the successes over the trials. It plays the same role for a proportion test that the sample mean plays for a mean test.

13. One of these is false

Two truths and a lie

All three concern the table.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Two of the three rows use a normal distribution
  • C. Only two questions are needed to select a row
  • B. The sample size determines which row applies

Survives elimination: B

Why: The survivor is false. The sample size affects the degrees of freedom for a t and the validity check for a proportion, but it never selects the row. Section 8.2 established that the deciding question for means is whether sigma is known, not how much data there is.

14. Why only one t?

Prediction

Commit before reasoning.

Predict first

Why does only one of the three rows use a Student t?

  • Because only that row estimates a second quantity, the standard deviation
  • Because only that row has a small sample
  • Because only that row is about a mean
  • Because the t is harder to compute

Correct: Because only that row estimates a second quantity.

Why: The t exists to account for the extra uncertainty in estimating sigma with s. The sigma-known row does not estimate it, and the proportion row gets its spread free from p prime — so neither needs the correction. The third option is close but not enough: both mean rows are about a mean, and only one uses a t.

15. Assumptions for the mean tests

Section

Section 2

16. A random sample, and approximate normality

Concept

For a z-test of a single mean you take a simple random sample, the population is normally distributed or the sample is sufficiently large, and you know sigma — which, the book adds, in reality is rarely known. For a t-test the data should be a simple random sample from an approximately normal population, and s approximates sigma.

sufficiently large — The condition that lets chapter 7's central limit theorem supply normality for the sample mean when the population is not normal. The book notes a t-test will work even if the population is not approximately normal, provided the sample is sufficiently large.

\[ \text{random sample} \;+\; \text{normal population OR large } n \]

The book's parenthetical about sigma being rarely known in reality is the reason the second row is the ordinary case and the first is nearly a textbook curiosity. Quality control with a long-established process is one of the few settings where a population standard deviation genuinely is known; almost everywhere else it has to be estimated, and the t applies.

Figure (svg): The assumptions each of the three tests requires

The z-test's last requirement is the one that is rarely met in practice, which is why the t-test is the ordinary case.

OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 467 — the assumptions for the z-test and the t-test

17. What each test assumes

Picture it

Five conditions across the three tests.

Figure (svg): The assumptions each of the three tests requires

The z-test's last requirement is the one that is rarely met in practice, which is why the t-test is the ordinary case.

The first line applies to all three and is the one no amount of data repairs. Section 8.2 made the same point: a large biased sample estimates the wrong quantity more precisely, which is worse than being imprecise about the right one.

18. Worked example: checking the assumptions

Worked example

A t-test on fifteen measurements.

\[ n = 15, \; \sigma \text{ unknown, population believed normal} \]

Random sample?

Why: Must be established.

Population normal?

Why: Believed so.

\[ \text{matters at } n = 15 \]

Sigma known?

Why: No.

Choose the distribution

Why: By the second row.

\[ t\text{ with } 14\text{ df} \]

Figure (svg): The solution to Worked example checking the assumptions shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ T \sim t_{14} \]

Verify: confirm which assumption a larger sample would relieve

Why: The book says a t-test will work even if the population is not approximately normal when the sample is sufficiently large, so normality is the relieved condition. Randomness is not: it is a separate assumption, and the book states it separately for exactly that reason. At n equal to 15 the normality assumption is doing real work and should be checked against the data's shape.

OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 467

19. Which assumption is threatened?

Discrimination

Each situation threatens one condition.

Sort into buckets

Sort each by which assumption it puts at risk.

Normality
a sample of 8 from a strongly skewed population; a sample of 12 with two extreme outliers
Random sampling
volunteers recruited by advertisement; only subjects who completed the study; the first 20 patients on an alphabetical list
norm
The population's shape, or evidence of it in a small sample, is far from bell shaped.
rand
The selection mechanism systematically favours part of the population.

20. Worked example: when normality stops mattering

Worked example

The same test on a much larger sample.

\[ n = 200, \; \text{population plainly skewed} \]

Note the population's shape

Why: Skewed.

Recall the theorem

Why: Chapter 7.

Check the sample size

Why: Two hundred.

Conclude

Why: The test is valid.

Figure (svg): The solution to Worked example when normality stops mattering shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ n \text{ large} \;\Rightarrow\; \bar{X} \text{ approximately normal} \]

Verify: confirm which quantity the normality assumption is really about

Why: The assumption is not that the data are normal but that the SAMPLE MEAN's distribution is, and chapter 7 supplies that for a large sample whatever the population looks like. For a small sample there is no such rescue, so the population's own shape has to carry it — which is why the assumption is stated in terms of the population but only bites when n is small.

OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 467

21. Error analysis: four claims about the assumptions

Error analysis

Which correctly describe what the mean tests require?

Annotate

On: \( \begin{aligned} &(1)\; \text{a large sample repairs a biased selection} \\ &(2)\; \text{normality is needed regardless of } n \\ &(3)\; \text{a } z\text{-test may be used whenever } n > 30 \\ &(4)\; \text{normality matters most when } n \text{ is small} \end{aligned} \)

  • (1) is false and is the most consequential error here: bias is never repaired by sample size.
  • (2) is false: a sufficiently large sample makes the t-test work on a non-normal population.
  • (3) is the outdated rule section 8.2 replaced. The deciding question is whether sigma is known.
  • (4) is correct: without a large sample there is no central limit theorem to supply normality.

Errors (2) and (4) are opposite readings of the same condition, and the difference matters at the design stage — a small study needs its normality assumption checked, and a large one mostly does not.

22. One of these is false

Two truths and a lie

All three concern the mean tests.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Sigma is rarely known in reality
  • C. A large sample relieves the normality assumption
  • B. Random sampling follows from a large sample size

Survives elimination: B

Why: The survivor is false. Randomness is a property of how the sample was selected, and no amount of additional data changes a biased selection mechanism — it simply pins down the biased quantity more tightly.

23. What is the assumption really about?

Prediction

Commit before reasoning.

Predict first

The t-test assumes an approximately normal population. Which quantity's distribution does this actually secure?

  • The distribution of the sample mean
  • The distribution of the individual observations
  • The distribution of the population parameter
  • The distribution of alpha

Correct: The distribution of the sample mean.

Why: The test statistic is built from the sample mean, so it is the mean's sampling distribution that must be normal. The population's shape is a way of securing that when the sample is small; when the sample is large, chapter 7 secures it directly and the population's shape stops mattering.

24. Choose the test

Faded example

A sample of 40 measurements, with s computed from the data and the population moderately skewed.

Fill in the blanks

\textt 39 \text___ ___ \text___

Why: A t with 39 degrees of freedom. The moderate skew is not a problem at n equal to 40, since the central limit theorem makes the sample mean approximately normal.

25. Assumptions for the proportion test

Section

Section 3

26. Binomial conditions, and a shape requirement

Concept

For a test of a single proportion you take a simple random sample and must meet the conditions for a binomial distribution: a certain number n of independent trials, outcomes of success or failure, and the same probability of success on each trial.

the binomial conditions — A fixed number of independent trials, two outcomes per trial, and a constant success probability — the three characteristics of section 4.3, restated as requirements for the test.

\[ X \sim B(n, p): \; n \text{ fixed}, \; \text{independent}, \; p \text{ constant} \]

These are section 4.3's three characteristics used as a checklist. The independence condition is the one that fails most often in practice — surveys of households, students in the same class, or patients at the same clinic all produce responses that are correlated rather than independent, and a test that assumes otherwise will understate its uncertainty.

Figure (svg): The assumptions each of the three tests requires

The z-test's last requirement is the one that is rarely met in practice, which is why the t-test is the ordinary case.

OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 467 — the conditions for a proportion test

27. The conditions in context

Picture it

The proportion test's requirements among the others.

Figure (svg): The assumptions each of the three tests requires

The z-test's last requirement is the one that is rarely met in practice, which is why the t-test is the ordinary case.

The last two lines both belong to the proportion test, and they do different jobs: the binomial conditions make the model right, and the np and nq condition makes the normal approximation to that model adequate. Both have to hold, and the next idea is about the second.

28. Worked example: checking the binomial conditions

Worked example

A survey of 100 first-time brides.

\[ 100 \text{ brides surveyed}; \; 53 \text{ are younger than their grooms} \]

A fixed number of trials

Why: One hundred surveyed.

\[ n = 100 \]

Two outcomes each

Why: Younger, or not.

The same probability each

Why: If randomly selected.

Independent

Why: One answer does not affect another.

Figure (svg): The solution to Worked example checking the binomial conditions shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ X \sim B(100, p) \]

Verify: confirm which condition is most at risk in survey work

Why: Independence is the fragile one. If the sample had been drawn from a single social circle, or if couples influenced one another's answers, the responses would be correlated and the effective information would be less than 100 independent observations suggests. The test would then report more certainty than the data supports — an error that no calculation inside the test can detect.

OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 467

29. Do the binomial conditions hold?

Sorting

Check for a fixed n, two outcomes, constant p and independence.

Sort into buckets

Sort each survey.

Conditions plausibly hold
500 randomly selected adults, each asked once; 100 independently sampled brides; 50 randomly selected patients from a national register
Independence is doubtful
all members of 20 randomly chosen households; every student in three randomly chosen classes
ok
Each respondent is selected independently, so their answers are plausibly independent too.
no
Respondents are grouped, and people in a group tend to resemble one another.

The two doubtful cases share a shape: sampling clusters and then taking everyone inside them. It is a common and often sensible design for practical reasons, but it needs methods that account for the clustering rather than the plain binomial test.

30. Worked example: a case where independence fails

Worked example

A survey design that breaks the model.

\[ \text{20 households surveyed, all members of each asked} \]

Count the responses

Why: Perhaps 60 people.

\[ n\text{ looks like } 60 \]

Ask about independence

Why: Household members agree.

Say what breaks

Why: The independence condition.

\[ \text{not } 60\text{ independent trials} \]

Note the consequence

Why: Overstated precision.

Figure (svg): The solution to Worked example a case where independence fails shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{correlated responses} \;\Rightarrow\; \text{effective } n < 60 \]

Verify: confirm the direction of the error this causes

Why: Correlated observations carry less information than independent ones, so the true standard error is larger than the formula gives — which makes the interval too narrow and the p-value too small. Every consequence runs toward overstating the evidence, which is why the independence condition is worth checking against the survey design rather than assumed.

OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 467

31. Trap: counting responses instead of independent trials

Trap

The trap

\[ 60 \text{ responses from } 20 \text{ households} \;\Rightarrow\; n = 60 \]

Count every answer collected

Why: Each is a data point.

\[ \text{but members of a household are not independent} \]

The effective number of independent observations is closer to 20 than to 60.

The fix

\[ \text{ask how many INDEPENDENT trials there were} \]

Count independent units, not responses

Why: Independence is one of the three binomial conditions.

This is a design problem rather than an arithmetic one, and it cannot be fixed at the analysis stage by any adjustment in this book — methods for clustered data exist but are beyond it. The practical lesson is that the conditions have to be checked when the survey is planned, because by the time the data arrive the damage is done.

32. One of these is false

Two truths and a lie

All three concern the proportion test's conditions.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The trials must be independent
  • C. Each trial must have the same probability of success
  • B. Correlated responses make the test conservative

Survives elimination: B

Why: The survivor is false and has the direction backwards. Correlated responses carry less information than the count suggests, so the computed standard error is too small — making intervals too narrow and p-values too small. The error overstates the evidence rather than being cautious about it.

33. What does clustering do?

Prediction

Commit before reasoning.

Predict first

A survey samples whole households rather than individuals. What happens to the true uncertainty?

  • It is larger than the formula reports
  • It is smaller than the formula reports
  • It is unchanged
  • It cannot be assessed

Correct: Larger than the formula reports.

Why: Members of a household give similar answers, so sixty responses from twenty households carry less information than sixty independent ones. The formula uses the raw count and therefore understates the real variability — reporting more precision than the design supports.

34. The three conditions

Fill the middle

What a proportion test requires of its data.

Fill in the blanks

n \textprobability of success ___ \text___

Why: The same probability of success on every trial — section 4.3's third characteristic, restated here as a requirement for the test to be valid.

35. The np and nq condition

Section

Section 4

36. Both products above five

Concept

The shape of the binomial distribution needs to be similar to the shape of the normal distribution. To ensure this, the quantities np and nq must both be greater than five. Then the binomial distribution of a sample proportion can be approximated by the normal distribution.

np and nq greater than five — The condition that makes the normal approximation to the binomial adequate. Both products are required, since a proportion near zero fails on np and one near one fails on nq.

\[ np > 5 \quad\text{and}\quad nq > 5 \]

The reason both are needed is symmetry. A binomial with a small np is bunched against zero and strongly right-skewed; one with a small nq is bunched against n and left-skewed. Either way the distribution is lopsided and bounded on one side, while a normal is symmetric and unbounded — so the approximation fails at exactly the place a test would be looking, which is the tail.

Figure (svg): Two binomial stem plots side by side, the left one strongly skewed with a small np and the right one symmetric with a larger np

The condition is what lets a discrete, bounded binomial be approximated by a continuous, unbounded normal.

OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 467 — the np and nq requirement

37. Why both products matter

Picture it

Two binomials with the same n and very different shapes.

Figure (svg): Two binomial stem plots side by side, the left one strongly skewed with a small np and the right one symmetric with a larger np

The condition is what lets a discrete, bounded binomial be approximated by a continuous, unbounded normal.

The left distribution has np equal to 2 and is plainly skewed, with a hard floor at zero that no normal curve respects. The right has np equal to 10 and is close to symmetric. Fitting a normal to the left one would misstate both tails, and the right tail is usually where a p-value is computed.

38. Worked example: checking the condition

Worked example

The brides survey of Example 9.17.

\[ n = 100, \; p = 0.50 \text{ from } H_0 \]

Compute np

Why: One hundred times a half.

\[ 50 \]

Compute nq

Why: The same.

\[ 50 \]

Compare with five

Why: Both far above.

Conclude

Why: Use the normal.

Figure (svg): The solution to Worked example checking the condition shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ np = nq = 50 > 5 \]

Verify: confirm which p is used in the check

Why: The p used is the one from the null hypothesis, not the sample proportion — because the test computes everything under the assumption that the null is true. Using p prime instead would usually give a similar answer here, but for a proportion near the boundary the two checks can disagree, and the null's value is the one the test actually needs.

OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 467

39. Check the condition

Faded example

A test of p = 0.30 on a sample of 50.

Fill in the blanks

np = 15, \quad nq = 35, \quad \text___

Why: Both 15 and 35 exceed five comfortably, so the normal approximation to the binomial is adequate and the test may proceed.

40. Worked example: a case where the condition fails

Worked example

A small sample with a proportion near zero.

\[ n = 40, \; p = 0.05 \text{ from } H_0 \]

Compute np

Why: Forty times 0.05.

\[ 2 \]

Compare with five

Why: Below it.

Compute nq

Why: Forty times 0.95.

\[ 38 \]

Conclude

Why: One product fails.

Figure (svg): The solution to Worked example a case where the condition fails shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ np = 2 \not> 5 \]

Verify: confirm what the failure looks like in practice

Why: Section 8.3's transfer problem showed the symptom: an interval running to a negative endpoint, which is impossible for a proportion. The same underlying failure affects a test, where the p-value computed from a normal can be badly wrong in the tail. The remedies are the plus-four adjustment for an interval, or an exact binomial calculation for a test.

OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 467

41. Trap: checking only one of the two products

Trap

The trap

\[ n = 40, \; p = 0.95: \quad nq = 2, \text{ but } np = 38 > 5 \]

Check np and stop, since it comfortably passes

Why: The condition names np first.

\[ nq = 2, \text{ so the distribution is skewed the other way} \]

A proportion near one is bunched against the upper limit, and a normal fits it no better than it fits one bunched against zero.

The fix

\[ \text{check BOTH: } np > 5 \text{ and } nq > 5 \]

Compute both products every time

Why: The condition is symmetric in p and q.

Checking only the first product passes every proportion above about 0.125 at n equal to 40, including 0.99 — where the approximation is hopeless. The two products are equally binding, and computing both costs one extra multiplication.

42. Does the condition hold?

Discrimination

Compute both products for each.

Sort into buckets

Sort each setup.

Both products exceed five
n = 100, p = 0.50; n = 50, p = 0.30; n = 500, p = 0.02
At least one fails
n = 40, p = 0.05; n = 30, p = 0.97
ok
Both np and nq are above five, so the binomial is close enough to symmetric.
no
One product is at or below five, so the distribution is bunched against a boundary.

43. One of these is false

Two truths and a lie

All three concern the condition.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Both products must exceed five
  • C. The p used is the one from the null hypothesis
  • B. A proportion below 0.1 can never satisfy the condition

Survives elimination: B

Why: The survivor is false. A proportion of 0.02 with n equal to 500 gives np of 10, which passes comfortably. The condition constrains the PRODUCT rather than the proportion, so any p can satisfy it given a large enough sample.

44. How large a sample?

Estimation

A null hypothesis states p = 0.04.

Predict first

Roughly what sample size does the condition require?

  • At least about 125
  • At least about 25
  • At least about 5
  • Any size

Correct: At least about 125.

Why: np exceeds five when n exceeds 5 over 0.04, which is 125; nq is then comfortably large. Small proportions demand large samples for the approximation to hold, which is why studies of rare events need far more data than their headline rates suggest.

45. Choosing correctly, and what goes wrong

Section

Section 5

46. The consequences of picking the wrong row

Concept

Selecting the wrong distribution does not produce an obviously wrong answer. It produces a p-value that is systematically too small or too large, with nothing in the output to indicate the error.

the direction of the error — Using a z where a t is needed makes the p-value too small and the evidence look stronger than it is. Using a normal where the binomial is too skewed distorts the tail, which is exactly where the p-value is computed.

\[ t_{df} > z \;\Longrightarrow\; \text{a z gives too small a } p\text{-value} \]

That the errors run one way is worth dwelling on. A z used in place of a t always understates the p-value, so it always makes results look more significant than they are — the same one-sided bias section 8.2 found for confidence intervals, where a z interval was always too narrow. Systematic errors that flatter the researcher are the ones worth guarding against hardest.

Figure (svg): A standard normal curve drawn over a t curve with few degrees of freedom, the t sitting lower in the centre and higher in the tails

Choosing the wrong one shifts every p-value, and always in the direction that overstates the evidence.

OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, pp. 466-467 — the table and the assumptions together

47. The two curves

Picture it

A standard normal and a t with four degrees of freedom.

Figure (svg): A standard normal curve drawn over a t curve with few degrees of freedom, the t sitting lower in the centre and higher in the tails

Choosing the wrong one shifts every p-value, and always in the direction that overstates the evidence.

The gap is largest in the tails, which is precisely where a p-value is computed — so the choice between them matters most exactly where it is being used. At four degrees of freedom the two-tailed p-value for a statistic of 2.5 is about 0.067 from the t and 0.012 from the normal, which straddles the conventional threshold.

48. Worked example: the cost of the wrong curve

Worked example

The same test statistic read from both distributions.

\[ t = 2.5, \; \text{df} = 4 \]

From the t

Why: Two-tailed, 4 df.

\[ \text{about } 0.067 \]

From the normal

Why: Two-tailed.

\[ \text{about } 0.012 \]

Compare with 0.05

Why: They straddle it.

Say which is right

Why: Sigma was estimated.

Figure (svg): The solution to Worked example the cost of the wrong curve shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ p_t \approx 0.067 \quad\text{against}\quad p_z \approx 0.012 \]

Verify: confirm the error always runs in the same direction

Why: The t has heavier tails, so the same statistic always leaves more area beyond it under a t than under a normal — the z p-value is smaller every time, never larger. That means the mistake always makes the evidence look stronger, and it is largest for small samples where the evidence is weakest to begin with.

OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, pp. 466-467

49. Which row, and is it valid?

Sorting

Select the distribution and check the conditions.

Sort into buckets

Sort each by whether the standard test applies.

The standard test applies
a mean, sigma known, n = 40, population normal; a mean, s from the data, n = 60; a proportion, n = 500, null p = 0.30
A condition fails
a proportion, n = 40, null p = 0.05; a proportion from all members of 15 households
ok
The right row applies and every condition it requires is satisfied.
no
Either np is below five, or the trials are not independent.

The two failures need different remedies: (b) needs an exact binomial method or the plus-four adjustment, while (d) needs a design that accounts for clustering. Neither is repaired by choosing a different curve from the table.

50. Worked example: a full selection

Worked example

Working through the two questions and the checks.

\[ \text{a claim about a proportion}; \; n = 100, \; p_0 = 0.50 \]

Question one

Why: Mean or proportion?

Select the row

Why: The third.

Check the binomial conditions

Why: Independent, two outcomes.

Check np and nq

Why: Both 50.

Figure (svg): The solution to Worked example a full selection shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P' \sim N\left(p, \sqrt{\tfrac{pq}{n}}\right) \]

Verify: confirm nothing about the sample size changed the row

Why: The sample of 100 entered only the np and nq check, never the choice of curve — a proportion test uses a normal at n equal to 20 and at n equal to 2,000 alike. What changes with n is whether the approximation is good enough, and that is what the products test. Keeping the two questions separate prevents the small-sample instinct from reaching for a t.

OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, pp. 466-467

51. Trap: treating the checks as formalities

Trap

The trap

\[ \text{the conditions are probably fine, so proceed} \]

Skip the checks and compute the test

Why: They usually pass.

\[ \text{a failed condition produces a plausible wrong answer} \]

Nothing in the output shows that an assumption was violated, so the error survives every subsequent step.

The fix

\[ \text{state the row, then check each condition explicitly} \]

Write the checks down, since they cost seconds and catch silent errors

Why: Failures do not announce themselves.

The book states the assumptions in a section of their own precisely because they are invisible afterwards. A test on clustered data, or on a proportion with np of 2, returns a number that looks exactly like a valid one — and the only place the problem could have been caught was before the arithmetic began.

52. One of these is false

Two truths and a lie

All three concern getting the choice wrong.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Using a z where a t is needed understates the p-value
  • C. A violated assumption does not show up in the output
  • B. Choosing the wrong distribution produces an obviously invalid answer

Survives elimination: B

Why: The survivor is false, and its falsity is why this section exists. The wrong choice returns a perfectly ordinary-looking p-value between zero and one. Nothing about it signals the error, which is the definition of a silent failure.

53. Which way does the error run?

Prediction

Commit before reasoning.

Predict first

A researcher uses a z when a t was required. What happens to their p-value?

  • It is too small, making the evidence look stronger
  • It is too large, making the evidence look weaker
  • It is unaffected
  • It could go either way

Correct: Too small, making the evidence look stronger.

Why: The t is heavier in the tails than the normal, so the area beyond any given statistic is larger under a t — meaning the correct p-value is larger than the one a z gives. The mistake therefore always flatters the result, and it does so most at the small sample sizes where the correction was most needed.

54. Explain why the checks matter

Explain it

A classmate says the assumption checks are box-ticking, since the formula works either way.

Discussion prompt

In two sentences or fewer, answer them.

Hint: Ask what a violated assumption would look like in the output.

Answer:

The formula does return a number either way, and that is exactly the problem — a p-value computed on clustered data or a badly skewed binomial looks identical to a valid one.

The checks are the only stage at which the error is visible, because after them the arithmetic proceeds identically whether the conditions held or not.

55. The three tests and their requirements

Comparison

Fill the blanks. Two questions select the row; the conditions decide whether it is legitimate.

Comparison matrix

TestDistributionIts own condition
Mean, sigma knownnormalsigma genuinely known
Mean, sigma unknownStudent t, with n - 1 dfpopulation approximately normal, or n large
Proportionnormalbinomial conditions, and np and nq above five
All threea simple random samplenever repaired by a larger sample

The last row is the one that carries across every test in the book. Distributional conditions can be relieved by more data; a biased selection cannot, and the extra data only makes the wrong answer more precise.

56. Selecting the distribution, in order

Pattern

Five steps, and the first two settle the row.

  1. Ask whether the parameter is a mean or a proportion.
  2. If a mean, ask whether the population standard deviation is genuinely known — a normal if so, a t with n minus one degrees of freedom if not.
  3. Confirm the sample is a simple random sample, which no test can do without.
  4. For a mean, confirm the population is approximately normal or the sample is large enough for the central limit theorem.
  5. For a proportion, confirm the binomial conditions and check that np and nq both exceed five.

Use the null hypothesis's value of p in the np and nq check, since the whole test is computed under the assumption that the null is true.

OpenStax Introductory Business Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3 Probability Distribution Needed for Hypothesis Testing

57. Check yourself 1 of 3

Check

Selecting the distribution.

Check your understanding

A test concerns a mean, and the sample standard deviation is computed from 25 observations. Which distribution?

  • A. A Student t with 24 degrees of freedom (correct)
  • B. A normal distribution
  • C. A Student t with 25 degrees of freedom
  • D. A binomial distribution

Answer: A

Why: The parameter is a mean and sigma is estimated by s, so a t applies with n minus one degrees of freedom.

Why B tempts people
A normal would be right only if sigma were genuinely known.
Why C tempts people
The degrees of freedom are n minus one, not n.
Why D tempts people
The binomial belongs to proportion tests, not to tests about a mean.

58. Check yourself 2 of 3

Check

The proportion condition.

Check your understanding

A test of p = 0.05 uses a sample of 40. Does the normal approximation apply?

  • A. No: np is only 2, which is below five (correct)
  • B. Yes: nq is 38, comfortably above five
  • C. Yes: the sample exceeds 30
  • D. No: proportions always need a t

Answer: A

Why: Both products must exceed five, and np is 2. The binomial is bunched against zero and a normal fits it badly.

Why B tempts people
Checking only one product passes almost any proportion; both are required.
Why C tempts people
The thirty-observation rule is not the condition here; the products are.
Why D tempts people
Proportion tests never use a t — their small-sample remedy is different.

59. Check yourself 3 of 3

Check

Assumptions.

Check your understanding

Which assumption is NOT relieved by collecting a larger sample?

  • A. That the sample is randomly selected (correct)
  • B. That the population is approximately normal
  • C. That np and nq exceed five
  • D. All are relieved by more data

Answer: A

Why: A larger biased sample estimates the wrong quantity more precisely. Randomness is a property of the selection, not of the amount of data.

Why B tempts people
The central limit theorem relieves this one as the sample grows.
Why C tempts people
Both products grow with n, so a larger sample helps here too.
Why D tempts people
Two of the three are relieved by more data; the randomness assumption is not.

60. Where this shows up outside the textbook

Real world

A hospital audits whether its rate of a rare surgical complication exceeds the national benchmark of 2 percent. It reviews 80 randomly selected cases, finds 4 complications, and runs a one-proportion z-test, obtaining a p-value of 0.16 and concluding no evidence of a problem.

Discussion prompt

Assess whether the test was appropriate, and say what should have been done instead.

Hint: Check the condition before questioning the conclusion.

Answer:

The test was not appropriate, and the check fails before any arithmetic. Under the null the proportion is 0.02, so np is 80 times 0.02, which is 1.6 — well below five. The binomial here is strongly right-skewed and bunched against zero, and a normal curve does not describe it.

\[ np = (80)(0.02) = 1.6 \;\not>\; 5, \qquad nq = 78.4 \]

The p-value of 0.16 is therefore unreliable, and its error is not in a predictable direction. For a discrete, skewed distribution approximated by a continuous symmetric one, the tail area can be substantially off either way — so the conclusion cannot be trusted even though the number looks ordinary.

The right approach is an exact binomial test, which computes the probability of seeing 4 or more complications in 80 cases when the true rate is 0.02 directly from the binomial rather than through any approximation. That calculation needs no shape condition because it uses the actual distribution, and it is the standard method for rare events.

Two further points belong in a careful answer. Rare-event auditing needs far more cases than intuition suggests — at a 2 percent benchmark the condition would need about 250 cases before a normal approximation became defensible, which is a design consideration rather than an analysis one. And the audit should also state what rate it had the power to detect: with 80 cases it would miss even a substantially elevated complication rate much of the time, so a non-significant result here is close to uninformative regardless of which test produced it.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

Why must np and nq BOTH exceed five for a proportion test?

  • Because the sample must be large
  • Because a small product on either side leaves the binomial bunched against a boundary and badly skewed
  • Because the formula divides by them
  • Because five is the minimum sample size

Correct: Because either one being small makes the binomial skewed.

\[ np > 5 \;\text{ and }\; nq > 5 \;\Longrightarrow\; \text{the binomial is close to symmetric} \]

Why: A small np bunches the distribution against zero and a small nq bunches it against n, and in both cases the shape is lopsided and bounded on one side — which a symmetric, unbounded normal cannot match, least of all in the tail where a p-value is computed. Checking only one product passes almost any proportion, which is why both are required.

62. Explain it to someone a year behind you

Explain it

They used a t distribution for a proportion test because their sample was only 30.

Discussion prompt

In two sentences or fewer, correct them.

Hint: Ask what the t is correcting for.

Answer:

The t exists to account for estimating a standard deviation with s, and a proportion has no separate standard deviation to estimate — its spread comes straight from p prime.

A proportion test always uses a normal, and its small-sample safeguard is the np and nq check rather than a change of curve.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Selecting the right row from the parameter and whether sigma is known
  • Stating the assumptions each mean test requires
  • Checking the binomial conditions on a survey design
  • Applying the np and nq check, using the right value of p

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the first, two questions settle it: mean or proportion, then is sigma known. For the second, a random sample always, and normality that a large sample relieves. For the third, look for independence, which clustering breaks. For the fourth, compute both products using the null's p. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Fifteen minutes.

Draw it

At the top, redraw the book's table with its four columns — type of test, parameter, point estimate and distribution — filling in all three rows. Beside it, write the two questions that select a row, in order, and note that nothing else is needed. Below that, list the assumptions in three groups: what all three tests require, what the two mean tests require, and what the proportion test requires. Mark clearly which single assumption is not relieved by a larger sample. In the middle of the page, draw two binomial stem plots side by side for n equal to 20: one with p equal to 0.1 and one with p equal to 0.5, and write np under each. Say in one sentence why a normal curve fits the second and not the first, and write the condition in full with both products. At the bottom, draw a standard normal and a t with four degrees of freedom on the same axes, shade the region beyond 2.5 on each, and write the two p-values — then one sentence on which is correct when sigma was estimated and which direction the error runs.

Check your two stem plots by confirming the left one has a visible floor at zero that the curve would have to cross, since that is the concrete reason the approximation fails. Check your bottom figure by confirming the t's shaded area is the larger of the two, which is what makes a z's p-value always too small.

65. What you can do now

Recap

Five things, and the last is the check chapter 8 left implicit.

If you seeThen
A mean with sigma givenA normal distribution
A mean with s computed from dataA t with n minus one degrees of freedom
A proportion, any sample sizeA normal, then check np and nq
A small proportion sampleCheck the products; do not reach for a t
Clustered or grouped responsesIndependence fails; the plain test does not apply
np or nq at or below fiveUse an exact binomial method instead
A biased sampleNo sample size and no distribution repairs it

Section 9.4 supplies the last piece: what to compute once the distribution is chosen. The p-value measures how unlikely the observed sample would be if the null were true, and comparing it against a preset significance level turns that measurement into a decision.

OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, pp. 466-467 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 466-467
  2. OpenStax Introductory Business Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108