A short section that does one job: it says which probability distribution each of the three one-sample tests uses, and what has to be true for that choice to be legitimate. A test of a mean when the population standard deviation is known uses a normal distribution; a test of a mean when it is unknown and estimated by the sample standard deviation uses a Student t; and a test of a proportion uses a normal distribution built on the binomial. The three rows are chapter 8's three confidence intervals in the same order and for the same reasons. What the section adds is an explicit statement of the assumptions behind each: a simple random sample in every case, approximate normality of the population for the two mean tests, and for a proportion the binomial conditions together with the requirement that n times p and n times q both exceed five, so that the binomial is close enough in shape to the normal for the approximation to hold.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 9 — Hypothesis Testing with One Sample
Probability Distribution Needed for Hypothesis Testing
Objectives
Five outcomes, and the last is the condition the earlier chapters left implicit.
OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, pp. 466-467 — the section these objectives are drawn from
Warm-up
Chapter 8 built three confidence intervals: a z for a mean with sigma known, a t for a mean without it, and a z for a proportion.
Discussion prompt
A hypothesis test asks a different question from a confidence interval. Does it need a different set of distributions?
Hint: What does each procedure actually compute an area under?
Answer:
No. A confidence interval starts from a point estimate and works outward to find the values consistent with it; a test starts from a claimed value and works outward to see whether the point estimate is consistent with THAT. Both operate on the same sampling distribution.
So the three rows of this section's table are exactly chapter 8's three cases, and the deciding questions are the same: is the parameter a mean or a proportion, and if a mean, is sigma known?
What this section adds is the assumptions, stated in one place for the first time — including one condition for proportions that chapter 8 never spelled out, and which turns out to matter.
Concept
Particular distributions are associated with various types of hypothesis testing. A test for a mean with the population standard deviation known uses the normal distribution; a test for a mean with it unknown uses the Student t; and a test for proportions uses the normal distribution.
the three one-sample tests — A mean with sigma known, a mean with sigma unknown, and a proportion. The parameter and the point estimate differ across them, and the distribution follows from the pair.
\[ \bar{x} \to N \text{ or } t; \qquad p' \to N \]
It is worth noticing that two of the three rows use a normal distribution, and for quite different reasons. The first uses it because sigma is known, so nothing beyond the mean is being estimated; the third uses it because a proportion's spread follows from its own point estimate, so again nothing extra is estimated. Only the middle row estimates a second quantity, and only it needs a t.
Figure (svg): A four-column table listing the three one-sample tests with their parameters, point estimates and distributions
OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 466
Section
Section 1
Concept
Each row of the book's table names three things: the population parameter being tested, the sample statistic that estimates it, and the probability distribution used to compute the test. The first two determine the third.
the three rows — Mean with sigma known, estimated by the sample mean, using a normal. Mean with sigma unknown, estimated by the sample mean, using a t. Proportion, estimated by the sample proportion, using a normal.
\[ (\mu, \bar{x}, N); \quad (\mu, \bar{x}, t); \quad (p, p', N) \]
The table is worth learning as a whole rather than as three separate facts, because the questions that select a row are asked in a fixed order. First, is the parameter a mean or a proportion — that settles rows one and two against row three. Then, for a mean, is sigma known — that settles row one against row two. Two questions, three answers, and no other information is needed.
Figure (svg): A four-column table listing the three one-sample tests with their parameters, point estimates and distributions
OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 466 — the table of tests and distributions
Picture it
Three tests, with what each uses.
Figure (svg): A four-column table listing the three one-sample tests with their parameters, point estimates and distributions
The parameter column contains only mu and p, and the point estimate column only x-bar and p prime — which is a reminder from section 9.1 that hypotheses concern parameters while the data supply statistics. Every test in this chapter and the next has that same two-column structure.
Worked example
Three problems, asked in the fixed order.
\[ \text{(a) } \sigma = 0.8 \text{ given}; \; \text{(b) only data}; \; \text{(c) } 53 \text{ of } 100 \]
Case (a)
Why: A mean, sigma known.
Case (b)
Why: A mean, sigma unknown.
Case (c)
Why: A proportion.
Note the order of questions
Why: Parameter, then sigma.
Figure (svg): The solution to Worked example selecting a row shown as a ladder of expressions, one row per legal move
\[ N, \quad t, \quad N \]
Verify: confirm the two questions are enough
Why: Nothing else was needed: not the sample size, not the confidence level, not the direction of the alternative. Sample size affects the degrees of freedom in case (b) and the validity check in case (c), but it never changes which row applies. Keeping the selection to two questions makes it hard to get wrong.
OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 466
Sorting
Ask about the parameter first, then about sigma.
Sort into buckets
Sort each test.
Both proportion problems use a normal regardless of their sample sizes, which is the distinction the previous trap turns on. Sample size matters for proportions through the np and nq check rather than through the choice of curve.
Worked example
Rows one and three both use a normal, for different reasons.
\[ \text{row 1: } \sigma \text{ known}; \quad \text{row 3: a proportion} \]
Row one
Why: Sigma is given.
Row three
Why: The spread comes from p prime.
Row two
Why: s estimates sigma.
Conclude
Why: Only row two needs a t.
Figure (svg): The solution to Worked example why two rows share a distribution shown as a ladder of expressions, one row per legal move
\[ \text{one estimate} \to N; \quad \text{two estimates} \to t \]
Verify: confirm this matches section 8.2's account
Why: Section 8.2 introduced the t precisely because s varies from sample to sample, adding a second source of uncertainty that a z cannot account for. A proportion has no such second source, since the binomial's spread is determined by its probability — which is why no t distribution appears anywhere in the proportion row, however small the sample.
OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, pp. 466-467
Trap
\[ n = 25 \text{ for a proportion} \;\Rightarrow\; \text{use } t_{24} \]
Apply section 8.2's small-sample rule
Why: A t handles small samples.
\[ \text{but nothing beyond } p' \text{ is being estimated} \]
The t corrects for an estimated standard deviation, and a proportion has none to estimate.
\[ \text{use a normal, and check } np > 5 \text{ and } nq > 5 \]
Use a normal, with the binomial condition as the small-sample safeguard
Why: That is what the third row says.
Small samples are a real problem for proportions, but the remedy is different: the np and nq check, and if it fails, the plus-four adjustment of section 8.3 or an exact binomial method. Reaching for a t is a reasonable instinct applied to the wrong row.
Fill the middle
For a hypothesis test about a population proportion.
Fill in the blanks
\textproportion ___
Why: The sample proportion, p prime — the successes over the trials. It plays the same role for a proportion test that the sample mean plays for a mean test.
Two truths and a lie
All three concern the table.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The sample size affects the degrees of freedom for a t and the validity check for a proportion, but it never selects the row. Section 8.2 established that the deciding question for means is whether sigma is known, not how much data there is.
Prediction
Commit before reasoning.
Predict first
Why does only one of the three rows use a Student t?
Correct: Because only that row estimates a second quantity.
Why: The t exists to account for the extra uncertainty in estimating sigma with s. The sigma-known row does not estimate it, and the proportion row gets its spread free from p prime — so neither needs the correction. The third option is close but not enough: both mean rows are about a mean, and only one uses a t.
Section
Section 2
Concept
For a z-test of a single mean you take a simple random sample, the population is normally distributed or the sample is sufficiently large, and you know sigma — which, the book adds, in reality is rarely known. For a t-test the data should be a simple random sample from an approximately normal population, and s approximates sigma.
sufficiently large — The condition that lets chapter 7's central limit theorem supply normality for the sample mean when the population is not normal. The book notes a t-test will work even if the population is not approximately normal, provided the sample is sufficiently large.
\[ \text{random sample} \;+\; \text{normal population OR large } n \]
The book's parenthetical about sigma being rarely known in reality is the reason the second row is the ordinary case and the first is nearly a textbook curiosity. Quality control with a long-established process is one of the few settings where a population standard deviation genuinely is known; almost everywhere else it has to be estimated, and the t applies.
Figure (svg): The assumptions each of the three tests requires
OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 467 — the assumptions for the z-test and the t-test
Picture it
Five conditions across the three tests.
Figure (svg): The assumptions each of the three tests requires
The first line applies to all three and is the one no amount of data repairs. Section 8.2 made the same point: a large biased sample estimates the wrong quantity more precisely, which is worse than being imprecise about the right one.
Worked example
A t-test on fifteen measurements.
\[ n = 15, \; \sigma \text{ unknown, population believed normal} \]
Random sample?
Why: Must be established.
Population normal?
Why: Believed so.
\[ \text{matters at } n = 15 \]
Sigma known?
Why: No.
Choose the distribution
Why: By the second row.
\[ t\text{ with } 14\text{ df} \]
Figure (svg): The solution to Worked example checking the assumptions shown as a ladder of expressions, one row per legal move
\[ T \sim t_{14} \]
Verify: confirm which assumption a larger sample would relieve
Why: The book says a t-test will work even if the population is not approximately normal when the sample is sufficiently large, so normality is the relieved condition. Randomness is not: it is a separate assumption, and the book states it separately for exactly that reason. At n equal to 15 the normality assumption is doing real work and should be checked against the data's shape.
OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 467
Discrimination
Each situation threatens one condition.
Sort into buckets
Sort each by which assumption it puts at risk.
Worked example
The same test on a much larger sample.
\[ n = 200, \; \text{population plainly skewed} \]
Note the population's shape
Why: Skewed.
Recall the theorem
Why: Chapter 7.
Check the sample size
Why: Two hundred.
Conclude
Why: The test is valid.
Figure (svg): The solution to Worked example when normality stops mattering shown as a ladder of expressions, one row per legal move
\[ n \text{ large} \;\Rightarrow\; \bar{X} \text{ approximately normal} \]
Verify: confirm which quantity the normality assumption is really about
Why: The assumption is not that the data are normal but that the SAMPLE MEAN's distribution is, and chapter 7 supplies that for a large sample whatever the population looks like. For a small sample there is no such rescue, so the population's own shape has to carry it — which is why the assumption is stated in terms of the population but only bites when n is small.
OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 467
Error analysis
Which correctly describe what the mean tests require?
Annotate
On: \( \begin{aligned} &(1)\; \text{a large sample repairs a biased selection} \\ &(2)\; \text{normality is needed regardless of } n \\ &(3)\; \text{a } z\text{-test may be used whenever } n > 30 \\ &(4)\; \text{normality matters most when } n \text{ is small} \end{aligned} \)
Errors (2) and (4) are opposite readings of the same condition, and the difference matters at the design stage — a small study needs its normality assumption checked, and a large one mostly does not.
Two truths and a lie
All three concern the mean tests.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. Randomness is a property of how the sample was selected, and no amount of additional data changes a biased selection mechanism — it simply pins down the biased quantity more tightly.
Prediction
Commit before reasoning.
Predict first
The t-test assumes an approximately normal population. Which quantity's distribution does this actually secure?
Correct: The distribution of the sample mean.
Why: The test statistic is built from the sample mean, so it is the mean's sampling distribution that must be normal. The population's shape is a way of securing that when the sample is small; when the sample is large, chapter 7 secures it directly and the population's shape stops mattering.
Faded example
A sample of 40 measurements, with s computed from the data and the population moderately skewed.
Fill in the blanks
\textt 39 \text___ ___ \text___
Why: A t with 39 degrees of freedom. The moderate skew is not a problem at n equal to 40, since the central limit theorem makes the sample mean approximately normal.
Section
Section 3
Concept
For a test of a single proportion you take a simple random sample and must meet the conditions for a binomial distribution: a certain number n of independent trials, outcomes of success or failure, and the same probability of success on each trial.
the binomial conditions — A fixed number of independent trials, two outcomes per trial, and a constant success probability — the three characteristics of section 4.3, restated as requirements for the test.
\[ X \sim B(n, p): \; n \text{ fixed}, \; \text{independent}, \; p \text{ constant} \]
These are section 4.3's three characteristics used as a checklist. The independence condition is the one that fails most often in practice — surveys of households, students in the same class, or patients at the same clinic all produce responses that are correlated rather than independent, and a test that assumes otherwise will understate its uncertainty.
Figure (svg): The assumptions each of the three tests requires
OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 467 — the conditions for a proportion test
Picture it
The proportion test's requirements among the others.
Figure (svg): The assumptions each of the three tests requires
The last two lines both belong to the proportion test, and they do different jobs: the binomial conditions make the model right, and the np and nq condition makes the normal approximation to that model adequate. Both have to hold, and the next idea is about the second.
Worked example
A survey of 100 first-time brides.
\[ 100 \text{ brides surveyed}; \; 53 \text{ are younger than their grooms} \]
A fixed number of trials
Why: One hundred surveyed.
\[ n = 100 \]
Two outcomes each
Why: Younger, or not.
The same probability each
Why: If randomly selected.
Independent
Why: One answer does not affect another.
Figure (svg): The solution to Worked example checking the binomial conditions shown as a ladder of expressions, one row per legal move
\[ X \sim B(100, p) \]
Verify: confirm which condition is most at risk in survey work
Why: Independence is the fragile one. If the sample had been drawn from a single social circle, or if couples influenced one another's answers, the responses would be correlated and the effective information would be less than 100 independent observations suggests. The test would then report more certainty than the data supports — an error that no calculation inside the test can detect.
OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 467
Sorting
Check for a fixed n, two outcomes, constant p and independence.
Sort into buckets
Sort each survey.
The two doubtful cases share a shape: sampling clusters and then taking everyone inside them. It is a common and often sensible design for practical reasons, but it needs methods that account for the clustering rather than the plain binomial test.
Worked example
A survey design that breaks the model.
\[ \text{20 households surveyed, all members of each asked} \]
Count the responses
Why: Perhaps 60 people.
\[ n\text{ looks like } 60 \]
Ask about independence
Why: Household members agree.
Say what breaks
Why: The independence condition.
\[ \text{not } 60\text{ independent trials} \]
Note the consequence
Why: Overstated precision.
Figure (svg): The solution to Worked example a case where independence fails shown as a ladder of expressions, one row per legal move
\[ \text{correlated responses} \;\Rightarrow\; \text{effective } n < 60 \]
Verify: confirm the direction of the error this causes
Why: Correlated observations carry less information than independent ones, so the true standard error is larger than the formula gives — which makes the interval too narrow and the p-value too small. Every consequence runs toward overstating the evidence, which is why the independence condition is worth checking against the survey design rather than assumed.
OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 467
Trap
\[ 60 \text{ responses from } 20 \text{ households} \;\Rightarrow\; n = 60 \]
Count every answer collected
Why: Each is a data point.
\[ \text{but members of a household are not independent} \]
The effective number of independent observations is closer to 20 than to 60.
\[ \text{ask how many INDEPENDENT trials there were} \]
Count independent units, not responses
Why: Independence is one of the three binomial conditions.
This is a design problem rather than an arithmetic one, and it cannot be fixed at the analysis stage by any adjustment in this book — methods for clustered data exist but are beyond it. The practical lesson is that the conditions have to be checked when the survey is planned, because by the time the data arrive the damage is done.
Two truths and a lie
All three concern the proportion test's conditions.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and has the direction backwards. Correlated responses carry less information than the count suggests, so the computed standard error is too small — making intervals too narrow and p-values too small. The error overstates the evidence rather than being cautious about it.
Prediction
Commit before reasoning.
Predict first
A survey samples whole households rather than individuals. What happens to the true uncertainty?
Correct: Larger than the formula reports.
Why: Members of a household give similar answers, so sixty responses from twenty households carry less information than sixty independent ones. The formula uses the raw count and therefore understates the real variability — reporting more precision than the design supports.
Fill the middle
What a proportion test requires of its data.
Fill in the blanks
n \textprobability of success ___ \text___
Why: The same probability of success on every trial — section 4.3's third characteristic, restated here as a requirement for the test to be valid.
Section
Section 4
Concept
The shape of the binomial distribution needs to be similar to the shape of the normal distribution. To ensure this, the quantities np and nq must both be greater than five. Then the binomial distribution of a sample proportion can be approximated by the normal distribution.
np and nq greater than five — The condition that makes the normal approximation to the binomial adequate. Both products are required, since a proportion near zero fails on np and one near one fails on nq.
\[ np > 5 \quad\text{and}\quad nq > 5 \]
The reason both are needed is symmetry. A binomial with a small np is bunched against zero and strongly right-skewed; one with a small nq is bunched against n and left-skewed. Either way the distribution is lopsided and bounded on one side, while a normal is symmetric and unbounded — so the approximation fails at exactly the place a test would be looking, which is the tail.
Figure (svg): Two binomial stem plots side by side, the left one strongly skewed with a small np and the right one symmetric with a larger np
OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 467 — the np and nq requirement
Picture it
Two binomials with the same n and very different shapes.
Figure (svg): Two binomial stem plots side by side, the left one strongly skewed with a small np and the right one symmetric with a larger np
The left distribution has np equal to 2 and is plainly skewed, with a hard floor at zero that no normal curve respects. The right has np equal to 10 and is close to symmetric. Fitting a normal to the left one would misstate both tails, and the right tail is usually where a p-value is computed.
Worked example
The brides survey of Example 9.17.
\[ n = 100, \; p = 0.50 \text{ from } H_0 \]
Compute np
Why: One hundred times a half.
\[ 50 \]
Compute nq
Why: The same.
\[ 50 \]
Compare with five
Why: Both far above.
Conclude
Why: Use the normal.
Figure (svg): The solution to Worked example checking the condition shown as a ladder of expressions, one row per legal move
\[ np = nq = 50 > 5 \]
Verify: confirm which p is used in the check
Why: The p used is the one from the null hypothesis, not the sample proportion — because the test computes everything under the assumption that the null is true. Using p prime instead would usually give a similar answer here, but for a proportion near the boundary the two checks can disagree, and the null's value is the one the test actually needs.
OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 467
Faded example
A test of p = 0.30 on a sample of 50.
Fill in the blanks
np = 15, \quad nq = 35, \quad \text___
Why: Both 15 and 35 exceed five comfortably, so the normal approximation to the binomial is adequate and the test may proceed.
Worked example
A small sample with a proportion near zero.
\[ n = 40, \; p = 0.05 \text{ from } H_0 \]
Compute np
Why: Forty times 0.05.
\[ 2 \]
Compare with five
Why: Below it.
Compute nq
Why: Forty times 0.95.
\[ 38 \]
Conclude
Why: One product fails.
Figure (svg): The solution to Worked example a case where the condition fails shown as a ladder of expressions, one row per legal move
\[ np = 2 \not> 5 \]
Verify: confirm what the failure looks like in practice
Why: Section 8.3's transfer problem showed the symptom: an interval running to a negative endpoint, which is impossible for a proportion. The same underlying failure affects a test, where the p-value computed from a normal can be badly wrong in the tail. The remedies are the plus-four adjustment for an interval, or an exact binomial calculation for a test.
OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, p. 467
Trap
\[ n = 40, \; p = 0.95: \quad nq = 2, \text{ but } np = 38 > 5 \]
Check np and stop, since it comfortably passes
Why: The condition names np first.
\[ nq = 2, \text{ so the distribution is skewed the other way} \]
A proportion near one is bunched against the upper limit, and a normal fits it no better than it fits one bunched against zero.
\[ \text{check BOTH: } np > 5 \text{ and } nq > 5 \]
Compute both products every time
Why: The condition is symmetric in p and q.
Checking only the first product passes every proportion above about 0.125 at n equal to 40, including 0.99 — where the approximation is hopeless. The two products are equally binding, and computing both costs one extra multiplication.
Discrimination
Compute both products for each.
Sort into buckets
Sort each setup.
Two truths and a lie
All three concern the condition.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. A proportion of 0.02 with n equal to 500 gives np of 10, which passes comfortably. The condition constrains the PRODUCT rather than the proportion, so any p can satisfy it given a large enough sample.
Estimation
A null hypothesis states p = 0.04.
Predict first
Roughly what sample size does the condition require?
Correct: At least about 125.
Why: np exceeds five when n exceeds 5 over 0.04, which is 125; nq is then comfortably large. Small proportions demand large samples for the approximation to hold, which is why studies of rare events need far more data than their headline rates suggest.
Section
Section 5
Concept
Selecting the wrong distribution does not produce an obviously wrong answer. It produces a p-value that is systematically too small or too large, with nothing in the output to indicate the error.
the direction of the error — Using a z where a t is needed makes the p-value too small and the evidence look stronger than it is. Using a normal where the binomial is too skewed distorts the tail, which is exactly where the p-value is computed.
\[ t_{df} > z \;\Longrightarrow\; \text{a z gives too small a } p\text{-value} \]
That the errors run one way is worth dwelling on. A z used in place of a t always understates the p-value, so it always makes results look more significant than they are — the same one-sided bias section 8.2 found for confidence intervals, where a z interval was always too narrow. Systematic errors that flatter the researcher are the ones worth guarding against hardest.
Figure (svg): A standard normal curve drawn over a t curve with few degrees of freedom, the t sitting lower in the centre and higher in the tails
OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, pp. 466-467 — the table and the assumptions together
Picture it
A standard normal and a t with four degrees of freedom.
Figure (svg): A standard normal curve drawn over a t curve with few degrees of freedom, the t sitting lower in the centre and higher in the tails
The gap is largest in the tails, which is precisely where a p-value is computed — so the choice between them matters most exactly where it is being used. At four degrees of freedom the two-tailed p-value for a statistic of 2.5 is about 0.067 from the t and 0.012 from the normal, which straddles the conventional threshold.
Worked example
The same test statistic read from both distributions.
\[ t = 2.5, \; \text{df} = 4 \]
From the t
Why: Two-tailed, 4 df.
\[ \text{about } 0.067 \]
From the normal
Why: Two-tailed.
\[ \text{about } 0.012 \]
Compare with 0.05
Why: They straddle it.
Say which is right
Why: Sigma was estimated.
Figure (svg): The solution to Worked example the cost of the wrong curve shown as a ladder of expressions, one row per legal move
\[ p_t \approx 0.067 \quad\text{against}\quad p_z \approx 0.012 \]
Verify: confirm the error always runs in the same direction
Why: The t has heavier tails, so the same statistic always leaves more area beyond it under a t than under a normal — the z p-value is smaller every time, never larger. That means the mistake always makes the evidence look stronger, and it is largest for small samples where the evidence is weakest to begin with.
OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, pp. 466-467
Sorting
Select the distribution and check the conditions.
Sort into buckets
Sort each by whether the standard test applies.
The two failures need different remedies: (b) needs an exact binomial method or the plus-four adjustment, while (d) needs a design that accounts for clustering. Neither is repaired by choosing a different curve from the table.
Worked example
Working through the two questions and the checks.
\[ \text{a claim about a proportion}; \; n = 100, \; p_0 = 0.50 \]
Question one
Why: Mean or proportion?
Select the row
Why: The third.
Check the binomial conditions
Why: Independent, two outcomes.
Check np and nq
Why: Both 50.
Figure (svg): The solution to Worked example a full selection shown as a ladder of expressions, one row per legal move
\[ P' \sim N\left(p, \sqrt{\tfrac{pq}{n}}\right) \]
Verify: confirm nothing about the sample size changed the row
Why: The sample of 100 entered only the np and nq check, never the choice of curve — a proportion test uses a normal at n equal to 20 and at n equal to 2,000 alike. What changes with n is whether the approximation is good enough, and that is what the products test. Keeping the two questions separate prevents the small-sample instinct from reaching for a t.
OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, pp. 466-467
Trap
\[ \text{the conditions are probably fine, so proceed} \]
Skip the checks and compute the test
Why: They usually pass.
\[ \text{a failed condition produces a plausible wrong answer} \]
Nothing in the output shows that an assumption was violated, so the error survives every subsequent step.
\[ \text{state the row, then check each condition explicitly} \]
Write the checks down, since they cost seconds and catch silent errors
Why: Failures do not announce themselves.
The book states the assumptions in a section of their own precisely because they are invisible afterwards. A test on clustered data, or on a proportion with np of 2, returns a number that looks exactly like a valid one — and the only place the problem could have been caught was before the arithmetic began.
Two truths and a lie
All three concern getting the choice wrong.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and its falsity is why this section exists. The wrong choice returns a perfectly ordinary-looking p-value between zero and one. Nothing about it signals the error, which is the definition of a silent failure.
Prediction
Commit before reasoning.
Predict first
A researcher uses a z when a t was required. What happens to their p-value?
Correct: Too small, making the evidence look stronger.
Why: The t is heavier in the tails than the normal, so the area beyond any given statistic is larger under a t — meaning the correct p-value is larger than the one a z gives. The mistake therefore always flatters the result, and it does so most at the small sample sizes where the correction was most needed.
Explain it
A classmate says the assumption checks are box-ticking, since the formula works either way.
Discussion prompt
In two sentences or fewer, answer them.
Hint: Ask what a violated assumption would look like in the output.
Answer:
The formula does return a number either way, and that is exactly the problem — a p-value computed on clustered data or a badly skewed binomial looks identical to a valid one.
The checks are the only stage at which the error is visible, because after them the arithmetic proceeds identically whether the conditions held or not.
Comparison
Fill the blanks. Two questions select the row; the conditions decide whether it is legitimate.
Comparison matrix
| Test | Distribution | Its own condition |
|---|---|---|
| Mean, sigma known | normal | sigma genuinely known |
| Mean, sigma unknown | Student t, with n - 1 df | population approximately normal, or n large |
| Proportion | normal | binomial conditions, and np and nq above five |
| All three | a simple random sample | never repaired by a larger sample |
The last row is the one that carries across every test in the book. Distributional conditions can be relieved by more data; a biased selection cannot, and the extra data only makes the wrong answer more precise.
Pattern
Five steps, and the first two settle the row.
Use the null hypothesis's value of p in the np and nq check, since the whole test is computed under the assumption that the null is true.
OpenStax Introductory Business Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3 Probability Distribution Needed for Hypothesis Testing
Check
Selecting the distribution.
Check your understanding
A test concerns a mean, and the sample standard deviation is computed from 25 observations. Which distribution?
Answer: A
Why: The parameter is a mean and sigma is estimated by s, so a t applies with n minus one degrees of freedom.
Check
The proportion condition.
Check your understanding
A test of p = 0.05 uses a sample of 40. Does the normal approximation apply?
Answer: A
Why: Both products must exceed five, and np is 2. The binomial is bunched against zero and a normal fits it badly.
Check
Assumptions.
Check your understanding
Which assumption is NOT relieved by collecting a larger sample?
Answer: A
Why: A larger biased sample estimates the wrong quantity more precisely. Randomness is a property of the selection, not of the amount of data.
Real world
A hospital audits whether its rate of a rare surgical complication exceeds the national benchmark of 2 percent. It reviews 80 randomly selected cases, finds 4 complications, and runs a one-proportion z-test, obtaining a p-value of 0.16 and concluding no evidence of a problem.
Discussion prompt
Assess whether the test was appropriate, and say what should have been done instead.
Hint: Check the condition before questioning the conclusion.
Answer:
The test was not appropriate, and the check fails before any arithmetic. Under the null the proportion is 0.02, so np is 80 times 0.02, which is 1.6 — well below five. The binomial here is strongly right-skewed and bunched against zero, and a normal curve does not describe it.
\[ np = (80)(0.02) = 1.6 \;\not>\; 5, \qquad nq = 78.4 \]
The p-value of 0.16 is therefore unreliable, and its error is not in a predictable direction. For a discrete, skewed distribution approximated by a continuous symmetric one, the tail area can be substantially off either way — so the conclusion cannot be trusted even though the number looks ordinary.
The right approach is an exact binomial test, which computes the probability of seeing 4 or more complications in 80 cases when the true rate is 0.02 directly from the binomial rather than through any approximation. That calculation needs no shape condition because it uses the actual distribution, and it is the standard method for rare events.
Two further points belong in a careful answer. Rare-event auditing needs far more cases than intuition suggests — at a 2 percent benchmark the condition would need about 250 cases before a normal approximation became defensible, which is a design consideration rather than an analysis one. And the audit should also state what rate it had the power to detect: with 80 cases it would miss even a substantially elevated complication rate much of the time, so a non-significant result here is close to uninformative regardless of which test produced it.
Commit first
Answer, then rate your confidence honestly.
Predict first
Why must np and nq BOTH exceed five for a proportion test?
Correct: Because either one being small makes the binomial skewed.
\[ np > 5 \;\text{ and }\; nq > 5 \;\Longrightarrow\; \text{the binomial is close to symmetric} \]
Why: A small np bunches the distribution against zero and a small nq bunches it against n, and in both cases the shape is lopsided and bounded on one side — which a symmetric, unbounded normal cannot match, least of all in the tail where a p-value is computed. Checking only one product passes almost any proportion, which is why both are required.
Explain it
They used a t distribution for a proportion test because their sample was only 30.
Discussion prompt
In two sentences or fewer, correct them.
Hint: Ask what the t is correcting for.
Answer:
The t exists to account for estimating a standard deviation with s, and a proportion has no separate standard deviation to estimate — its spread comes straight from p prime.
A proportion test always uses a normal, and its small-sample safeguard is the np and nq check rather than a change of curve.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, two questions settle it: mean or proportion, then is sigma known. For the second, a random sample always, and normality that a large sample relieves. For the third, look for independence, which clustering breaks. For the fourth, compute both products using the null's p. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top, redraw the book's table with its four columns — type of test, parameter, point estimate and distribution — filling in all three rows. Beside it, write the two questions that select a row, in order, and note that nothing else is needed. Below that, list the assumptions in three groups: what all three tests require, what the two mean tests require, and what the proportion test requires. Mark clearly which single assumption is not relieved by a larger sample. In the middle of the page, draw two binomial stem plots side by side for n equal to 20: one with p equal to 0.1 and one with p equal to 0.5, and write np under each. Say in one sentence why a normal curve fits the second and not the first, and write the condition in full with both products. At the bottom, draw a standard normal and a t with four degrees of freedom on the same axes, shade the region beyond 2.5 on each, and write the two p-values — then one sentence on which is correct when sigma was estimated and which direction the error runs.
Check your two stem plots by confirming the left one has a visible floor at zero that the curve would have to cross, since that is the concrete reason the approximation fails. Check your bottom figure by confirming the t's shaded area is the larger of the two, which is what makes a z's p-value always too small.
Recap
Five things, and the last is the check chapter 8 left implicit.
| If you see | Then |
|---|---|
| A mean with sigma given | A normal distribution |
| A mean with s computed from data | A t with n minus one degrees of freedom |
| A proportion, any sample size | A normal, then check np and nq |
| A small proportion sample | Check the products; do not reach for a t |
| Clustered or grouped responses | Independence fails; the plain test does not apply |
| np or nq at or below five | Use an exact binomial method instead |
| A biased sample | No sample size and no distribution repairs it |
Section 9.4 supplies the last piece: what to compute once the distribution is chosen. The p-value measures how unlikely the observed sample would be if the null were true, and comparing it against a preset significance level turns that measurement into a decision.
OpenStax Introductory Statistics 2e, §9.3 Probability Distribution Needed for Hypothesis Testing §9.3, pp. 466-467 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.