Section 12.3 computed a correlation of 0.6631 from eleven students and left open whether that is enough to conclude anything. The correlation coefficient tells us about the strength and direction of the linear relationship, but the reliability of the linear model also depends on how many observed data points are in the sample — so r and n have to be looked at together. This section performs a hypothesis test of the significance of the correlation coefficient, with a null that the population correlation rho equals zero and a two-tailed alternative that it does not. The book gives two methods, a p-value from a t statistic on n minus two degrees of freedom and a table of critical values at a fixed 5 percent level, and calls them equivalent. They are in fact the same test written twice: every critical value in the table equals the two-tailed t critical value divided by the square root of df plus that value squared.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 12 — Linear Regression and Correlation
Testing the Significance of the Correlation Coefficient
Objectives
Six outcomes, and the second explains why the section exists at all.
OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 631-635 — the section these objectives are drawn from
Warm-up
Section 12.3 produced r = 0.6631 from eleven students; chapter 9 tested claims about parameters.
Discussion prompt
If two variables were completely unrelated in the population, would a sample of eleven give a correlation of exactly zero?
Hint: Ask what any sample statistic does around its parameter.
Answer:
No. A sample correlation varies from sample to sample just as a sample mean does, so even with no relationship at all in the population a sample of eleven will produce some non-zero r — sometimes a fairly large one, purely by chance.
So the question is not whether r differs from zero, but whether it differs by more than chance would readily produce at this sample size. That is a hypothesis test, and its parameter is the population correlation coefficient.
\[ H_0: \rho = 0 \qquad\text{against}\qquad H_a: \rho \ne 0 \]
The sample size enters directly. Eleven points can look convincingly linear by accident far more easily than a hundred can, which is why the book insists on looking at the value of r and the sample size n together.
Concept
The correlation coefficient tells us about the strength and direction of the linear relationship, but the reliability of the linear model also depends on how many observed data points are in the sample. We perform a hypothesis test of the significance of the correlation coefficient to decide whether the linear relationship in the sample data is strong enough to use to model the relationship in the population.
rho against r — Rho is the population correlation coefficient and is unknown; r is the sample correlation coefficient, computed from the data, and is our estimate of it.
\[ H_0: \rho = 0, \qquad H_a: \rho \ne 0, \qquad \alpha = 0.05 \]
The book fixes the significance level for the whole chapter: we will always use a significance level of 5 percent. That is not a statistical necessity but a consequence of the tool — the table of critical values provided assumes 5 percent, and a different level would need a different table. The p-value method carries no such restriction.
Figure (svg): A card showing that the same correlation means different things at different sample sizes
OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 631-632
Section
Section 1
Concept
The sample data are used to compute r, the correlation coefficient for the sample. Because we have only sample data, we cannot calculate the population correlation coefficient, so r is our estimate of the unknown rho. Whether an observed r is convincing depends on how many points produced it.
significant — If the test concludes that the correlation coefficient is significantly different from zero, we say the correlation coefficient is significant, and the regression line may be used to model the relationship in the population.
\[ \text{same } r, \text{ different } n \;\Longrightarrow\; \text{different verdicts} \]
The dependence runs in the direction intuition suggests but more sharply than most people expect. At six points a correlation of 0.776 is not significant; at nineteen a correlation of 0.567 is. Doubling the sample roughly halves the correlation needed, so sample size buys a great deal here.
Figure (svg): A card showing that the same correlation means different things at different sample sizes
OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 631-634 — we need to look at both r and the sample size together
Picture it
Examples 12.7, 12.9 and 12.10a.
Figure (svg): A card showing that the same correlation means different things at different sample sizes
The middle panel is significant with a correlation lower than the left panel's, which is not. Reading r alone would rank them the other way round, which is precisely the error the section prevents.
Worked example
A correlation of 0.776 computed from six data points.
\[ r = 0.776, \; n = 6 \]
Degrees of freedom
Why: Six minus two.
\[ 4 \]
Look up the critical values
Why: At df 4.
\[ -0.811\text{ and } 0.811 \]
Compare
Why: 0.776 sits between.
Conclude
Why: Not significant.
Figure (svg): The solution to Worked example Example 12.9, a large r that fails shown as a ladder of expressions, one row per legal move
\[ -0.811 < 0.776 < 0.811 \]
Verify: confirm the same verdict by the p-value route
Why: The t statistic is 0.776 times the root of 4 over the root of 1 minus 0.602, which is 1.552 over 0.631, or 2.46. On four degrees of freedom the two-tailed p-value is 0.070 — above 0.05, so the same do-not-reject. The two methods agree because they are the same test, which is the subject of a later idea in this lesson.
OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 633-634
Sorting
Each gives a correlation, a sample size and the critical value.
Sort into buckets
Sort by the verdict.
These are Examples 12.7 through 12.10. Item (b) has a larger correlation than item (c) and fails, because it rests on six points rather than nine.
Worked example
A correlation of -0.567 computed from nineteen data points.
\[ r = -0.567, \; n = 19 \]
Degrees of freedom
Why: Nineteen minus two.
\[ 17 \]
Look up the critical value
Why: At df 17.
\[ -0.456 \]
Compare
Why: -0.567 is further out.
Conclude
Why: Significant.
Figure (svg): The solution to Worked example Example 12.10a, a smaller r that passes shown as a ladder of expressions, one row per legal move
\[ -0.567 < -0.456 \]
Verify: confirm the comparison is about distance from zero
Why: For a negative correlation the test asks whether r falls BELOW the negative critical value, which is the mirror of asking whether a positive r exceeds the positive one. Both amount to asking whether the magnitude exceeds 0.456, and stating it that way avoids the sign errors that comparing negative numbers invites.
OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, p. 634
Trap
\[ r = 0.776 > 0.567 \;\Rightarrow\; \text{the first is the stronger finding} \]
Rank two results by their correlations
Why: A larger correlation is a tighter relationship.
\[ \text{but } n = 6 \text{ against } n = 19 \]
The larger correlation comes from six points and is not significant; the smaller comes from nineteen and is.
\[ \text{compare each } r \text{ against its own critical value} \]
Read r and n together, as the book insists
Why: The threshold moves with the sample size.
Both statements can be true at once: the first sample shows a tighter relationship AND provides weaker evidence that any relationship exists. Strength and evidence are different quantities, and conflating them is the most common misreading of a correlation.
Two truths and a lie
All three concern sample size.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and Example 12.9 is the counterexample: 0.776 on six data points is not significant, because the critical value at four degrees of freedom is 0.811.
Estimation
The critical value is 0.811 at n = 6 and 0.456 at n = 19.
Predict first
Roughly what happens to the critical value as the sample grows?
Correct: It falls steadily.
Why: More data makes a given correlation harder to attribute to chance, so the bar drops. Tripling the sample from 6 to 19 nearly halves the required correlation, and at n = 100 the critical value is about 0.197 — which is why large studies routinely report significant correlations that explain very little variation.
Prediction
Commit before reasoning.
Predict first
Example 12.10d has r = 0 with n = 5. What is the verdict, and does n matter?
Correct: Not significant, whatever the sample size.
Why: The book says it directly: no matter what the degrees of freedom are, r equal to zero lies between the two critical values. Zero is the null's own value, so it can never be evidence against the null — which is why Try It 12.10 poses the same question at n = 100 and gets the same answer.
Section
Section 2
Concept
The null hypothesis is that the population correlation coefficient rho equals zero; the alternative is that it does not. In words, the null says there is not a significant linear relationship between x and y in the population, and the alternative says there is.
what significant means here — That the correlation coefficient is significantly different from zero. If so, the regression line may be used to model the linear relationship between x and y in the population.
\[ H_0: \rho = 0, \qquad H_a: \rho \ne 0 \]
The alternative names no direction, so the test is two-tailed throughout. That is a deliberate choice: the question being asked is whether ANY linear relationship exists, and a relationship running either way answers it. It also means the p-value is the combined area in both tails, which is where a factor of two enters the arithmetic.
Figure (svg): A card giving the hypotheses in symbols and in words
OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 631-632 — the hypotheses in symbols and in words
Picture it
Symbols above, words below.
Figure (svg): A card giving the hypotheses in symbols and in words
This is the chapter's first null with a Greek parameter, and it restores chapter 9's structure after three sections of description. What it does not restore is a choice of tail.
Worked example
The exam data, where the test rejects.
\[ \text{reject } H_0 \]
Name the verdict
Why: Reject.
Say why
Why: r differs from zero.
Name the variables
Why: In context.
Add what it licenses
Why: With a linear plot.
Figure (svg): The solution to Worked example stating the conclusion in words shown as a ladder of expressions, one row per legal move
\[ \text{reject } H_0: \rho = 0 \]
Verify: confirm the second clause is doing real work
Why: The book's wording always gives the reason — because the correlation coefficient is significantly different from zero — rather than stopping at the claim. That matters because a reader needs to know what was tested: not that the variables are related in some general sense, but that their LINEAR association is more than chance would produce at this sample size.
OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 632-633
Faded example
For a test of the significance of a correlation coefficient.
Fill in the blanks
H_0: \rho = 0, \qquad H_a: \rho ≠ 0
Why: Two-tailed throughout, because the question is whether any linear relationship exists rather than whether it runs in a particular direction.
Worked example
Example 12.9, where 0.776 on six points is not significant.
\[ \text{do not reject } H_0 \]
Name the verdict
Why: Do not reject.
State it carefully
Why: Insufficient evidence.
What it forbids
Why: Using the line.
Note the reason
Why: Only six points.
Figure (svg): The solution to Worked example the conclusion when the test fails shown as a ladder of expressions, one row per legal move
\[ \text{do not reject}: \; \text{no prediction} \]
Verify: confirm what failing to reject does not establish
Why: It does not establish that rho is zero. With six observations the test can only detect correlations above 0.811, so a real population correlation of 0.7 would usually go undetected here. Chapter 9's asymmetry applies unchanged: failing to reject means the data are consistent with the null, not that the null is true.
OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 631-634
Error analysis
Which are correct?
Annotate
On: \( \begin{aligned} &(1)\; H_0: \rho = 0, \; H_a: \rho \ne 0 \\ &(2)\; H_0: r = 0, \; H_a: r \ne 0 \\ &(3)\; H_0: \rho = 0, \; H_a: \rho > 0 \\ &(4)\; H_0: \rho \ne 0, \; H_a: \rho = 0 \end{aligned} \)
Error (2) is the same confusion chapter 11's variance test invited between sigma and s, and the fix is the same: the Greek letter is the unknown being claimed about, and the Roman letter is what the sample supplies.
Matching
Match each symbol to what it denotes.
Match the pairs
Why: The degrees of freedom are n minus two because the line uses two estimates — an intercept and a slope — which is the same accounting section 12.6 gives for dividing the SSE by n minus two.
Two truths and a lie
All three concern the hypotheses.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The sample correlation is known — it was computed in section 12.3 — so there is nothing to hypothesise about it. The unknown is rho, the correlation in the population the sample came from.
Prediction
Commit before reasoning.
Predict first
Why is this test two-tailed rather than one-tailed?
Correct: Any direction answers the question.
Why: The purpose is to decide whether the regression line may be used at all, and a strong negative relationship licenses that just as a strong positive one does. Since the alternative covers both, the p-value is the combined area in both tails — which is where the factor of two in the calculator command comes from.
Section
Section 3
Concept
The p-value is calculated using a t distribution with n minus two degrees of freedom. The test statistic is r times the square root of n minus two, divided by the square root of one minus r squared, and it has the same sign as r. The p-value is the combined area in both tails.
the test statistic — Built entirely from r and n, which is the section's whole point: the two are combined into one number whose distribution is known when rho is zero.
\[ t = \frac{r\sqrt{n-2}}{\sqrt{1-r^2}}, \qquad \text{df} = n - 2 \]
The formula makes the dependence on both quantities visible. A larger r raises the numerator and lowers the denominator, so the statistic grows quickly with the correlation; a larger n raises the numerator through the square root. Both push the statistic further into the tails, which is why either a stronger correlation or a bigger sample can produce significance.
Figure (svg): A t distribution on nine degrees of freedom with both tails shaded beyond plus and minus 2.66
OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, p. 632 — the calculation notes and the p-value method
Picture it
Both tails beyond 2.66 on nine degrees of freedom.
Figure (svg): A t distribution on nine degrees of freedom with both tails shaded beyond plus and minus 2.66
The shaded area totals 0.026, which is below 0.05 — so the correlation is significant and, since section 12.2 confirmed a linear pattern, the line may be used to predict final exam scores.
Worked example
The line of best fit is ŷ = -173.51 + 4.83x with r = 0.6631 and n = 11. Can the line be used for prediction?
\[ r = 0.6631, \; n = 11 \]
Degrees of freedom
Why: Eleven minus two.
\[ 9 \]
The numerator
Why: 0.6631 times root 9.
\[ 1.989 \]
The denominator
Why: Root of 1 minus 0.4397.
\[ 0.7485 \]
The statistic
Why: Divide.
\[ 2.6576 \]
Both tails on df 9
Why: Twice one tail.
\[ 0.026 \]
Figure (svg): The solution to Worked example the exam data by the p-value method shown as a ladder of expressions, one row per legal move
\[ t = 2.6576, \quad p = 0.026 < 0.05 \]
Verify: confirm the denominator is what section 12.3 computed
Why: One minus r squared is one minus 0.4397, which is 0.5603 — exactly the share of variation the line does NOT explain. So the test statistic is built from the explained and unexplained shares directly, and a line explaining more of the variation produces a smaller denominator and a larger statistic. The connection to r-squared is not a coincidence but the structure of the formula.
OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 632-633
Faded example
A correlation of 0.708 from nine data points, so r-squared is 0.5013.
Fill in the blanks
t = \frac7}}}2.65} = \frac______ = ___
Why: That is Example 12.10b, whose two-tailed p-value on seven degrees of freedom is 0.033 — below 0.05, agreeing with the critical-value verdict of significant.
Worked example
The calculator command the book gives is 2 times tcdf of the absolute t.
\[ 2 \times \text{tcdf}(|t|, 10^{99}, n-2) \]
The command's inner part
Why: Area beyond |t|.
For the exam data
Why: Beyond 2.6576 on df 9.
\[ 0.0131 \]
Double it
Why: The other tail matches.
\[ 0.0262 \]
Why
Why: The alternative is two-sided.
Figure (svg): The solution to Worked example why the area is doubled shown as a ladder of expressions, one row per legal move
\[ p = 2(0.0131) = 0.0262 \]
Verify: confirm the doubling is not optional
Why: Reporting the one-tailed 0.0131 would nearly halve the p-value and overstate the evidence. The t distribution is symmetric, so the two tails are equal and the factor is exactly two — but the reason to include it is the alternative's wording, not the symmetry. A one-tailed alternative would use a single tail even on the same symmetric curve.
OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, p. 632
Trap
\[ p = P(t > 2.6576) = 0.0131 \]
Take the area beyond the statistic
Why: It is where the evidence lies.
\[ \text{but } H_a: \rho \ne 0 \text{ covers both sides} \]
A correlation of -0.6631 would be equally strong evidence, so the lower tail counts too.
\[ p = 2 \times P(t > 2.6576) = 0.026 \]
Double the tail, because the alternative is two-sided
Why: The book's own calculator command does exactly this.
Here the decision survives either way, since both 0.0131 and 0.026 fall below 0.05. But a one-tailed 0.03 would double to 0.06 and reverse the verdict, so the habit matters more than this example shows.
Two truths and a lie
All three concern the p-value method.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The book states that the p-value is the combined area in both tails, and its calculator command multiplies the one-tailed area by two — because the alternative covers correlations of either sign.
Prediction
Commit before reasoning.
Predict first
The statistic divides by the square root of one minus r squared. What is that quantity?
Correct: The unexplained share.
Why: Section 12.3 defined one minus r-squared as the percent of variation in y not explained by variation in x. So the statistic is a ratio of explained to unexplained influence, scaled by the sample size — which is why a line explaining more of the variation yields a larger statistic.
Estimation
A correlation of 0.2 from 500 observations.
Predict first
Would that be significant at 5 percent?
Correct: Yes, comfortably.
Why: At 498 degrees of freedom the critical correlation is about 0.088, so 0.2 clears it easily. The correlation still explains only 4 percent of the variation, which is the tension section 12.3's r-squared was introduced to reveal: significance and importance are different questions.
Section
Section 4
Concept
The 95 percent critical values of the sample correlation coefficient table can be used to decide whether the computed value of r is significant. Compare r to the appropriate critical value for n minus two degrees of freedom: if r is not between the positive and negative critical values, then the correlation coefficient is significant.
the critical value — The smallest correlation, at a given degrees of freedom, that would be significant at 5 percent. It equals the two-tailed t critical value divided by the square root of df plus that value squared.
\[ r^* = \frac{t^*}{\sqrt{\text{df} + (t^*)^2}} \]
The identity is worth checking once. At nine degrees of freedom the two-tailed 5 percent t critical value is 2.262, and 2.262 divided by the root of 9 plus 5.117 is 0.602 — exactly the value the book quotes for the exam data. Every critical value the section uses reproduces the same way, which is what makes the two methods the same test rather than merely agreeing ones.
Figure (svg): A card showing that the critical-value table is the p-value method precomputed
OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 633-634 — the critical-value method and its worked cases
Picture it
How the table is produced.
Figure (svg): A card showing that the critical-value table is the p-value method precomputed
Because the table fixes t* at a 5 percent level, it can only answer that question — which is exactly the restriction the book states, and why the p-value method is the more flexible of the two.
Worked example
Two applications of the table.
\[ r = 0.801, n = 10; \quad r = 0.6631, n = 11 \]
First: df
Why: Ten minus two.
\[ 8 \]
Its critical values
Why: From the table.
\[ -0.632\text{ and } 0.632 \]
Compare
Why: 0.801 exceeds 0.632.
Second: df
Why: Eleven minus two.
\[ 9 \]
Compare against 0.602
Why: 0.6631 exceeds it.
Figure (svg): The solution to Worked example Example 12.7 and the exam data shown as a ladder of expressions, one row per legal move
\[ 0.801 > 0.632; \qquad 0.6631 > 0.602 \]
Verify: confirm the second margin is narrow, and what that implies
Why: The exam data's correlation exceeds its critical value by only 0.061 — a small margin, matching the p-value of 0.026 that sits only half a level below 0.05. Had there been one fewer student, the critical value at eight degrees of freedom would be 0.632 and the same correlation would fail. That fragility is worth reporting alongside the verdict.
OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 633-634
Faded example
A correlation is computed from 14 data points.
Fill in the blanks
\text14 = 12 - 2 = ___
Why: That is Examples 12.8 and 12.10c, both at twelve degrees of freedom with a critical value of 0.532 — one significant at -0.624 and one not at 0.134.
Worked example
Checking the table entry for nine degrees of freedom.
\[ \text{df} = 9 \]
The two-tailed t critical value
Why: At 5 percent, df 9.
\[ 2.262 \]
Square it
Why: Add to df.
\[ 9 + 5.117 = 14.117 \]
Take the root
Why: Of 14.117.
\[ 3.757 \]
Divide
Why: 2.262 over 3.757.
\[ 0.602 \]
Figure (svg): The solution to Worked example deriving a critical value shown as a ladder of expressions, one row per legal move
\[ r^* = \frac{2.262}{\sqrt{14.117}} = 0.602 \]
Verify: confirm the derivation by inverting the test statistic
Why: Setting the t statistic equal to t* and solving for r gives precisely this expression, which is why the two methods can never disagree. It also explains the book's restriction that a different significance level would need a different table: changing alpha changes t*, and every entry with it.
OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 632-634
Error analysis
Which are correct?
Annotate
On: \( \begin{aligned} &(1)\; \text{look up the critical value at df} = n - 2 \\ &(2)\; \text{look it up at df} = n - 1 \\ &(3)\; r \text{ outside the two critical values means significant} \\ &(4)\; \text{the table works at any significance level} \end{aligned} \)
Error (2) shifts the threshold in the direction of finding significance too readily, which makes it the more dangerous kind of mistake — it produces conclusions rather than obvious errors.
Two truths and a lie
All three concern the two methods.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The book notes that using the p-value method you could choose any appropriate significance level; only the table carries the restriction, because its entries were computed at one fixed alpha.
Sorting
Each pairs a correlation with its critical value.
Sort into buckets
Sort by whether r falls outside the critical values.
Items (c), (d) and (e) are Try Its 12.9, 12.8 and 12.7. Comparing magnitudes rather than signed values makes the negative case no harder than the positive one.
Prediction
Commit before reasoning.
Predict first
What fixes the table's significance level?
Correct: Each entry came from a t critical value at a fixed alpha.
Why: The entry at df is t* over the root of df plus t* squared, and t* depends on the significance level. Changing alpha to 1 percent would raise every t* and every entry with it — which is exactly why the book says a different level would need different tables not provided in this textbook.
Section
Section 5
Concept
If r is significant and the scatter plot shows a linear trend, the line can be used to predict the value of y for values of x within the domain of observed x values. If r is not significant, or if the scatter plot does not show a linear trend, the line should not be used for prediction. Even when both hold, the line may not be reliable outside the observed domain.
the three notes — Significance alone is not enough — the plot must also show a linear trend — and neither is enough to license prediction outside the range of observed x values.
\[ r \text{ significant} \;\wedge\; \text{linear trend} \;\Rightarrow\; \text{predict within } [x_{\min}, x_{\max}] \]
The first condition keeps section 12.2 in play. A significant correlation says the linear component of a relationship is real; it does not say the relationship is linear, which is exactly the situation Example 12.14's Consumer Price Index data presents — a significant r accompanying a pattern that the book says a curve would model better.
Figure (svg): The assumptions underlying the test of significance
OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 631-635 — the three notes, and the assumptions
Picture it
The book's own list.
Figure (svg): The assumptions underlying the test of significance
Assumptions two and three describe a picture: at every x, the y values are normally distributed about the line with the same spread. That is a strong claim, and section 12.3's residuals plot is the practical way to look for violations of it.
Worked example
The correlation is significant and section 12.2 confirmed a linear pattern. Third-exam scores ran from 65 to 75.
\[ \text{predict at } x = 73 \text{ or } x = 50? \]
Check significance
Why: p = 0.026.
Check the plot
Why: Linear trend.
Is 73 in range
Why: Between 65 and 75.
Is 50 in range
Why: Below 65.
Figure (svg): The solution to Worked example what the exam result licenses shown as a ladder of expressions, one row per legal move
\[ 65 \le 73 \le 75; \qquad 50 < 65 \]
Verify: confirm why the domain restriction is separate from significance
Why: The test says the linear relationship is real among students scoring 65 to 75. Nothing in the data speaks to students scoring 50, and the line's behaviour there rests entirely on an assumption that the same straight relationship continues — which no evidence supports. The book's own reminder makes exactly this point, and section 12.5 gives it a name.
OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 632-634
Sorting
Each describes a situation.
Sort into buckets
Sort by whether prediction is licensed.
Items (b), (c) and (d) each fail a different condition, which is why the book states three separate notes rather than one.
Worked example
Example 12.14's CPI data has a significant correlation of 0.8694 and a visibly bending pattern.
\[ r \text{ significant, pattern curved} \]
Check significance
Why: 0.8694 beats 0.532.
Check the plot
Why: It bends.
Apply the second note
Why: One condition fails.
What is needed
Why: A curve.
Figure (svg): The solution to Worked example significance with the wrong shape shown as a ladder of expressions, one row per legal move
\[ \text{significant} \;\nRightarrow\; \text{linear} \]
Verify: confirm the book reaches the same conclusion
Why: Its note says that although the correlation coefficient is significant, the pattern in the scatterplot indicates that a curve would be a more appropriate model, and that a statistician should prefer other methods. Then it adds the general lesson: in addition to doing the calculations, it is always important to look at the scatterplot when deciding whether a linear model is appropriate.
OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 631-634
Trap
\[ p = 0.004 \;\Rightarrow\; \text{use the line} \]
Take a significant correlation as clearance to predict
Why: The test was passed.
\[ \text{but the plot may bend, and } x \text{ may be out of range} \]
Two further conditions have to hold, and neither is tested by the p-value.
\[ \text{significant } \wedge \text{ linear trend } \wedge \; x \text{ in range} \]
Check all three before predicting
Why: The book states them as three separate notes.
The three fail in different ways and are caught by different means: significance by the test, linearity by the scatter plot and residuals plot, and range by comparing x against the observed minimum and maximum. Only the first is a calculation, which is why the other two are the ones usually skipped.
Two truths and a lie
All three concern the assumptions.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and reverses the third assumption, which is that the standard deviations of the y values about the line are EQUAL for each value of x. Spread that grows with x is a violation, and section 12.3 names it: the residuals should not consistently increase as x increases.
Prediction
Commit before reasoning.
Predict first
Which of the five assumptions does a residuals plot help check?
Correct: Linearity, equal spread and independence.
Why: Section 12.3 says a residuals plot should appear random with no pattern and no outliers, and should show constant error variance. A curve in the residuals violates linearity, a widening band violates equal spread, and any systematic pattern violates independence — three assumptions in one picture.
Explain it
A classmate has a significant correlation and concludes the relationship must be linear.
Discussion prompt
In two sentences or fewer, correct them.
Hint: Ask what the test actually measures.
Answer:
The test measures whether the LINEAR component of the relationship is more than chance would produce, which a curved relationship can easily pass since a rising curve has a strong linear component.
Example 12.14 is exactly that case: a significant correlation of 0.8694 alongside a pattern the book says a curve would model better.
Comparison
Fill the blanks. They are one test written twice.
Comparison matrix
| p-value method | Critical-value method | |
|---|---|---|
| What is computed | a t statistic from r and n | nothing: r is looked up against a table |
| Degrees of freedom | n - 2 | n - 2 |
| Decision rule | reject if p is below alpha | significant if r is outside the critical values |
| Significance levels | any | 5 percent only |
The last row is the only genuine difference, and it follows from the third: a table of critical values has to fix alpha before it can list anything, while a p-value is computed after the fact and can be compared against any level.
Pattern
Six steps, and the last two are conditions rather than calculations.
The two methods cannot disagree, since each critical value is the t critical value solved for r at a 5 percent level.
OpenStax Introductory Business Statistics 2e, §13.2 Testing the Significance of the Correlation Coefficient §13.2 Testing the Significance of the Correlation Coefficient
Check
Degrees of freedom.
Check your understanding
A correlation is computed from 20 paired observations. What degrees of freedom does the test use?
Answer: A
Why: The degrees of freedom are n minus two, because a regression line uses two estimates: an intercept and a slope.
Check
The verdict.
Check your understanding
A correlation of 0.55 comes from 8 data points, where the critical value is 0.707. What follows?
Answer: A
Why: 0.55 lies between -0.707 and 0.707, so it is not significantly different from zero and the line should not be used.
Check
What significance licenses.
Check your understanding
A correlation is significant and the scatter plot shows a clear curve. May the line be used to predict?
Answer: A
Why: The book's note is explicit: if the scatter plot does not show a linear trend, the line should not be used for prediction, whatever the p-value says.
Real world
A pharmaceutical analysis of 4,000 patients reports a correlation of 0.06 between a drug's dose and symptom improvement, with a p-value of 0.0001, and the press release states that the study found a highly significant relationship between dose and improvement.
Discussion prompt
Is the statistical claim correct, and what should a reader take from it?
Hint: Compute how much of the variation the correlation explains.
Answer:
The statistical claim is correct and almost meaningless. With 4,000 patients the critical correlation at 5 percent is about 0.031, so 0.06 clears it comfortably and the tiny p-value is exactly what the arithmetic gives. The linear relationship is real.
\[ r^2 = 0.06^2 = 0.0036 \;\Rightarrow\; 0.36\% \text{ of the variation explained} \]
Dose explains about a third of one percent of the variation in improvement, leaving 99.6 percent to everything else. The phrase highly significant describes the p-value, which measures confidence that the effect is non-zero — not its size. At this sample size a correlation far too small to matter clinically is detected with near-certainty.
This is the mirror image of Example 12.9. There, a correlation of 0.776 from six points was a strong relationship with weak evidence; here 0.06 from four thousand is negligible with overwhelming evidence. Strength and evidence move independently, and a p-value reports only the second.
What a reader should take: the drug's effect on symptoms is almost certainly not zero, and is very small. The honest report gives r-squared alongside the p-value, states the effect in the outcome's own units, and lets a clinician judge whether an effect that size is worth anything. Reporting significance alone, at this sample size, tells the reader nothing they can act on.
Commit first
Answer, then rate your confidence honestly.
Predict first
Why is a correlation of 0.776 from 6 points not significant while 0.567 from 19 points is?
Correct: The critical value falls as the sample grows.
\[ r^* = \frac{t^*}{\sqrt{\text{df} + (t^*)^2}}: \quad 0.811 \text{ at df } 4, \quad 0.456 \text{ at df } 17 \]
Why: At four degrees of freedom the threshold is 0.811 and at seventeen it is 0.456. A small sample produces large correlations by chance quite easily, so a great deal of correlation is needed to rule chance out; a larger sample makes a modest correlation hard to explain that way. Strength and evidence are different quantities, and the test measures the second.
Explain it
They looked up the critical value at n minus one degrees of freedom instead of n minus two.
Discussion prompt
In two sentences or fewer, correct them.
Hint: Ask how many quantities the line estimates.
Answer:
A regression line estimates two things, an intercept and a slope, so two degrees of freedom are spent and the test uses n minus two.
Using n minus one gives a critical value that is slightly too small, which errs in the direction of declaring a correlation significant when it is not.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, small samples produce large correlations by chance. For the second, r times root n minus two over root one minus r squared, then double. For the third, each entry is a t critical value solved for r at one fixed alpha. For the fourth, significant, linear, and inside the range. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top, write the hypotheses in symbols and beneath each its wording in words, marking clearly that rho is the unknown population correlation and r the computed sample one. Underneath, write the test statistic as r times the root of n minus two over the root of one minus r squared, note that the degrees of freedom are n minus two, and work the exam data through it: 0.6631 times 3 over the root of 0.5603, giving 2.6576. Beside that, draw a t curve on nine degrees of freedom with both tails shaded beyond plus and minus 2.66, labelled with a combined area of 0.026. In the middle, write the critical-value method and the identity that generates it — critical r equals t star over the root of df plus t star squared — then verify it once at nine degrees of freedom: 2.262 over the root of 14.117 is 0.602. Below, make a five-row table of the book's cases with columns for r, n, df, critical value and verdict, and circle the row where 0.776 fails while a smaller correlation elsewhere passes. At the bottom, write the three conditions for using the line in a box: r significant, a linear trend in the plot, and x inside the observed range.
Check your critical-value derivation by trying it at four degrees of freedom, where t star is 2.776 and the answer should be 0.811 — the value that made Example 12.9 fail. Check your table by confirming that the verdicts do not simply follow the order of the correlations.
Recap
Six things, and two of them are conditions the p-value does not check.
| If you see | Then |
|---|---|
| A correlation to test | Degrees of freedom are n minus two |
| A t statistic computed | Double the tail: the alternative is two-sided |
| r outside the critical values | Significant at 5 percent |
| r between them | Not significant; do not use the line |
| r equal to zero | Never significant, at any sample size |
| A significant r with a curved plot | Do not use the line: a curve is needed |
| A prediction outside the observed x range | Not licensed, however significant r is |
| A large sample and a tiny r | Significant and unimportant; report r-squared too |
Section 12.5 uses the line for what it was fitted for. With the correlation established as significant and the pattern confirmed linear, predictions can be made — and the section draws the line between interpolation, which is licensed, and extrapolation, which is not.
OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 631-635 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.