12.4 Testing the Significance of the Correlation Coefficient

Section 12.3 computed a correlation of 0.6631 from eleven students and left open whether that is enough to conclude anything. The correlation coefficient tells us about the strength and direction of the linear relationship, but the reliability of the linear model also depends on how many observed data points are in the sample — so r and n have to be looked at together. This section performs a hypothesis test of the significance of the correlation coefficient, with a null that the population correlation rho equals zero and a two-tailed alternative that it does not. The book gives two methods, a p-value from a t statistic on n minus two degrees of freedom and a table of critical values at a fixed 5 percent level, and calls them equivalent. They are in fact the same test written twice: every critical value in the table equals the two-tailed t critical value divided by the square root of df plus that value squared.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 12.4 Testing the Significance of the Correlation Coefficient

Title

Statistics · Chapter 12 — Linear Regression and Correlation

Testing the Significance of the Correlation Coefficient

2. By the end of this lesson you can

Objectives

Six outcomes, and the second explains why the section exists at all.

OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 631-635 — the section these objectives are drawn from

3. What you already have

Warm-up

Section 12.3 produced r = 0.6631 from eleven students; chapter 9 tested claims about parameters.

Discussion prompt

If two variables were completely unrelated in the population, would a sample of eleven give a correlation of exactly zero?

Hint: Ask what any sample statistic does around its parameter.

Answer:

No. A sample correlation varies from sample to sample just as a sample mean does, so even with no relationship at all in the population a sample of eleven will produce some non-zero r — sometimes a fairly large one, purely by chance.

So the question is not whether r differs from zero, but whether it differs by more than chance would readily produce at this sample size. That is a hypothesis test, and its parameter is the population correlation coefficient.

\[ H_0: \rho = 0 \qquad\text{against}\qquad H_a: \rho \ne 0 \]

The sample size enters directly. Eleven points can look convincingly linear by accident far more easily than a hundred can, which is why the book insists on looking at the value of r and the sample size n together.

4. Is r large enough, given n?

Concept

The correlation coefficient tells us about the strength and direction of the linear relationship, but the reliability of the linear model also depends on how many observed data points are in the sample. We perform a hypothesis test of the significance of the correlation coefficient to decide whether the linear relationship in the sample data is strong enough to use to model the relationship in the population.

rho against r — Rho is the population correlation coefficient and is unknown; r is the sample correlation coefficient, computed from the data, and is our estimate of it.

\[ H_0: \rho = 0, \qquad H_a: \rho \ne 0, \qquad \alpha = 0.05 \]

The book fixes the significance level for the whole chapter: we will always use a significance level of 5 percent. That is not a statistical necessity but a consequence of the tool — the table of critical values provided assumes 5 percent, and a different level would need a different table. The p-value method carries no such restriction.

Figure (svg): A card showing that the same correlation means different things at different sample sizes

Examples 12.7, 12.9 and 12.10a: the correlation alone never settles the question.

OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 631-632

5. Why the sample size matters

Section

Section 1

6. r and n together

Concept

The sample data are used to compute r, the correlation coefficient for the sample. Because we have only sample data, we cannot calculate the population correlation coefficient, so r is our estimate of the unknown rho. Whether an observed r is convincing depends on how many points produced it.

significant — If the test concludes that the correlation coefficient is significantly different from zero, we say the correlation coefficient is significant, and the regression line may be used to model the relationship in the population.

\[ \text{same } r, \text{ different } n \;\Longrightarrow\; \text{different verdicts} \]

The dependence runs in the direction intuition suggests but more sharply than most people expect. At six points a correlation of 0.776 is not significant; at nineteen a correlation of 0.567 is. Doubling the sample roughly halves the correlation needed, so sample size buys a great deal here.

Figure (svg): A card showing that the same correlation means different things at different sample sizes

Examples 12.7, 12.9 and 12.10a: the correlation alone never settles the question.

OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 631-634 — we need to look at both r and the sample size together

7. Three cases

Picture it

Examples 12.7, 12.9 and 12.10a.

Figure (svg): A card showing that the same correlation means different things at different sample sizes

Examples 12.7, 12.9 and 12.10a: the correlation alone never settles the question.

The middle panel is significant with a correlation lower than the left panel's, which is not. Reading r alone would rank them the other way round, which is precisely the error the section prevents.

8. Worked example: Example 12.9, a large r that fails

Worked example

A correlation of 0.776 computed from six data points.

\[ r = 0.776, \; n = 6 \]

Degrees of freedom

Why: Six minus two.

\[ 4 \]

Look up the critical values

Why: At df 4.

\[ -0.811\text{ and } 0.811 \]

Compare

Why: 0.776 sits between.

Conclude

Why: Not significant.

Figure (svg): The solution to Worked example Example 12.9, a large r that fails shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ -0.811 < 0.776 < 0.811 \]

Verify: confirm the same verdict by the p-value route

Why: The t statistic is 0.776 times the root of 4 over the root of 1 minus 0.602, which is 1.552 over 0.631, or 2.46. On four degrees of freedom the two-tailed p-value is 0.070 — above 0.05, so the same do-not-reject. The two methods agree because they are the same test, which is the subject of a later idea in this lesson.

OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 633-634

9. Significant at 5 percent?

Sorting

Each gives a correlation, a sample size and the critical value.

Sort into buckets

Sort by the verdict.

Significant
r = 0.801, n = 10, critical 0.632; r = 0.708, n = 9, critical 0.666; r = -0.624, n = 14, critical 0.532
Not significant
r = 0.776, n = 6, critical 0.811; r = 0.134, n = 14, critical 0.532
yes
The magnitude of r exceeds the critical value.
no
r lies between the positive and negative critical values.

These are Examples 12.7 through 12.10. Item (b) has a larger correlation than item (c) and fails, because it rests on six points rather than nine.

10. Worked example: Example 12.10a, a smaller r that passes

Worked example

A correlation of -0.567 computed from nineteen data points.

\[ r = -0.567, \; n = 19 \]

Degrees of freedom

Why: Nineteen minus two.

\[ 17 \]

Look up the critical value

Why: At df 17.

\[ -0.456 \]

Compare

Why: -0.567 is further out.

Conclude

Why: Significant.

Figure (svg): The solution to Worked example Example 12.10a, a smaller r that passes shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ -0.567 < -0.456 \]

Verify: confirm the comparison is about distance from zero

Why: For a negative correlation the test asks whether r falls BELOW the negative critical value, which is the mirror of asking whether a positive r exceeds the positive one. Both amount to asking whether the magnitude exceeds 0.456, and stating it that way avoids the sign errors that comparing negative numbers invites.

OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, p. 634

11. Trap: judging significance from r alone

Trap

The trap

\[ r = 0.776 > 0.567 \;\Rightarrow\; \text{the first is the stronger finding} \]

Rank two results by their correlations

Why: A larger correlation is a tighter relationship.

\[ \text{but } n = 6 \text{ against } n = 19 \]

The larger correlation comes from six points and is not significant; the smaller comes from nineteen and is.

The fix

\[ \text{compare each } r \text{ against its own critical value} \]

Read r and n together, as the book insists

Why: The threshold moves with the sample size.

Both statements can be true at once: the first sample shows a tighter relationship AND provides weaker evidence that any relationship exists. Strength and evidence are different quantities, and conflating them is the most common misreading of a correlation.

12. One of these is false

Two truths and a lie

All three concern sample size.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. A larger sample makes a smaller correlation significant
  • C. r estimates the unknown population correlation rho
  • B. A correlation above 0.7 is always significant

Survives elimination: B

Why: The survivor is false, and Example 12.9 is the counterexample: 0.776 on six data points is not significant, because the critical value at four degrees of freedom is 0.811.

13. How much does n buy?

Estimation

The critical value is 0.811 at n = 6 and 0.456 at n = 19.

Predict first

Roughly what happens to the critical value as the sample grows?

  • It falls steadily, so weaker correlations become detectable
  • It rises, so stronger correlations are needed
  • It stays about the same
  • It becomes negative

Correct: It falls steadily.

Why: More data makes a given correlation harder to attribute to chance, so the bar drops. Tripling the sample from 6 to 19 nearly halves the required correlation, and at n = 100 the critical value is about 0.197 — which is why large studies routinely report significant correlations that explain very little variation.

14. What does r = 0 give?

Prediction

Commit before reasoning.

Predict first

Example 12.10d has r = 0 with n = 5. What is the verdict, and does n matter?

  • Not significant, and no sample size would change it
  • Significant, since zero is a definite value
  • Not significant, but a larger n would make it so
  • The test cannot be run

Correct: Not significant, whatever the sample size.

Why: The book says it directly: no matter what the degrees of freedom are, r equal to zero lies between the two critical values. Zero is the null's own value, so it can never be evidence against the null — which is why Try It 12.10 poses the same question at n = 100 and gets the same answer.

15. The hypotheses

Section

Section 2

16. A Greek parameter, and no direction

Concept

The null hypothesis is that the population correlation coefficient rho equals zero; the alternative is that it does not. In words, the null says there is not a significant linear relationship between x and y in the population, and the alternative says there is.

what significant means here — That the correlation coefficient is significantly different from zero. If so, the regression line may be used to model the linear relationship between x and y in the population.

\[ H_0: \rho = 0, \qquad H_a: \rho \ne 0 \]

The alternative names no direction, so the test is two-tailed throughout. That is a deliberate choice: the question being asked is whether ANY linear relationship exists, and a relationship running either way answers it. It also means the p-value is the combined area in both tails, which is where a factor of two enters the arithmetic.

Figure (svg): A card giving the hypotheses in symbols and in words

The first null in the chapter with a Greek parameter, and it is tested two-tailed throughout.

OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 631-632 — the hypotheses in symbols and in words

17. The hypotheses

Picture it

Symbols above, words below.

Figure (svg): A card giving the hypotheses in symbols and in words

The first null in the chapter with a Greek parameter, and it is tested two-tailed throughout.

This is the chapter's first null with a Greek parameter, and it restores chapter 9's structure after three sections of description. What it does not restore is a choice of tail.

18. Worked example: stating the conclusion in words

Worked example

The exam data, where the test rejects.

\[ \text{reject } H_0 \]

Name the verdict

Why: Reject.

Say why

Why: r differs from zero.

Name the variables

Why: In context.

Add what it licenses

Why: With a linear plot.

Figure (svg): The solution to Worked example stating the conclusion in words shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{reject } H_0: \rho = 0 \]

Verify: confirm the second clause is doing real work

Why: The book's wording always gives the reason — because the correlation coefficient is significantly different from zero — rather than stopping at the claim. That matters because a reader needs to know what was tested: not that the variables are related in some general sense, but that their LINEAR association is more than chance would produce at this sample size.

OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 632-633

19. Write the hypotheses

Faded example

For a test of the significance of a correlation coefficient.

Fill in the blanks

H_0: \rho = 0, \qquad H_a: \rho ≠ 0

Why: Two-tailed throughout, because the question is whether any linear relationship exists rather than whether it runs in a particular direction.

20. Worked example: the conclusion when the test fails

Worked example

Example 12.9, where 0.776 on six points is not significant.

\[ \text{do not reject } H_0 \]

Name the verdict

Why: Do not reject.

State it carefully

Why: Insufficient evidence.

What it forbids

Why: Using the line.

Note the reason

Why: Only six points.

Figure (svg): The solution to Worked example the conclusion when the test fails shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{do not reject}: \; \text{no prediction} \]

Verify: confirm what failing to reject does not establish

Why: It does not establish that rho is zero. With six observations the test can only detect correlations above 0.811, so a real population correlation of 0.7 would usually go undetected here. Chapter 9's asymmetry applies unchanged: failing to reject means the data are consistent with the null, not that the null is true.

OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 631-634

21. Error analysis: four statements of the hypotheses

Error analysis

Which are correct?

Annotate

On: \( \begin{aligned} &(1)\; H_0: \rho = 0, \; H_a: \rho \ne 0 \\ &(2)\; H_0: r = 0, \; H_a: r \ne 0 \\ &(3)\; H_0: \rho = 0, \; H_a: \rho > 0 \\ &(4)\; H_0: \rho \ne 0, \; H_a: \rho = 0 \end{aligned} \)

  • (1) is the book's own statement, and it is two-tailed.
  • (2) uses r, the SAMPLE correlation. Hypotheses are always about a population parameter; r is the evidence.
  • (3) is one-tailed. The book's alternative names no direction, since the question is whether any linear relationship exists.
  • (4) reverses them. The null must be the equality, since only a specific value generates a distribution for the statistic.

Error (2) is the same confusion chapter 11's variance test invited between sigma and s, and the fix is the same: the Greek letter is the unknown being claimed about, and the Roman letter is what the sample supplies.

22. Symbol to meaning

Matching

Match each symbol to what it denotes.

Match the pairs

  • l1. rho
  • l2. r
  • l3. n - 2
  • l4. alpha
  • r1. the population correlation coefficient, unknown
  • r2. the sample correlation coefficient, computed
  • r3. the degrees of freedom for the test
  • r4. the significance level, fixed at 0.05 in this chapter

Why: The degrees of freedom are n minus two because the line uses two estimates — an intercept and a slope — which is the same accounting section 12.6 gives for dividing the SSE by n minus two.

23. One of these is false

Two truths and a lie

All three concern the hypotheses.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The test is two-tailed
  • C. The hypotheses concern the population, not the sample
  • B. The alternative is that the sample correlation differs from zero

Survives elimination: B

Why: The survivor is false. The sample correlation is known — it was computed in section 12.3 — so there is nothing to hypothesise about it. The unknown is rho, the correlation in the population the sample came from.

24. Why two-tailed?

Prediction

Commit before reasoning.

Predict first

Why is this test two-tailed rather than one-tailed?

  • The question is whether ANY linear relationship exists, and either direction answers it
  • Because r can be negative
  • Because the t distribution is symmetric
  • Because alpha is 0.05

Correct: Any direction answers the question.

Why: The purpose is to decide whether the regression line may be used at all, and a strong negative relationship licenses that just as a strong positive one does. Since the alternative covers both, the p-value is the combined area in both tails — which is where the factor of two in the calculator command comes from.

25. The p-value method

Section

Section 3

26. A t statistic on n minus two degrees of freedom

Concept

The p-value is calculated using a t distribution with n minus two degrees of freedom. The test statistic is r times the square root of n minus two, divided by the square root of one minus r squared, and it has the same sign as r. The p-value is the combined area in both tails.

the test statistic — Built entirely from r and n, which is the section's whole point: the two are combined into one number whose distribution is known when rho is zero.

\[ t = \frac{r\sqrt{n-2}}{\sqrt{1-r^2}}, \qquad \text{df} = n - 2 \]

The formula makes the dependence on both quantities visible. A larger r raises the numerator and lowers the denominator, so the statistic grows quickly with the correlation; a larger n raises the numerator through the square root. Both push the statistic further into the tails, which is why either a stronger correlation or a bigger sample can produce significance.

Figure (svg): A t distribution on nine degrees of freedom with both tails shaded beyond plus and minus 2.66

The test is always two-tailed, so the p-value is the combined area — twice the area beyond 2.66 alone.

OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, p. 632 — the calculation notes and the p-value method

27. The exam data's p-value

Picture it

Both tails beyond 2.66 on nine degrees of freedom.

Figure (svg): A t distribution on nine degrees of freedom with both tails shaded beyond plus and minus 2.66

The test is always two-tailed, so the p-value is the combined area — twice the area beyond 2.66 alone.

The shaded area totals 0.026, which is below 0.05 — so the correlation is significant and, since section 12.2 confirmed a linear pattern, the line may be used to predict final exam scores.

28. Worked example: the exam data by the p-value method

Worked example

The line of best fit is ŷ = -173.51 + 4.83x with r = 0.6631 and n = 11. Can the line be used for prediction?

\[ r = 0.6631, \; n = 11 \]

Degrees of freedom

Why: Eleven minus two.

\[ 9 \]

The numerator

Why: 0.6631 times root 9.

\[ 1.989 \]

The denominator

Why: Root of 1 minus 0.4397.

\[ 0.7485 \]

The statistic

Why: Divide.

\[ 2.6576 \]

Both tails on df 9

Why: Twice one tail.

\[ 0.026 \]

Figure (svg): The solution to Worked example the exam data by the p-value method shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ t = 2.6576, \quad p = 0.026 < 0.05 \]

Verify: confirm the denominator is what section 12.3 computed

Why: One minus r squared is one minus 0.4397, which is 0.5603 — exactly the share of variation the line does NOT explain. So the test statistic is built from the explained and unexplained shares directly, and a line explaining more of the variation produces a smaller denominator and a larger statistic. The connection to r-squared is not a coincidence but the structure of the formula.

OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 632-633

29. Compute the statistic

Faded example

A correlation of 0.708 from nine data points, so r-squared is 0.5013.

Fill in the blanks

t = \frac7}}}2.65} = \frac______ = ___

Why: That is Example 12.10b, whose two-tailed p-value on seven degrees of freedom is 0.033 — below 0.05, agreeing with the critical-value verdict of significant.

30. Worked example: why the area is doubled

Worked example

The calculator command the book gives is 2 times tcdf of the absolute t.

\[ 2 \times \text{tcdf}(|t|, 10^{99}, n-2) \]

The command's inner part

Why: Area beyond |t|.

For the exam data

Why: Beyond 2.6576 on df 9.

\[ 0.0131 \]

Double it

Why: The other tail matches.

\[ 0.0262 \]

Why

Why: The alternative is two-sided.

Figure (svg): The solution to Worked example why the area is doubled shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ p = 2(0.0131) = 0.0262 \]

Verify: confirm the doubling is not optional

Why: Reporting the one-tailed 0.0131 would nearly halve the p-value and overstate the evidence. The t distribution is symmetric, so the two tails are equal and the factor is exactly two — but the reason to include it is the alternative's wording, not the symmetry. A one-tailed alternative would use a single tail even on the same symmetric curve.

OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, p. 632

31. Trap: reporting the one-tailed area

Trap

The trap

\[ p = P(t > 2.6576) = 0.0131 \]

Take the area beyond the statistic

Why: It is where the evidence lies.

\[ \text{but } H_a: \rho \ne 0 \text{ covers both sides} \]

A correlation of -0.6631 would be equally strong evidence, so the lower tail counts too.

The fix

\[ p = 2 \times P(t > 2.6576) = 0.026 \]

Double the tail, because the alternative is two-sided

Why: The book's own calculator command does exactly this.

Here the decision survives either way, since both 0.0131 and 0.026 fall below 0.05. But a one-tailed 0.03 would double to 0.06 and reverse the verdict, so the habit matters more than this example shows.

32. One of these is false

Two truths and a lie

All three concern the p-value method.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The degrees of freedom are n minus two
  • C. The statistic has the same sign as r
  • B. The p-value is the area in one tail

Survives elimination: B

Why: The survivor is false. The book states that the p-value is the combined area in both tails, and its calculator command multiplies the one-tailed area by two — because the alternative covers correlations of either sign.

33. What is in the denominator?

Prediction

Commit before reasoning.

Predict first

The statistic divides by the square root of one minus r squared. What is that quantity?

  • The share of variation the line does NOT explain
  • The share it does explain
  • The standard error of the slope
  • The sample size

Correct: The unexplained share.

Why: Section 12.3 defined one minus r-squared as the percent of variation in y not explained by variation in x. So the statistic is a ratio of explained to unexplained influence, scaled by the sample size — which is why a line explaining more of the variation yields a larger statistic.

34. Predict the verdict

Estimation

A correlation of 0.2 from 500 observations.

Predict first

Would that be significant at 5 percent?

  • Yes, comfortably: the critical value at that sample size is under 0.09
  • No: 0.2 is a weak correlation
  • Exactly borderline
  • Impossible to say

Correct: Yes, comfortably.

Why: At 498 degrees of freedom the critical correlation is about 0.088, so 0.2 clears it easily. The correlation still explains only 4 percent of the variation, which is the tension section 12.3's r-squared was introduced to reveal: significance and importance are different questions.

35. The critical-value method

Section

Section 4

36. The same test, solved for r

Concept

The 95 percent critical values of the sample correlation coefficient table can be used to decide whether the computed value of r is significant. Compare r to the appropriate critical value for n minus two degrees of freedom: if r is not between the positive and negative critical values, then the correlation coefficient is significant.

the critical value — The smallest correlation, at a given degrees of freedom, that would be significant at 5 percent. It equals the two-tailed t critical value divided by the square root of df plus that value squared.

\[ r^* = \frac{t^*}{\sqrt{\text{df} + (t^*)^2}} \]

The identity is worth checking once. At nine degrees of freedom the two-tailed 5 percent t critical value is 2.262, and 2.262 divided by the root of 9 plus 5.117 is 0.602 — exactly the value the book quotes for the exam data. Every critical value the section uses reproduces the same way, which is what makes the two methods the same test rather than merely agreeing ones.

Figure (svg): A card showing that the critical-value table is the p-value method precomputed

The book calls the two methods equivalent. They are the same test: one solves for a p-value, the other for r.

OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 633-634 — the critical-value method and its worked cases

37. Two methods, one test

Picture it

How the table is produced.

Figure (svg): A card showing that the critical-value table is the p-value method precomputed

The book calls the two methods equivalent. They are the same test: one solves for a p-value, the other for r.

Because the table fixes t* at a 5 percent level, it can only answer that question — which is exactly the restriction the book states, and why the p-value method is the more flexible of the two.

38. Worked example: Example 12.7 and the exam data

Worked example

Two applications of the table.

\[ r = 0.801, n = 10; \quad r = 0.6631, n = 11 \]

First: df

Why: Ten minus two.

\[ 8 \]

Its critical values

Why: From the table.

\[ -0.632\text{ and } 0.632 \]

Compare

Why: 0.801 exceeds 0.632.

Second: df

Why: Eleven minus two.

\[ 9 \]

Compare against 0.602

Why: 0.6631 exceeds it.

Figure (svg): The solution to Worked example Example 12.7 and the exam data shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 0.801 > 0.632; \qquad 0.6631 > 0.602 \]

Verify: confirm the second margin is narrow, and what that implies

Why: The exam data's correlation exceeds its critical value by only 0.061 — a small margin, matching the p-value of 0.026 that sits only half a level below 0.05. Had there been one fewer student, the critical value at eight degrees of freedom would be 0.632 and the same correlation would fail. That fragility is worth reporting alongside the verdict.

OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 633-634

39. Find the degrees of freedom

Faded example

A correlation is computed from 14 data points.

Fill in the blanks

\text14 = 12 - 2 = ___

Why: That is Examples 12.8 and 12.10c, both at twelve degrees of freedom with a critical value of 0.532 — one significant at -0.624 and one not at 0.134.

40. Worked example: deriving a critical value

Worked example

Checking the table entry for nine degrees of freedom.

\[ \text{df} = 9 \]

The two-tailed t critical value

Why: At 5 percent, df 9.

\[ 2.262 \]

Square it

Why: Add to df.

\[ 9 + 5.117 = 14.117 \]

Take the root

Why: Of 14.117.

\[ 3.757 \]

Divide

Why: 2.262 over 3.757.

\[ 0.602 \]

Figure (svg): The solution to Worked example deriving a critical value shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ r^* = \frac{2.262}{\sqrt{14.117}} = 0.602 \]

Verify: confirm the derivation by inverting the test statistic

Why: Setting the t statistic equal to t* and solving for r gives precisely this expression, which is why the two methods can never disagree. It also explains the book's restriction that a different significance level would need a different table: changing alpha changes t*, and every entry with it.

OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 632-634

41. Error analysis: four uses of the table

Error analysis

Which are correct?

Annotate

On: \( \begin{aligned} &(1)\; \text{look up the critical value at df} = n - 2 \\ &(2)\; \text{look it up at df} = n - 1 \\ &(3)\; r \text{ outside the two critical values means significant} \\ &(4)\; \text{the table works at any significance level} \end{aligned} \)

  • (1) is correct: a regression uses two estimates, so two degrees of freedom are spent.
  • (2) borrows the rule from a one-sample t test, and gives a critical value that is slightly too small.
  • (3) is the book's own criterion, stated as: if r is not between the positive and negative critical values, then r is significant.
  • (4) is false, and the book says so: the table provided assumes a significance level of 5 percent.

Error (2) shifts the threshold in the direction of finding significance too readily, which makes it the more dangerous kind of mistake — it produces conclusions rather than obvious errors.

42. One of these is false

Two truths and a lie

All three concern the two methods.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The two methods always give the same verdict
  • C. The table assumes a 5 percent significance level
  • B. The p-value method is also restricted to 5 percent

Survives elimination: B

Why: The survivor is false. The book notes that using the p-value method you could choose any appropriate significance level; only the table carries the restriction, because its entries were computed at one fixed alpha.

43. Inside or outside?

Sorting

Each pairs a correlation with its critical value.

Sort into buckets

Sort by whether r falls outside the critical values.

Outside: significant
r = 0.6631, critical 0.602; r = -0.7204, critical 0.707; r = 0.6501, critical 0.576
Inside: not significant
r = 0.776, critical 0.811; r = 0.5204, critical 0.666
out
The magnitude of r exceeds the critical value.
in
The magnitude falls short, so r lies between the two critical values.

Items (c), (d) and (e) are Try Its 12.9, 12.8 and 12.7. Comparing magnitudes rather than signed values makes the negative case no harder than the positive one.

44. Why can the table only do 5 percent?

Prediction

Commit before reasoning.

Predict first

What fixes the table's significance level?

  • Each entry was computed from the t critical value at that alpha
  • The sample sizes it covers
  • The number of decimal places
  • The two-tailed alternative

Correct: Each entry came from a t critical value at a fixed alpha.

Why: The entry at df is t* over the root of df plus t* squared, and t* depends on the significance level. Changing alpha to 1 percent would raise every t* and every entry with it — which is exactly why the book says a different level would need different tables not provided in this textbook.

45. What significance licenses, and what it assumes

Section

Section 5

46. Three conditions, and five assumptions

Concept

If r is significant and the scatter plot shows a linear trend, the line can be used to predict the value of y for values of x within the domain of observed x values. If r is not significant, or if the scatter plot does not show a linear trend, the line should not be used for prediction. Even when both hold, the line may not be reliable outside the observed domain.

the three notes — Significance alone is not enough — the plot must also show a linear trend — and neither is enough to license prediction outside the range of observed x values.

\[ r \text{ significant} \;\wedge\; \text{linear trend} \;\Rightarrow\; \text{predict within } [x_{\min}, x_{\max}] \]

The first condition keeps section 12.2 in play. A significant correlation says the linear component of a relationship is real; it does not say the relationship is linear, which is exactly the situation Example 12.14's Consumer Price Index data presents — a significant r accompanying a pattern that the book says a curve would model better.

Figure (svg): The assumptions underlying the test of significance

The book's own list. Section 12.3's residuals plot is the practical check on the first, third and fourth.

OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 631-635 — the three notes, and the assumptions

47. The five assumptions

Picture it

The book's own list.

Figure (svg): The assumptions underlying the test of significance

The book's own list. Section 12.3's residuals plot is the practical check on the first, third and fourth.

Assumptions two and three describe a picture: at every x, the y values are normally distributed about the line with the same spread. That is a strong claim, and section 12.3's residuals plot is the practical way to look for violations of it.

48. Worked example: what the exam result licenses

Worked example

The correlation is significant and section 12.2 confirmed a linear pattern. Third-exam scores ran from 65 to 75.

\[ \text{predict at } x = 73 \text{ or } x = 50? \]

Check significance

Why: p = 0.026.

Check the plot

Why: Linear trend.

Is 73 in range

Why: Between 65 and 75.

Is 50 in range

Why: Below 65.

Figure (svg): The solution to Worked example what the exam result licenses shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 65 \le 73 \le 75; \qquad 50 < 65 \]

Verify: confirm why the domain restriction is separate from significance

Why: The test says the linear relationship is real among students scoring 65 to 75. Nothing in the data speaks to students scoring 50, and the line's behaviour there rests entirely on an assumption that the same straight relationship continues — which no evidence supports. The book's own reminder makes exactly this point, and section 12.5 gives it a name.

OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 632-634

49. May the line be used to predict?

Sorting

Each describes a situation.

Sort into buckets

Sort by whether prediction is licensed.

Prediction licensed
r significant, plot linear, x inside the observed range; r significant, plot linear, x at the edge of the range
Not licensed
r significant, plot clearly curved; r not significant, plot linear; r significant, plot linear, x far outside the range
yes
All three conditions hold: significant, linear, and within the observed domain.
no
At least one of the three conditions fails.

Items (b), (c) and (d) each fail a different condition, which is why the book states three separate notes rather than one.

50. Worked example: significance with the wrong shape

Worked example

Example 12.14's CPI data has a significant correlation of 0.8694 and a visibly bending pattern.

\[ r \text{ significant, pattern curved} \]

Check significance

Why: 0.8694 beats 0.532.

Check the plot

Why: It bends.

Apply the second note

Why: One condition fails.

What is needed

Why: A curve.

Figure (svg): The solution to Worked example significance with the wrong shape shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{significant} \;\nRightarrow\; \text{linear} \]

Verify: confirm the book reaches the same conclusion

Why: Its note says that although the correlation coefficient is significant, the pattern in the scatterplot indicates that a curve would be a more appropriate model, and that a statistician should prefer other methods. Then it adds the general lesson: in addition to doing the calculations, it is always important to look at the scatterplot when deciding whether a linear model is appropriate.

OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 631-634

51. Trap: treating significance as sufficient

Trap

The trap

\[ p = 0.004 \;\Rightarrow\; \text{use the line} \]

Take a significant correlation as clearance to predict

Why: The test was passed.

\[ \text{but the plot may bend, and } x \text{ may be out of range} \]

Two further conditions have to hold, and neither is tested by the p-value.

The fix

\[ \text{significant } \wedge \text{ linear trend } \wedge \; x \text{ in range} \]

Check all three before predicting

Why: The book states them as three separate notes.

The three fail in different ways and are caught by different means: significance by the test, linearity by the scatter plot and residuals plot, and range by comparing x against the observed minimum and maximum. Only the first is a calculation, which is why the other two are the ones usually skipped.

52. One of these is false

Two truths and a lie

All three concern the assumptions.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The y values at each x are assumed normally distributed about the line
  • C. The residual errors are assumed mutually independent
  • B. The spread of y is assumed to grow as x grows

Survives elimination: B

Why: The survivor is false and reverses the third assumption, which is that the standard deviations of the y values about the line are EQUAL for each value of x. Spread that grows with x is a violation, and section 12.3 names it: the residuals should not consistently increase as x increases.

53. What does the residuals plot check?

Prediction

Commit before reasoning.

Predict first

Which of the five assumptions does a residuals plot help check?

  • Linearity, equal spread, and independence of the errors
  • Only the random sampling
  • Only normality
  • None of them

Correct: Linearity, equal spread and independence.

Why: Section 12.3 says a residuals plot should appear random with no pattern and no outliers, and should show constant error variance. A curve in the residuals violates linearity, a widening band violates equal spread, and any systematic pattern violates independence — three assumptions in one picture.

54. Explain the limit

Explain it

A classmate has a significant correlation and concludes the relationship must be linear.

Discussion prompt

In two sentences or fewer, correct them.

Hint: Ask what the test actually measures.

Answer:

The test measures whether the LINEAR component of the relationship is more than chance would produce, which a curved relationship can easily pass since a rising curve has a strong linear component.

Example 12.14 is exactly that case: a significant correlation of 0.8694 alongside a pattern the book says a curve would model better.

55. The two methods

Comparison

Fill the blanks. They are one test written twice.

Comparison matrix

p-value methodCritical-value method
What is computeda t statistic from r and nnothing: r is looked up against a table
Degrees of freedomn - 2n - 2
Decision rulereject if p is below alphasignificant if r is outside the critical values
Significance levelsany5 percent only

The last row is the only genuine difference, and it follows from the third: a table of critical values has to fix alpha before it can list anything, while a p-value is computed after the fact and can be compared against any level.

56. Testing a correlation for significance, in order

Pattern

Six steps, and the last two are conditions rather than calculations.

  1. State the hypotheses: rho equals zero against rho not equal to zero, two-tailed.
  2. Compute the degrees of freedom as n minus two.
  3. Either compute the t statistic and double its tail area, or look up the critical value for those degrees of freedom.
  4. Reject if the p-value is below 0.05, or equivalently if r lies outside the critical values.
  5. Check the scatter plot shows a linear trend before using the line at all.
  6. Predict only for x values inside the observed range.

The two methods cannot disagree, since each critical value is the t critical value solved for r at a 5 percent level.

OpenStax Introductory Business Statistics 2e, §13.2 Testing the Significance of the Correlation Coefficient §13.2 Testing the Significance of the Correlation Coefficient

57. Check yourself 1 of 3

Check

Degrees of freedom.

Check your understanding

A correlation is computed from 20 paired observations. What degrees of freedom does the test use?

  • A. 18 (correct)
  • B. 19
  • C. 20
  • D. 2

Answer: A

Why: The degrees of freedom are n minus two, because a regression line uses two estimates: an intercept and a slope.

Why B tempts people
That is n minus one, the rule for a one-sample t test on a mean.
Why C tempts people
The sample size itself is never the degrees of freedom.
Why D tempts people
Two is the number of estimates spent, not what remains.

58. Check yourself 2 of 3

Check

The verdict.

Check your understanding

A correlation of 0.55 comes from 8 data points, where the critical value is 0.707. What follows?

  • A. Not significant: the line should not be used for prediction (correct)
  • B. Significant, since 0.55 is a moderate correlation
  • C. Significant, since it is positive
  • D. The population correlation is zero

Answer: A

Why: 0.55 lies between -0.707 and 0.707, so it is not significantly different from zero and the line should not be used.

Why B tempts people
A moderate correlation from eight points is exactly what chance readily produces.
Why C tempts people
Sign has no bearing; the comparison is against the magnitude.
Why D tempts people
Failing to reject never establishes the null, only that the data are consistent with it.

59. Check yourself 3 of 3

Check

What significance licenses.

Check your understanding

A correlation is significant and the scatter plot shows a clear curve. May the line be used to predict?

  • A. No: a linear trend is a separate requirement (correct)
  • B. Yes: significance is what matters
  • C. Yes, but only outside the observed range
  • D. Only if the sample exceeds 30

Answer: A

Why: The book's note is explicit: if the scatter plot does not show a linear trend, the line should not be used for prediction, whatever the p-value says.

Why B tempts people
Significance shows the linear component is real, not that the relationship is linear.
Why C tempts people
Prediction outside the observed range is the least reliable use, not the most.
Why D tempts people
No sample size makes a curved relationship linear.

60. Where this shows up outside the textbook

Real world

A pharmaceutical analysis of 4,000 patients reports a correlation of 0.06 between a drug's dose and symptom improvement, with a p-value of 0.0001, and the press release states that the study found a highly significant relationship between dose and improvement.

Discussion prompt

Is the statistical claim correct, and what should a reader take from it?

Hint: Compute how much of the variation the correlation explains.

Answer:

The statistical claim is correct and almost meaningless. With 4,000 patients the critical correlation at 5 percent is about 0.031, so 0.06 clears it comfortably and the tiny p-value is exactly what the arithmetic gives. The linear relationship is real.

\[ r^2 = 0.06^2 = 0.0036 \;\Rightarrow\; 0.36\% \text{ of the variation explained} \]

Dose explains about a third of one percent of the variation in improvement, leaving 99.6 percent to everything else. The phrase highly significant describes the p-value, which measures confidence that the effect is non-zero — not its size. At this sample size a correlation far too small to matter clinically is detected with near-certainty.

This is the mirror image of Example 12.9. There, a correlation of 0.776 from six points was a strong relationship with weak evidence; here 0.06 from four thousand is negligible with overwhelming evidence. Strength and evidence move independently, and a p-value reports only the second.

What a reader should take: the drug's effect on symptoms is almost certainly not zero, and is very small. The honest report gives r-squared alongside the p-value, states the effect in the outcome's own units, and lets a clinician judge whether an effect that size is worth anything. Reporting significance alone, at this sample size, tells the reader nothing they can act on.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

Why is a correlation of 0.776 from 6 points not significant while 0.567 from 19 points is?

  • Because the second is negative
  • Because the critical value falls as the sample grows, so less correlation is needed
  • Because 0.567 is closer to zero
  • An error: the larger correlation must be the significant one

Correct: The critical value falls as the sample grows.

\[ r^* = \frac{t^*}{\sqrt{\text{df} + (t^*)^2}}: \quad 0.811 \text{ at df } 4, \quad 0.456 \text{ at df } 17 \]

Why: At four degrees of freedom the threshold is 0.811 and at seventeen it is 0.456. A small sample produces large correlations by chance quite easily, so a great deal of correlation is needed to rule chance out; a larger sample makes a modest correlation hard to explain that way. Strength and evidence are different quantities, and the test measures the second.

62. Explain it to someone a year behind you

Explain it

They looked up the critical value at n minus one degrees of freedom instead of n minus two.

Discussion prompt

In two sentences or fewer, correct them.

Hint: Ask how many quantities the line estimates.

Answer:

A regression line estimates two things, an intercept and a slope, so two degrees of freedom are spent and the test uses n minus two.

Using n minus one gives a critical value that is slightly too small, which errs in the direction of declaring a correlation significant when it is not.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Explaining why the sample size matters as much as r
  • Computing the t statistic and doubling its tail
  • Using the critical-value table and knowing why it is fixed at 5 percent
  • Stating the three conditions before a line may be used to predict

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the first, small samples produce large correlations by chance. For the second, r times root n minus two over root one minus r squared, then double. For the third, each entry is a t critical value solved for r at one fixed alpha. For the fourth, significant, linear, and inside the range. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Fifteen minutes.

Draw it

At the top, write the hypotheses in symbols and beneath each its wording in words, marking clearly that rho is the unknown population correlation and r the computed sample one. Underneath, write the test statistic as r times the root of n minus two over the root of one minus r squared, note that the degrees of freedom are n minus two, and work the exam data through it: 0.6631 times 3 over the root of 0.5603, giving 2.6576. Beside that, draw a t curve on nine degrees of freedom with both tails shaded beyond plus and minus 2.66, labelled with a combined area of 0.026. In the middle, write the critical-value method and the identity that generates it — critical r equals t star over the root of df plus t star squared — then verify it once at nine degrees of freedom: 2.262 over the root of 14.117 is 0.602. Below, make a five-row table of the book's cases with columns for r, n, df, critical value and verdict, and circle the row where 0.776 fails while a smaller correlation elsewhere passes. At the bottom, write the three conditions for using the line in a box: r significant, a linear trend in the plot, and x inside the observed range.

Check your critical-value derivation by trying it at four degrees of freedom, where t star is 2.776 and the answer should be 0.811 — the value that made Example 12.9 fail. Check your table by confirming that the verdicts do not simply follow the order of the correlations.

65. What you can do now

Recap

Six things, and two of them are conditions the p-value does not check.

If you seeThen
A correlation to testDegrees of freedom are n minus two
A t statistic computedDouble the tail: the alternative is two-sided
r outside the critical valuesSignificant at 5 percent
r between themNot significant; do not use the line
r equal to zeroNever significant, at any sample size
A significant r with a curved plotDo not use the line: a curve is needed
A prediction outside the observed x rangeNot licensed, however significant r is
A large sample and a tiny rSignificant and unimportant; report r-squared too

Section 12.5 uses the line for what it was fitted for. With the correlation established as significant and the pattern confirmed linear, predictions can be made — and the section draws the line between interpolation, which is licensed, and extrapolation, which is not.

OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient §12.4, pp. 631-635 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §12.4 Testing the Significance of the Correlation Coefficient — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 631-635
  2. OpenStax Introductory Business Statistics 2e, §13.2 Testing the Significance of the Correlation Coefficient — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108