The chapter's central section. Data rarely fit a straight line exactly, so the question is which of the many lines that miss should be used — and the answer is the one minimising the sum of the squared vertical distances from the points to the line. Each such distance is a residual, positive where the line underestimates and negative where it overestimates, and the sum of their squares is the SSE. Two properties of the resulting line are worth more than its formulas: it always passes through the point of the two means, and its slope is the correlation coefficient times the ratio of the two standard deviations — which is why r and the slope always share a sign. The section closes with r itself, measuring the strength and direction of the linear association, and with r-squared, the share of the variation in y that the line explains. For the running example that share is 44 percent, which is a useful corrective to reading a correlation of 0.66 as a strong result.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 12 — Linear Regression and Correlation
The Regression Equation
Objectives
Six outcomes, and two of them are properties that check the other four.
OpenStax Introductory Statistics 2e, §12.3 The Regression Equation §12.3, pp. 623-631 — the section these objectives are drawn from
Warm-up
Section 12.2 confirmed the exam scores show a linear pattern; section 12.1 wrote lines exactly.
Discussion prompt
If each of you were to fit a line to those eleven points by eye, would you all draw the same one? What would decide which is best?
Hint: Ask what makes one line better than another when none passes through every point.
Answer:
No — everyone would draw something slightly different, which is exactly the book's point in its collaborative exercise on pinky-finger lengths and heights. By eye there is no single answer, so a criterion is needed.
The natural criterion is how badly a line misses. For each point, the vertical distance from the point to the line is that point's miss, and a good line should make those misses small overall.
\[ \varepsilon_i = y_i - \hat{y}_i, \qquad \text{SSE} = \sum_{i=1}^{11} \varepsilon_i^2 \]
Squaring before adding is what makes it work, for the same reason chapter 2 squared deviations and chapter 11 squared departures: misses above and below would otherwise cancel, and a terrible line could score zero. The line minimising that sum is the least-squares line, and calculus produces exactly one of them.
Concept
Data rarely fit a straight line exactly, so usually you must be satisfied with rough predictions. The line of best fit, or least-squares line, is the one for which the sum of the squared errors is as small as possible. Any other line you might choose would have a higher SSE.
the least-squares line — The line minimising the sum of squared vertical distances from the observed points. Using calculus, the values of a and b that make the SSE a minimum can be determined, and they are unique.
\[ \hat{y} = a + bx \quad\text{minimising}\quad \text{SSE} = \sum (y_i - \hat{y}_i)^2 \]
The hat matters. The book writes y-hat for the estimated value of y — the value obtained from the regression line — and notes it is not generally equal to the y from the data. Keeping the two apart is what makes a residual expressible at all, since a residual is precisely the difference between them.
Figure (svg): A scatter plot of eleven exam scores with a fitted line and a dashed vertical segment from each point to the line
OpenStax Introductory Statistics 2e, §12.3 The Regression Equation §12.3, pp. 623-627
Section
Section 1
Concept
The term y minus y-hat is called the error or residual. It is not an error in the sense of a mistake: its absolute value measures the vertical distance between the actual value of y and the estimated value. If the observed point lies above the line the residual is positive and the line underestimates; if below, the residual is negative and the line overestimates.
the residual — Observed minus predicted, written epsilon. Squaring each and adding gives the sum of squared errors, and the criterion for the best fit line is that this sum is minimised.
\[ \varepsilon_i = y_i - \hat{y}_i; \qquad \text{SSE} = \sum \varepsilon_i^2 \]
The sign convention runs in the direction that is easy to reverse. A POSITIVE residual means the point is above the line, which means the line predicted too LOW — so a positive residual accompanies an underestimate. Saying it aloud in that order once is worth more than reading it three times.
Figure (svg): A card defining the residual and the sign convention attached to it
OpenStax Introductory Statistics 2e, §12.3 The Regression Equation §12.3, p. 625 — the residual, its sign, and the SSE
Picture it
The eleven exam points with their vertical misses.
Figure (svg): A scatter plot of eleven exam scores with a fitted line and a dashed vertical segment from each point to the line
Green segments are points above the line and pink ones below. The least-squares line is the unique line making the total of their squared lengths as small as possible; tilting or shifting it any amount increases that total.
Worked example
For the student who scored 65 on the third exam and 175 on the final, with the fitted line ŷ = -173.51 + 4.83x.
\[ (65, 175) \]
Predict from the line
Why: Substitute 65.
\[ 140.45 \]
Subtract from observed
Why: 175 minus 140.45.
\[ 34.55 \]
Read the sign
Why: Positive.
Say what that means
Why: Predicted too low.
Figure (svg): The solution to Worked example one residual shown as a ladder of expressions, one row per legal move
\[ \varepsilon = 175 - 140.45 = 34.55 \]
Verify: confirm against the value the full-precision line gives
Why: Using the unrounded coefficients, -173.51336 plus 4.827394 times 65 gives 140.267, so the residual is 34.73 rather than 34.55. The book's own table, computed from the rounded line, lists 35. All three describe the same point; the differences come entirely from how much of the slope and intercept was carried. Section 12.6 uses 35 and the exact value 34.73, and both exceed the outlier threshold.
OpenStax Introductory Statistics 2e, §12.3 The Regression Equation §12.3, pp. 625-628
Faded example
A student scored 75 on the third exam and 198 on the final; the line predicts 188.6.
Fill in the blanks
\varepsilon = 198 - 188.6 = 9.4, \textabove ___ \text___
Why: A positive residual puts the point above the line, so the line underestimated this student. The exact residual is 9.46 and the book's rounded table gives 9.
Worked example
Suppose the residuals for four points were 10, -10, 10 and -10.
\[ 10, -10, 10, -10 \]
Add them directly
Why: Ten minus ten twice.
\[ 0 \]
What that suggests
Why: A perfect fit.
Square, then add
Why: Four hundreds.
\[ 400 \]
What that captures
Why: Every miss counts.
Figure (svg): The solution to Worked example why the misses are squared shown as a ladder of expressions, one row per legal move
\[ \sum \varepsilon_i = 0 \quad\text{against}\quad \sum \varepsilon_i^2 = 400 \]
Verify: confirm the plain sum is always zero, not just here
Why: For any least-squares line the residuals sum to exactly zero, which is a consequence of how the intercept is chosen. So the plain sum carries no information about any fitted line and cannot distinguish a good one from a bad one. Squaring is what makes the criterion meaningful — the same reason chapter 2's variance squared its deviations.
OpenStax Introductory Statistics 2e, §12.3 The Regression Equation §12.3, pp. 625-627
Trap
\[ \varepsilon > 0 \;\Rightarrow\; \text{the line OVERestimates} \]
Reason that a positive residual means the line is too high
Why: Positive sounds like too much.
\[ \varepsilon = y - \hat{y}, \text{ so } y > \hat{y} \]
A positive residual means the OBSERVED value exceeds the prediction, so the prediction was too low.
\[ \varepsilon > 0 \;\Rightarrow\; y > \hat{y} \;\Rightarrow\; \text{the line UNDERestimates} \]
Read the subtraction in the order it is written
Why: Observed minus predicted.
The reliable check is the picture. A positive residual puts the point ABOVE the line, and a line drawn below a point has predicted too little for it. Reading the sign off the plot rather than off the word takes the ambiguity out entirely.
Sorting
Each is a residual from the exam data.
Sort into buckets
Sort by where the point lies.
These are the first five exact residuals from Example 12.6. Six of the eleven are positive and five negative, and all eleven sum to zero — as they do for any least-squares line.
Two truths and a lie
All three concern residuals.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The residuals of any least-squares line already sum to exactly zero, so that sum cannot distinguish between lines at all. The squares are what the criterion minimises.
Prediction
Commit before reasoning.
Predict first
Two lines are fitted to the same data and one has a smaller SSE. What follows?
Correct: It sits closer to the points overall.
Why: SSE measures total squared vertical miss, so a smaller value means a better overall fit. The book states the criterion as a uniqueness claim: any other line you might choose would have a higher SSE than the best-fit line.
Section
Section 2
Concept
The best fit line always passes through the point given by the two sample means. Its slope can be written as r times the standard deviation of the y values divided by the standard deviation of the x values, where r is the correlation coefficient.
the point of means — The point whose coordinates are the mean of the x values and the mean of the y values. Every least-squares line passes through it exactly, whatever the data.
\[ (\bar{x}, \bar{y}) \text{ on the line}; \qquad b = r\,\frac{s_y}{s_x} \]
Both properties are checks that take one line of arithmetic and catch most errors. If a reported line does not reproduce the mean of y when the mean of x is substituted, something is wrong; and if the slope's sign disagrees with the correlation's, something is wrong. Neither check requires refitting anything.
Figure (svg): A card giving the two properties of the least-squares line
OpenStax Introductory Statistics 2e, §12.3 The Regression Equation §12.3, p. 626 — the line passes through the point of means; b = r sy over sx
Picture it
With the exam data's numbers.
Figure (svg): A card giving the two properties of the least-squares line
The slope identity is the more useful of the two, because it links this section to the next. The correlation coefficient and the slope are the same quantity in different units, which is why testing whether r differs from zero is the same as testing whether the slope does.
Worked example
The exam data has mean third-exam score 69.18 and mean final score 160.45.
\[ \bar{x} = 69.18, \; \bar{y} = 160.45 \]
Substitute the mean of x
Why: Into the line.
\[ -173.513 + 4.8274(69.18) \]
Compute the product
Why: The slope term.
\[ 333.97 \]
Add the intercept
Why: Subtract 173.513.
\[ 160.45 \]
Compare with the mean of y
Why: The same.
Figure (svg): The solution to Worked example checking the point of means shown as a ladder of expressions, one row per legal move
\[ \hat{y}(69.18) = 160.45 = \bar{y} \]
Verify: confirm this is a property rather than a coincidence
Why: It holds for every least-squares line because the intercept is defined as the mean of y minus the slope times the mean of x — so substituting the mean of x returns the mean of y by construction. That makes it a genuine check on arithmetic rather than on the data: any reported intercept and slope can be tested against it in seconds.
OpenStax Introductory Statistics 2e, §12.3 The Regression Equation §12.3, p. 626
Faded example
A data set has r = -0.8, sy = 12 and sx = 4.
Fill in the blanks
b = -0.8 \times \frac3-2.4 = -0.8 \times ___ = ___
Why: The negative correlation forces a negative slope, since the ratio of standard deviations is always positive.
Worked example
The exam data has r = 0.6631, sy = 20.80 and sx = 2.857.
\[ b = r\,s_y/s_x \]
The ratio of spreads
Why: 20.80 over 2.857.
\[ 7.280 \]
Multiply by r
Why: 0.6631 times 7.280.
\[ 4.827 \]
Compare with the calculator
Why: LinRegTTest.
\[ 4.8273 \]
Note the sign
Why: r positive.
Figure (svg): The solution to Worked example the slope from r shown as a ladder of expressions, one row per legal move
\[ b = 0.6631 \times \frac{20.80}{2.857} = 4.827 \]
Verify: confirm why the two signs can never disagree
Why: Standard deviations are never negative, so the ratio is always positive and the sign of b comes entirely from the sign of r. That is why the book can state as a separate fact that the sign of r is the same as the sign of the slope: it is not an observation about data but a consequence of this identity.
OpenStax Introductory Statistics 2e, §12.3 The Regression Equation §12.3, pp. 626-630
Error analysis
Which are correct?
Annotate
On: \( \begin{aligned} &(1)\; \text{it passes through the point of the two means} \\ &(2)\; \text{it passes through at least two of the data points} \\ &(3)\; \text{its slope and } r \text{ always share a sign} \\ &(4)\; \text{it is unique for a given data set} \end{aligned} \)
Error (2) is worth naming because the point of means is so often not a data point either. The line is anchored by a computed point rather than by any observation, which is exactly what lets it ignore individual points in favour of the overall pattern.
Two truths and a lie
All three concern the two properties.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The slope is r SCALED by the ratio of the two standard deviations, and they are equal only in the special case where those spreads happen to match. For the exam data r is 0.6631 while the slope is 4.827.
Prediction
Commit before reasoning.
Predict first
Why does every least-squares line pass through the point of means?
Correct: The intercept is defined that way.
Why: Once the slope is determined, the intercept is chosen to be the mean of y minus the slope times the mean of x — which is exactly the condition that the line pass through that point. So the property is built into the construction and holds for any data at all.
Estimation
A data set has r = 0.5, with y spread about ten times as wide as x.
Predict first
Roughly what slope should be expected?
Correct: About 5.
Why: The slope is r times the ratio of spreads, so 0.5 times 10 gives 5. This is why a slope's size says little about the strength of a relationship on its own — it depends on the units of both variables, while r does not.
Section
Section 3
Concept
The slope of the line tells us how the dependent variable changes for every one unit increase in the independent variable, on average. It is important to interpret the slope in the context of the situation, and you should be able to write a sentence interpreting it in plain English.
on average — The two words that separate a fitted line from an exact one. Individual points depart from the line, so the slope describes the trend rather than what happens to any one observation.
\[ b = 4.83: \quad +1 \text{ third-exam point} \;\to\; +4.83 \text{ final points, on average} \]
This is section 12.1's interpretation habit with one phrase added, and that phrase carries the whole difference between the two chapters. Svetlana really did earn exactly fifteen more dollars per hour; no student's final score rises by exactly 4.83 points per third-exam point. The line describes the average behaviour of a scattered cloud.
Figure (svg): Fitting and reporting a regression line
OpenStax Introductory Statistics 2e, §12.3 The Regression Equation §12.3, p. 628 — interpretation of the slope, and the worked sentence
Picture it
Six steps, of which the arithmetic is one.
Figure (svg): Fitting and reporting a regression line
Steps one and six are the book's own reminders: plot first, and predict only within the observed range. The interpretation in steps four and five is what a calculator cannot supply.
Worked example
For the third-exam and final-exam data, the fitted slope is 4.83.
\[ b = 4.83 \]
Units of x
Why: Third exam.
\[ \text{points out of } 80 \]
Units of y
Why: Final exam.
\[ \text{points out of } 200 \]
Units of b
Why: y per x.
Add the qualifier
Why: It is a trend.
Figure (svg): The solution to Worked example the book's own interpretation shown as a ladder of expressions, one row per legal move
\[ b = 4.83 \text{ final points per third-exam point} \]
Verify: confirm the magnitude is plausible given the two scales
Why: The third exam is out of 80 and the final out of 200, a ratio of 2.5 — so a slope near 2.5 would mean the two exams track proportionally. The fitted 4.83 is nearly twice that, meaning the final exam spreads students out considerably more than the third does. That is consistent with the two standard deviations, 20.80 against 2.857, and the slope identity makes the connection explicit.
OpenStax Introductory Statistics 2e, §12.3 The Regression Equation §12.3, p. 628
Faded example
For the dive data, the fitted slope is -1.11 minutes per foot.
Fill in the blanks
\text1.11 on average \text___ ___
Why: The negative sign is carried by the word costs, and the qualifier records that this is a trend across the six observations rather than an exact rule.
Worked example
The fitted intercept is -173.51.
\[ a = -173.51 \]
Read it literally
Why: y when x is zero.
\[ a\text{ third-exam score of } 0 \]
Check the data range
Why: Scores run 65 to 75.
Check plausibility
Why: A negative final score.
Conclude
Why: Placement only.
Figure (svg): The solution to Worked example what the intercept means here shown as a ladder of expressions, one row per legal move
\[ x = 0 \text{ lies far outside } [65, 75] \]
Verify: confirm the line is still sound despite this
Why: The intercept is far outside the data because the x values are clustered between 65 and 75 — so the line has to be extended a long way left to reach the vertical axis, and small changes in slope swing that endpoint dramatically. Within the observed range the line predicts sensibly, which is exactly the point section 12.5 makes about interpolation and extrapolation.
OpenStax Introductory Statistics 2e, §12.3 The Regression Equation §12.3, pp. 628-629
Trap
\[ +1 \text{ third-exam point} \;\Rightarrow\; +4.83 \text{ final points} \]
State the slope as an exact consequence
Why: That is what it meant in section 12.1.
\[ \text{but the residuals run from } -19 \text{ to } +35 \]
Individual students depart from the line by far more than 4.83 points, so no student's score moves by exactly that amount.
\[ +1 \text{ point} \;\Rightarrow\; +4.83 \text{ final points, ON AVERAGE} \]
State the slope as a trend across the group
Why: The line describes the cloud, not any point in it.
The residual spread makes the size of the qualification concrete: the standard deviation of the residuals is 16.4 final-exam points, against a slope of 4.83 per third-exam point. So the scatter around the line is more than three times the effect of a one-point change in x, and treating the slope as exact would badly overstate what it predicts about an individual.
Two truths and a lie
All three concern the slope's meaning.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The residuals for these eleven students range from about -19 to +35, so individual final scores depart from the line by far more than the slope. The slope summarises the trend across the group.
Prediction
Commit before reasoning.
Predict first
The final exam is rescaled from 200 points to 100. What happens to the slope and to r?
Correct: The slope halves and r is unchanged.
Why: The slope carries units of y per x, so halving the y scale halves it. The correlation is a pure number with no units and is unaffected by rescaling either variable — which is why r can be compared across studies while slopes cannot.
Explain it
A classmate writes that a student who scores one point higher on the third exam will score 4.83 points higher on the final.
Discussion prompt
In two sentences or fewer, correct them.
Hint: Ask how far the actual points sit from the line.
Answer:
The slope describes the average trend across the eleven students, not what happens to any one of them — the observed points sit as much as 35 final-exam points away from the line.
The correct sentence adds on average, which is the difference between a fitted line and the exact ones in section 12.1.
Section
Section 4
Concept
The correlation coefficient r, developed by Karl Pearson in the early 1900s, provides a measure of the strength and direction of the linear association between the independent variable x and the dependent variable y. Its value always lies between minus one and one.
what r measures — The strength and direction of a LINEAR association. Values close to minus one or one indicate a stronger linear relationship; a value of zero indicates there is likely no linear correlation.
\[ -1 \le r \le 1 \]
The book attaches a caution to r equal to zero that is easy to skip: it is important to view the scatterplot, because data that exhibit a curved or horizontal pattern may have a correlation of zero. A correlation near zero rules out a straight-line association and rules out nothing else, which is why section 12.2 came first.
Figure (svg): A number line from minus one to one marking perfect negative, no correlation, perfect positive, and the exam data at 0.6631
OpenStax Introductory Statistics 2e, §12.3 The Regression Equation §12.3, pp. 629-630 — the correlation coefficient, its value and its sign
Picture it
What the value and the sign tell you.
Figure (svg): A number line from minus one to one marking perfect negative, no correlation, perfect positive, and the exam data at 0.6631
The exam data's 0.6631 sits between the middle and the positive extreme — a real association, and not a tight one. The lower panel is the caution that keeps a value near zero from being over-read.
Worked example
The calculator reports r = 0.663 for the third-exam and final-exam scores.
\[ r = 0.6631 \]
Read the sign
Why: Positive.
Check against the slope
Why: Also positive.
Read the magnitude
Why: Between 0 and 1.
Say what it does not settle
Why: Whether to use the line.
\[ \text{section } 12.4 \]
Figure (svg): The solution to Worked example reading r for the exam data shown as a ladder of expressions, one row per legal move
\[ r = 0.6631 > 0, \text{ and } b = 4.83 > 0 \]
Verify: confirm the two signs agree, as they must
Why: The slope identity forces it: b equals r times a ratio of standard deviations, both of which are positive. So a positive r and a negative slope can never occur together, and finding them in a computer output means the two came from different data or different columns. It is a free consistency check on any regression output.
OpenStax Introductory Statistics 2e, §12.3 The Regression Equation §12.3, pp. 629-630
Sorting
Each is a correlation coefficient.
Sort into buckets
Sort by the strength of the linear relationship.
Strength depends on distance from zero, not on sign: the negative values here are as strong as the positive ones. Item (b) is the exam data, and it is the one the chapter spends most time on despite being the weakest of the strong ones.
Worked example
Consider points symmetric about a vertical axis, following a clear U shape.
\[ \text{a symmetric curve} \]
Left half
Why: y falls as x rises.
Right half
Why: y rises as x rises.
Add them
Why: They cancel.
But the pattern
Why: Tight and clear.
Figure (svg): The solution to Worked example why r near zero is not no relationship shown as a ladder of expressions, one row per legal move
\[ r \approx 0 \text{ with a tight curve} \]
Verify: confirm the book flags exactly this case
Why: It says that if r is zero there is likely no linear correlation, and immediately adds that it is important to view the scatterplot, because data exhibiting a curved or horizontal pattern may have a correlation of zero. Both of section 12.2's warnings reappear here — the curve and the horizontal line — as reasons never to report r without having looked at the plot.
OpenStax Introductory Statistics 2e, §12.3 The Regression Equation §12.3, p. 630
Trap
\[ r = 0.94 \;\Rightarrow\; x \text{ causes } y \]
Take a tight linear association as a causal finding
Why: The relationship is undeniably strong.
\[ \text{but the book states the rule directly} \]
Strong correlation does not suggest that x causes y or y causes x. We say correlation does not imply causation.
\[ r = 0.94: \text{ a strong LINEAR ASSOCIATION} \]
Report association, and treat cause as a separate question
Why: It depends on the design, not the number.
Section 12.2's condition is the practical form of this: a regression line is only worth fitting if one variable helps explain or predict the other, and that judgement is made from the situation. No value of r, however extreme, converts an observational association into a causal one.
Two truths and a lie
All three concern r.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. r measures LINEAR association only, so a tight U-shaped or horizontal pattern can produce a correlation near zero while a strong relationship is plainly visible in the plot.
Faded example
State the range of the correlation coefficient.
Fill in the blanks
-1 \le r \le 1
Why: Both endpoints mean every observed point lies exactly on a straight line, which the book notes will not generally happen in the real world.
Prediction
Commit before reasoning.
Predict first
A dataset has r exactly equal to 1. What follows?
Correct: All the points lie on a rising straight line.
Why: The book states this directly: at r equal to one or minus one, all of the original data points lie on a straight line. The slope can be anything positive, since r describes how tightly the points follow a line rather than how steep it is.
Section
Section 5
Concept
The variable r-squared is called the coefficient of determination and is the square of the correlation coefficient, usually stated as a percent. It represents the percent of variation in the dependent variable y that can be explained by variation in the independent variable x using the regression line.
r-squared — The square of r, read as a percent. Its complement, one minus r-squared, is the percent of variation in y NOT explained by variation in x — seen as the scattering of the points about the line.
\[ r^2 = 0.6631^2 = 0.4397 \approx 44\% \]
Squaring is what makes a correlation legible. A value of 0.6631 sounds like a solid two thirds of something; squared it becomes 44 percent, with 56 percent of the variation in final exam grades left unexplained by third exam grades. That deflation is systematic — any r below one squares to something smaller — and it is why r-squared is the number usually reported alongside a regression.
Figure (svg): A bar split into a smaller explained portion and a larger unexplained portion, labelled forty-four and fifty-six percent
OpenStax Introductory Statistics 2e, §12.3 The Regression Equation §12.3, pp. 630-631 — the coefficient of determination and its interpretation
Picture it
The exam data's variation, split.
Figure (svg): A bar split into a smaller explained portion and a larger unexplained portion, labelled forty-four and fifty-six percent
The unexplained share is what the scatter about the line represents. A student's final exam score depends on much besides their third exam score, and this is the quantitative form of that observation.
Worked example
The correlation is 0.6631.
\[ r = 0.6631 \]
Square it
Why: 0.6631 squared.
\[ 0.4397 \]
As a percent
Why: Multiply by 100.
\[ \text{about } 44 \% \]
State what is explained
Why: Variation in y.
Take the complement
Why: One minus 0.44.
\[ 56 \%\text{ not} \]
Figure (svg): The solution to Worked example interpreting r-squared for the exam data shown as a ladder of expressions, one row per legal move
\[ r^2 = 0.4397; \qquad 1 - r^2 = 0.5603 \]
Verify: confirm the calculator's reported value
Why: LinRegTTest reports r-squared as 0.43969, which squares back to a correlation of 0.66309 — the 0.663 shown on the same screen. Reading both from the output and checking that one is the square of the other catches a mistyped data value quickly, since a single wrong entry would change them inconsistently.
OpenStax Introductory Statistics 2e, §12.3 The Regression Equation §12.3, pp. 629-631
Faded example
A study reports a correlation of 0.8.
Fill in the blanks
r^2 = 0.64, \text64 ___ \text___
Why: A correlation of 0.8 sounds close to complete and explains under two thirds of the variation, leaving 36 percent to everything else.
Worked example
Four correlations and the shares of variation they explain.
\[ r = 0.3, 0.5, 0.7, 0.9 \]
r = 0.3
Why: Squared.
\[ 9 \% \]
r = 0.5
Why: Squared.
\[ 25 \% \]
r = 0.7
Why: Squared.
\[ 49 \% \]
r = 0.9
Why: Squared.
\[ 81 \% \]
Figure (svg): The solution to Worked example how r-squared deflates a correlation shown as a ladder of expressions, one row per legal move
\[ r^2 < r \text{ whenever } 0 < r < 1 \]
Verify: confirm where the halfway point falls
Why: Explaining half the variation requires a correlation of about 0.707, since 0.707 squared is 0.5. So a correlation described as strong at 0.7 accounts for slightly less than half of what it is meant to explain — which is why reporting r alone tends to flatter a relationship and why the book gives both numbers for the running example.
OpenStax Introductory Statistics 2e, §12.3 The Regression Equation §12.3, pp. 630-631
Error analysis
Which are correct?
Annotate
On: \( \begin{aligned} &(1)\; 44\% \text{ of the variation in } y \text{ is explained by variation in } x \\ &(2)\; \text{the line is correct } 44\% \text{ of the time} \\ &(3)\; 56\% \text{ of the variation in } y \text{ is not explained} \\ &(4)\; \text{the correlation is } 0.44 \end{aligned} \)
Error (4) runs in both directions and is worth guarding against, since 0.44 and 0.66 are both plausible-looking correlations. Checking that r-squared is the smaller of the two settles which is which.
Two truths and a lie
All three concern r-squared.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false for any correlation strictly between minus one and one, since squaring a number smaller than one in magnitude makes it smaller. That deflation is the reason r-squared is worth reporting.
Estimation
Explaining half the variation means r-squared equal to 0.5.
Predict first
What correlation is that?
Correct: About 0.71.
Why: The square root of 0.5 is 0.707. So a correlation has to reach about 0.71 before the line accounts for as much of y's variation as it leaves unaccounted for — a useful benchmark when reading reported correlations.
Prediction
Commit before reasoning.
Predict first
For the exam data, what is the 56 percent that the line does not explain?
Correct: The scattering of the points about the line.
Why: The book says exactly this: the unexplained percent can be seen as the scattering of the observed data points about the regression line. Final exam scores vary for many reasons besides the third exam score, and that variation shows up as vertical spread.
Comparison
Fill the blanks. Both describe the same relationship, and only one is comparable across studies.
Comparison matrix
| Slope b | Correlation r | |
|---|---|---|
| Units | y-units per x-unit | none |
| Range | any real number | between -1 and 1 |
| If y is rescaled | changes | unchanged |
| Sign | the same as r's | the same as b's |
The third row is why r travels between studies and a slope does not. The fourth is guaranteed by the identity b equals r times a ratio of standard deviations, both of which are positive.
Pattern
Six steps, and only the second is arithmetic a calculator does.
Compute the slope a second way as r times the ratio of standard deviations; agreement confirms both numbers at once.
OpenStax Introductory Business Statistics 2e, §13.4 The Regression Equation §13.4 The Regression Equation
Check
Residuals.
Check your understanding
An observed point lies above the regression line. What does that mean?
Answer: A
Why: The residual is observed minus predicted, so a point above the line has a positive residual — and a prediction below the observation is an underestimate.
Check
The fitted line.
Check your understanding
Which point does every least-squares line pass through?
Answer: A
Why: The intercept is defined so that substituting the mean of x returns the mean of y, so the property holds for any data set.
Check
The coefficient of determination.
Check your understanding
A regression has r = 0.5. What percent of the variation in y does the line explain?
Answer: A
Why: The coefficient of determination is r squared, and 0.5 squared is 0.25, or 25 percent.
Real world
A newspaper reports that a study of 200 towns found a correlation of 0.62 between the number of public libraries per capita and residents' median income, and concludes that building libraries raises incomes, noting that the relationship explains 62 percent of the variation in income.
Discussion prompt
Find the statistical errors and say what the study does support.
Hint: Check the reported percentage, then ask what a correlation can establish.
Answer:
The percentage is wrong. The share of variation explained is r-squared, not r, so 0.62 squared gives 0.3844 — about 38 percent, not 62. The article has reported the correlation as though it were the coefficient of determination, which overstates the relationship's explanatory reach by two thirds.
\[ r = 0.62 \;\Rightarrow\; r^2 = 0.3844 \approx 38\%, \qquad 1 - r^2 \approx 62\% \]
The two numbers have in fact been swapped, since 62 percent is very nearly the share of variation the line does NOT explain. That coincidence makes the error easy to miss and easy to check: r-squared is always the smaller of the two, so a reported explained share exceeding the correlation is a reliable warning sign.
The causal conclusion is not supported at all, and the book states the rule plainly: correlation does not imply causation. Wealthier towns can afford more libraries, so the arrow may run the other way entirely; and both may follow from population density, education levels, or municipal budgets. With 200 towns observed rather than assigned, none of these can be separated.
What the study supports: towns with more libraries per capita tend to have higher median incomes, a moderate positive linear association accounting for roughly 38 percent of the variation in income. What it does not support: that building a library would raise anyone's income. Section 12.2's condition applies — a regression line here predicts within the observed towns without licensing any claim about intervening.
Commit first
Answer, then rate your confidence honestly.
Predict first
What does the least-squares line minimise?
Correct: The sum of squared vertical distances.
\[ \text{SSE} = \sum (y_i - \hat{y}_i)^2 \quad \text{minimised}; \qquad \sum (y_i - \hat{y}_i) = 0 \text{ always} \]
Why: That sum is the SSE, and the book states the criterion as a uniqueness claim: any other line you might choose would have a higher SSE than the best fit line. The plain sum of residuals cannot be the criterion because it is exactly zero for every least-squares line, so it cannot distinguish between them. Squaring is what makes misses above and below both count.
Explain it
They report a correlation of 0.66 and say the third exam explains 66 percent of the variation in final exam scores.
Discussion prompt
In two sentences or fewer, correct them.
Hint: Ask which quantity is the percent of variation.
Answer:
The percent of variation explained is r SQUARED, not r, so 0.6631 squared gives about 44 percent rather than 66.
The remaining 56 percent is the scatter of the points about the line — everything besides the third exam score that affects a final grade.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, observed minus predicted, and positive means the line predicted too low. For the second, substitute the mean of x and expect the mean of y. For the third, y-units per x-unit, on average. For the fourth, r-squared is the smaller one and is the percent of variation explained. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Eighteen minutes — this is the chapter's central section.
Draw it
At the top left, plot the eleven exam points roughly, draw the fitted line ŷ = -173.51 + 4.83x through them, and drop a vertical dashed segment from three or four points to the line. Label one such segment as a residual and write beside it that observed minus predicted is positive when the point is above the line, meaning the line underestimated. Beneath, write the least-squares criterion: the sum of the squared residuals is minimised, and any other line has a higher SSE. In the middle, box the two properties — the line passes through the point 69.18 comma 160.45, and the slope is r times sy over sx, which is 0.6631 times 20.80 over 2.857, giving 4.827. Verify the first by substituting 69.18 and getting 160.45. To the right, draw a number line from minus one to one, mark minus one, zero and one with what each means, and mark 0.6631 on it. At the bottom, draw a bar split at 44 percent, labelling the left part explained and the right part not explained, and write the sentence in full: approximately 44 percent of the variation in final exam grades can be explained by variation in third exam grades. Finish with the slope interpretation in a sentence, underlining the words on average.
Check your slope calculation by confirming both routes agree — the calculator's 4.8273 and r times the ratio of spreads. Check your bar by confirming the explained portion is smaller than the correlation would suggest, which is true of every correlation below one.
Recap
Six things, and two of them check the others.
| If you see | Then |
|---|---|
| A point above the line | A positive residual: the line underestimates |
| A point below the line | A negative residual: the line overestimates |
| A fitted line to check | Substitute the mean of x and expect the mean of y |
| A slope and an r with opposite signs | An error: the identity forbids it |
| A slope to interpret | y-units per one x-unit, on average |
| r near zero | No LINEAR correlation; look at the plot before saying more |
| A correlation to report | Give r-squared too, as the percent of variation explained |
| A strong correlation | Association, never causation |
Section 12.4 asks the question this section left open. A correlation of 0.6631 was computed from only eleven students, and the reliability of a linear model depends on the sample size as well as on r — so the next step is a hypothesis test of whether the correlation is significantly different from zero.
OpenStax Introductory Statistics 2e, §12.3 The Regression Equation §12.3, pp. 623-631 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.