Handles data that trends linearly without lying on a line. Builds the scatter plot, fits a line by eye and then by least-squares regression, and introduces the correlation coefficient — including its two standard misreadings, that its sign measures quality and that a strong value implies causation, neither of which any calculation will catch.
Subject: Precalculus · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Precalculus · Chapter 2 — Linear Functions
§2.4 Fitting Linear Models to Data, pp. 247-262
Objectives
Five things, each one you can check yourself on paper.
OpenStax, Precalculus, §2.4 Fitting Linear Models to Data §2.4, pp. 247-262 — the pages these objectives are drawn from
Warm-up
Every model so far has passed exactly through its data. Real measurements do not cooperate.
Discussion prompt
You plot eight measurements and they trend clearly upward but no line passes through all eight. Is a linear model wrong here?
Hint: What might account for the scatter, apart from the relationship not being linear?
Answer:
Not necessarily. The underlying relationship may be linear while the measurements carry error, or the situation may have small influences the model does not track. A cloud of points around a line is what a linear relationship looks like once measured.
So the question changes from 'which line passes through the data' — usually none — to 'which line passes closest to it'. That is a different and harder question, and it needs a definition of 'closest' before it has an answer.
This section supplies that definition and the number that reports how well the answer did. Both are worth having, and both are routinely misread, which is why the interpretation gets as much attention here as the procedure.
Concept
When no line fits exactly, the best line is the one whose vertical misses are collectively smallest. Squaring each miss before adding them makes the answer unique and penalises large misses heavily.
residual — The vertical distance from a data point to the fitted line: the observed output minus the output the line predicts. A residual is positive when the point lies above the line and negative when below.
\[ \text{minimise } \sum (y_i - \hat{y}_i)^2 \]
Squaring does two jobs. It makes every miss count positively, so that a point far above and a point far below do not cancel, and it makes large misses count disproportionately, so the fitted line will not tolerate one badly missed point in order to fit the others slightly better.
Figure (svg): A scatter plot of points that trend upward without lying on a line, with a straight line drawn through the middle of them and short vertical segments showing each point's distance from the line
OpenStax, Precalculus, §2.4 Fitting Linear Models to Data §2.4, pp. 247-251
Section
Section 1
Concept
A scatter plot shows each data pair as a point. Its shape decides whether a linear model is appropriate at all, and no calculation can substitute for looking.
The habit of plotting first is worth more than any technique in the section. Almost every misuse of regression begins with fitting a line to data whose plot, if anyone had looked at it, would obviously have shown a curve, two clusters, or a single outlier dragging everything.
Figure (svg): A scatter plot of points that trend upward without lying on a line, with a straight line drawn through the middle of them and short vertical segments showing each point's distance from the line
OpenStax, Precalculus, §2.4 Fitting Linear Models to Data §2.4, pp. 247-252
Picture it
Each vertical segment is one residual — how far that point misses the line.
Figure (svg): A scatter plot of points that trend upward without lying on a line, with a straight line drawn through the middle of them and short vertical segments showing each point's distance from the line
A good fit has short segments, roughly half of them upward and half downward. All the points on one side means the line is misplaced even if the segments are short.
Worked example
The first question is whether a line is appropriate, not which line.
\[ \text{A plot shows points rising steeply, levelling off, and then flat. Fit a line?} \]
Describe the shape
Why: Rising then flattening is a curve, not a line.
Ask what a line would do
Why: It would over-predict early and under-predict late.
Check the residual pattern
Why: They would be negative, then positive, then negative.
Conclude
Why: A line is the wrong model here.
Figure (svg): The solution to Worked example judge a plot before fitting shown as a ladder of expressions, one row per legal move
\[ \text{A line is inappropriate: the residuals would show a pattern.} \]
Verify: say what the residual pattern means
Why: Residuals that follow a pattern rather than scattering randomly are the signature of a wrong model shape. Random scatter means the line has captured the trend and only noise is left; a pattern means there is structure the line failed to capture. This is worth more than the correlation coefficient as a diagnostic, and it costs only a glance.
OpenStax, Precalculus, §2.4 Fitting Linear Models to Data §2.4, pp. 248-250
Sorting
Judge from the described shape of the plot.
Sort into buckets
Sort each described scatter plot.
Worked example
A hand-drawn line is enough for an estimate.
\[ \text{A fitted line passes near } (2,5) \text{ and } (8,14). \text{ Estimate the output at } 5. \]
Compute the slope from the two line points
Why: Not from data points, from the drawn line.
\[ \frac{9}{6} = 1.5 \]
Use point-slope
Why: With either point on the line.
\[ y - 5 = 1.5(x - 2) \]
Expand
Why: Distribute and collect.
\[ y = 1.5 x + 2 \]
Evaluate at the wanted input
Why: Substitute 5.
\[ y = 9.5 \]
Figure (svg): The solution to Worked example estimate with an eyeballed line shown as a ladder of expressions, one row per legal move
\[ y = 1.5x+2, \quad y(5)=9.5 \]
Verify: note what kind of answer this is
Why: The input 5 lies between 2 and 8, so this is an interpolation and reasonably safe. The word 'about' is doing real work: the line was drawn by eye, so a second person would draw a slightly different line and get a slightly different estimate. Reporting an eyeballed estimate to four decimal places would claim a precision the method does not have.
OpenStax, Precalculus, §2.4 Fitting Linear Models to Data §2.4, pp. 250-252
Trap
\[ r = 0.12 \;\Longrightarrow\; \text{the two quantities are unrelated} \]
Compute the correlation coefficient and read its size
Why: The value is near zero, so no relationship is inferred.
The two quantities are reported as having nothing to do with each other.
A correlation near zero means no LINEAR relationship. The data may follow a strong curve, and a symmetric curve produces a correlation of almost exactly zero while being nearly deterministic.
Only the scatter plot distinguishes these two situations, and the number gives no hint which one you have.
Plot the data before computing anything, and look at it again afterwards. The correlation coefficient answers one narrow question, and it answers it whether or not that was the question worth asking.
Prediction
A straight line is fitted to data that actually follows a U-shaped curve.
Predict first
What do the residuals look like?
Correct: They follow a pattern: positive, negative, then positive.
Why: A line through a U misses above at both ends and below in the middle, so the residuals change sign systematically. Random scatter is what a correct model shape produces; a pattern in the residuals is the clearest evidence that the shape is wrong, and it is visible even when the correlation coefficient looks acceptable.
Faded example
An eyeballed line passes near the points (1, 4) and (7, 16).
Fill in the blanks
m = \frac22 = ___, \qquad y - 4 = ___(x-1) \;\Longrightarrow\; y = 2x + ___
Why: The slope is 12 over 6, which is 2, and point-slope with the first point expands to 2x plus 2. These are points chosen on the drawn line rather than data points, which is the distinction that matters: the line need not pass through any actual observation.
Step zero
You are handed a table of paired measurements and asked to model them.
Discussion prompt
What do you do before any computation, and what are you looking for?
Hint: What can a picture tell you that a number cannot?
Answer:
Plot the data. Before any slope, any regression, any correlation, put the points on axes and look at them.
You are looking for three things: whether the trend is straight or curved, whether there are outliers far from the rest, and whether the points form one cloud or several clusters. Each of those changes what you should do next, and none is visible in the table.
Skipping this step is how regression gets misused. The procedure will produce a line and a correlation coefficient for any data whatsoever, including data with no relationship at all, and neither output complains about being asked.
Section
Section 2
Concept
The least-squares line is the one minimising the sum of the squared residuals. That criterion picks out exactly one line for any data set, which is what makes it a procedure rather than a judgement.
In practice the line is produced by a calculator or a spreadsheet, which is entirely reasonable — the formulas are not illuminating. What is worth knowing is what is being minimised, because that explains the method's one notable weakness: a single outlier can move the line a long way, since its residual is squared along with everyone else's.
Figure (svg): A scatter plot of points that trend upward without lying on a line, with a straight line drawn through the middle of them and short vertical segments showing each point's distance from the line
OpenStax, Precalculus, §2.4 Fitting Linear Models to Data §2.4, pp. 252-257
Picture it
Each pink segment is a residual, and the criterion adds their squares.
Figure (svg): A scatter plot of points that trend upward without lying on a line, with a straight line drawn through the middle of them and short vertical segments showing each point's distance from the line
Moving the line to shorten some segments lengthens others. The least-squares line is the position where no adjustment reduces the total, which is a genuine balance rather than a compromise.
Worked example
Observed minus predicted, one point at a time.
\[ \text{For the line } y=2x+1 \text{ and the point } (3,8), \text{ find the residual.} \]
Compute the prediction
Why: Substitute the input into the line.
\[ 2(3) + 1 = 7 \]
Subtract in the right order
Why: Observed minus predicted.
\[ 8 - 7 \]
Read the value
Why: A positive residual.
\[ \text{residual } = 1 \]
Interpret the sign
Why: Positive means the point sits above the line.
Figure (svg): The solution to Worked example compute residuals for a proposed line shown as a ladder of expressions, one row per legal move
\[ 8-7 = 1: \text{ the point is } 1 \text{ above the line} \]
Verify: check the sign convention
Why: Observed minus predicted is the standard order, and it makes a positive residual mean the observation exceeded the prediction. Reversing the subtraction would flip every sign, which does not change the sum of squares but does reverse every interpretation — so the order is worth fixing once and keeping.
OpenStax, Precalculus, §2.4 Fitting Linear Models to Data §2.4, pp. 253-254
Faded example
The line is y equal to 3x minus 2, and a data point sits at (4, 8).
Fill in the blanks
\text10 = -2, \qquad \text___ = 8 - ___ = ___
Why: The line predicts 12 minus 2, which is 10, and the observation is 8. Observed minus predicted gives negative 2, so the point sits 2 below the line. The negative sign is informative rather than an error, and it is what makes squaring necessary before the residuals are added.
Worked example
Once the line is produced, it is used exactly like any other linear model.
\[ \text{A regression gives } y=1.18x+2.2 \text{ from data with } x \text{ from } 1 \text{ to } 8. \text{ Predict at } 5 \text{ and } 30. \]
Evaluate at the first input
Why: Substitute 5.
\[ 5.9 + 2.2 = 8.1 \]
Classify that prediction
Why: Five lies inside the data range.
Evaluate at the second
Why: Substitute 30.
\[ 35.4 + 2.2 = 37.6 \]
Classify it
Why: Thirty is far beyond the data.
Figure (svg): The solution to Worked example use a regression line shown as a ladder of expressions, one row per legal move
\[ y(5)=8.1 \text{ (safe)}; \quad y(30)=37.6 \text{ (extrapolation)} \]
Verify: notice the regression changed nothing about this
Why: The interpolation and extrapolation distinction from §2.3 applies exactly as before. A regression line is a linear model like any other, and being computed by a well-defined procedure gives it no extra authority outside the range of the data it was fitted to.
OpenStax, Precalculus, §2.4 Fitting Linear Models to Data §2.4, pp. 255-257
Error analysis
A student computes how far a point misses the fitted line.
Annotate
On: \( \text{residual} = \text{perpendicular distance from the point to the line} \)
Residuals are vertical because regression predicts the output from the input, and the input is treated as known. Fitting by perpendicular distance is a different and less common technique with different uses.
Prediction
One data point is moved far above the rest, with its input unchanged.
Predict first
What happens to the least-squares line?
Correct: It moves noticeably towards the outlier.
Why: The criterion squares each residual, so a point far from the line contributes enormously to the total, and the line can reduce the total substantially by moving towards it. This sensitivity is the method's main weakness and the reason for looking at the plot: an outlier that is a measurement error will distort the fit badly and the correlation coefficient will not report it.
Socratic
Several ways of combining the misses could have been chosen.
Discussion prompt
Why not simply add the residuals, or add their absolute values?
Hint: What happens to a point above the line and a point below it?
Answer:
Adding them plainly fails because positive and negative residuals cancel. A line missing one point by ten above and another by ten below would score zero, tying with a perfect fit.
Adding absolute values works and is a genuine alternative, used under the name least absolute deviations. It is less sensitive to outliers, which is sometimes an advantage.
Squaring is preferred mainly because it is smooth, which makes the minimising line computable by a formula rather than by search, and because it has clean statistical properties. The trade-off is exactly the outlier sensitivity of the previous probe — the method's strength and its weakness are the same feature.
Matching
The vocabulary of fitting is worth keeping straight.
Match the pairs
Why: These four cover the section's machinery. Note that an outlier is defined relative to the pattern rather than by any threshold, which is why it is identified by looking at the plot rather than by a calculation — and why the plot cannot be skipped.
Section
Section 3
Concept
The correlation coefficient measures how closely the data clusters about a straight line. Its sign gives the direction of the trend and its size gives the strength of it.
\[ -1 \le r \le 1 \]
The two pieces of information are genuinely independent, and treating the number as a single quality score is what produces the classic error of calling a correlation of negative 0.95 worse than one of positive 0.6. The first describes an almost perfect falling relationship; the second describes a weak rising one.
Figure (svg): Four scatter plots with different correlation coefficients: strong positive, weak positive, strong negative, and no correlation, each labelled with an approximate value of r
OpenStax, Precalculus, §2.4 Fitting Linear Models to Data §2.4, pp. 257-260
Picture it
Sign across, strength within — the four plots vary both independently.
Figure (svg): Four scatter plots with different correlation coefficients: strong positive, weak positive, strong negative, and no correlation, each labelled with an approximate value of r
The third plot has the second-largest absolute value and the most negative sign. It is the second-best fit of the four, and reading its sign as a verdict would rank it last.
Worked example
Separate the sign from the size before judging.
\[ \text{Which is the better linear fit: } r=-0.93 \text{ or } r=0.58? \]
Take absolute values
Why: This is what measures strength.
\[ 0.93\text{ and } 0.58 \]
Compare them
Why: The larger absolute value is the better fit.
\[ 0.93 > 0.58 \]
Read the signs separately
Why: They describe direction, not quality.
State both facts
Why: Strength and direction together.
Figure (svg): Four scatter plots with different correlation coefficients: strong positive, weak positive, strong negative, and no correlation, each labelled with an approximate value of r
\[ |-0.93| > |0.58|: \text{ the first is the stronger linear fit.} \]
Verify: say what each describes
Why: The first describes points lying tightly along a falling line — an excellent linear relationship in which one quantity decreases as the other increases. The second describes a loose rising trend with substantial scatter. Comparing the raw numbers would have ranked them the other way round, which is the error this example exists to prevent.
OpenStax, Precalculus, §2.4 Fitting Linear Models to Data §2.4, pp. 258-259
Ranking
Strongest first, using absolute values.
Put in order
Why: The absolute values are 0.98, 0.72, 0.31 and 0.05, which is the order given. The strongest fit has the most negative coefficient, which is the point: the sign says which way the trend runs and contributes nothing to how tightly the points cluster.
Worked example
The number answers a narrower question than it appears to.
\[ \text{Data lies almost exactly on a symmetric U, and } r \approx 0. \text{ Explain.} \]
Recall what r measures
Why: Closeness to a STRAIGHT line.
Consider the left half
Why: Output falls as input rises.
Consider the right half
Why: Output rises as input rises.
Combine them
Why: The two halves cancel.
Figure (svg): A scatter plot with a clear curved trend and a straight line fitted through it, showing that a low correlation coefficient does not mean there is no relationship
\[ r \approx 0: \text{ no LINEAR trend, despite a near-perfect relationship} \]
Verify: state what would have caught it
Why: The scatter plot, immediately. The relationship is nearly deterministic — knowing the input almost determines the output — and the coefficient reports none of that because it is not the question r asks. This is the strongest argument in the section for plotting before computing.
OpenStax, Precalculus, §2.4 Fitting Linear Models to Data §2.4, pp. 259-260
Trap
\[ r_1 = -0.97, \quad r_2 = 0.40 \;\Longrightarrow\; \text{the second model is better} \]
Compare the two numbers directly
Why: Positive 0.40 is greater than negative 0.97, so the second is ranked higher.
The model with the positive coefficient is chosen as the better fit.
The first is far better. Its absolute value is 0.97 against 0.40, so its points cluster much more tightly about their line.
The minus sign means the trend falls rather than rises. A falling relationship is not a worse relationship, and often it is the one the situation predicts.
Compare absolute values for strength and read signs for direction. They are two separate readings of one number, and combining them into a single ranking is the error.
Prediction
A data set has a correlation coefficient of 0.02.
Predict first
What can you conclude?
Correct: There is no linear trend, but there may be a strong curved one.
Why: The coefficient measures linear association only. A symmetric curve produces a coefficient near zero while the relationship is nearly perfect, as the U-shaped example showed. Concluding that the quantities are unrelated goes beyond what the number says, and only the scatter plot can settle it.
Discrimination
Each question is answered by one of the two pieces of information in r.
Sort into buckets
Sort each question by which part of r answers it.
Counterexample
A classmate claims a correlation coefficient near zero means the two quantities have nothing to do with each other.
Discussion prompt
Give a relationship that is nearly perfect and yet has a correlation near zero.
Hint: What shape makes the rising and falling halves cancel?
Answer:
Any symmetric curve. Take the squaring rule on inputs from negative 5 to 5: knowing the input determines the output exactly, so the relationship is as strong as a relationship can be.
But the coefficient is near zero, because the left half falls and the right half rises and the two linear trends cancel. There is no straight-line trend at all, and that is precisely what r reports.
So the honest statement is that r near zero means no linear association, not no association. The claim as stated is false, and the counterexample is not exotic — it is the second toolkit function from §1.2.
Section
Section 4
Concept
A strong correlation between two quantities has at least three possible explanations, and no number computed from the data distinguishes them.
The lurking variable is the one worth watching for, because it produces genuinely strong correlations between quantities with no direct connection at all. Ice cream sales and drowning deaths correlate strongly, and neither causes the other — hot weather causes both.
Figure (svg): Three explanations for a strong correlation between two quantities: one causes the other, the other causes the one, or a third quantity causes both, illustrated with a lurking-variable diagram
OpenStax, Precalculus, §2.4 Fitting Linear Models to Data §2.4, pp. 260-262
Picture it
Each diagram produces the same scatter plot.
Figure (svg): Three explanations for a strong correlation between two quantities: one causes the other, the other causes the one, or a third quantity causes both, illustrated with a lurking-variable diagram
Nothing in the data distinguishes them, because the data records only that the two move together. Which arrows are real is a question about the world.
Worked example
When two quantities correlate with no plausible direct link, look for a third.
\[ \text{Ice cream sales and drowning deaths correlate strongly. Explain.} \]
Test the first explanation
Why: Does eating ice cream cause drowning?
Test the second
Why: Do drownings cause ice cream sales?
Look for a third quantity
Why: Something that would raise both.
Check it explains both
Why: Heat raises ice cream sales and swimming alike.
Figure (svg): Three explanations for a strong correlation between two quantities: one causes the other, the other causes the one, or a third quantity causes both, illustrated with a lurking-variable diagram
\[ \text{Temperature causes both; neither causes the other.} \]
Verify: check the correlation is genuine
Why: The correlation is real and would be reproduced by any honest analysis — it is not a fluke or an artefact. That is what makes the example instructive: a strong, reproducible correlation with no causal link between the two quantities, which is exactly the situation the warning is about.
OpenStax, Precalculus, §2.4 Fitting Linear Models to Data §2.4, pp. 261-261
Sorting
Distinguish claims the data supports from claims it does not.
Sort into buckets
Sort each conclusion drawn from a strong correlation.
Worked example
Sometimes both causal directions are plausible and the data cannot separate them.
\[ \text{Exercise hours and reported wellbeing correlate. Which causes which?} \]
Consider the first direction
Why: Exercise plausibly improves wellbeing.
Consider the second
Why: Feeling well plausibly makes people exercise.
Consider a third quantity
Why: Good health could raise both.
Ask what would settle it
Why: A study that assigns exercise rather than observing it.
Figure (svg): The solution to Worked example when the direction is ambiguous shown as a ladder of expressions, one row per legal move
\[ \text{Correlation alone cannot decide; a controlled study is required.} \]
Verify: say what an experiment adds
Why: Randomly assigning who exercises breaks the link between exercise and any pre-existing difference between participants, so a difference in outcome afterwards can be attributed to the exercise itself. That is what a controlled experiment buys and what observational correlation cannot: the ability to rule out the second and third explanations by design.
OpenStax, Precalculus, §2.4 Fitting Linear Models to Data §2.4, pp. 261-262
Error analysis
A report examines school funding and test scores.
Annotate
On: \( r = 0.86 \;\Longrightarrow\; \text{funding causes higher scores} \)
State the association, then state what would be needed to establish a cause. The gap between the two is not a technicality — it is the whole difference between an observation and an experiment.
Prediction
Two quantities correlate, and neither plausibly causes the other.
Predict first
What is the most likely explanation?
Correct: A third quantity influencing both.
Why: The lurking variable is the standard explanation when neither direct link is plausible, and it accounts for a great many strong real-world correlations. The correlation itself is usually perfectly correct, which is what makes the situation interesting rather than a mistake to be found.
Two truths and a lie
Two of these are true of correlation and one is false.
Eliminate the wrong options
One of these claims is wrong.
Survives elimination: B
Why: B is the false claim, and no threshold would make it true. There is no value of the correlation coefficient that establishes a cause, because the coefficient measures association and every one of the three causal pictures produces association. Only the study's design can distinguish them.
Real world
Correlational findings are reported as causal claims constantly.
Discussion prompt
A headline says a habit is 'linked to' an outcome. What has probably been measured, and what should you ask?
Hint: What does 'linked to' translate to in this section's vocabulary?
Answer:
'Linked to' almost always means correlated, measured in an observational study where people were watched rather than assigned. That is an association and nothing more.
The questions to ask are the three explanations: could the outcome cause the habit rather than the reverse, and is there a third factor — income, age, health — plausibly driving both?
The follow-up question is whether the study was observational or controlled. If participants were randomly assigned, the causal claim may hold; if they were merely observed, the honest reading is association. This section's vocabulary is enough to ask both questions precisely, which is most of what it is for.
Section
Section 5
Concept
A regression line is a linear model, subject to everything §2.3 said about domains and extrapolation, plus the further caution that it was fitted to noisy data.
The temptation a formal procedure creates is to trust its output more than an eyeballed line, and the procedure is indeed more consistent. But it is minimising a criterion, not verifying that a line was the right shape, and it will produce a confident-looking answer for data it should have refused.
Figure (svg): A scatter plot with a clear curved trend and a straight line fitted through it, showing that a low correlation coefficient does not mean there is no relationship
OpenStax, Precalculus, §2.4 Fitting Linear Models to Data §2.4, pp. 252-262
Picture it
A regression line has been fitted to data that is obviously curved.
Figure (svg): A scatter plot with a clear curved trend and a straight line fitted through it, showing that a low correlation coefficient does not mean there is no relationship
The procedure ran without complaint and produced a line. Nothing in its output says the shape was wrong; only the plot does.
Worked example
Two questions: is the shape right, and is the input inside the data?
\[ \text{A fit has } r=0.94 \text{ on data from } 2 \text{ to } 9. \text{ Trust the prediction at } 25? \]
Check the fit quality
Why: The coefficient is high.
Check where the input sits
Why: Twenty five is far outside 2 to 9.
Ask what r covers
Why: It describes the fit over the DATA range only.
Conclude
Why: The number is computable but not trustworthy.
Figure (svg): The solution to Worked example decide whether to trust a prediction shown as a ladder of expressions, one row per legal move
\[ \text{Compute it if asked, but state that it is an unsupported extrapolation.} \]
Verify: say what r actually described
Why: The coefficient reports how tightly the observed points cluster about the line, across inputs from 2 to 9. It contains no information whatever about inputs near 25, where no data was collected. A high correlation and a far extrapolation are entirely compatible, and combining them is the most common misuse of a good fit.
OpenStax, Precalculus, §2.4 Fitting Linear Models to Data §2.4, pp. 257-260
Sorting
Some facts support a prediction and others are irrelevant to it.
Sort into buckets
Sort each fact by whether it makes a prediction at a distant input more trustworthy.
Worked example
Investigate first; remove only with a reason and a note.
\[ \text{One point sits far from the rest and pulls the line noticeably. What should be done?} \]
Check for a recording error
Why: Was it entered or measured wrongly?
Check whether it is genuine
Why: Some outliers are real and important.
Fit both ways if unsure
Why: With and without the point.
Report what was done
Why: Removal must always be stated.
Figure (svg): The solution to Worked example handle an outlier shown as a ladder of expressions, one row per legal move
\[ \text{Investigate, fit both ways, and state any removal explicitly.} \]
Verify: consider what silent removal amounts to
Why: Dropping inconvenient points without saying so makes the fit look better than the evidence supports, and it is indistinguishable from manipulating the result. Reporting both fits costs one extra line and lets the reader judge — which is the whole point of reporting rather than merely concluding.
OpenStax, Precalculus, §2.4 Fitting Linear Models to Data §2.4, pp. 254-256
Trap
\[ r=0.98 \text{ on } x \in [1,10] \;\Longrightarrow\; \text{the prediction at } x=200 \text{ is reliable} \]
Note the excellent fit and apply the line at a distant input
Why: The correlation is very high, so the model is treated as trustworthy everywhere.
The prediction far outside the data range is reported without qualification.
The correlation describes the fit over the data range only. It measures how closely the observed points lie to the line, and there are no observed points near 200.
A perfect fit on ten data points still says nothing about what happens twenty times further out, where the underlying relationship may have changed completely.
Fit quality and extrapolation risk are independent. A high correlation makes interpolation more trustworthy and does nothing at all for extrapolation, which is limited by the range of the evidence rather than by the tightness of the fit.
Prediction
A line has been fitted to data that genuinely is linear, with measurement noise.
Predict first
What should the residuals look like?
Correct: Randomly scattered above and below zero, with no pattern.
Why: When the model's shape is right, what remains is noise, and noise has no pattern. Any structure in the residuals — a curve, a trend, a fan — indicates the model has failed to capture something. Checking the residual plot is a better diagnostic than the correlation coefficient, because it reveals the shape problems r is blind to.
Elimination
Several practices are proposed for using a fitted model. Rule out the responsible ones.
Eliminate the wrong options
One of these is NOT responsible practice.
Survives elimination: B
Why: B is the irresponsible practice. Removing inconvenient data silently makes the fit appear better supported than it is and is indistinguishable from manipulating the result. An outlier may legitimately be removed after investigation, but the removal and its reason must always be reported.
Explain it
A high correlation coefficient is persuasive, and it should not persuade of everything.
Discussion prompt
Explain to a classmate what a correlation of 0.97 does and does not entitle them to claim.
Hint: Distinguish fit, prediction range, and causation.
Answer:
It does entitle them to say the points lie tightly along a rising straight line, and that predictions within the data range should be good.
It does not entitle them to predict far outside that range, because the coefficient describes only where the data is. And it does not establish that one quantity causes the other, because every one of the three causal pictures produces the same tight cluster.
The clean summary: a high correlation is a statement about fit, not about reach and not about cause. Those are three separate questions, and the single number answers only the first.
Comparison
Fill the blanks from memory. Each answers one question and is silent on the others.
Comparison matrix
| the scatter plot | the regression line | the correlation coefficient | |
|---|---|---|---|
| reveals the shape | yes | no | no |
| reveals outliers | yes | no | no |
| gives a prediction | roughly | yes | no |
| measures fit quality | roughly, by eye | no | yes |
| establishes causation | no | no | no |
The last row is the one worth committing to memory, and the first column is why the plot is never optional: it is the only tool in the table that can tell you a line was the wrong idea.
Pattern
Six steps, and the first and last are the ones a calculator will not prompt you for.
Steps 1 and 4 both amount to looking at a picture, and both catch the failures that every number in the procedure is blind to.
OpenStax Algebra and Trigonometry 2e, §4.3 Fitting Linear Models to Data §4.3
Check
Size for strength, sign for direction.
Check your understanding
Which correlation indicates the strongest linear relationship?
Answer: A
Why: Strength is the absolute value, and 0.91 is the largest of 0.91, 0.63, 0.44 and 0.20. The minus sign means the trend falls, which describes the direction rather than the quality of the fit.
Check
Observed minus predicted.
Check your understanding
A line predicts 14 at an input where the observed value was 11. What is the residual?
Answer: A
Why: Observed minus predicted is 11 minus 14, which is negative 3. A negative residual means the observation fell short of the prediction, so the point sits below the line.
Check
What can and cannot be concluded.
Check your understanding
Two quantities have a correlation of 0.88. What is the strongest justified conclusion?
Answer: A
Why: A correlation measures how closely the data clusters about a straight line, so a strong association is exactly what it establishes. Any causal claim requires a study designed to rule out reverse causation and lurking variables.
Real world
Regression is the most widely used statistical tool there is, and its two standard misuses are equally widespread.
Discussion prompt
A company finds that employees who attend an optional training course perform better, with a correlation of 0.7, and makes the course mandatory. What has been assumed?
Hint: Who chose to attend the course, and what else might be true of them?
Answer:
It assumes the training causes the performance, when the course was optional — so the people who attended chose to. Motivated or already-strong employees are more likely to volunteer, which is a lurking variable.
Making the course mandatory changes exactly the thing that produced the correlation. The volunteers may improve because they were the kind of people who improve, and the rest may not.
What would settle it is random assignment: send a randomly chosen half and compare. That breaks the link between attending and whatever else distinguished the volunteers, which is the one thing a correlation on observational data can never do.
Commit first
State your confidence along with your answer.
Predict first
Data lies almost exactly on a downward-opening parabola. What is its correlation coefficient likely to be?
Correct: Near zero, because there is no linear trend.
Why: The coefficient measures linear association only. A symmetric parabola rises on one side and falls on the other, so the two trends cancel and the value comes out near zero despite an almost perfect relationship. It is only exactly zero when the data is perfectly symmetric about the vertex, which is why the answer is 'near' rather than 'exactly'.
Explain it
The test of understanding is being able to say why, not just what.
Discussion prompt
Explain to a classmate why looking at the scatter plot matters even when you have the correlation coefficient.
Hint: What can a picture show that a single number cannot?
Answer:
A single number compresses the whole data set, and compression loses things. The coefficient reports one narrow fact — how close the points are to a straight line — and is silent on everything else.
The plot shows what the number cannot: whether the trend is curved, whether one outlier is dragging the fit, whether there are two clusters masquerading as one trend. All three produce coefficients that look unremarkable.
The clean argument is that the coefficient answers the question you asked, and the plot tells you whether it was the right question. That is why the plot comes first and gets looked at again afterwards.
Exit ticket
One honest answer, so the next chapter can start in the right place.
Predict first
Which idea from this lesson would you most want to see again?
Correct: Any of these is a legitimate answer; the useful one is the honest one.
Why: There is no correct choice here. The third produces wrong answers on assessments and the fourth produces wrong conclusions everywhere else, which makes both worth the time. The second is the least examinable and the most useful for understanding when the procedure can be trusted.
Connect it up
One page, drawn from memory, closes the chapter.
Draw it
Sketch four small scatter plots: one strong positive, one strong negative, one weak, and one strongly curved. Label each with roughly what its correlation coefficient would be. Beneath them, write the three explanations for a correlation between two quantities, and note which of them the coefficient can distinguish.
If your curved plot is labelled with a coefficient near zero, and none of the three explanations is marked as distinguishable, you have the two ideas this section exists to protect against.
Recap
Five things, and two of them are judgements rather than calculations.
| if you remember one thing | it should be this |
|---|---|
| about the procedure | plot the data first, and look again at the residuals afterwards |
| about the coefficient | sign is direction, size is strength, and they are independent |
| about r near zero | it means no LINEAR trend, not no relationship |
| about conclusions | association is what you have; causation needs a designed study |
Chapter 3 leaves linear functions behind for polynomials and rational functions, where the graph can turn, break, and run off to infinity — and where the questions this chapter answered in one step take a section each.
OpenStax, Precalculus, §2.4 Fitting Linear Models to Data §2.4, pp. 247-262 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Precalculus — $55/session, free consultation.