Scatter plots and the three kinds of correlation, the correlation coefficient r, approximating a best-fitting line by hand in four steps, using a line of fit to predict, and finding the regression line with a graphing calculator.
Subject: Algebra 2 · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Algebra 2 · Chapter 2 — Linear Equations and Functions
Draw Scatter Plots and Best-Fitting Lines
Objectives
Five outcomes. The fourth is where the honest limits of a model live.
McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 113-119 — the lesson these objectives are drawn from
Warm-up
Lesson 2.5 tested data against a model that had to pass through every point. Real data almost never does.
Discussion prompt
You measure eight pairs of values and none of them lies on any single straight line. Does that mean no linear model is worth writing? What would you look at before deciding?
Hint: Think about the difference between the points being scattered and their being patternless.
Answer:
No. Measurement always carries noise, so exact agreement would be more suspicious than disagreement. What matters is whether the points trend in one direction and cluster near some line, not whether any of them sits on it.
This lesson gives you two ways to answer that: a picture, and a single number between negative one and one that summarises how close the cluster is to a line.
Concept
A scatter plot is a graph of data pairs. When y tends to rise as x rises the data has positive correlation; when y tends to fall it has negative correlation; when the points show no pattern there is approximately none. The word tends is doing the work.
scatter plot — A graph of a set of data pairs, plotted as points with no line joining them.
Correlation is a statement about the cloud as a whole. Individual points may run the other way without changing the verdict.
Figure (svg): Three scatter plots side by side showing positive correlation, negative correlation, and no correlation
McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 113-113
Section
Section 1
Concept
Plot the pairs and look. If the cloud slopes upward the correlation is positive, if downward it is negative, and if it has no discernible slope there is approximately none. No arithmetic is needed for this first judgement.
positive correlation — A tendency for y to increase as x increases. Negative correlation is the tendency for y to decrease as x increases.
The word approximately in approximately no correlation is deliberate. Real data never gives an exactly patternless cloud, so the verdict is always a judgement.
Figure (svg): Three scatter plots side by side showing positive correlation, negative correlation, and no correlation
McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 113-113 — Describe correlation
Picture it
The same number of points, arranged three ways.
Figure (svg): Three scatter plots side by side showing positive correlation, negative correlation, and no correlation
The first two clouds each have a direction you could describe to someone over the phone. The third does not, and that is the whole distinction.
Worked example
Example 1. Two scatter plots from the same period, 1995 to 2003.
\[ \text{Cellular subscribers against service regions; and cellular subscribers against corded phone sales.} \]
Read the first plot's direction
Why: As the number of subscribers rose, the number of service regions also tended to rise.
Name it
Why: Rising together is positive correlation.
Read the second plot's direction
Why: As subscribers rose, corded phone sales tended to fall.
Name it
Why: One rising while the other falls is negative correlation.
Figure (svg): The solution to Worked example describe two correlations shown as a ladder of expressions, one row per algebraic move
\[ \text{first: positive correlation} \qquad \text{second: negative correlation} \]
Verify: say whether each direction makes sense
Why: More subscribers plausibly needs more service regions, so a positive correlation is what you would expect. And mobile phones replacing landlines would push corded sales down as subscribers rise, so a negative correlation fits too. Data that contradicts a sensible expectation is worth re-checking before it is believed.
McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 113-113
Sorting
Judge each from what you know of the two quantities.
Sort into buckets
Sort each pair of quantities by the correlation you would expect.
Predicting the correlation before plotting is a genuine check: a plot that contradicts a strong expectation usually means an error in the data or in the plotting.
Worked example
Exercise 6 asks how to do this without drawing anything.
\[ \text{Oil production } y \text{ in thousands of barrels, } x \text{ years after } 1994: \; 6660, 6560, 6470, 6450, 6250, 5880, 5820, 5800, 5750. \]
Check the inputs are in increasing order
Why: They are: zero through eight. If they were not, sorting them first is the only way the next step means anything.
\[ x\text{ rises } 0\text{ to } 8 \]
Read down the output column
Why: The values fall throughout, from 6660 to 5750.
Name the correlation
Why: One rising while the other falls is negative correlation.
Judge the strength
Why: The fall is steady, with no reversals at all, so the points will lie close to a falling line.
Figure (svg): The solution to Worked example read correlation from a table shown as a ladder of expressions, one row per algebraic move
\[ \text{negative correlation, and strong} \]
Verify: check for any reversal in the sequence
Why: Reading the nine values in order, each is smaller than the one before it — there is no point where the sequence turns back up. A perfectly monotonic column is stronger evidence than a merely downward-trending one, which is why the verdict is strong rather than just negative.
McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 117-117
Trap
\[ \text{Subscribers up, corded phone sales down: negative correlation.} \]
Conclude that buying a mobile phone causes corded phone sales to fall
Why: A relationship in the data is read as a mechanism in the world.
The data cannot distinguish that from a third factor — falling landline prices, changing housing patterns — moving both.
\[ \text{Subscribers up, corded phone sales down: negative correlation.} \]
State the correlation, and separate it from any claim about cause
Why: Correlation says the two quantities moved oppositely. It does not say which, if either, moved the other.
Here a causal story is plausible, and it may even be right — but the plausibility comes from knowing about phones, not from the scatter plot. The plot alone supports only the correlation.
Prediction
Commit before you reason.
Predict first
A data set has a clear positive correlation. One new point is added that sits well below the trend. What happens to the correlation?
Correct: It stays positive but weakens.
This is why the number of data points matters as much as their spread. A correlation computed from four points is a much weaker claim than the same number computed from forty.
Why: Correlation describes a tendency across the whole cloud, so one contrary point cannot reverse a clear trend — but it does make the cloud less tight around any line, which is what weakening means. How much depends on how far out the point is and how many points there already are: with eight points one outlier matters considerably, and with eight hundred it barely registers.
Explain it
A classmate says a positive correlation is broken because one point goes the wrong way.
Discussion prompt
In three sentences, explain what positive correlation actually claims and why a single contrary point does not refute it. Give them one everyday example where the tendency is obvious but exceptions are common.
Hint: Compare it with a rule that must hold for every case.
Answer:
Positive correlation says the two quantities tend to rise together across the data as a whole, not that every single pair obeys it. It is a statement about the cloud, so one point going the other way is an exception, not a counterexample.
Height and weight is the everyday case: taller people weigh more on average, and yet you can easily find a tall light person and a short heavy one without the tendency being wrong.
Two truths and a lie
All three are about correlation.
Eliminate the wrong options
Two of these are true. Knock those out and keep the false one.
Survives elimination: C
Why: The survivor is the false one. Correlation measures how the two quantities moved together; it says nothing about mechanism. A third factor may drive both, the causation may run the other way, or the association may be coincidence — and no amount of strength in the correlation distinguishes those cases.
Section
Section 2
Concept
The correlation coefficient r measures how well a line fits the data. Near one, the points lie close to a rising line; near negative one, close to a falling line; near zero, they lie close to no line at all.
correlation coefficient — A number r from negative one to one measuring how well a line fits a set of data pairs. Its sign gives the direction and its size gives the strength.
\[ -1 \leq r \leq 1 \]
Two facts in one number: the sign is the direction of the trend, and how close the size is to one is the strength of it.
Figure (svg): A number line from negative one to one showing what values of the correlation coefficient mean
McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 114-114 — Correlation coefficients
Picture it
What each part of the scale looks like.
Figure (svg): A number line from negative one to one showing what values of the correlation coefficient mean
Note that r near zero does not mean the two quantities are unrelated — only that no straight line describes them. A perfect parabola can have an r of nearly zero.
Worked example
Example 2. Choose from negative one, negative one half, zero, one half, and one.
\[ \text{(a) a clear but weak downward cloud; (b) a patternless cloud; (c) a tight upward cloud.} \]
Judge the first plot's direction and strength
Why: It slopes downward, so r is negative; but the points are loosely scattered, so it is not near negative one.
\[ \text{between } 0\text{ and } -1 \]
Choose the closest option
Why: Negative one half is the only negative option that is not extreme. The actual value is about negative 0.46.
\[ r = -0.5 \]
Judge the second plot
Why: No direction at all, so r is near zero. The actual value is about negative 0.02.
\[ r = 0 \]
Judge the third plot
Why: A tight upward cloud, so r is near positive one. The actual value is about 0.98.
\[ r = 1 \]
Figure (svg): The solution to Worked example estimate r from three plots shown as a ladder of expressions, one row per algebraic move
\[ r \approx -0.5, \quad r \approx 0, \quad r \approx 1 \]
Verify: check each estimate against the stated actual value
Why: The three actual values are -0.46, -0.02 and 0.98, and each estimate is the nearest of the five options. Notice the second is not exactly zero — real data almost never is — which is why the question asks which value r is CLOSEST to rather than what it equals.
McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 114-114
Matching
Four clouds, four coefficients.
Match the pairs
Why: Two features must be read separately: the direction, which fixes the sign, and the tightness, which fixes how close the size is to one. A loose downward cloud and a tight downward cloud share a sign and differ in magnitude, which is exactly the distinction between the second and fourth rows.
Worked example
Guided Practice 4's data, judged before any calculator is reached for.
\[ \text{Oil production falls steadily from } 6660 \text{ to } 5750 \text{ over nine years.} \]
Read the direction
Why: Every value is lower than the one before, so the trend is downward and r is negative.
Judge the strength from the reversals
Why: There are none: the sequence never turns back up. That is a very tight trend.
Check the steadiness of the steps
Why: The drops are 100, 90, 20, 200, 370, 60, 20, 50 — uneven, so the fit is good but not perfect.
Estimate
Why: Strongly negative but not exactly negative one. The computed value is about negative 0.97.
\[ r\text{ near } -1 \]
Figure (svg): The solution to Worked example estimate r for the oil data shown as a ladder of expressions, one row per algebraic move
\[ r \approx -0.97 \]
Verify: reconcile the two observations
Why: No reversals argues for r very near negative one, while uneven step sizes argue for slightly less. Both are true, and the computed value of -0.97 sits exactly where those two observations predict: extremely strong, but not perfect.
McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 117-117
Trap
\[ r \approx 0 \text{ for a data set} \]
Conclude that x and y are unrelated
Why: The coefficient is read as measuring relationship in general rather than linear relationship in particular.
But a set of points lying exactly on a symmetric parabola can have an r of essentially zero, and nothing about that data is unrelated.
\[ r \approx 0 \text{ for a data set} \]
Conclude that no LINE fits the data well
Why: That is all r measures. A strong non-linear relationship can produce an r near zero.
\[ (-2,4), (-1,1), (0,0), (1,1), (2,4) \;\Longrightarrow\; r = 0 \]
Those five points lie exactly on y equals x squared — a perfect relationship with zero linear correlation. Always look at the plot, not only at r.
Discrimination
Each statement about r tells you something. Say what.
Sort into buckets
Sort each fact about r by what it establishes.
Estimation
Eight points that rise steadily with only slight scatter.
Predict first
Roughly what is r?
Correct: About 0.98.
The alternative-fuel vehicle data in Example 3 gives r of about 0.993 for precisely this reason: eight points rising with very little scatter.
Why: Rising means the sign is positive, and steadily with only slight scatter means the magnitude is near one. Half is what a visibly loose cloud gives, not a tight one, and Example 2 makes exactly that comparison: its loose cloud came out at -0.46 while its tight one came out at 0.98. Judging the sign first and the magnitude second turns this into two easy questions instead of one hard one.
Counterexample
A classmate proposes a rule about r.
\[ r \approx 0 \;\Longrightarrow\; x \text{ and } y \text{ have nothing to do with each other} \]
Discussion prompt
Find five data pairs with an r of essentially zero whose two quantities are nevertheless perfectly related. Then say precisely what r does and does not measure.
Hint: Try a symmetric shape centred on the vertical axis.
Answer:
\[ (-2,4), \; (-1,1), \; (0,0), \; (1,1), \; (2,4) \]
Those points lie exactly on y equals x squared, so y is completely determined by x — and yet the cloud has no upward or downward slope at all, so r is zero.
What r measures is how well a straight line fits. It is silent about every other kind of relationship, which is why the plot must always be looked at alongside the number.
Section
Section 3
Concept
Draw the scatter plot, sketch a line that follows the trend with about as many points above it as below, choose two points on your line, and write the equation through them. Those two points do not have to be data points.
best-fitting line — The line that lies as close as possible to all the data points. A hand-sketched approximation to it is called a line of fit.
Choosing points on your drawn line rather than two data points is what makes the fit use the whole cloud instead of just two members of it.
Figure (svg): A scatter plot of alternative-fuel vehicle data with a fitted line through two chosen points, one of which is not a data point
McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 115-115 — Approximating a Best-Fitting Line
Picture it
Example 3: alternative-fuelled vehicles in use, in thousands, x years after 1997.
Figure (svg): A scatter plot of alternative-fuel vehicle data with a fitted line through two chosen points, one of which is not a data point
The chosen points were (1, 300), which is not a data point at all, and (7, 548), which happens to be one. Both lie on the sketched line, which is the only requirement.
Worked example
Example 3, all four steps.
\[ \text{Data } (0,280), (1,295), (2,322), (3,395), (4,425), (5,471), (6,511), (7,548). \text{ Fit a line.} \]
Plot the eight points
Why: They rise steadily with mild scatter, so a line is a reasonable model.
Sketch a line through the trend
Why: Aim for about as many points above as below, rather than through any particular pair.
Choose two points ON THE LINE
Why: The point (1, 300) is on the sketched line but is not in the data; (7, 548) happens to be both.
\[ (1, 300)\text{ and } (7, 548) \]
Find the slope of the line through them
Why: Five hundred and forty-eight minus 300 is 248, over 7 minus 1, which is 6.
\[ m = \frac{248}{6} = 41.3 \]
Write the equation with point-slope form
Why: Using (1, 300): y minus 300 equals 41.3 times x minus 1, which simplifies.
\[ y = 41.3 x + 259 \]
Figure (svg): The solution to Worked example approximate the best-fitting line shown as a ladder of expressions, one row per algebraic move
\[ y \approx 41.3x + 259 \]
Verify: count the points above and below the line
Why: Substituting each x into the model gives 259, 300, 342, 383, 424, 465, 507, 548 against the data's 280, 295, 322, 395, 425, 471, 511, 548. Three data points sit above the line, four below, and one on it — a balanced fit, which is what step two asked for.
McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 115-115
Ranking
Approximating a best-fitting line.
Put in order
Why: The plot has to exist before a line can be sketched onto it, and the line has to exist before points can be chosen from it — that ordering is what prevents the endpoints error. Writing the equation is Lesson 2.4's two-point method, unchanged. The check is not in the textbook's four steps but is worth adding: it is the only way to notice that a line is systematically high or low.
Worked example
Guided Practice 4a. The oil production data, fitted by the same four steps.
\[ \text{Oil production, thousands of barrels, } x \text{ years after } 1994: \; 6660, 6560, 6470, 6450, 6250, 5880, 5820, 5800, 5750. \]
Plot and sketch
Why: The points fall steadily; a line drawn through the middle of the cloud has some above and some below.
Choose two points on the sketched line
Why: Taking (1, 6560) and (7, 5800), both of which are data points that lie close to the trend.
\[ (1, 6560)\text{ and } (7, 5800) \]
Find the slope
Why: Five thousand eight hundred minus 6560 is negative 760, over 6.
\[ m = -126.7 \]
Write the equation
Why: Using (1, 6560): y minus 6560 equals negative 126.7 times x minus 1.
\[ y = -126.7 x + 6687 \]
Figure (svg): The solution to Worked example fit a falling line shown as a ladder of expressions, one row per algebraic move
\[ y \approx -126.7x + 6687 \]
Verify: check the fit at both ends of the data
Why: At x equal to 0 the model gives 6687 against the measured 6660, and at x equal to 8 it gives 5673 against the measured 5750. Both are within about one percent, and the errors fall on opposite sides, which is the signature of a balanced fit rather than one that drifts.
McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 117-117
Error analysis
A student fits a line to the vehicle data by joining the first and last data points.
Annotate
On: \( m = \frac{548 - 280}{7 - 0} = 38.3 \;\Longrightarrow\; y = 38.3x + 280 \)
The endpoints are two data points among many, and they carry no special authority. A line of fit is supposed to be influenced by every point.
Fill the middle
Example 3's two chosen points.
Fill in the blanks
m = \frac7 - 1___} = \frac______ \approx 41.3
Why: The numerator used the second point's y value first, so the denominator must use the second point's x value first: 7 minus 1, which is 6. That consistency rule is Lesson 2.2's, unchanged. Note that one of these two points, (1,300), is not in the data at all — it was read off the sketched line.
Definition probe
Step three of the procedure.
Sort into buckets
Sort each point by whether it may be used to write the equation of the line of fit.
Explain it to yourself
It would certainly be quicker.
\[ y = 38.3x + 280 \quad \text{versus} \quad y = 41.3x + 259 \]
Discussion prompt
Explain what is lost by fitting a line through the first and last data points instead of sketching a line through the whole cloud. When would the two methods agree, and what does that tell you about when the shortcut is safe?
Hint: Think about what happens if one of the two endpoints is unusual.
Answer:
The endpoints method gives all the influence to two points and none to the rest, so an unusual first or last measurement drags the whole line with it. The six middle points, which carry most of the evidence, are simply ignored.
The two methods agree when the endpoints happen to sit on the overall trend. That is exactly the case where the shortcut is safe — and you can only know it by drawing the plot, which is the work the shortcut was trying to avoid.
Section
Section 4
Concept
Once you have a model, predicting is one substitution. What matters is whether the input lies inside the range of the data, which is interpolation, or outside it, which is extrapolation and carries much more risk.
\[ y = 41.3(13) + 259 \approx 796 \]
The arithmetic is identical either way. Only the confidence differs, and it differs a great deal.
Figure (svg): The fitted line extended beyond the data to a prediction thirteen years after 1997
McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 116-116 — Use a line of fit to make a prediction
Picture it
Example 4: predicting 2010 from data covering 1997 to 2004.
Figure (svg): The fitted line extended beyond the data to a prediction thirteen years after 1997
Thirteen years after 1997 is six years past the last measurement — nearly doubling the span the model was built on. The model does not know that, and will answer just as confidently either way.
Worked example
Example 4, using the line from Example 3.
\[ \text{Using } y = 41.3x + 259, \text{ predict the number of vehicles in } 2010. \]
Convert the year into the model's input
Why: The model counts years after 1997, and 2010 is thirteen years after.
\[ x = 13 \]
Substitute
Why: Forty-one point three times thirteen is 536.9.
\[ y = 536.9 + 259 \]
Add and round sensibly
Why: The data was given to three figures, so three figures is honest.
\[ y = 796 \]
State the answer with its unit
Why: The values were in thousands, so this is about 796 thousand vehicles.
\[ \text{about } 796, 000 \]
Figure (svg): The solution to Worked example predict 2010 shown as a ladder of expressions, one row per algebraic move
\[ y = 41.3(13) + 259 \approx 796 \text{ thousand vehicles} \]
Verify: check that the size is consistent with the trend
Why: The data ended at 548 thousand in 2004, and the model adds about 41 thousand a year for six more years, which is about 248 thousand. Five hundred and forty-eight plus 248 is 796 — the same answer reached by extending the trend rather than by substituting, which confirms the arithmetic.
McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 116-116
Sorting
The vehicle data covers x from 0 to 7.
Sort into buckets
Sort each prediction by whether it lies inside or outside the data.
Extrapolation is not forbidden — it is often the whole point of building a model — but it should be labelled, and its risk grows with distance from the data.
Worked example
The same model, used for a year that lies within the measured range.
\[ \text{Using } y = 41.3x + 259, \text{ estimate the figure for } 2001, \text{ and compare with the data.} \]
Convert the year
Why: Two thousand and one is four years after 1997.
\[ x = 4 \]
Substitute
Why: Forty-one point three times four is 165.2.
\[ y = 165.2 + 259 \]
Compute
Why: Four hundred and twenty-four thousand.
\[ y = 424 \]
Compare with the measured value
Why: The data records 425 thousand for that year, so the model is one thousand low.
\[ \text{data says } 425 \]
Figure (svg): The solution to Worked example predict inside the data shown as a ladder of expressions, one row per algebraic move
\[ y(4) = 424 \quad \text{against the measured } 425 \]
Verify: quantify the error and compare it with the extrapolation
Why: The model is off by 1 in 425, about a quarter of one percent, which is excellent — and it should be, because this is interpolation inside the data the model was fitted to. No such check is available for the 2010 prediction, which is precisely what makes extrapolation riskier.
McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 116-116
Error analysis
A student predicts 2010 from the vehicle model.
Annotate
On: \( y = 41.3(2010) + 259 = 83\,272 \text{ thousand vehicles} \)
Every model comes with a definition of its input. An answer that is absurdly large or small almost always means that definition was skipped.
Edge cases
The vehicle model, y equals 41.3x plus 259.
Discussion prompt
What does the model predict for the year 2100? Is that a believable number of alternative-fuelled vehicles? Say what feature of a linear model makes long-range predictions unreliable for quantities like this.
Hint: First convert 2100 into the model's input.
Answer:
\[ x = 103 \;\Longrightarrow\; y = 41.3(103) + 259 \approx 4514 \text{ thousand} \]
About four and a half million — which is not absurd on its own, but the model assumes the SAME 41 thousand vehicles are added every single year for a century, regardless of population, technology or saturation.
A linear model has a constant rate built into it, so it can never level off or accelerate. Real adoption curves do both, which is why Chapter 7's exponential and logistic shapes exist.
Real world
The oil production model, y equals negative 126.7x plus 6687, with x years after 1994 and y in thousands of barrels per day.
Discussion prompt
Predict production in 2009, say whether it is interpolation or extrapolation, and give one specific reason the prediction might be badly wrong.
Hint: The data ran from 1994 to 2002.
Answer:
\[ x = 15 \;\Longrightarrow\; y = -126.7(15) + 6687 \approx 4787 \text{ thousand barrels per day} \]
This is extrapolation: the data covers x from 0 to 8, and 15 is nearly twice as far out as the data reaches.
It could be badly wrong because oil production responds to price and to new extraction technology, neither of which moves steadily. The shale boom that began around 2008 reversed the decline entirely — a linear model fitted to the 1990s could not possibly have anticipated it, and would have kept predicting a fall.
Commit first
Answer, then rate your confidence honestly.
Predict first
A model fits its data with r equal to 0.99. How confident should you be in a prediction ten years beyond the data?
Correct: Cautious — r measures how well the line fits the data you have, and says nothing about what happens outside it.
The honest way to report a prediction is with both facts: how well the model fits what was measured, and how far outside that range the prediction sits.
Why: A correlation coefficient is computed entirely from the observed pairs. It cannot know whether the trend continues, because no evidence about the future is in the calculation. A model can fit a decade of data essentially perfectly and still be wrong about the next year, if the underlying situation changes. The number of data points matters too, but it does not repair the basic problem with extrapolation.
Section
Section 5
Concept
A graphing calculator's linear regression feature finds the best-fitting line using every data point, and reports the correlation coefficient alongside it. The result is reproducible: two people get the same line from the same data.
\[ y = 40.9x + 263, \quad r \approx 0.993 \]
Enter the inputs in one list and the outputs in another, then run the linear regression command. If r does not appear, turn diagnostics on.
Figure (svg): Two columns comparing the hand-sketched line of fit with the calculator's regression line
McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 116-116 — Use a graphing calculator to find a best-fitting line
Picture it
The same eight data points, fitted two ways.
Figure (svg): Two columns comparing the hand-sketched line of fit with the calculator's regression line
The two lines differ by about one percent in slope and give predictions one part in eight hundred apart. A careful hand fit is not a poor substitute — it is a good approximation to the same thing.
Worked example
Example 5. The same vehicle data, entered and computed.
\[ \text{Find the regression line for the eight vehicle data pairs.} \]
Enter the data into two lists
Why: Years since 1997 in the first list, vehicle counts in the second, in matching order.
\[ L 1: 0..7, L 2: 280..548 \]
Run the linear regression command
Why: The calculator reports a slope of about 40.869 and an intercept of about 262.83.
\[ a = 40.87, b = 262.83 \]
Round sensibly
Why: The data has three significant figures, so the model should not claim more.
\[ y = 40.9 x + 263 \]
Read the correlation coefficient
Why: The calculator reports r of about 0.993, a very strong positive correlation.
\[ r = 0.993 \]
Figure (svg): The solution to Worked example run the regression shown as a ladder of expressions, one row per algebraic move
\[ y = 40.9x + 263, \quad r \approx 0.993 \]
Verify: plot the line with the scatter plot
Why: Graphing the regression equation over the scatter plot shows it running through the middle of the cloud with points on both sides throughout — the same balance the hand sketch aimed for. An r of 0.993 predicts exactly that appearance, and seeing it confirms no data was mistyped.
McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 116-116
Comparison
Fill the blanks. Both methods answer the same question differently.
Comparison matrix
| Feature | Sketched by hand | Linear regression |
|---|---|---|
| Which points influence it | all of them, through your eye | all of them, by computation |
| Reproducible by someone else | no - it depends on your sketch | yes, exactly |
| Reports a correlation coefficient | no | yes |
| Result on the vehicle data | y = 41.3x + 259 | y = 40.9x + 263 |
| Needs equipment | no | yes |
The hand method's weakness is reproducibility, not accuracy. On well-behaved data a careful sketch lands within a percent or two of the computed line.
Worked example
Guided Practice 4c in structure: repeat a prediction with the regression line.
\[ \text{Predict } 2010 \text{ with } y = 40.9x + 263, \text{ and compare with the hand fit's } 796. \]
Substitute thirteen into the regression model
Why: Forty point nine times thirteen is 531.7.
\[ y = 531.7 + 263 \]
Compute
Why: Seven hundred and ninety-five thousand approximately.
\[ y = 795 \]
Compare with the hand fit
Why: The hand fit gave 796; the difference is one thousand out of nearly eight hundred thousand.
\[ 796\text{ against } 795 \]
Judge the agreement
Why: Roughly one part in eight hundred, far smaller than the uncertainty in extrapolating six years past the data.
Figure (svg): The solution to Worked example compare the two predictions shown as a ladder of expressions, one row per algebraic move
\[ y(13) = 40.9(13) + 263 \approx 795 \]
Verify: ask which source of error dominates
Why: The two methods differ by about 0.1 percent, while the risk in extrapolating six years beyond the data is far larger than that. So the choice of fitting method is not what limits this prediction — the extrapolation is. Knowing which error dominates tells you where extra care is worth spending.
McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 116-116
Error analysis
A student enters the vehicle data with the second list sorted independently and gets a suspicious result.
Annotate
On: \( \text{L1: } 0,1,2,3,4,5,6,7 \qquad \text{L2: } 548,511,471,425,395,322,295,280 \)
Always check the sign of the reported slope against the direction you can see in the data. A sign flip is the usual signature of mismatched lists.
Ranking
Finding and checking a regression line.
Put in order
Why: Entering the data has to come first, and the regression cannot run without it. Rounding comes before graphing so that the equation you plot is the one you will actually quote. The last two steps are the check, and they are the ones most often skipped — a mistyped value can shift the line without producing any error message, and only the plot reveals it.
Prediction
Commit before reasoning.
Predict first
The oil production data falls steadily with no reversals across nine years. What will the regression report for r?
Correct: About -0.97.
\[ y \approx -129.8x + 6702, \quad r \approx -0.97 \]
Why: Falling makes the sign negative, and no reversals across nine points makes the magnitude close to one. Half would correspond to a visibly loose cloud, and zero to no trend at all — neither matches a sequence where every value is lower than the last. Predicting r before running the regression is a real check: a reported value with the wrong sign means the lists were mismatched.
Socratic
One question, and nothing else on this slide.
\[ y = 40.9x + 263 \]
Discussion prompt
The calculator calls this the best-fitting line. Best according to what standard? Consider two candidate lines that each miss the data by the same total amount, one missing every point slightly and one matching six points exactly while missing two badly. Which should count as better, and why might reasonable people disagree?
Hint: Think about what you want the line for.
Answer:
Regression minimises the sum of the SQUARED vertical distances, which punishes large misses much more heavily than small ones. So it prefers the line that misses everything slightly over the one with two bad misses.
Squaring is a choice, not a law. Minimising the plain distances instead gives a different line, one less disturbed by outliers, and for data with a few bad measurements that is arguably the better answer.
Reasonable people disagree because the right standard depends on whether the outliers are errors, which you want ignored, or real, which you want accounted for. The calculator cannot know which, so it applies one convention and leaves the judgement to you.
Comparison
Fill the blanks. Each stage answers a different question.
Comparison matrix
| Stage | Question it answers | Tool |
|---|---|---|
| Scatter plot | is there a trend, and which way? | your eyes |
| Correlation coefficient | how tightly do the points hug a line? | the number r |
| Line of fit | what line describes the trend? | sketch, then two points on it |
| Prediction | what value at a new input? | substitute into the model |
| Regression | what is the best line, exactly? | a graphing calculator |
The first two stages decide whether the last three are worth doing at all. Fitting a line to a patternless cloud produces an equation and no information.
Pattern
One routine takes data to a usable model.
Step three is where this lesson differs from every earlier one. Everywhere else two points determined the line exactly; here they are read off a line that was drawn from all the data.
OpenStax Algebra and Trigonometry 2e, §4.3 Fitting Linear Models to Data §4.3
Check
Estimating r from a description.
Check your understanding
A scatter plot shows a clear but fairly loose downward cloud. Which value is r closest to?
Answer: A
Why: Downward makes r negative, and loose means the magnitude is well short of one. Example 2's matching plot had an actual value of about -0.46, which rounds to the -0.5 option.
Check
Fitting by hand. Which two points may be used?
Check your understanding
In step 3 of approximating a best-fitting line, which two points should you choose?
Answer: A
Why: The points must lie on the sketched line, which was drawn using every data point. They need not be data points at all — Example 3 uses (1, 300), which is not in the data.
Check
Predicting. Convert the input first.
Check your understanding
Using y = 41.3x + 259, where x is years after 1997 and y is thousands of vehicles, predict the figure for 2010.
Answer: A
Why: 2010 is 13 years after 1997, so x is 13. Then 41.3 times 13 is 536.9, and adding 259 gives about 796 thousand.
Real world
A study reports a strong positive correlation between the number of fire engines sent to a fire and the amount of damage the fire causes.
Discussion prompt
Describe the correlation, say what a naive causal reading would conclude, and explain what is really going on. Then say what extra variable would need to be measured to make sense of the data.
Hint: Ask what decides how many engines are sent.
Answer:
The correlation is positive and probably strong: bigger numbers of engines go with bigger damage figures. A naive reading concludes that sending engines causes damage, and therefore that fewer should be sent.
What is really happening is that a third variable — the size of the fire — drives both. Big fires get more engines and cause more damage; the engines are a response to the size, not a cause of the damage.
Measuring fire size, or damage per unit of fire size, would separate the two. This is the classic illustration of why correlation does not establish causation, and it is exactly the trap from Section 1 in a case where the naive conclusion is actively harmful.
Commit first
Answer, then rate your confidence honestly.
Predict first
Two data sets both have r equal to 0.9. Does that mean their scatter plots look equally tight?
Correct: Roughly, but the slopes may differ completely.
This is why a model needs both numbers reported: the slope says how much y changes per unit of x, and r says how much to trust that the relationship is linear at all.
Why: The coefficient r measures how closely the points cluster around their line, not how steep that line is. Two data sets can both have r equal to 0.9 with one rising gently and the other almost vertically. The clouds look similarly tight around their own lines, which is what r captures, but the pictures can look very different overall. Number of points affects how much confidence r deserves rather than what it measures.
Explain it
They have only ever seen graphs where the points lie exactly on a line.
Discussion prompt
In four sentences or fewer, explain why real data does not, what a line of fit is for, and the one thing they should check before trusting a prediction made from one.
Hint: The check is about where the prediction sits relative to the data.
Answer:
Real measurements carry noise, so points scatter around a trend rather than sitting on it. A line of fit is the line that follows that trend as closely as possible, and it lets you estimate values you did not measure.
Before trusting a prediction, check whether the input lies inside the range of the data. Inside is interpolation and is usually safe; outside is extrapolation, and the further out you go the less the data has to say about it.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For estimating r, read the sign from the direction and the size from the tightness, as two separate questions. For choosing points, take them off the line you drew, not from the data. For predicting, reread the variable definition before substituting anything. For correlation and cause, always ask what third variable could be driving both. Do five of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
Across the top of a page draw three small scatter plots showing positive correlation, negative correlation and none, and write beneath each the value of r you would estimate. Underneath, draw one larger scatter plot of at least eight points of your own invention that show a clear trend with visible scatter. Sketch a line of fit through it, balancing points above and below, then mark two points ON YOUR LINE — making at least one of them not a data point — and work out the equation through them, showing the slope calculation and the point-slope step. Beside the plot, draw a vertical dashed line where your data ends and label everything to the right of it extrapolation. At the bottom, use your equation to predict one value inside the data and one well outside it, and write one sentence about how much you trust each.
If both predictions got the same level of trust in your last sentence, reread Section 4. The arithmetic is identical for both; the confidence is not.
Recap
Five things, and the last two are about knowing what a model cannot tell you.
| If you see | Then |
|---|---|
| A cloud sloping upward | Positive correlation, r positive |
| A cloud with no direction | r near zero: no LINE fits |
| A sketched line of fit | Take your two points off it |
| A prediction beyond the data | Label it extrapolation |
| A strong correlation | Ask what third variable could cause both |
Lesson 2.7 leaves data behind and returns to exact functions, taking the absolute value function from Chapter 1 and moving it around the plane.
McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 113-119 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Algebra 2 — $55/session, free consultation.