2.6 Scatter Plots and Best-Fitting Lines

Scatter plots and the three kinds of correlation, the correlation coefficient r, approximating a best-fitting line by hand in four steps, using a line of fit to predict, and finding the regression line with a graphing calculator.

Subject: Algebra 2 · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Lesson 2.6 Scatter Plots and Best-Fitting Lines

Title

Algebra 2 · Chapter 2 — Linear Equations and Functions

Draw Scatter Plots and Best-Fitting Lines

2. By the end of this lesson you can

Objectives

Five outcomes. The fourth is where the honest limits of a model live.

McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 113-119 — the lesson these objectives are drawn from

3. What you already have

Warm-up

Lesson 2.5 tested data against a model that had to pass through every point. Real data almost never does.

Discussion prompt

You measure eight pairs of values and none of them lies on any single straight line. Does that mean no linear model is worth writing? What would you look at before deciding?

Hint: Think about the difference between the points being scattered and their being patternless.

Answer:

No. Measurement always carries noise, so exact agreement would be more suspicious than disagreement. What matters is whether the points trend in one direction and cluster near some line, not whether any of them sits on it.

This lesson gives you two ways to answer that: a picture, and a single number between negative one and one that summarises how close the cluster is to a line.

4. A cloud can have a direction

Concept

A scatter plot is a graph of data pairs. When y tends to rise as x rises the data has positive correlation; when y tends to fall it has negative correlation; when the points show no pattern there is approximately none. The word tends is doing the work.

scatter plot — A graph of a set of data pairs, plotted as points with no line joining them.

Correlation is a statement about the cloud as a whole. Individual points may run the other way without changing the verdict.

Figure (svg): Three scatter plots side by side showing positive correlation, negative correlation, and no correlation

Correlation describes a tendency across the whole cloud, not a rule that every point obeys.

McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 113-113

5. Scatter plots and correlation

Section

Section 1

6. Three verdicts, read from the shape

Concept

Plot the pairs and look. If the cloud slopes upward the correlation is positive, if downward it is negative, and if it has no discernible slope there is approximately none. No arithmetic is needed for this first judgement.

positive correlation — A tendency for y to increase as x increases. Negative correlation is the tendency for y to decrease as x increases.

The word approximately in approximately no correlation is deliberate. Real data never gives an exactly patternless cloud, so the verdict is always a judgement.

Figure (svg): Three scatter plots side by side showing positive correlation, negative correlation, and no correlation

Correlation describes a tendency across the whole cloud, not a rule that every point obeys.

McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 113-113 — Describe correlation

7. The three shapes

Picture it

The same number of points, arranged three ways.

Figure (svg): Three scatter plots side by side showing positive correlation, negative correlation, and no correlation

Correlation describes a tendency across the whole cloud, not a rule that every point obeys.

The first two clouds each have a direction you could describe to someone over the phone. The third does not, and that is the whole distinction.

8. Worked example: describe two correlations

Worked example

Example 1. Two scatter plots from the same period, 1995 to 2003.

\[ \text{Cellular subscribers against service regions; and cellular subscribers against corded phone sales.} \]

Read the first plot's direction

Why: As the number of subscribers rose, the number of service regions also tended to rise.

Name it

Why: Rising together is positive correlation.

Read the second plot's direction

Why: As subscribers rose, corded phone sales tended to fall.

Name it

Why: One rising while the other falls is negative correlation.

Figure (svg): The solution to Worked example describe two correlations shown as a ladder of expressions, one row per algebraic move

The whole solution at once: each drop is one legal move.

\[ \text{first: positive correlation} \qquad \text{second: negative correlation} \]

Verify: say whether each direction makes sense

Why: More subscribers plausibly needs more service regions, so a positive correlation is what you would expect. And mobile phones replacing landlines would push corded sales down as subscribers rise, so a negative correlation fits too. Data that contradicts a sensible expectation is worth re-checking before it is believed.

McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 113-113

9. Which correlation?

Sorting

Judge each from what you know of the two quantities.

Sort into buckets

Sort each pair of quantities by the correlation you would expect.

Positive
hours studied and test score; height and arm span
Negative
car age and resale value; outdoor temperature and heating cost
Approximately none
shoe size and favourite colour code
pos
Both quantities tend to rise together. Height and arm span in particular are close to proportional in most people, so the correlation is strong as well as positive.
neg
One rises as the other falls. Older cars are worth less, and warmer weather costs less to heat — in both cases the cloud slopes downward.
none
There is no reason for the two to move together, and a favourite-colour code is an arbitrary label rather than a measurement. A scatter plot of these would be a patternless cloud.

Predicting the correlation before plotting is a genuine check: a plot that contradicts a strong expectation usually means an error in the data or in the plotting.

10. Worked example: read correlation from a table

Worked example

Exercise 6 asks how to do this without drawing anything.

\[ \text{Oil production } y \text{ in thousands of barrels, } x \text{ years after } 1994: \; 6660, 6560, 6470, 6450, 6250, 5880, 5820, 5800, 5750. \]

Check the inputs are in increasing order

Why: They are: zero through eight. If they were not, sorting them first is the only way the next step means anything.

\[ x\text{ rises } 0\text{ to } 8 \]

Read down the output column

Why: The values fall throughout, from 6660 to 5750.

Name the correlation

Why: One rising while the other falls is negative correlation.

Judge the strength

Why: The fall is steady, with no reversals at all, so the points will lie close to a falling line.

Figure (svg): The solution to Worked example read correlation from a table shown as a ladder of expressions, one row per algebraic move

The whole solution at once: each drop is one legal move.

\[ \text{negative correlation, and strong} \]

Verify: check for any reversal in the sequence

Why: Reading the nine values in order, each is smaller than the one before it — there is no point where the sequence turns back up. A perfectly monotonic column is stronger evidence than a merely downward-trending one, which is why the verdict is strong rather than just negative.

McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 117-117

11. Trap: reading correlation as cause

Trap

The trap

\[ \text{Subscribers up, corded phone sales down: negative correlation.} \]

Conclude that buying a mobile phone causes corded phone sales to fall

Why: A relationship in the data is read as a mechanism in the world.

The data cannot distinguish that from a third factor — falling landline prices, changing housing patterns — moving both.

The fix

\[ \text{Subscribers up, corded phone sales down: negative correlation.} \]

State the correlation, and separate it from any claim about cause

Why: Correlation says the two quantities moved oppositely. It does not say which, if either, moved the other.

Here a causal story is plausible, and it may even be right — but the plausibility comes from knowing about phones, not from the scatter plot. The plot alone supports only the correlation.

12. Does one point decide it?

Prediction

Commit before you reason.

Predict first

A data set has a clear positive correlation. One new point is added that sits well below the trend. What happens to the correlation?

  • It becomes negative
  • It stays positive but weakens
  • It becomes exactly zero
  • It is unaffected

Correct: It stays positive but weakens.

This is why the number of data points matters as much as their spread. A correlation computed from four points is a much weaker claim than the same number computed from forty.

Why: Correlation describes a tendency across the whole cloud, so one contrary point cannot reverse a clear trend — but it does make the cloud less tight around any line, which is what weakening means. How much depends on how far out the point is and how many points there already are: with eight points one outlier matters considerably, and with eight hundred it barely registers.

13. Explain the word tends

Explain it

A classmate says a positive correlation is broken because one point goes the wrong way.

Discussion prompt

In three sentences, explain what positive correlation actually claims and why a single contrary point does not refute it. Give them one everyday example where the tendency is obvious but exceptions are common.

Hint: Compare it with a rule that must hold for every case.

Answer:

Positive correlation says the two quantities tend to rise together across the data as a whole, not that every single pair obeys it. It is a statement about the cloud, so one point going the other way is an exception, not a counterexample.

Height and weight is the everyday case: taller people weigh more on average, and yet you can easily find a tall light person and a short heavy one without the tendency being wrong.

14. One of these claims is false

Two truths and a lie

All three are about correlation.

Eliminate the wrong options

Two of these are true. Knock those out and keep the false one.

  • A. A negative correlation means y tends to fall as x rises
  • B. Correlation can be judged from a table without drawing the plot
  • C. A strong correlation between two quantities shows that one causes the other

Survives elimination: C

Why: The survivor is the false one. Correlation measures how the two quantities moved together; it says nothing about mechanism. A third factor may drive both, the causation may run the other way, or the association may be coincidence — and no amount of strength in the correlation distinguishes those cases.

15. The correlation coefficient

Section

Section 2

16. One number, from negative one to one

Concept

The correlation coefficient r measures how well a line fits the data. Near one, the points lie close to a rising line; near negative one, close to a falling line; near zero, they lie close to no line at all.

correlation coefficient — A number r from negative one to one measuring how well a line fits a set of data pairs. Its sign gives the direction and its size gives the strength.

\[ -1 \leq r \leq 1 \]

Two facts in one number: the sign is the direction of the trend, and how close the size is to one is the strength of it.

Figure (svg): A number line from negative one to one showing what values of the correlation coefficient mean

One number carries two facts: its sign is the direction of the trend and its size is how tightly the points hug a line.

McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 114-114 — Correlation coefficients

17. The whole range of r

Picture it

What each part of the scale looks like.

Figure (svg): A number line from negative one to one showing what values of the correlation coefficient mean

One number carries two facts: its sign is the direction of the trend and its size is how tightly the points hug a line.

Note that r near zero does not mean the two quantities are unrelated — only that no straight line describes them. A perfect parabola can have an r of nearly zero.

18. Worked example: estimate r from three plots

Worked example

Example 2. Choose from negative one, negative one half, zero, one half, and one.

\[ \text{(a) a clear but weak downward cloud; (b) a patternless cloud; (c) a tight upward cloud.} \]

Judge the first plot's direction and strength

Why: It slopes downward, so r is negative; but the points are loosely scattered, so it is not near negative one.

\[ \text{between } 0\text{ and } -1 \]

Choose the closest option

Why: Negative one half is the only negative option that is not extreme. The actual value is about negative 0.46.

\[ r = -0.5 \]

Judge the second plot

Why: No direction at all, so r is near zero. The actual value is about negative 0.02.

\[ r = 0 \]

Judge the third plot

Why: A tight upward cloud, so r is near positive one. The actual value is about 0.98.

\[ r = 1 \]

Figure (svg): The solution to Worked example estimate r from three plots shown as a ladder of expressions, one row per algebraic move

The whole solution at once: each drop is one legal move.

\[ r \approx -0.5, \quad r \approx 0, \quad r \approx 1 \]

Verify: check each estimate against the stated actual value

Why: The three actual values are -0.46, -0.02 and 0.98, and each estimate is the nearest of the five options. Notice the second is not exactly zero — real data almost never is — which is why the question asks which value r is CLOSEST to rather than what it equals.

McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 114-114

19. Description to value of r

Matching

Four clouds, four coefficients.

Match the pairs

  • l1. tight cloud sloping up
  • l2. loose cloud sloping down
  • l3. no discernible pattern
  • l4. tight cloud sloping down
  • r1. r close to 1
  • r2. r close to -0.5
  • r3. r close to 0
  • r4. r close to -1

Why: Two features must be read separately: the direction, which fixes the sign, and the tightness, which fixes how close the size is to one. A loose downward cloud and a tight downward cloud share a sign and differ in magnitude, which is exactly the distinction between the second and fourth rows.

20. Worked example: estimate r for the oil data

Worked example

Guided Practice 4's data, judged before any calculator is reached for.

\[ \text{Oil production falls steadily from } 6660 \text{ to } 5750 \text{ over nine years.} \]

Read the direction

Why: Every value is lower than the one before, so the trend is downward and r is negative.

Judge the strength from the reversals

Why: There are none: the sequence never turns back up. That is a very tight trend.

Check the steadiness of the steps

Why: The drops are 100, 90, 20, 200, 370, 60, 20, 50 — uneven, so the fit is good but not perfect.

Estimate

Why: Strongly negative but not exactly negative one. The computed value is about negative 0.97.

\[ r\text{ near } -1 \]

Figure (svg): The solution to Worked example estimate r for the oil data shown as a ladder of expressions, one row per algebraic move

The whole solution at once: each drop is one legal move.

\[ r \approx -0.97 \]

Verify: reconcile the two observations

Why: No reversals argues for r very near negative one, while uneven step sizes argue for slightly less. Both are true, and the computed value of -0.97 sits exactly where those two observations predict: extremely strong, but not perfect.

McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 117-117

21. Trap: reading r near zero as no relationship

Trap

The trap

\[ r \approx 0 \text{ for a data set} \]

Conclude that x and y are unrelated

Why: The coefficient is read as measuring relationship in general rather than linear relationship in particular.

But a set of points lying exactly on a symmetric parabola can have an r of essentially zero, and nothing about that data is unrelated.

The fix

\[ r \approx 0 \text{ for a data set} \]

Conclude that no LINE fits the data well

Why: That is all r measures. A strong non-linear relationship can produce an r near zero.

\[ (-2,4), (-1,1), (0,0), (1,1), (2,4) \;\Longrightarrow\; r = 0 \]

Those five points lie exactly on y equals x squared — a perfect relationship with zero linear correlation. Always look at the plot, not only at r.

22. Direction, strength, or both?

Discrimination

Each statement about r tells you something. Say what.

Sort into buckets

Sort each fact about r by what it establishes.

Direction only
r is negative; r is positive
Strength only
the size of r is near 1; the size of r is near 0
Both
r = -0.97
dir
The sign alone says whether the cloud rises or falls, and nothing about how tightly it does so. A negative r could be -0.1 or -0.99.
str
The magnitude alone says how close the points lie to a line, and nothing about which way that line slopes. Both 0.9 and -0.9 are strong.
both
A specific numerical value carries the sign and the magnitude together, so it fixes the direction and the strength at once.

23. Estimate before computing

Estimation

Eight points that rise steadily with only slight scatter.

Predict first

Roughly what is r?

  • About 0.98
  • About 0.5
  • About 0
  • About -0.98

Correct: About 0.98.

The alternative-fuel vehicle data in Example 3 gives r of about 0.993 for precisely this reason: eight points rising with very little scatter.

Why: Rising means the sign is positive, and steadily with only slight scatter means the magnitude is near one. Half is what a visibly loose cloud gives, not a tight one, and Example 2 makes exactly that comparison: its loose cloud came out at -0.46 while its tight one came out at 0.98. Judging the sign first and the magnitude second turns this into two easy questions instead of one hard one.

24. Break a plausible claim

Counterexample

A classmate proposes a rule about r.

\[ r \approx 0 \;\Longrightarrow\; x \text{ and } y \text{ have nothing to do with each other} \]

Discussion prompt

Find five data pairs with an r of essentially zero whose two quantities are nevertheless perfectly related. Then say precisely what r does and does not measure.

Hint: Try a symmetric shape centred on the vertical axis.

Answer:

\[ (-2,4), \; (-1,1), \; (0,0), \; (1,1), \; (2,4) \]

Those points lie exactly on y equals x squared, so y is completely determined by x — and yet the cloud has no upward or downward slope at all, so r is zero.

What r measures is how well a straight line fits. It is silent about every other kind of relationship, which is why the plot must always be looked at alongside the number.

25. Fitting a line by hand

Section

Section 3

26. Four steps, and the third is the surprising one

Concept

Draw the scatter plot, sketch a line that follows the trend with about as many points above it as below, choose two points on your line, and write the equation through them. Those two points do not have to be data points.

best-fitting line — The line that lies as close as possible to all the data points. A hand-sketched approximation to it is called a line of fit.

Choosing points on your drawn line rather than two data points is what makes the fit use the whole cloud instead of just two members of it.

Figure (svg): A scatter plot of alternative-fuel vehicle data with a fitted line through two chosen points, one of which is not a data point

The two points you use to write the equation need not be data points; they only have to lie on the line you drew.

McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 115-115 — Approximating a Best-Fitting Line

27. The four steps on real data

Picture it

Example 3: alternative-fuelled vehicles in use, in thousands, x years after 1997.

Figure (svg): A scatter plot of alternative-fuel vehicle data with a fitted line through two chosen points, one of which is not a data point

The two points you use to write the equation need not be data points; they only have to lie on the line you drew.

The chosen points were (1, 300), which is not a data point at all, and (7, 548), which happens to be one. Both lie on the sketched line, which is the only requirement.

28. Worked example: approximate the best-fitting line

Worked example

Example 3, all four steps.

\[ \text{Data } (0,280), (1,295), (2,322), (3,395), (4,425), (5,471), (6,511), (7,548). \text{ Fit a line.} \]

Plot the eight points

Why: They rise steadily with mild scatter, so a line is a reasonable model.

Sketch a line through the trend

Why: Aim for about as many points above as below, rather than through any particular pair.

Choose two points ON THE LINE

Why: The point (1, 300) is on the sketched line but is not in the data; (7, 548) happens to be both.

\[ (1, 300)\text{ and } (7, 548) \]

Find the slope of the line through them

Why: Five hundred and forty-eight minus 300 is 248, over 7 minus 1, which is 6.

\[ m = \frac{248}{6} = 41.3 \]

Write the equation with point-slope form

Why: Using (1, 300): y minus 300 equals 41.3 times x minus 1, which simplifies.

\[ y = 41.3 x + 259 \]

Figure (svg): The solution to Worked example approximate the best-fitting line shown as a ladder of expressions, one row per algebraic move

The whole solution at once: each drop is one legal move.

\[ y \approx 41.3x + 259 \]

Verify: count the points above and below the line

Why: Substituting each x into the model gives 259, 300, 342, 383, 424, 465, 507, 548 against the data's 280, 295, 322, 395, 425, 471, 511, 548. Three data points sit above the line, four below, and one on it — a balanced fit, which is what step two asked for.

McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 115-115

29. Order the four steps

Ranking

Approximating a best-fitting line.

Put in order

  1. Draw a scatter plot of the data
  2. Sketch the line that follows the trend, balancing points above and below
  3. Choose two points that lie on the sketched line
  4. Write the equation of the line through those two points
  5. Check by comparing the model's values with the data

Why: The plot has to exist before a line can be sketched onto it, and the line has to exist before points can be chosen from it — that ordering is what prevents the endpoints error. Writing the equation is Lesson 2.4's two-point method, unchanged. The check is not in the textbook's four steps but is worth adding: it is the only way to notice that a line is systematically high or low.

30. Worked example: fit a falling line

Worked example

Guided Practice 4a. The oil production data, fitted by the same four steps.

\[ \text{Oil production, thousands of barrels, } x \text{ years after } 1994: \; 6660, 6560, 6470, 6450, 6250, 5880, 5820, 5800, 5750. \]

Plot and sketch

Why: The points fall steadily; a line drawn through the middle of the cloud has some above and some below.

Choose two points on the sketched line

Why: Taking (1, 6560) and (7, 5800), both of which are data points that lie close to the trend.

\[ (1, 6560)\text{ and } (7, 5800) \]

Find the slope

Why: Five thousand eight hundred minus 6560 is negative 760, over 6.

\[ m = -126.7 \]

Write the equation

Why: Using (1, 6560): y minus 6560 equals negative 126.7 times x minus 1.

\[ y = -126.7 x + 6687 \]

Figure (svg): The solution to Worked example fit a falling line shown as a ladder of expressions, one row per algebraic move

The whole solution at once: each drop is one legal move.

\[ y \approx -126.7x + 6687 \]

Verify: check the fit at both ends of the data

Why: At x equal to 0 the model gives 6687 against the measured 6660, and at x equal to 8 it gives 5673 against the measured 5750. Both are within about one percent, and the errors fall on opposite sides, which is the signature of a balanced fit rather than one that drifts.

McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 117-117

31. Find the error: two data points used instead of two line points

Error analysis

A student fits a line to the vehicle data by joining the first and last data points.

Annotate

On: \( m = \frac{548 - 280}{7 - 0} = 38.3 \;\Longrightarrow\; y = 38.3x + 280 \)

  • The arithmetic is correct, and the resulting line does pass through two genuine data points.
  • But it uses only two of the eight, ignoring the six in between. Step three asks for two points on the SKETCHED line, which is drawn using all of them.
  • The consequence shows in the residuals: this line runs below the data through the middle of the range - at x equal to 3 it gives 395, which happens to match, but at x equal to 5 it gives 472 against 471 and at x equal to 1 it gives 318 against 295.
  • The calculator's regression line is y = 40.9x + 263, and the hand fit of 41.3x + 259 is much closer to that than the endpoints line's 38.3x + 280. Joining the extremes is a different method, and a worse one.

The endpoints are two data points among many, and they carry no special authority. A line of fit is supposed to be influenced by every point.

32. Complete the slope

Fill the middle

Example 3's two chosen points.

Fill in the blanks

m = \frac7 - 1___} = \frac______ \approx 41.3

Why: The numerator used the second point's y value first, so the denominator must use the second point's x value first: 7 minus 1, which is 6. That consistency rule is Lesson 2.2's, unchanged. Note that one of these two points, (1,300), is not in the data at all — it was read off the sketched line.

33. Which points may be used?

Definition probe

Step three of the procedure.

Sort into buckets

Sort each point by whether it may be used to write the equation of the line of fit.

May be used
(1, 300), read off the sketched line; (7, 548), a data point that lies on the line; (5, 465), read off the line between two data points
May not be used
(3, 395), a data point that lies above the line; (0, 280), a data point that lies above the line
ok
Each of these lies on the sketched line, which is the only requirement. Whether it also happens to be a data point is irrelevant — the middle one is both, and that is a coincidence rather than a qualification.
no
These are data points that do NOT lie on the sketched line. Using one would force the line to pass through it, which changes the line you carefully drew and throws away the balance you achieved in step two.

34. Why not just use two data points?

Explain it to yourself

It would certainly be quicker.

\[ y = 38.3x + 280 \quad \text{versus} \quad y = 41.3x + 259 \]

Discussion prompt

Explain what is lost by fitting a line through the first and last data points instead of sketching a line through the whole cloud. When would the two methods agree, and what does that tell you about when the shortcut is safe?

Hint: Think about what happens if one of the two endpoints is unusual.

Answer:

The endpoints method gives all the influence to two points and none to the rest, so an unusual first or last measurement drags the whole line with it. The six middle points, which carry most of the evidence, are simply ignored.

The two methods agree when the endpoints happen to sit on the overall trend. That is exactly the case where the shortcut is safe — and you can only know it by drawing the plot, which is the work the shortcut was trying to avoid.

35. Predicting from the line

Section

Section 4

36. Substitute, and know which side of the data you are on

Concept

Once you have a model, predicting is one substitution. What matters is whether the input lies inside the range of the data, which is interpolation, or outside it, which is extrapolation and carries much more risk.

\[ y = 41.3(13) + 259 \approx 796 \]

The arithmetic is identical either way. Only the confidence differs, and it differs a great deal.

Figure (svg): The fitted line extended beyond the data to a prediction thirteen years after 1997

Predicting inside the data is interpolation and is usually safe; predicting beyond it is extrapolation and is not.

McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 116-116 — Use a line of fit to make a prediction

37. Inside the data, and beyond it

Picture it

Example 4: predicting 2010 from data covering 1997 to 2004.

Figure (svg): The fitted line extended beyond the data to a prediction thirteen years after 1997

Predicting inside the data is interpolation and is usually safe; predicting beyond it is extrapolation and is not.

Thirteen years after 1997 is six years past the last measurement — nearly doubling the span the model was built on. The model does not know that, and will answer just as confidently either way.

38. Worked example: predict 2010

Worked example

Example 4, using the line from Example 3.

\[ \text{Using } y = 41.3x + 259, \text{ predict the number of vehicles in } 2010. \]

Convert the year into the model's input

Why: The model counts years after 1997, and 2010 is thirteen years after.

\[ x = 13 \]

Substitute

Why: Forty-one point three times thirteen is 536.9.

\[ y = 536.9 + 259 \]

Add and round sensibly

Why: The data was given to three figures, so three figures is honest.

\[ y = 796 \]

State the answer with its unit

Why: The values were in thousands, so this is about 796 thousand vehicles.

\[ \text{about } 796, 000 \]

Figure (svg): The solution to Worked example predict 2010 shown as a ladder of expressions, one row per algebraic move

The whole solution at once: each drop is one legal move.

\[ y = 41.3(13) + 259 \approx 796 \text{ thousand vehicles} \]

Verify: check that the size is consistent with the trend

Why: The data ended at 548 thousand in 2004, and the model adds about 41 thousand a year for six more years, which is about 248 thousand. Five hundred and forty-eight plus 248 is 796 — the same answer reached by extending the trend rather than by substituting, which confirms the arithmetic.

McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 116-116

39. Interpolation or extrapolation?

Sorting

The vehicle data covers x from 0 to 7.

Sort into buckets

Sort each prediction by whether it lies inside or outside the data.

Interpolation
x = 4, the year 2001; x = 2.5, mid-1999; x = 7, the year 2004
Extrapolation
x = 13, the year 2010; x = -3, the year 1994
in
The input lies between the smallest and largest values the data covers, so the model is describing a region it was actually fitted on. The endpoint at x equal to 7 counts as inside, since it is a measured value.
out
The input lies beyond the data in one direction or the other. Going backwards to 1994 is just as much an extrapolation as going forwards to 2010, and it is easy to forget that.

Extrapolation is not forbidden — it is often the whole point of building a model — but it should be labelled, and its risk grows with distance from the data.

40. Worked example: predict inside the data

Worked example

The same model, used for a year that lies within the measured range.

\[ \text{Using } y = 41.3x + 259, \text{ estimate the figure for } 2001, \text{ and compare with the data.} \]

Convert the year

Why: Two thousand and one is four years after 1997.

\[ x = 4 \]

Substitute

Why: Forty-one point three times four is 165.2.

\[ y = 165.2 + 259 \]

Compute

Why: Four hundred and twenty-four thousand.

\[ y = 424 \]

Compare with the measured value

Why: The data records 425 thousand for that year, so the model is one thousand low.

\[ \text{data says } 425 \]

Figure (svg): The solution to Worked example predict inside the data shown as a ladder of expressions, one row per algebraic move

The whole solution at once: each drop is one legal move.

\[ y(4) = 424 \quad \text{against the measured } 425 \]

Verify: quantify the error and compare it with the extrapolation

Why: The model is off by 1 in 425, about a quarter of one percent, which is excellent — and it should be, because this is interpolation inside the data the model was fitted to. No such check is available for the 2010 prediction, which is precisely what makes extrapolation riskier.

McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 116-116

41. Find the error: the calendar year substituted directly

Error analysis

A student predicts 2010 from the vehicle model.

Annotate

On: \( y = 41.3(2010) + 259 = 83\,272 \text{ thousand vehicles} \)

  • The arithmetic is right: 41.3 times 2010 really is about 83,013, and adding 259 gives 83,272.
  • But the model's input is defined as years AFTER 1997, not the calendar year itself. Substituting 2010 asks the model about the year 3-thousand-and-something.
  • The size gives it away instantly: 83 million alternative-fuelled vehicles when the most recent measurement was 548 thousand - a factor of more than a hundred and fifty in six years.
  • Corrected: x = 2010 - 1997 = 13, giving about 796 thousand. Reading the variable definition before substituting is what prevents this, and it is why Lesson 2.4 insisted on writing that definition down as step one.

Every model comes with a definition of its input. An answer that is absurdly large or small almost always means that definition was skipped.

42. Push the prediction far out

Edge cases

The vehicle model, y equals 41.3x plus 259.

Discussion prompt

What does the model predict for the year 2100? Is that a believable number of alternative-fuelled vehicles? Say what feature of a linear model makes long-range predictions unreliable for quantities like this.

Hint: First convert 2100 into the model's input.

Answer:

\[ x = 103 \;\Longrightarrow\; y = 41.3(103) + 259 \approx 4514 \text{ thousand} \]

About four and a half million — which is not absurd on its own, but the model assumes the SAME 41 thousand vehicles are added every single year for a century, regardless of population, technology or saturation.

A linear model has a constant rate built into it, so it can never level off or accelerate. Real adoption curves do both, which is why Chapter 7's exponential and logistic shapes exist.

43. Predict, and say how much to trust it

Real world

The oil production model, y equals negative 126.7x plus 6687, with x years after 1994 and y in thousands of barrels per day.

Discussion prompt

Predict production in 2009, say whether it is interpolation or extrapolation, and give one specific reason the prediction might be badly wrong.

Hint: The data ran from 1994 to 2002.

Answer:

\[ x = 15 \;\Longrightarrow\; y = -126.7(15) + 6687 \approx 4787 \text{ thousand barrels per day} \]

This is extrapolation: the data covers x from 0 to 8, and 15 is nearly twice as far out as the data reaches.

It could be badly wrong because oil production responds to price and to new extraction technology, neither of which moves steadily. The shale boom that began around 2008 reversed the decline entirely — a linear model fitted to the 1990s could not possibly have anticipated it, and would have kept predicting a fall.

44. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

A model fits its data with r equal to 0.99. How confident should you be in a prediction ten years beyond the data?

  • Very confident — a high r means accurate predictions
  • Cautious — r measures fit to the data you have, not beyond it
  • Not confident at all — a high r is meaningless
  • It depends only on how many data points there were

Correct: Cautious — r measures how well the line fits the data you have, and says nothing about what happens outside it.

The honest way to report a prediction is with both facts: how well the model fits what was measured, and how far outside that range the prediction sits.

Why: A correlation coefficient is computed entirely from the observed pairs. It cannot know whether the trend continues, because no evidence about the future is in the calculation. A model can fit a decade of data essentially perfectly and still be wrong about the next year, if the underlying situation changes. The number of data points matters too, but it does not repair the basic problem with extrapolation.

45. Regression on a calculator

Section

Section 5

46. The same idea, computed rather than sketched

Concept

A graphing calculator's linear regression feature finds the best-fitting line using every data point, and reports the correlation coefficient alongside it. The result is reproducible: two people get the same line from the same data.

\[ y = 40.9x + 263, \quad r \approx 0.993 \]

Enter the inputs in one list and the outputs in another, then run the linear regression command. If r does not appear, turn diagnostics on.

Figure (svg): Two columns comparing the hand-sketched line of fit with the calculator's regression line

Two methods, the same data, predictions one part in eight hundred apart.

McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 116-116 — Use a graphing calculator to find a best-fitting line

47. Hand fit against regression

Picture it

The same eight data points, fitted two ways.

Figure (svg): Two columns comparing the hand-sketched line of fit with the calculator's regression line

Two methods, the same data, predictions one part in eight hundred apart.

The two lines differ by about one percent in slope and give predictions one part in eight hundred apart. A careful hand fit is not a poor substitute — it is a good approximation to the same thing.

48. Worked example: run the regression

Worked example

Example 5. The same vehicle data, entered and computed.

\[ \text{Find the regression line for the eight vehicle data pairs.} \]

Enter the data into two lists

Why: Years since 1997 in the first list, vehicle counts in the second, in matching order.

\[ L 1: 0..7, L 2: 280..548 \]

Run the linear regression command

Why: The calculator reports a slope of about 40.869 and an intercept of about 262.83.

\[ a = 40.87, b = 262.83 \]

Round sensibly

Why: The data has three significant figures, so the model should not claim more.

\[ y = 40.9 x + 263 \]

Read the correlation coefficient

Why: The calculator reports r of about 0.993, a very strong positive correlation.

\[ r = 0.993 \]

Figure (svg): The solution to Worked example run the regression shown as a ladder of expressions, one row per algebraic move

The whole solution at once: each drop is one legal move.

\[ y = 40.9x + 263, \quad r \approx 0.993 \]

Verify: plot the line with the scatter plot

Why: Graphing the regression equation over the scatter plot shows it running through the middle of the cloud with points on both sides throughout — the same balance the hand sketch aimed for. An r of 0.993 predicts exactly that appearance, and seeing it confirms no data was mistyped.

McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 116-116

49. Hand fit against regression

Comparison

Fill the blanks. Both methods answer the same question differently.

Comparison matrix

FeatureSketched by handLinear regression
Which points influence itall of them, through your eyeall of them, by computation
Reproducible by someone elseno - it depends on your sketchyes, exactly
Reports a correlation coefficientnoyes
Result on the vehicle datay = 41.3x + 259y = 40.9x + 263
Needs equipmentnoyes

The hand method's weakness is reproducibility, not accuracy. On well-behaved data a careful sketch lands within a percent or two of the computed line.

50. Worked example: compare the two predictions

Worked example

Guided Practice 4c in structure: repeat a prediction with the regression line.

\[ \text{Predict } 2010 \text{ with } y = 40.9x + 263, \text{ and compare with the hand fit's } 796. \]

Substitute thirteen into the regression model

Why: Forty point nine times thirteen is 531.7.

\[ y = 531.7 + 263 \]

Compute

Why: Seven hundred and ninety-five thousand approximately.

\[ y = 795 \]

Compare with the hand fit

Why: The hand fit gave 796; the difference is one thousand out of nearly eight hundred thousand.

\[ 796\text{ against } 795 \]

Judge the agreement

Why: Roughly one part in eight hundred, far smaller than the uncertainty in extrapolating six years past the data.

Figure (svg): The solution to Worked example compare the two predictions shown as a ladder of expressions, one row per algebraic move

The whole solution at once: each drop is one legal move.

\[ y(13) = 40.9(13) + 263 \approx 795 \]

Verify: ask which source of error dominates

Why: The two methods differ by about 0.1 percent, while the risk in extrapolating six years beyond the data is far larger than that. So the choice of fitting method is not what limits this prediction — the extrapolation is. Knowing which error dominates tells you where extra care is worth spending.

McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 116-116

51. Find the error: lists entered out of order

Error analysis

A student enters the vehicle data with the second list sorted independently and gets a suspicious result.

Annotate

On: \( \text{L1: } 0,1,2,3,4,5,6,7 \qquad \text{L2: } 548,511,471,425,395,322,295,280 \)

  • Both lists contain the correct eight numbers, so nothing was mistyped and nothing is missing.
  • But the second list has been reversed, so each year is now paired with the wrong count: year 0 with 548 instead of 280.
  • The regression reports a slope of about -40.9 and an r of about -0.993 - the same magnitudes as before with both signs flipped.
  • The tell is the sign: the data plainly rises over time, so a negative slope is impossible. Corrected, the pairs must stay in their original order, giving y = 40.9x + 263 with r = 0.993.

Always check the sign of the reported slope against the direction you can see in the data. A sign flip is the usual signature of mismatched lists.

52. Order the calculator steps

Ranking

Finding and checking a regression line.

Put in order

  1. Enter the inputs in one list and the outputs in another, in matching order
  2. Run the linear regression command
  3. Round the reported slope and intercept to match the data's precision
  4. Make a scatter plot of the data
  5. Graph the regression equation over the scatter plot to see how well it fits

Why: Entering the data has to come first, and the regression cannot run without it. Rounding comes before graphing so that the equation you plot is the one you will actually quote. The last two steps are the check, and they are the ones most often skipped — a mistyped value can shift the line without producing any error message, and only the plot reveals it.

53. What will r be?

Prediction

Commit before reasoning.

Predict first

The oil production data falls steadily with no reversals across nine years. What will the regression report for r?

  • About -0.97
  • About 0.97
  • About -0.5
  • About 0

Correct: About -0.97.

\[ y \approx -129.8x + 6702, \quad r \approx -0.97 \]

Why: Falling makes the sign negative, and no reversals across nine points makes the magnitude close to one. Half would correspond to a visibly loose cloud, and zero to no trend at all — neither matches a sequence where every value is lower than the last. Predicting r before running the regression is a real check: a reported value with the wrong sign means the lists were mismatched.

54. What does best-fitting actually mean?

Socratic

One question, and nothing else on this slide.

\[ y = 40.9x + 263 \]

Discussion prompt

The calculator calls this the best-fitting line. Best according to what standard? Consider two candidate lines that each miss the data by the same total amount, one missing every point slightly and one matching six points exactly while missing two badly. Which should count as better, and why might reasonable people disagree?

Hint: Think about what you want the line for.

Answer:

Regression minimises the sum of the SQUARED vertical distances, which punishes large misses much more heavily than small ones. So it prefers the line that misses everything slightly over the one with two bad misses.

Squaring is a choice, not a law. Minimising the plain distances instead gives a different line, one less disturbed by outliers, and for data with a few bad measurements that is arguably the better answer.

Reasonable people disagree because the right standard depends on whether the outliers are errors, which you want ignored, or real, which you want accounted for. The calculator cannot know which, so it applies one convention and leaves the judgement to you.

55. Fitting a line to data, at a glance

Comparison

Fill the blanks. Each stage answers a different question.

Comparison matrix

StageQuestion it answersTool
Scatter plotis there a trend, and which way?your eyes
Correlation coefficienthow tightly do the points hug a line?the number r
Line of fitwhat line describes the trend?sketch, then two points on it
Predictionwhat value at a new input?substitute into the model
Regressionwhat is the best line, exactly?a graphing calculator

The first two stages decide whether the last three are worth doing at all. Fitting a line to a patternless cloud produces an equation and no information.

56. The procedure, in order

Pattern

One routine takes data to a usable model.

  1. Plot the pairs and look. Name the correlation as positive, negative or approximately none, and stop here if there is no trend.
  2. Estimate r from the picture: its sign from the direction, its size from how tightly the points cluster.
  3. Sketch the line that follows the trend with about as many points above as below, then choose TWO points on that sketched line — not two data points.
  4. Write the equation through those two points using the slope formula and point-slope form, exactly as in Lesson 2.4.
  5. Predict by substituting, stating whether the input is inside the data or beyond it, and check the model against the data you already have.

Step three is where this lesson differs from every earlier one. Everywhere else two points determined the line exactly; here they are read off a line that was drawn from all the data.

OpenStax Algebra and Trigonometry 2e, §4.3 Fitting Linear Models to Data §4.3

57. Check yourself 1 of 3

Check

Estimating r from a description.

Check your understanding

A scatter plot shows a clear but fairly loose downward cloud. Which value is r closest to?

  • A. -0.5 (correct)
  • B. -1
  • C. 0
  • D. 0.5

Answer: A

Why: Downward makes r negative, and loose means the magnitude is well short of one. Example 2's matching plot had an actual value of about -0.46, which rounds to the -0.5 option.

Why B tempts people
Negative one would mean the points lie essentially on a straight falling line, with almost no scatter. The word loose rules that out.
Why C tempts people
Zero would mean no discernible direction at all, but the cloud was described as clearly downward.
Why D tempts people
The sign is wrong. A downward cloud gives a negative coefficient; only a rising one gives a positive.

58. Check yourself 2 of 3

Check

Fitting by hand. Which two points may be used?

Check your understanding

In step 3 of approximating a best-fitting line, which two points should you choose?

  • A. Any two points that lie on the line you sketched (correct)
  • B. The first and last data points
  • C. The two data points closest to the middle
  • D. Any two data points, as long as they are far apart

Answer: A

Why: The points must lie on the sketched line, which was drawn using every data point. They need not be data points at all — Example 3 uses (1, 300), which is not in the data.

Why B tempts people
This gives all the influence to two points and ignores the rest. On the vehicle data it produces a slope of 38.3 against the regression's 40.9.
Why C tempts people
The middle points are no more authoritative than any others, and forcing the line through them abandons the balance achieved by sketching.
Why D tempts people
Being far apart makes the slope less sensitive to reading error, which is genuinely useful — but only among points ON the line. Two data points off the line change the line.

59. Check yourself 3 of 3

Check

Predicting. Convert the input first.

Check your understanding

Using y = 41.3x + 259, where x is years after 1997 and y is thousands of vehicles, predict the figure for 2010.

  • A. About 796 thousand (correct)
  • B. About 83,272 thousand
  • C. About 672 thousand
  • D. About 537 thousand

Answer: A

Why: 2010 is 13 years after 1997, so x is 13. Then 41.3 times 13 is 536.9, and adding 259 gives about 796 thousand.

Why B tempts people
The calendar year 2010 was substituted instead of 13. The size gives it away: 83 million vehicles six years after a measurement of 548 thousand.
Why C tempts people
The input was taken as 10 rather than 13, as if the model counted years after 2000. The definition says years after 1997.
Why D tempts people
The intercept was omitted, leaving only 41.3 times 13. The constant term is part of the model and cannot be dropped.

60. Where this shows up outside the textbook

Real world

A study reports a strong positive correlation between the number of fire engines sent to a fire and the amount of damage the fire causes.

Discussion prompt

Describe the correlation, say what a naive causal reading would conclude, and explain what is really going on. Then say what extra variable would need to be measured to make sense of the data.

Hint: Ask what decides how many engines are sent.

Answer:

The correlation is positive and probably strong: bigger numbers of engines go with bigger damage figures. A naive reading concludes that sending engines causes damage, and therefore that fewer should be sent.

What is really happening is that a third variable — the size of the fire — drives both. Big fires get more engines and cause more damage; the engines are a response to the size, not a cause of the damage.

Measuring fire size, or damage per unit of fire size, would separate the two. This is the classic illustration of why correlation does not establish causation, and it is exactly the trap from Section 1 in a case where the naive conclusion is actively harmful.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

Two data sets both have r equal to 0.9. Does that mean their scatter plots look equally tight?

  • Yes — r measures tightness, so equal r means equal appearance
  • Roughly, but the slopes may differ completely
  • No — r says nothing about the plots
  • Only if both have the same number of points

Correct: Roughly, but the slopes may differ completely.

This is why a model needs both numbers reported: the slope says how much y changes per unit of x, and r says how much to trust that the relationship is linear at all.

Why: The coefficient r measures how closely the points cluster around their line, not how steep that line is. Two data sets can both have r equal to 0.9 with one rising gently and the other almost vertically. The clouds look similarly tight around their own lines, which is what r captures, but the pictures can look very different overall. Number of points affects how much confidence r deserves rather than what it measures.

62. Explain it to someone a year behind you

Explain it

They have only ever seen graphs where the points lie exactly on a line.

Discussion prompt

In four sentences or fewer, explain why real data does not, what a line of fit is for, and the one thing they should check before trusting a prediction made from one.

Hint: The check is about where the prediction sits relative to the data.

Answer:

Real measurements carry noise, so points scatter around a trend rather than sitting on it. A line of fit is the line that follows that trend as closely as possible, and it lets you estimate values you did not measure.

Before trusting a prediction, check whether the input lies inside the range of the data. Inside is interpolation and is usually safe; outside is extrapolation, and the further out you go the less the data has to say about it.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Estimating r from a scatter plot
  • Choosing the two points for a hand-fitted line
  • Converting a year into the model's input before predicting
  • Saying what a correlation does and does not prove

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For estimating r, read the sign from the direction and the size from the tightness, as two separate questions. For choosing points, take them off the line you drew, not from the data. For predicting, reread the variable definition before substituting anything. For correlation and cause, always ask what third variable could be driving both. Do five of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Fifteen minutes.

Draw it

Across the top of a page draw three small scatter plots showing positive correlation, negative correlation and none, and write beneath each the value of r you would estimate. Underneath, draw one larger scatter plot of at least eight points of your own invention that show a clear trend with visible scatter. Sketch a line of fit through it, balancing points above and below, then mark two points ON YOUR LINE — making at least one of them not a data point — and work out the equation through them, showing the slope calculation and the point-slope step. Beside the plot, draw a vertical dashed line where your data ends and label everything to the right of it extrapolation. At the bottom, use your equation to predict one value inside the data and one well outside it, and write one sentence about how much you trust each.

If both predictions got the same level of trust in your last sentence, reread Section 4. The arithmetic is identical for both; the confidence is not.

65. What you can do now

Recap

Five things, and the last two are about knowing what a model cannot tell you.

If you seeThen
A cloud sloping upwardPositive correlation, r positive
A cloud with no directionr near zero: no LINE fits
A sketched line of fitTake your two points off it
A prediction beyond the dataLabel it extrapolation
A strong correlationAsk what third variable could cause both

Lesson 2.7 leaves data behind and returns to exact functions, taking the absolute value function from Chapter 1 and moving it around the plane.

McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines §2.6, pp. 113-119 — everything on these slides traces back here

Sources

  1. McDougal Littell Algebra 2 (Texas Edition), Ch. 2 Linear Equations and Functions — Lesson 2.6 Draw Scatter Plots and Best-Fitting Lines — Larson, Boswell, Kanold & Stiff, McDougal Littell / Houghton Mifflin, 2007, pp. 113-119
  2. OpenStax Algebra and Trigonometry 2e, §4.3 Fitting Linear Models to Data

Want this taught 1-on-1? Alexander tutors Algebra 2 — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108