A short section that finally uses the line for what it was fitted for. Everything needed was established earlier — section 12.2 confirmed a linear pattern, section 12.3 fitted the least-squares line, section 12.4 established that its correlation is significant — so predicting is a substitution. The section's real content is the boundary around where that substitution may be trusted. Predicting for an x inside the range of observed x values is interpolation and is licensed; predicting outside that range is extrapolation and is not, however significant the correlation. The book demonstrates rather than asserts this: substituting a third-exam score of 90, well above the observed maximum of 75, gives a predicted final exam score of 261.19 when the largest score the final exam can carry is 200. The prediction is also a prediction of a mean, not of what any individual will score.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 12 — Linear Regression and Correlation
Prediction
Objectives
Five outcomes, and three of them are checks made before any substitution.
OpenStax Introductory Statistics 2e, §12.5 Prediction §12.5, pp. 635-636 — the section these objectives are drawn from
Warm-up
The line is fitted and its correlation is significant. Third-exam scores in the sample ran from 65 to 75.
Discussion prompt
Two students are asked about: one scored 73 on the third exam, the other 90. Can the line predict both final scores?
Hint: Compare each score against the range the data actually covered.
Answer:
The first, yes. Seventy-three lies between 65 and 75, so the sample contains students on both sides of it and the line has evidence behind it there. Substituting gives about 179.
The second, no — and not because the arithmetic fails. Substituting 90 gives 261.19, which is a perfectly ordinary calculation and a worthless answer, since the final exam is scored out of 200.
\[ \hat{y}(90) = -173.51 + 4.83(90) = 261.19 > 200 \]
Nothing in the data concerns students scoring 90 on the third exam, because none did. The line continues past 75 as a matter of algebra; the evidence for it stops there. That distinction is the whole of this section.
Concept
We examined the scatterplot and showed that the correlation coefficient is significant, and we found the equation of the best-fit line. We can now use the least-squares regression line for prediction — substituting an x value between the smallest and largest observed and reading off the predicted mean y.
interpolation and extrapolation — The process of predicting inside the observed x values is called interpolation; predicting outside them is called extrapolation. Only the first is licensed by the data.
\[ x_{\min} \le x \le x_{\max} \;\Longrightarrow\; \text{predict } \hat{y} = a + bx \]
The conditions are cumulative. Section 12.2 asked whether the pattern is linear, section 12.4 asked whether the correlation is significant, and this section adds whether the x lies in range. Failing any one of the three blocks the prediction, and only the first two involve any statistics — the third is a comparison of three numbers.
Figure (svg): A scatter plot with the fitted line drawn solid across the observed range and dashed beyond it, with a horizontal line marking the maximum possible score of two hundred
OpenStax Introductory Statistics 2e, §12.5 Prediction §12.5, pp. 635-636
Section
Section 1
Concept
Suppose you want to estimate the mean final exam score of statistics students who received 73 on the third exam. The exam scores range from 65 to 75, and since 73 is between them, substitute x equal to 73 into the equation.
what is predicted — The MEAN value of y for that x, not the value any individual will take. The book's wording is that such students will earn that grade on the final exam, on average.
\[ \hat{y} = -173.51 + 4.83(73) = 179.08 \]
The two words on average carry the same weight here as they did for the slope in section 12.3, and for the same reason. The residuals in this data run from about -19 to +35, so an individual student scoring 73 could plausibly finish anywhere from the low 160s to the low 200s. The line locates the centre of that spread, not its extent.
Figure (svg): A card explaining what a predicted value is a prediction of
OpenStax Introductory Statistics 2e, §12.5 Prediction §12.5, p. 635 — the prediction at 73, and its wording
Picture it
What the number 179.08 is a statement about.
Figure (svg): A card explaining what a predicted value is a prediction of
Reporting a prediction without the qualifier invites the reader to expect that score from a particular student, which the data do not support. The residual spread is the honest measure of how far an individual may land from it.
Worked example
The fitted line is ŷ = -173.51 + 4.83x, and third exam scores ran from 65 to 75.
\[ x = 73 \]
Check the range
Why: Between 65 and 75.
Multiply
Why: 4.83 times 73.
\[ 352.59 \]
Add the intercept
Why: Subtract 173.51.
\[ 179.08 \]
State it in context
Why: With the qualifier.
Figure (svg): The solution to Worked example the book's prediction at 73 shown as a ladder of expressions, one row per legal move
\[ \hat{y}(73) = 179.08 \]
Verify: confirm the value against the unrounded line
Why: The calculator's own coefficients are -173.513 and 4.8273, and those give 178.89 rather than 179.08. Both figures come from the book's own pages: it quotes the rounded line in the prediction and the fuller one in the LinRegTTest output. A fifth of a point changes nothing here, but the pattern is worth noticing — a rounded slope multiplied by a number near 73 accumulates a visible difference.
OpenStax Introductory Statistics 2e, §12.5 Prediction §12.5, p. 635
Faded example
Using ŷ = -173.51 + 4.83x at a third exam score of 70.
Fill in the blanks
\hat338.1 = -173.51 + 4.83(70) = -173.51 + 164.59 = ___
Why: One student did score 70, finishing with 163 — about 1.6 points below the prediction, one of the smallest residuals in the data.
Worked example
What would you predict the final exam score to be for a student who scored 66 on the third exam?
\[ x = 66 \]
Check the range
Why: 65 to 75.
Multiply
Why: 4.83 times 66.
\[ 318.78 \]
Add the intercept
Why: Subtract 173.51.
\[ 145.27 \]
Compare with the data
Why: Two students scored near 66.
\[ 126\text{ and } 175 \]
Figure (svg): The solution to Worked example Example 12.11a, at 66 shown as a ladder of expressions, one row per legal move
\[ \hat{y}(66) = 145.27 \]
Verify: confirm the prediction sits sensibly among nearby observations
Why: The two students who scored 65 and 66 on the third exam finished with 175 and 126 — one well above the prediction and one well below, averaging 150.5 against the predicted 145.3. That is exactly what a prediction of a mean should look like: near the centre of nearby outcomes and matching none of them. It also shows how wide the individual spread is at a single x.
OpenStax Introductory Statistics 2e, §12.5 Prediction §12.5, pp. 635-636
Trap
\[ \hat{y}(73) = 179 \;\Rightarrow\; \text{this student will score } 179 \]
Take the predicted value as a forecast for one person
Why: It was computed from that person's x.
\[ \text{but the residuals run from } -19 \text{ to } +35 \]
Individual students at a given third-exam score finish across a wide band, and the line marks its centre.
\[ \text{students scoring } 73 \text{ average about } 179 \text{ on the final} \]
Report the prediction as a mean, with the qualifier
Why: That is the book's own wording.
The two students who scored 65 and 66 make the point concretely: they finished on 175 and 126, forty-nine points apart at nearly the same third-exam score. Any prediction at that x is a statement about their average, and about neither of them individually.
Two truths and a lie
All three concern what is predicted.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. Two students in this sample scored 65 and 66 on the third exam and finished 49 points apart, so no single prediction could be right for both. The line gives the centre of the outcomes at each x.
Prediction
Commit before reasoning.
Predict first
The standard deviation of the residuals is 16.4. Roughly how far from the prediction might an individual student fall?
Correct: Commonly 15 to 20, occasionally over 30.
Why: The residual standard deviation measures the typical vertical miss, so most individuals fall within about one of them and a few fall two or more away — which is exactly what section 12.6's outlier rule of two standard deviations formalises. The observed residuals here run from -19.1 to +34.7.
Estimation
Using ŷ = -173.51 + 4.83x at the observed maximum of 75.
Predict first
Roughly what final score does the line predict?
Correct: About 189.
Why: Four point eight three times 75 is 362.25, less 173.51 gives 188.74. The student who actually scored 75 finished with 198, about nine points above the line — the largest third-exam score in the sample and a positive residual.
Section
Section 2
Concept
The process of predicting inside of the observed x values is called interpolation. The process of predicting outside of the observed x values is called extrapolation. You should not use the line to predict for x values outside the domain of the sample data.
the domain — The interval from the smallest to the largest observed x value. For the exam data that is 65 to 75, so 73 is interpolation and 90 is extrapolation.
\[ [x_{\min}, x_{\max}] = [65, 75] \]
The reason is that a straight line is a model fitted to a region, not a law. Within the observed range there are data on both sides of any x, so the line's behaviour there is constrained by evidence. Outside it, the line's value rests entirely on the assumption that the same straight relationship continues — an assumption the data cannot support because they never reached there.
Figure (svg): A card contrasting interpolation with extrapolation
OpenStax Introductory Statistics 2e, §12.5 Prediction §12.5, p. 636 — the note defining interpolation and extrapolation
Picture it
One licensed, one not.
Figure (svg): A card contrasting interpolation with extrapolation
Section 12.4's third note said the same thing before this section named it: even when r is significant and the plot is linear, the line may not be appropriate or reliable outside the domain of observed x values.
Worked example
What would you predict the final exam score to be for a student who scored 90 on the third exam?
\[ x = 90 \]
Check the range
Why: Data run 65 to 75.
\[ 90\text{ is outside} \]
Give the verdict
Why: Extrapolation.
Substitute anyway
Why: 4.83 times 90.
\[ 434.7 \]
Add the intercept
Why: Subtract 173.51.
\[ 261.19 \]
Compare with the maximum
Why: Final is out of 200.
Figure (svg): The solution to Worked example Example 12.11b, the demonstration at 90 shown as a ladder of expressions, one row per legal move
\[ 261.19 > 200 \]
Verify: confirm what makes this demonstration convincing
Why: It does not appeal to a rule; it shows the model producing an impossible answer. The final exam is scored out of 200, so any predicted value above that is self-evidently wrong — and the line reaches it at a third exam score of about 77, barely beyond the observed maximum of 75. Extrapolation fails quickly here, not eventually.
OpenStax Introductory Statistics 2e, §12.5 Prediction §12.5, p. 636
Sorting
The observed third-exam scores run from 65 to 75.
Sort into buckets
Sort each prediction point.
Item (d) is the book's own second warning: you should not use the line to predict the final exam score for a student who earned 50 on the third exam, because 50 is not within the domain of the sample data.
Worked example
Finding where the fitted line first predicts an impossible score.
\[ -173.51 + 4.83x = 200 \]
Add the intercept back
Why: 200 plus 173.51.
\[ 373.51 \]
Divide by the slope
Why: 373.51 over 4.83.
\[ 77.3 \]
Compare with the data
Why: Maximum observed 75.
Read the lesson
Why: Little margin.
Figure (svg): The solution to Worked example how far the line stays possible shown as a ladder of expressions, one row per legal move
\[ x \approx 77.3 \]
Verify: confirm this explains why the book chose 90 rather than 78
Why: At 78 the prediction would be 203.23 — impossible, but only just, and easy to dismiss as rounding. At 90 it is 261.19, which no reader can mistake for a near miss. The point is that the failure begins almost immediately outside the data; the book simply chose an x where it is unmistakable.
OpenStax Introductory Statistics 2e, §12.5 Prediction §12.5, p. 636
Error analysis
The exam data covers x from 65 to 75. Which arguments hold?
Annotate
On: \( \begin{aligned} &(1)\; \text{the correlation is significant, so the line holds everywhere} \\ &(2)\; 90 \text{ is only 15 past the data, so it is nearly inside} \\ &(3)\; \text{the arithmetic works, so the answer is valid} \\ &(4)\; \text{nothing in the data concerns } x = 90 \end{aligned} \)
Argument (3) is the one worth naming, because a calculator never refuses. Every gate in this section has to be applied by the person, since the arithmetic will produce a number under any conditions at all.
Faded example
For the exam data, whose third-exam scores are 65, 66, 67, 67, 69, 69, 70, 71, 71, 71 and 75.
Fill in the blanks
\text65 [75, ___]
Why: Only these endpoints matter, not how the values are distributed between them. Any x in that interval is interpolation, and any x outside it is extrapolation.
Two truths and a lie
All three concern the boundary.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. A wider sample widens the DOMAIN, so more x values become interpolation — but predicting beyond whatever range was observed remains unsupported, because the data still say nothing about what happens there.
Prediction
Commit before reasoning.
Predict first
The prediction exceeds 200 by x = 77.3, only just past the data. Why so soon?
Correct: A steep slope over a narrow range.
Why: The observed x values span only ten points while the slope is 4.83 final-exam points per third-exam point, so the line climbs nearly 50 points across the data and keeps climbing at the same rate. A narrow domain and a steep slope together mean the line leaves plausible territory almost as soon as it leaves the data.
Section
Section 3
Concept
A prediction is licensed only if the scatter plot shows a linear trend, the correlation coefficient is significant, and the x value lies within the domain of the observed x values. Section 12.4 stated all three; this section applies them.
the three gates — Linearity from section 12.2, significance from section 12.4, and range from this section. All three must hold, and the arithmetic checks none of them.
\[ \text{linear} \;\wedge\; r \text{ significant} \;\wedge\; x \in [x_{\min}, x_{\max}] \]
Each gate fails in a different way and is caught by a different means. Linearity is judged from the scatter plot and the residuals plot, significance from a hypothesis test, and range from comparing three numbers. Only the second is a calculation, which is why the other two are the ones usually skipped.
Figure (svg): Making a prediction from a regression line
OpenStax Introductory Statistics 2e, §12.5 Prediction §12.5, pp. 635-636 — section 12.4's three notes, applied
Picture it
Six steps, of which one is arithmetic.
Figure (svg): Making a prediction from a regression line
The substitution comes last for a reason. A calculator will perform it under any conditions whatever, so the checks have to be made before there is a number to be persuaded by.
Worked example
For the exam data at x = 73.
\[ \text{may we predict?} \]
Linear pattern
Why: Section 12.2's plot.
r significant
Why: p = 0.026 < 0.05.
In range
Why: 65 to 75 contains 73.
Substitute
Why: All gates passed.
\[ 179.08 \]
Figure (svg): The solution to Worked example running all three gates shown as a ladder of expressions, one row per legal move
\[ \hat{y}(73) = 179.08 \]
Verify: confirm the order in which the gates were checked
Why: Linearity first, because a curved pattern makes the whole model wrong and nothing later can repair it. Significance second, because an insignificant correlation means the line describes noise. Range last, because it concerns a particular prediction rather than the model as a whole — the first two are settled once for the data, and the third is re-checked for every x.
OpenStax Introductory Statistics 2e, §12.5 Prediction §12.5, pp. 635-636
Matching
Match each condition to how it is verified.
Match the pairs
Why: Only the second is a calculation. The other three are things to look at or compare, which is why they are the ones that get skipped when a calculator is doing the work.
Worked example
Data relate hours per week practising a musical instrument to math test scores, with fitted line ŷ = 72.5 + 2.8x. Predict for a student practising five hours a week.
\[ \hat{y} = 72.5 + 2.8x, \; x = 5 \]
Multiply
Why: 2.8 times 5.
\[ 14 \]
Add the intercept
Why: Plus 72.5.
\[ 86.5 \]
Check the units
Why: A test score.
Note what is missing
Why: The domain.
Figure (svg): The solution to Worked example Try It 12.11 shown as a ladder of expressions, one row per legal move
\[ \hat{y}(5) = 72.5 + 2.8(5) = 86.5 \]
Verify: confirm what this problem does not supply
Why: It gives the fitted line without the data, so the domain of observed x values is unknown and the third gate cannot be checked. Five hours a week is plausible for such a study, but the problem as stated does not establish it — which is a reminder that a reported equation alone is not enough to predict responsibly. The range has to travel with it.
OpenStax Introductory Statistics 2e, §12.5 Prediction §12.5, p. 636
Trap
\[ \hat{y} = 72.5 + 2.8x \;\Rightarrow\; \text{predict at any } x \]
Treat a published equation as a general formula
Why: It is stated without qualification.
\[ \text{but the domain is not stated} \]
Twenty hours a week would give 128.5 on a test likely scored out of 100.
\[ \text{ask for the range of } x \text{ before substituting} \]
Treat the equation as valid only over the region it was fitted to
Why: The domain is part of the result.
This is why a reported regression should always carry the range of its data alongside the coefficients. Without it a reader cannot tell interpolation from extrapolation, and the equation alone will produce confident nonsense — as it does here at twenty hours.
Two truths and a lie
All three concern the conditions.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The test asks only whether the linear component is more than chance would produce. It says nothing about whether the pattern is genuinely straight, and nothing at all about which x values were observed.
Faded example
A study's x values run from 12 to 40, and a prediction is wanted at x = 45.
Fill in the blanks
45 > 40, \textextrapolation ___
Why: Five past the maximum is still extrapolation. The rule has no tolerance band — the question is simply whether the data reached that far.
Explain it
A classmate has a significant correlation and a linear plot, and predicts at an x well beyond the largest observed.
Discussion prompt
In two sentences or fewer, tell them what is wrong.
Hint: Ask what the data say about that region.
Answer:
Significance and linearity establish that the line describes the region the data covered; they say nothing about a region where nothing was observed.
That is extrapolation, and the book's own demonstration is that it can produce a predicted exam score of 261 on a test marked out of 200.
Section
Section 4
Concept
The book predicts using the rounded line ŷ = -173.51 + 4.83x, while its own calculator output gives a equal to -173.513 and b equal to 4.8273. The two produce slightly different predictions, and the gap widens as x grows.
carrying precision — Rounding coefficients before substituting introduces an error that the multiplication amplifies. The slope is multiplied by x, so a rounding in the third decimal place shows up in the first when x is near 70.
\[ 4.83 \text{ against } 4.8273: \quad \text{a gap of } 0.0027x \]
The size is predictable. The slope was rounded by 0.0027, and multiplying that by an x near 73 gives about 0.20 — which is exactly the discrepancy observed. Nothing here changes a conclusion, but the same habit applied to a steeper slope or a larger x would, and section 12.6 shows the book's SSE differing by 16 for the same reason.
Figure (svg): A table comparing predictions from the rounded and unrounded regression lines
OpenStax Introductory Statistics 2e, §12.5 Prediction §12.5, pp. 635-636 — the calculator output against the quoted equation
Picture it
At three values of x.
Figure (svg): A table comparing predictions from the rounded and unrounded regression lines
Every difference here is under a quarter of a point, and every one runs the same direction — the rounded slope is slightly larger, so it predicts slightly higher, and increasingly so as x grows.
Worked example
Comparing the two lines at x = 73.
\[ 4.83 \text{ against } 4.8273 \]
Difference in slope
Why: 4.83 minus 4.8273.
\[ 0.0027 \]
Difference in intercept
Why: -173.51 against -173.513.
\[ 0.003 \]
Slope effect at 73
Why: 0.0027 times 73.
\[ 0.197 \]
Total
Why: Plus the intercept gap.
\[ \text{about } 0.20 \]
Figure (svg): The solution to Worked example locating the discrepancy shown as a ladder of expressions, one row per legal move
\[ 0.0027(73) + 0.003 \approx 0.20 \]
Verify: confirm the prediction of the discrepancy against the values
Why: The rounded line gives 179.08 and the full-precision one 178.89, a difference of 0.19 — matching the 0.20 predicted from the rounding alone. Being able to account for a discrepancy exactly is what separates a rounding artefact from an arithmetic error, and it is worth doing whenever two computed values nearly agree.
OpenStax Introductory Statistics 2e, §12.5 Prediction §12.5, p. 635
Faded example
A slope of 3.7 is rounded from 3.6982, and a prediction is made at x = 50.
Fill in the blanks
\text0.0018 \approx 0.09 \times 50 = ___
Why: Under a tenth of a unit, which would be invisible in most reporting. The same rounding at x = 5000 would produce an error of 9.
Worked example
The same rounding, applied at a much larger x.
\[ x = 1000 \text{ with a slope rounded by } 0.0027 \]
The slope error
Why: As before.
\[ 0.0027 \]
Multiply by x
Why: 0.0027 times 1000.
\[ 2.7 \]
Compare with x = 73
Why: About 0.20.
Draw the rule
Why: Error grows with x.
Figure (svg): The solution to Worked example where rounding would matter shown as a ladder of expressions, one row per legal move
\[ \text{error} \approx (\Delta b) x \]
Verify: confirm this predicts the CPI example's behaviour
Why: Example 12.14 fits years against the Consumer Price Index with a slope of 1.6625 rounded to 1.662, and predicts at x = 1990. The rounding error is 0.0005 times 1990, about one unit — and indeed the book's 103.4 differs from the full-precision 103.9 by about half a unit. Where x is a year, even four-decimal rounding shows up in the answer.
OpenStax Introductory Statistics 2e, §12.5 Prediction §12.5, pp. 635-636
Error analysis
The rounded line gives 179.08 and the exact one 178.89. Which responses are sound?
Annotate
On: \( \begin{aligned} &(1)\; \text{one of the two must be an arithmetic error} \\ &(2)\; \text{the difference is explained by rounding the slope} \\ &(3)\; \text{report the exact value, noting the book's figure} \\ &(4)\; \text{the discrepancy grows with } x \end{aligned} \)
Response (1) is the instinct worth resisting. Two computed values that differ slightly are usually a precision question rather than a mistake, and the way to tell is to predict the size of the discrepancy from the rounding and see whether it matches.
Two truths and a lie
All three concern precision.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The intercept contributes its rounding error once, unchanged; the slope's error is multiplied by x. At x near 73 the slope's contribution here was 0.197 against the intercept's 0.003 — sixty times larger.
Prediction
Commit before reasoning.
Predict first
When the book's rounded prediction and the full-precision one differ slightly, what is best practice?
Correct: Report the full-precision value and note the book's.
Why: The full-precision value is the correct one, and noting the alternative lets a reader reconcile it with the printed source. Averaging invents a third number belonging to no line, and reporting only one without explanation leaves the discrepancy to surface later as an apparent error.
Estimation
A slope near 5 is used to predict at x values near 70.
Predict first
To keep the prediction accurate to about 0.01, how many decimals does the slope need?
Correct: Four.
Why: An error of 0.0001 in the slope becomes 0.007 at x = 70, which is under 0.01; three decimals would leave 0.07. Carrying four decimals is the usual practical answer for x values of this size, and carrying the full calculator value costs nothing.
Section
Section 5
Concept
A prediction should be stated in the context of the problem, in the units of the dependent variable, with the words on average — and, where it matters, with the range over which the line was fitted.
a complete report — The predicted value, what it is a prediction of, and the domain that licenses it. An equation quoted without its range cannot be used responsibly by a reader.
\[ \text{students scoring } 73 \;\to\; \text{about } 179 \text{ on the final, on average} \]
The habit connects back to section 12.1, where a slope and intercept had to be interpreted in complete sentences, and forward to any use of the result. A reported regression that omits the range of its x values leaves every reader unable to distinguish interpolation from extrapolation, which is the single most consequential judgement in using it.
Figure (svg): A card explaining what a predicted value is a prediction of
OpenStax Introductory Statistics 2e, §12.5 Prediction §12.5, pp. 635-636 — the book's own wording of its predictions
Picture it
The prediction, and what it is a prediction of.
Figure (svg): A card explaining what a predicted value is a prediction of
The right-hand panel is what the qualifier prevents. Without it a reader takes 179.08 as a forecast for a particular student, which the residual spread of this data plainly does not support.
Worked example
For the prediction at x = 73.
\[ \hat{y}(73) = 179 \]
Name the x
Why: In context.
\[ a\text{ third exam score of } 73 \]
Name the y
Why: In units.
Give the value
Why: Rounded sensibly.
\[ \text{about } 179 \]
Add the qualifier
Why: It is a mean.
Figure (svg): The solution to Worked example writing the sentence shown as a ladder of expressions, one row per legal move
\[ \hat{y}(73) \approx 179 \text{ points, on average} \]
Verify: confirm the rounding of the reported value is honest
Why: The final exam is scored in whole points and the residual standard deviation is 16.4, so reporting 179.08 implies a precision the data cannot support. Rounding to 179 — or even stating a range around it — matches what the model can actually distinguish. Quoting two decimals on a prediction whose typical miss is sixteen points overstates the result.
OpenStax Introductory Statistics 2e, §12.5 Prediction §12.5, p. 635
Two truths and a lie
All three concern reporting.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. With a residual standard deviation of 16.4 points, two decimal places on a prediction of 179.08 claim a precision three orders of magnitude finer than the model can distinguish.
Worked example
Everything a reader needs to use the exam regression.
\[ \text{the complete report} \]
The equation
Why: Both coefficients.
\[ -173.51 + 4.83 x \]
The strength
Why: r and r-squared.
\[ 0.6631\text{ and } 44 \% \]
The significance
Why: p and n.
\[ 0.026\text{ with } n = 11 \]
The domain
Why: The observed range.
\[ 65\text{ to } 75 \]
Figure (svg): The solution to Worked example what a report should carry shown as a ladder of expressions, one row per legal move
\[ \text{equation} + r^2 + p + [x_{\min}, x_{\max}] \]
Verify: confirm each item answers a question a reader will have
Why: The equation makes prediction possible; r-squared says how much of the variation it explains; the p-value with n says whether the relationship is more than chance at that sample size; and the domain says where predictions may be made at all. Omitting any one leaves a question a reader cannot answer from the rest, and the domain is the one most often left out.
OpenStax Introductory Statistics 2e, §12.5 Prediction §12.5, pp. 635-636
Trap
\[ \hat{y}(73) = 179.08 \text{ points} \]
Report every digit the calculator produced
Why: It is what the arithmetic gave.
\[ \text{but the residual standard deviation is } 16.4 \]
Two decimal places suggest a precision of a hundredth of a point on a prediction whose typical miss is sixteen points.
\[ \text{about } 179 \text{ points, on average} \]
Round to what the model can distinguish
Why: The residual spread sets the scale.
This is the same discipline chapter 2 applied to summary statistics and chapter 8 to confidence intervals: reported precision should reflect what the data support rather than what the arithmetic produced. A prediction accompanied by its residual standard deviation is more informative still, since it tells the reader how far individuals may fall from it.
Sorting
Each is a candidate item to report alongside a regression.
Sort into buckets
Sort by whether a reader needs it.
Entry order does not affect a least-squares fit at all — the same points in any sequence give the same line. The other four each answer a question the rest cannot.
Faded example
For a prediction of 86.5 on a math test from five hours of practice.
Fill in the blanks
\text86.5 on average, \; ___
Why: Both parts are needed. The value without the qualifier reads as a forecast for one student, which no regression prediction ever is.
Prediction
Commit before reasoning.
Predict first
Which item is most commonly missing from a reported regression, and what does its absence cost?
Correct: The range of observed x values.
Why: Equations, correlations and sample sizes are routinely reported; the domain rarely is. Without it a reader cannot tell whether their own x of interest falls inside the data, which is precisely the judgement this section is about — so the omission disables the most consequential check they could make.
Comparison
Fill the blanks. The arithmetic is identical; the standing is not.
Comparison matrix
| Interpolation | Extrapolation | |
|---|---|---|
| Where x lies | inside the observed range | outside it |
| Evidence behind it | data on both sides | none |
| The calculation | a substitution | the same substitution |
| Licensed? | yes, if r is significant and the plot linear | no, whatever r is |
The third row is why this section is needed at all. Since the arithmetic cannot distinguish the two cases, the person has to — and the book demonstrates the cost with a predicted score of 261 on a test marked out of 200.
Pattern
Six steps, and the substitution is the last of them.
Report the domain alongside the equation, so a later reader can apply step four without the original data.
OpenStax Introductory Business Statistics 2e, §13.6 Predicting with a Regression Equation §13.6 Predicting with a Regression Equation
Check
The boundary.
Check your understanding
A regression's x values run from 20 to 50. Predicting at x = 60 is called what?
Answer: A
Why: Sixty lies outside the observed range, which is extrapolation — and section 12.4's third note says a significant r does not make the line reliable there.
Check
What is predicted.
Check your understanding
A line predicts 179 at x = 73. What does that number estimate?
Answer: A
Why: The book's wording is that such students will earn that grade on average. Individuals scatter around the line by as much as the residual spread allows.
Check
The demonstration.
Check your understanding
Substituting x = 90 into the exam line gives 261.19. What does that show?
Answer: A
Why: The book substitutes 90 precisely to show how unreliable prediction becomes outside the observed x values: 261.19 exceeds the largest score the final exam can carry.
Real world
A regional planner fits a line to a town's population from 2010 to 2024, obtaining a significant correlation of 0.98 and a slope of 640 people per year. The report projects the population in 2075 by substituting into the same equation, and recommends water infrastructure sized for that figure.
Discussion prompt
Assess the fit and the projection separately, and say what should be done instead.
Hint: Compare the target year against the years the data covered.
Answer:
The fit is excellent within its range. A correlation of 0.98 over fifteen years explains about 96 percent of the variation in population, and with thirteen degrees of freedom that is overwhelmingly significant. For any year between 2010 and 2024 the line is well supported.
The projection is extrapolation, and by a long way. The data span fifteen years and the projection reaches fifty-one years past the last observation — more than three times the length of the observed record. The line predicts about 32,600 more people than in 2024, and nothing in the data speaks to any year after 2024.
\[ \text{observed } [2010, 2024]; \quad 2075 - 2024 = 51 \text{ years beyond} \]
The failure mode here is different from the exam example's and worse. A predicted exam score above 200 announces itself as impossible; a projected population is merely a large number, and nothing about it looks wrong. Populations also change regime — growth saturates, industries leave, policy shifts — so the assumption that one straight trend continues for half a century is precisely the assumption the data cannot test.
What should be done: report the fit for 2010 to 2024 with its range stated, and treat 2075 as a scenario rather than a prediction. Infrastructure planning at that horizon uses demographic models with explicit assumptions about births, deaths and migration, and sizes for a range of scenarios rather than a point estimate. The regression is evidence about the recent past; it is not a forecast, and a correlation of 0.98 does not make it one.
Commit first
Answer, then rate your confidence honestly.
Predict first
Why should a regression line not be used to predict outside the observed range of x values?
Correct: The data say nothing about that region.
\[ \hat{y}(90) = 261.19 > 200 = \text{the maximum possible score} \]
Why: The arithmetic works perfectly — that is the danger. Inside the observed range there are data on both sides of any x, so the line is constrained by evidence; outside it, the line continues only because a straight line continues, and the assumption that the same relationship holds there was never tested. The book shows the cost by predicting 261.19 on an exam marked out of 200.
Explain it
They have a significant correlation and are predicting at an x value well beyond the largest one observed, saying the p-value justifies it.
Discussion prompt
In two sentences or fewer, correct them.
Hint: Ask what region the test was about.
Answer:
The significance test says the linear relationship is real among the x values that were actually observed, and says nothing about a region where no data were collected.
That is extrapolation, and in the book's own example it predicts a final exam score of 261 on a test marked out of 200.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, linear pattern, significant r, and x in range. For the second, compare x against the smallest and largest observed. For the third, a mean at that x, not an individual. For the fourth, round to what the residual spread supports and say on average. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Twelve minutes — this is a short section.
Draw it
Draw the eleven exam points with the fitted line through them, and shade the vertical band between x equal to 65 and x equal to 75, labelling it the observed range. Extend the line to the right as a DASHED continuation past 75 out to x equal to 90, and draw a horizontal line at y equal to 200 labelled the maximum possible final exam score. Mark the point where the dashed line crosses that horizontal, and write beside it that it happens at about x equal to 77.3, barely past the data. Mark x equal to 73 inside the band with its prediction of 179, and x equal to 90 outside it with its prediction of 261. Beneath the plot, write the two words in boxes — interpolation for inside, extrapolation for outside — with a one-line note under each saying licensed and not licensed. To the right, list the three gates: a linear pattern, a significant correlation, and an x inside the range. At the bottom, write the reported sentence in full, underlining the words on average, and note that the rounded line gives 179.08 while the full-precision one gives 178.89.
Check your crossing point by solving 200 equals minus 173.51 plus 4.83x, which should give about 77.3. Check your dashed extension by confirming it leaves the plausible region almost immediately after leaving the data, which is the whole point of the picture.
Recap
Five things, and three of them happen before any arithmetic.
| If you see | Then |
|---|---|
| An x inside the observed range | Interpolation: substitute and report |
| An x outside it | Extrapolation: do not predict, whatever r is |
| A significant r with a curved plot | Do not predict at all |
| A prediction to report | State it as a mean, with on average |
| An equation quoted without its data range | Ask for the range before using it |
| A predicted value beyond what is possible | A sure sign of extrapolation |
| Coefficients rounded before substituting | An error of roughly the rounding times x |
Section 12.6 asks the question this chapter has so far deferred: which points should the line have been fitted to? Some observations sit far from the line, others sit far from the other x values, and both kinds can change a fitted line substantially.
OpenStax Introductory Statistics 2e, §12.5 Prediction §12.5, pp. 635-636 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.