The chapter's last section asks which points the line should have been fitted to. Two different kinds of point matter and the book keeps them apart: an outlier is far from the least-squares line vertically, meaning it has a large residual, while an influential point is far from the other observations horizontally and may have a big effect on the slope. The rough rule for flagging an outlier is any point further than two standard deviations from the line, where the standard deviation is that of the residuals, computed as the square root of the SSE divided by n minus two — dividing by n minus two because the regression model involves two estimates. For the running example that flags one point, the student who scored 65 on the third exam and 175 on the final. Deleting it changes the fitted line from a slope of 4.83 to 7.39 and the correlation from 0.6631 to 0.9121, which is why the book insists a deleted point be recorded and explained, or the results reported both ways.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 12 — Linear Regression and Correlation
Outliers
Objectives
Six outcomes, and the distinction in the first is the one that carries.
OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 636-643 — the section these objectives are drawn from
Warm-up
Section 12.3 fitted the line and defined residuals; section 12.5 predicted from it.
Discussion prompt
Among the eleven exam students, one scored 65 on the third exam and 175 on the final while another scored 66 and finished with 126. Should either change how the line is fitted?
Hint: Compute roughly what the line predicts for each, and compare.
Answer:
The line predicts about 140 at a third-exam score of 65 and about 145 at 66, so the first student is 35 points above the line and the second 19 points below it. Both are large misses, and the first is nearly twice the second.
Whether that matters needs a threshold, and the book supplies one: flag any point further than two standard deviations from the line, using the standard deviation of the RESIDUALS. Here that standard deviation is about 16.4, so the threshold is about 32.8.
\[ |\varepsilon| > 2s = 2(16.4) = 32.8 \]
By that rule the first point is flagged and the second is not. What to do next is a separate question — the book's answer is to examine the point rather than to delete it, since an outlier may be an error or may be genuinely informative.
Concept
Outliers are observed data points that are far from the least squares line, with large errors. Influential points are observed data points far from the other observed data points in the horizontal direction, which may have a big effect on the slope of the regression line.
the two kinds — An outlier is far vertically and has a large residual; an influential point is far horizontally and may swing the slope. A point can be either, both, or neither.
\[ \text{outlier}: |\varepsilon| \text{ large}; \qquad \text{influential}: x \text{ far from the other } x \]
Both need examining, and for different reasons. An outlier raises the question of whether the observation is correct; an influential point raises the question of whether the fitted line depends too heavily on one observation. The book's suggested test for the second is direct: remove it from the data set and see if the slope of the regression line is changed significantly.
Figure (svg): A card contrasting an outlier with an influential point
OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, p. 636
Section
Section 1
Concept
An outlier is far from the least squares line, so its residual is large. An influential point is far from the other observed data points in the horizontal direction, so it may exert a large effect on the slope. To begin to identify an influential point, remove it from the data set and see if the slope is changed significantly.
independent properties — Being far from the line and being far from the other x values are different things. The exam data contains a point of each kind, and they are different points.
\[ \text{vertical distance} \quad\text{against}\quad \text{horizontal position} \]
The exam data makes the distinction concrete. The point (75, 198) has a residual of only 9.5, well inside any outlier threshold — yet its third-exam score is the largest in the sample by four points, and removing it drops the fitted slope from 4.83 to 3.46. It is influential without being an outlier, which is exactly the case that a residual-based rule cannot detect.
Figure (svg): A card contrasting an outlier with an influential point
OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, p. 636 — the definitions of both, and the removal test
Picture it
With the exam data's example of each.
Figure (svg): A card contrasting an outlier with an influential point
The bottom panel is the part worth remembering. Checking residuals finds outliers and finds nothing about influence, so a point that quietly determines the slope can pass every check in this section unless it is looked for separately.
Worked example
The point (75, 198) has the largest third-exam score in the sample.
\[ (75, 198) \]
Its residual
Why: 198 minus 188.5.
\[ 9.46 \]
Against the threshold
Why: Two s is 32.8.
Its x position
Why: Next largest is 71.
Remove it and refit
Why: Slope changes.
\[ 4.83\text{ to } 3.46 \]
Figure (svg): The solution to Worked example an influential point that is not an outlier shown as a ladder of expressions, one row per legal move
\[ \varepsilon = 9.46 \ll 32.8; \quad b: 4.83 \to 3.46 \]
Verify: confirm why a distant x has this much leverage
Why: The other ten third-exam scores lie between 65 and 71, so this single point sits four beyond the rest and anchors the line's right end almost by itself. A least-squares fit minimises squared vertical misses, and a miss at a distant x can be reduced most cheaply by tilting the line — so points far from the other x values pull the slope hardest. That is the whole content of the word influential.
OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 636-639
Sorting
Each describes a point relative to the rest of the data.
Sort into buckets
Sort by which property it has.
Items (b) and (c) are both influential, and only one is an outlier — which is why sorting by residual alone answers half the question. Item (b) is (75, 198) and item (e) is (65, 175).
Worked example
The flagged point (65, 175) sits at the smallest third-exam score.
\[ (65, 175) \]
Its residual
Why: 175 minus 140.3.
\[ 34.73 \]
Against the threshold
Why: Exceeds 32.8.
Its x position
Why: Smallest in the sample.
Remove it and refit
Why: Slope changes.
\[ 4.83\text{ to } 7.39 \]
Figure (svg): The solution to Worked example a point that is both shown as a ladder of expressions, one row per legal move
\[ \varepsilon = 34.73 > 32.8; \quad b: 4.83 \to 7.39 \]
Verify: confirm why this combination is the most consequential
Why: A point that is far from the line AND far horizontally pulls the slope hardest of all, which is why deleting this one changes the fit so dramatically — the slope rises by more than half and the correlation climbs from 0.6631 to 0.9121. A large residual at a central x would barely move the line at all, since there would be little leverage to act through.
OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 637-639
Trap
\[ \text{small residual} \;\Rightarrow\; \text{the point does not matter} \]
Judge a point's importance by its distance from the line
Why: That is what the outlier rule measures.
\[ (75, 198): \; \varepsilon = 9.5 \text{ and yet } b: 4.83 \to 3.46 \]
A point can sit close to the line and still determine where the line goes, if its x is far from the others.
\[ \text{check residuals for outliers AND } x \text{ positions for influence} \]
Treat the two as separate checks
Why: The book defines them separately for this reason.
The reason a small residual can accompany large influence is circular in a useful way: an influential point pulls the line toward itself, so the fitted line ends up close to it — and its residual is small BECAUSE it had so much influence. The residual measures the outcome of the fit, not the point's importance to it.
Two truths and a lie
All three concern the two kinds.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and backwards in an instructive way. An influential point pulls the line toward itself, so its residual is often SMALL precisely because of the influence it exerted.
Prediction
Commit before reasoning.
Predict first
The book's suggested test for an influential point is what?
Correct: Remove it and refit.
Why: The book says to begin identifying an influential point by removing it from the data set and seeing if the slope of the regression line is changed significantly. There is no formula given; the test is direct and requires refitting.
Estimation
The full data gives a slope of 4.83; removing (75, 198) gives 3.46.
Predict first
Roughly what proportional change is that?
Correct: About 28 percent.
Why: The drop of 1.37 against 4.83 is 28 percent, from a point whose residual was under a third of the outlier threshold. That is a substantial change from one observation in eleven, and no residual-based check would have surfaced it.
Section
Section 2
Concept
The standard deviation of the residuals is calculated from the SSE as the square root of the SSE divided by n minus two. We divide by n minus two because the regression model involves two estimates.
s, the residual standard deviation — A measure of the typical vertical miss. The calculator reports it as s in the LinRegTTest output, and it is the quantity the outlier rule is stated in terms of.
\[ s = \sqrt{\frac{\text{SSE}}{n-2}} \]
The divisor is the same accounting section 12.4 used for its degrees of freedom. A regression estimates two quantities from the data, an intercept and a slope, so two degrees of freedom are spent and n minus two remain. That is the same principle behind chapter 8's n minus one for a single mean, applied to a model with one more parameter.
Figure (svg): A card reconciling the three values the book gives for the residual standard deviation
OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 638-639 — the formula for s and the note on dividing by n minus two
Picture it
For the same quantity, in the same example.
Figure (svg): A card reconciling the three values the book gives for the residual standard deviation
The first is the calculator's exact value and the second comes from the book's hand computation on rounded residuals. The third appears only as a threshold in one sentence and matches neither, but since the largest residual is 34.7 and the next is 19.1, all three flag the same single point.
Worked example
The book squares the residuals 35, 17, 16, 6, 19, 9, 3, 1, 10, 9 and 1, and adds them.
\[ \text{SSE} = 2440, \; n = 11 \]
The divisor
Why: n minus two.
\[ 9 \]
Divide
Why: 2440 over 9.
\[ 271.1 \]
Take the root
Why: Of 271.1.
\[ 16.47 \]
Double it
Why: The threshold.
\[ 32.94 \]
Figure (svg): The solution to Worked example computing s from the residuals shown as a ladder of expressions, one row per legal move
\[ s = \sqrt{\frac{2440}{9}} = 16.47 \]
Verify: confirm against the calculator's value and account for the gap
Why: LinRegTTest reports s = 16.412, from an exact SSE of 2424.30 rather than 2440. The difference comes entirely from the book's hand computation using residuals rounded to whole numbers, which inflates the SSE by about 16. Both give thresholds near 33, and since the second-largest residual is 19.1, no point sits between them — the conclusion is identical either way.
OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, p. 639
Faded example
A regression on 12 points has SSE = 900.
Fill in the blanks
s = \sqrt1018.97}} = 9.49, \qquad 2s = ___
Why: Any point whose residual exceeds about 19 in magnitude would be flagged as a potential outlier under the book's rough rule.
Worked example
Comparing with the divisors used earlier in the book.
\[ n-1 \text{ against } n-2 \]
A single mean
Why: One estimate.
\[ \text{divide by } n - 1 \]
A regression line
Why: Intercept and slope.
So the divisor
Why: Two fewer.
\[ n - 2 \]
Check against section 12.4
Why: Same count.
\[ d f = n - 2 \]
Figure (svg): The solution to Worked example why n minus two shown as a ladder of expressions, one row per legal move
\[ \text{df} = n - 2 = 9 \]
Verify: confirm the two uses of n minus two are the same quantity
Why: Section 12.4's significance test uses n minus two degrees of freedom for exactly this reason, and this section uses n minus two in the divisor for the same one. They are not two coincidental rules but one accounting applied twice: a fitted line costs two degrees of freedom, wherever that cost shows up.
OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, p. 638
Error analysis
With SSE = 2424.30 and n = 11, which are correct?
Annotate
On: \( \begin{aligned} &(1)\; s = \sqrt{2424.30/9} = 16.41 \\ &(2)\; s = \sqrt{2424.30/11} = 14.85 \\ &(3)\; s = \sqrt{2424.30/10} = 15.57 \\ &(4)\; s = 2424.30/9 = 269.4 \end{aligned} \)
Error (3) is the most tempting because n minus one is so familiar, and it understates s by about 5 percent — which lowers the outlier threshold and would flag points the correct rule does not.
Two truths and a lie
All three concern s.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The standard deviation of the y values here is 20.80, describing how much final scores vary overall; s is 16.41, describing how much they vary about the LINE. The second is smaller precisely because the line explains some of the variation.
Prediction
Commit before reasoning.
Predict first
For the exam data, sy is 20.80 and s is 16.41. Why is s smaller?
Correct: The line explains part of the variation.
Why: This is r-squared in a different guise. The line accounts for about 44 percent of y's variation, so what remains scattered around it is smaller than the total spread. For a perfect fit s would be zero, and for a useless one it would approach the spread of y itself.
Estimation
The exact SSE is 2424.30; the book's hand computation gives 2440.
Predict first
By roughly what percentage?
Correct: Under 1 percent.
Why: The gap of 15.7 against 2424.3 is about 0.6 percent, which moves s from 16.412 to 16.47 — a difference of 0.06. Rounding eleven residuals to whole numbers costs remarkably little here, which is why the book can compute by hand without changing any conclusion.
Section
Section 3
Concept
As a rough rule of thumb, flag any point located further than two standard deviations above or below the best-fit line. This can be done visually by drawing an extra pair of lines two standard deviations above and below the best-fit line, or numerically by calculating each residual and comparing it to twice the standard deviation.
the rule of thumb — The book's own description. It is a screening device rather than a test, and it produces candidates for examination rather than verdicts.
\[ |y - \hat{y}| \ge 2s \;\Longrightarrow\; \text{flag} \]
The two methods are the same rule seen two ways. The extra lines are the fitted line shifted vertically by two standard deviations, so a point outside them is exactly a point whose residual exceeds that amount. The book notes that on a calculator screen the graphical version can be hard to read when a point is close to a band — which is why the numerical version exists as a check.
Figure (svg): A scatter plot with the fitted line and two dashed parallel lines two residual standard deviations above and below, one point falling outside them
OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 636-639 — the graphical and numerical identification methods
Picture it
Two standard deviations above and below the fitted line.
Figure (svg): A scatter plot with the fitted line and two dashed parallel lines two residual standard deviations above and below, one point falling outside them
The flagged point is only just outside, which is the book's own observation: on the calculator screen it is just barely outside these lines. When a point is that close, the numerical comparison settles it.
Worked example
The book draws Y2 and Y3 alongside the best-fit line.
\[ s = 16.4 \]
The fitted line
Why: As Y1.
\[ -173.5 + 4.83 x \]
Two standard deviations
Why: Twice 16.4.
\[ 32.8 \]
The lower band
Why: Subtract it.
\[ Y 2 = -173.5 + 4.83 x - 32.8 \]
The upper band
Why: Add it.
\[ Y 3 = -173.5 + 4.83 x + 32.8 \]
Figure (svg): The solution to Worked example the graphical method shown as a ladder of expressions, one row per legal move
\[ Y_2, Y_3 = -173.5 + 4.83x \mp 2(16.4) \]
Verify: confirm the bands are parallel and why that matters
Why: Both have the same slope as the best-fit line, differing only in intercept — so the band has constant width across the whole range. That is consistent with section 12.4's third assumption, that the standard deviations of the y values about the line are equal for each value of x. A residuals plot that fanned outward would violate it, and the constant-width band would then be the wrong screening tool.
OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, p. 637
Faded example
A regression has s = 12, and a point has a residual of -26.
Fill in the blanks
2s = 24, \textexceeds |___| = 26 ___ \text___
Why: The point is flagged. The sign is irrelevant to the rule — the comparison uses the magnitude, so points below the line are flagged on the same terms as points above it.
Worked example
Comparing every residual against twice the standard deviation.
\[ 2s = 32.8 \]
The largest residual
Why: At (65, 175).
\[ 34.73 \]
Compare
Why: Exceeds 32.8.
The next largest
Why: At (66, 126).
\[ 19.09 \]
Compare
Why: Well below.
Figure (svg): The solution to Worked example the numerical method shown as a ladder of expressions, one row per legal move
\[ 34.73 > 32.8 > 19.09 \]
Verify: confirm the verdict is robust to which threshold is used
Why: The book's three candidate thresholds are 31.29, 32.8 and 32.94, and the two largest residuals are 34.7 and 19.1. Every threshold falls in the wide gap between them, so all three identify exactly the same point. When a screening rule's exact value is uncertain, checking whether the answer is sensitive to it is more useful than settling which value is right.
OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 639-640
Trap
\[ |\varepsilon| > 2s \;\Rightarrow\; \text{the point is wrong; delete it} \]
Read a flag as a verdict
Why: The rule produced a definite answer.
\[ \text{but the book calls it a rough rule of thumb} \]
It flags candidates for examination, not points for removal.
\[ |\varepsilon| > 2s \;\Rightarrow\; \text{examine this observation} \]
Investigate the flagged point before doing anything with it
Why: The key is to examine what causes it to be an outlier.
The book is explicit that an outlier may hold valuable information about the population and should remain included. In a sample of eleven, roughly one point would be expected beyond two standard deviations by chance alone under a normal model — so a single flag is close to what ordinary variation produces, and is not on its own evidence of anything wrong.
Two truths and a lie
All three concern flagging.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and reverses the book's guidance. A flagged point should be examined; it may be erroneous, or it may hold valuable information about the population and should remain included. The book deletes one only as a learning exercise, after supposing it was found to be an error.
Prediction
Commit before reasoning.
Predict first
Why do Y2 and Y3 have the same slope as the fitted line?
Correct: The threshold is a constant vertical distance.
Why: Two standard deviations is the same distance at every x, which is exactly section 12.4's third assumption — that the spread of y about the line is equal for each x. Bands of constant width encode that assumption, and if it failed the screening rule would be too strict at some x values and too lenient at others.
Estimation
Under the test's assumptions, residuals are roughly normal about the line.
Predict first
In a sample of 11, how many points would be expected beyond two standard deviations by chance?
Correct: About half a point.
Why: A normal distribution puts about 5 percent of its values beyond two standard deviations, and 5 percent of eleven is 0.55. So finding one flagged point in eleven is entirely ordinary and is not by itself evidence that anything is wrong — which is precisely why the book calls this a rough rule of thumb and asks for examination rather than action.
Section
Section 4
Concept
Outliers need to be examined closely. Sometimes they should not be included in the analysis of the data — it is possible that an outlier is a result of erroneous data. Other times an outlier may hold valuable information about the population under study and should remain included. The key is to examine carefully what causes a data point to be an outlier.
the reporting requirement — When outliers are deleted, the researcher should either record that data was deleted and why, or should provide results both with and without the deleted data.
\[ \text{examine} \to \text{correct, keep, or delete with a record} \]
The book adds a third option that is easy to overlook: if the data are erroneous and the correct values are known, the correction can simply be made — a student who actually scored 70 rather than 65 should have their value fixed rather than removed. Deletion is the last resort among three responses, not the default one.
Figure (svg): Handling a potential outlier
OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 636-640 — what outliers need, and the note on recording deletions
Picture it
Six steps, ending in a reporting obligation.
Figure (svg): Handling a potential outlier
The final step is the one most often skipped and the one that makes an analysis checkable. A reader who is told which point was removed and why can judge the decision; a reader given only the final line cannot.
Worked example
Having flagged (65, 175), what does the book do?
\[ (65, 175) \text{ flagged} \]
First response
Why: Re-examine the data.
If an error
Why: Fix or delete.
If correct
Why: Leave it in.
The book's course
Why: Supposes an error.
Figure (svg): The solution to Worked example the book's own decision shown as a ladder of expressions, one row per legal move
\[ \text{delete, having recorded why} \]
Verify: confirm the book flags its own artificiality
Why: It writes that for this problem we will suppose that we examined the data and found that this outlier data was an error, and adds a parenthetical reminder that we do not always delete an outlier. The supposition is stated rather than assumed, which is exactly the transparency its own reporting note asks for.
OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 639-640
Sorting
Each is a response to a flagged point.
Sort into buckets
Sort by whether the book supports it.
Item (e) is supported because of the second half: the book permits deletion provided the results are given both with and without the removed data, or the deletion is recorded with its reason.
Worked example
The points (1,5), (2,7), (2,6), (3,9), (4,12), (4,13), (5,18), (6,19), (7,12) and (7,21).
\[ n = 10 \]
Fit the line
Why: Least squares.
\[ 2.897 + 2.269 x \]
The residual standard deviation
Why: Root of SSE over 8.
\[ 3.063 \]
The threshold
Why: Twice it.
\[ 6.13 \]
The largest residual
Why: At (7, 12).
\[ -6.78:\text{ flagged} \]
Refit without it
Why: Nine points.
\[ 1.035 + 2.961 x \]
Figure (svg): The solution to Worked example applying the rule to Try It 12.13 shown as a ladder of expressions, one row per legal move
\[ \hat{y}(10) = 1.035 + 2.961(10) = 30.65 \]
Verify: confirm the flagged point is the one the picture suggests
Why: At x = 7 the other observation is 21 while this one is 12, so two points at the same x differ by nine — and the line, drawn through a generally rising pattern, passes well above 12. The residual of -6.78 exceeds the threshold of 6.13, confirming numerically what the plot shows. Note that the prediction at x = 10 is extrapolation, since the observed x values stop at 7.
OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, p. 640
Error analysis
Which are consistent with the book's guidance?
Annotate
On: \( \begin{aligned} &(1)\; \text{examine the observation for a recording error} \\ &(2)\; \text{delete it, since it worsens the fit} \\ &(3)\; \text{correct it, if the true value is known} \\ &(4)\; \text{delete it and report only the new line} \end{aligned} \)
Errors (2) and (4) compound each other. Deleting whatever fits poorly and reporting only the survivor produces a correlation that describes the deletion process rather than the population — and leaves a reader no way to detect it.
Two truths and a lie
All three concern handling.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false on two counts. Outliers cannot be identified before fitting, since the rule uses residuals from the fitted line — and even once identified, the book's first instruction is to examine rather than to remove.
Faded example
A researcher removes a flagged point from an analysis.
Fill in the blanks
\textwhy both ways, \text___ ___
Why: Either satisfies the requirement, and both leave a reader able to judge the decision rather than having to accept it.
Explain it
A classmate removes every point with a large residual, one at a time, until the correlation exceeds 0.95.
Discussion prompt
In two sentences or fewer, say what is wrong.
Hint: Ask what that procedure would do to unrelated data.
Answer:
Removing whatever fits poorly will raise the correlation of ANY dataset, including one with no relationship at all, so the resulting number measures the deletion process rather than the variables.
The book's rule is to examine a flagged point for a reason it might be wrong, not to delete points until the fit looks good.
Section
Section 5
Concept
Computing a new best-fit line using the ten remaining points gives a slope of 7.39 and an intercept of -355.19, with r equal to 0.9121. The new line is a stronger correlation than the original because 0.9121 is closer to one, so it fits the ten remaining data values better.
what moves — Everything: the equation, the correlation, the coefficient of determination, the SSE and every prediction made from the line.
\[ \hat{y} = -355.19 + 7.39x, \qquad r = 0.9121 \]
The correlation rising is not evidence that removing the point was right. Any least-squares fit improves when its worst-fitting observation is discarded, so a higher r is guaranteed by the procedure rather than earned by it. The justification for a deletion has to come from examining the observation, never from what happens to the fit afterwards.
Figure (svg): A table comparing the regression before and after the outlier is removed
OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 639-640 — the new line and correlation after deletion
Picture it
The same data, one point apart.
Figure (svg): A table comparing the regression before and after the outlier is removed
The SSE falls by more than two thirds and r-squared nearly doubles, from one observation in eleven. That magnitude is the argument for the reporting requirement: a reader shown only the second column would form a very different impression of the relationship.
Worked example
Using the new line based on the remaining ten points, what would a student who scores 73 on the third exam expect on the final?
\[ \hat{y} = -355.19 + 7.39x, \; x = 73 \]
Multiply
Why: 7.39 times 73.
\[ 539.47 \]
Add the intercept
Why: Subtract 355.19.
\[ 184.28 \]
The original prediction
Why: From section 12.5.
\[ 179.08 \]
The difference
Why: Subtract.
\[ 5.2\text{ points} \]
Figure (svg): The solution to Worked example Example 12.13, the new prediction shown as a ladder of expressions, one row per legal move
\[ \hat{y}(73) = 184.28 \quad\text{against}\quad 179.08 \]
Verify: confirm the difference is meaningful relative to the model's precision
Why: Five points on a 200-point exam is small, but the residual standard deviation has also fallen — from 16.4 to about 9.3 on the ten remaining points — so the new line is both different and more precise. Both changes come from one deleted observation, which is why the book's instruction to report results with and without it is not a formality.
OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, p. 640
Faded example
The refitted line is ŷ = -355.19 + 7.39x.
Fill in the blanks
\hat517.3(70) = -355.19 + 7.39(70) = -355.19 + 162.11 = ___
Why: The original line predicted 164.59 at the same x, so the two differ by only 2.5 points near the centre of the data — much less than the 5.2-point gap at x = 73, since the lines cross near the middle and diverge toward the edges.
Worked example
The Consumer Price Index data, with fourteen points and a fitted line ŷ = -3204 + 1.662x.
\[ s = 25.4 \]
The threshold
Why: Twice 25.4.
\[ 50.8 \]
Draw the bands
Why: Above and below.
\[ Y 2\text{ and } Y 3 \]
Check every point
Why: None outside.
The closest
Why: 1999.
Figure (svg): The solution to Worked example Example 12.14, when nothing is flagged shown as a ladder of expressions, one row per legal move
\[ \text{all } |\varepsilon| < 50.8 \]
Verify: confirm the more important caution in this example
Why: The book adds a note that matters more than the outlier check: although the correlation coefficient is significant, the pattern in the scatterplot indicates that a curve would be a more appropriate model than a line. So the data pass the outlier screen while failing the linearity condition — a reminder that the checks in this chapter are independent, and passing one says nothing about the others.
OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 641-642
Trap
\[ r: 0.6631 \to 0.9121 \;\Rightarrow\; \text{removing the point was right} \]
Take the better fit as evidence for the decision
Why: The improvement is dramatic.
\[ \text{but removing the worst point always improves the fit} \]
The same procedure applied to entirely unrelated data would also raise r, so the improvement proves nothing.
\[ \text{delete only on evidence about the OBSERVATION} \]
Justify from examination of the data point, then report both
Why: The book supposes an error was found, and says so.
The asymmetry is worth stating plainly: a deletion can be justified by finding a recording error, a measurement fault or a subject who does not belong to the population — all facts about the observation. It can never be justified by the effect on r, because that effect is guaranteed in advance.
Two truths and a lie
All three concern the effect of deletion.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. Removing the worst-fitting point raises r in any dataset whatsoever, so the improvement carries no information about whether the point deserved removal. That justification has to come from examining the observation itself.
Estimation
From 2424.3 on eleven points to 688.2 on ten.
Predict first
Roughly what fraction of the SSE did that one point contribute?
Correct: About 70 percent.
Why: The drop of 1736 against 2424 is 72 percent, from one observation in eleven. Its own squared residual was 1206, and removing it also let the line shift to fit the rest better — so a single point accounted for the majority of the total squared miss.
Prediction
Commit before reasoning.
Predict first
The CPI data has a significant correlation and no outliers. What does the book still object to?
Correct: The pattern bends.
Why: Its note says a statistician should prefer to use other methods to fit a curve to this data, and adds the general lesson that in addition to doing the calculations, it is always important to look at the scatterplot when deciding whether a linear model is appropriate. Passing the outlier screen and the significance test does not make a line the right model.
Comparison
Fill the blanks. Two different properties, and two different checks.
Comparison matrix
| Outlier | Influential point | |
|---|---|---|
| Far in which direction | vertically, from the line | horizontally, from the other points |
| How it is detected | a residual beyond two standard deviations | remove it and refit |
| What it threatens | the accuracy of one observation | the slope of the whole line |
| Its own residual | large, by definition | often small, because the line follows it |
The last row is why a residual check alone is not enough. An influential point pulls the line toward itself, so the very leverage that makes it matter also hides it from the outlier rule.
Pattern
Six steps, and the last is a reporting obligation rather than a calculation.
A better fit after deletion is guaranteed by the procedure and is never a justification for it.
OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 636-643
Check
The two kinds.
Check your understanding
A point has a small residual but its x value is far from every other. What is it?
Answer: A
Why: An outlier is far from the line vertically, which a small residual rules out. Being far horizontally is what makes a point influential, and the test is to remove it and see whether the slope changes.
Check
The standard deviation.
Check your understanding
A regression on 10 points has SSE = 512. What is the residual standard deviation?
Answer: A
Why: The divisor is n minus two, so 512 over 8 is 64, whose square root is 8.
Check
Handling.
Check your understanding
A flagged point is examined and the data are found to be correct. What should be done?
Answer: A
Why: The book says an outlier may hold valuable information about the population under study and should remain included. Being unusual is not a reason for removal.
Real world
A lab technician measures reaction rate against temperature at 20, 25, 30, 35, 40 and 95 degrees, fits a line, and reports a correlation of 0.97. A colleague notes that the 95-degree run had a small residual and was clearly consistent with the others, so nothing needs checking.
Discussion prompt
Assess the colleague's reasoning and say what should be done.
Hint: Ask where the 95-degree point sits relative to the other temperatures.
Answer:
The colleague has applied the outlier check and missed the influential-point one. A small residual rules out the point being an outlier and says nothing about influence — and this point's x value is 55 degrees beyond the next highest, while the other five span only 20 degrees in total.
The small residual is itself a symptom rather than reassurance. A point that far from the others exerts enormous leverage: the line is pulled toward it, so it ends up fitting closely BECAUSE it dominated the fit. The correlation of 0.97 is largely a statement that a five-point cluster and one distant point can be joined by a line, which two well-separated groups almost always can.
\[ \text{other } x: [20, 40]; \quad \text{this } x = 95, \; 55 \text{ beyond the next} \]
The book's own test applies directly: remove the 95-degree run and refit. If the slope changes substantially, the point is influential and the reported line describes it more than it describes the other five. That is a two-minute check and the colleague's argument gives no reason to skip it.
What should be done: refit without the 95-degree point and report both lines; collect runs at intermediate temperatures — 50, 65, 80 — so the gap is filled and the fit no longer rests on one observation; and check whether the reaction's behaviour is even linear across so wide a range, since many rates rise exponentially with temperature. The correlation is high; what it is high ABOUT is the open question.
Commit first
Answer, then rate your confidence honestly.
Predict first
Why can an influential point have a small residual?
Correct: It pulls the line toward itself.
\[ (75, 198): \; \varepsilon = 9.46 \text{ and yet } b: 4.83 \to 3.46 \text{ on removal} \]
Why: A least-squares fit minimises squared vertical misses, and a miss at an x far from the others is most cheaply reduced by tilting the whole line. So the point gets fitted closely, and its residual comes out small precisely because of the influence it exerted. The residual measures the outcome of the fit rather than the point's importance to it, which is why influence needs its own check.
Explain it
They deleted three points with the largest residuals, refitted, and report a correlation of 0.96 as their result.
Discussion prompt
In two sentences or fewer, say what is wrong.
Hint: Ask what that procedure does to data with no relationship.
Answer:
Removing the worst-fitting points raises the correlation of any dataset at all, so 0.96 describes the deletion procedure rather than the two variables.
A point may be removed only after examining it and finding a reason it is wrong, and the deletion has to be recorded with its reason or the results given both ways.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, far vertically against far horizontally. For the second, root of SSE over n minus two, then double it. For the third, examine, then correct or keep; delete only with a recorded reason. For the fourth, removing the worst point improves any fit. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top, write the two definitions side by side and box the words VERTICALLY and HORIZONTALLY: an outlier is far from the line vertically, an influential point is far from the other points horizontally. Beneath each, name the exam data's example — (65, 175) with residual 34.7 for the first, and (75, 198) with residual only 9.5 for the second — and note that removing the second changes the slope from 4.83 to 3.46. In the middle, plot the eleven points with the fitted line and two dashed parallel bands 32.8 above and below it, circling the single point that falls outside. Beside the plot, write s equals the root of SSE over n minus two, work it as the root of 2440 over 9 giving 16.47, and note the calculator's 16.412 from the exact SSE of 2424.3. At the bottom left, write the six-step procedure ending with the reporting requirement. At the bottom right, make a two-column before-and-after table: the equation, r, r-squared and the prediction at 73, showing 0.6631 becoming 0.9121 and 179.08 becoming 184.28 — and write one sentence saying that the rise in r is guaranteed by removing the worst point and therefore justifies nothing.
Check your bands by confirming they are parallel to the fitted line, since the threshold is the same vertical distance at every x. Check your before-and-after table by confirming every single row changed, which is the argument for reporting both.
Recap
Six things, and the distinction in the first governs the rest.
| If you see | Then |
|---|---|
| A residual beyond two standard deviations | Flag it and examine the observation |
| An x far from all the others | Remove it and refit to test for influence |
| A small residual at a distant x | Suspect influence, not accuracy |
| SSE and n | s is the root of SSE over n minus two |
| A flagged point with correct data | Keep it: it may be informative |
| A flagged point with a known correct value | Correct it rather than deleting it |
| A deletion | Record it with its reason, or report both ways |
| A higher r after deleting a point | Expected, and no justification at all |
That completes chapter 12, and with it the chapter's full sequence: look at the data, fit a line, test whether it means anything, predict within its range, and check which points it was answering to. Chapter 13 turns to comparing several population means at once, which needs a distribution the course has not yet met.
OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 636-643 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.