12.6 Outliers

The chapter's last section asks which points the line should have been fitted to. Two different kinds of point matter and the book keeps them apart: an outlier is far from the least-squares line vertically, meaning it has a large residual, while an influential point is far from the other observations horizontally and may have a big effect on the slope. The rough rule for flagging an outlier is any point further than two standard deviations from the line, where the standard deviation is that of the residuals, computed as the square root of the SSE divided by n minus two — dividing by n minus two because the regression model involves two estimates. For the running example that flags one point, the student who scored 65 on the third exam and 175 on the final. Deleting it changes the fitted line from a slope of 4.83 to 7.39 and the correlation from 0.6631 to 0.9121, which is why the book insists a deleted point be recorded and explained, or the results reported both ways.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 12.6 Outliers

Title

Statistics · Chapter 12 — Linear Regression and Correlation

Outliers

2. By the end of this lesson you can

Objectives

Six outcomes, and the distinction in the first is the one that carries.

OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 636-643 — the section these objectives are drawn from

3. What you already have

Warm-up

Section 12.3 fitted the line and defined residuals; section 12.5 predicted from it.

Discussion prompt

Among the eleven exam students, one scored 65 on the third exam and 175 on the final while another scored 66 and finished with 126. Should either change how the line is fitted?

Hint: Compute roughly what the line predicts for each, and compare.

Answer:

The line predicts about 140 at a third-exam score of 65 and about 145 at 66, so the first student is 35 points above the line and the second 19 points below it. Both are large misses, and the first is nearly twice the second.

Whether that matters needs a threshold, and the book supplies one: flag any point further than two standard deviations from the line, using the standard deviation of the RESIDUALS. Here that standard deviation is about 16.4, so the threshold is about 32.8.

\[ |\varepsilon| > 2s = 2(16.4) = 32.8 \]

By that rule the first point is flagged and the second is not. What to do next is a separate question — the book's answer is to examine the point rather than to delete it, since an outlier may be an error or may be genuinely informative.

4. Which points is the line answering to?

Concept

Outliers are observed data points that are far from the least squares line, with large errors. Influential points are observed data points far from the other observed data points in the horizontal direction, which may have a big effect on the slope of the regression line.

the two kinds — An outlier is far vertically and has a large residual; an influential point is far horizontally and may swing the slope. A point can be either, both, or neither.

\[ \text{outlier}: |\varepsilon| \text{ large}; \qquad \text{influential}: x \text{ far from the other } x \]

Both need examining, and for different reasons. An outlier raises the question of whether the observation is correct; an influential point raises the question of whether the fitted line depends too heavily on one observation. The book's suggested test for the second is direct: remove it from the data set and see if the slope of the regression line is changed significantly.

Figure (svg): A card contrasting an outlier with an influential point

The book defines both and the exam data illustrates both, which is what makes the distinction concrete rather than terminological.

OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, p. 636

5. Outliers and influential points

Section

Section 1

6. Far vertically, or far horizontally

Concept

An outlier is far from the least squares line, so its residual is large. An influential point is far from the other observed data points in the horizontal direction, so it may exert a large effect on the slope. To begin to identify an influential point, remove it from the data set and see if the slope is changed significantly.

independent properties — Being far from the line and being far from the other x values are different things. The exam data contains a point of each kind, and they are different points.

\[ \text{vertical distance} \quad\text{against}\quad \text{horizontal position} \]

The exam data makes the distinction concrete. The point (75, 198) has a residual of only 9.5, well inside any outlier threshold — yet its third-exam score is the largest in the sample by four points, and removing it drops the fitted slope from 4.83 to 3.46. It is influential without being an outlier, which is exactly the case that a residual-based rule cannot detect.

Figure (svg): A card contrasting an outlier with an influential point

The book defines both and the exam data illustrates both, which is what makes the distinction concrete rather than terminological.

OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, p. 636 — the definitions of both, and the removal test

7. Two ways to matter

Picture it

With the exam data's example of each.

Figure (svg): A card contrasting an outlier with an influential point

The book defines both and the exam data illustrates both, which is what makes the distinction concrete rather than terminological.

The bottom panel is the part worth remembering. Checking residuals finds outliers and finds nothing about influence, so a point that quietly determines the slope can pass every check in this section unless it is looked for separately.

8. Worked example: an influential point that is not an outlier

Worked example

The point (75, 198) has the largest third-exam score in the sample.

\[ (75, 198) \]

Its residual

Why: 198 minus 188.5.

\[ 9.46 \]

Against the threshold

Why: Two s is 32.8.

Its x position

Why: Next largest is 71.

Remove it and refit

Why: Slope changes.

\[ 4.83\text{ to } 3.46 \]

Figure (svg): The solution to Worked example an influential point that is not an outlier shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \varepsilon = 9.46 \ll 32.8; \quad b: 4.83 \to 3.46 \]

Verify: confirm why a distant x has this much leverage

Why: The other ten third-exam scores lie between 65 and 71, so this single point sits four beyond the rest and anchors the line's right end almost by itself. A least-squares fit minimises squared vertical misses, and a miss at a distant x can be reduced most cheaply by tilting the line — so points far from the other x values pull the slope hardest. That is the whole content of the word influential.

OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 636-639

9. Which kind of point?

Sorting

Each describes a point relative to the rest of the data.

Sort into buckets

Sort by which property it has.

Flagged as an outlier
a large residual at a central x value; a large residual at an extreme x value; a residual of 34.7 at the smallest x in the sample
Not an outlier
a small residual at an x far from all others; a small residual at a central x value
out
Its residual is large, which is what defines an outlier.
notout
Its residual is small, whatever its horizontal position.

Items (b) and (c) are both influential, and only one is an outlier — which is why sorting by residual alone answers half the question. Item (b) is (75, 198) and item (e) is (65, 175).

10. Worked example: a point that is both

Worked example

The flagged point (65, 175) sits at the smallest third-exam score.

\[ (65, 175) \]

Its residual

Why: 175 minus 140.3.

\[ 34.73 \]

Against the threshold

Why: Exceeds 32.8.

Its x position

Why: Smallest in the sample.

Remove it and refit

Why: Slope changes.

\[ 4.83\text{ to } 7.39 \]

Figure (svg): The solution to Worked example a point that is both shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \varepsilon = 34.73 > 32.8; \quad b: 4.83 \to 7.39 \]

Verify: confirm why this combination is the most consequential

Why: A point that is far from the line AND far horizontally pulls the slope hardest of all, which is why deleting this one changes the fit so dramatically — the slope rises by more than half and the correlation climbs from 0.6631 to 0.9121. A large residual at a central x would barely move the line at all, since there would be little leverage to act through.

OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 637-639

11. Trap: assuming a large residual means a large effect on the line

Trap

The trap

\[ \text{small residual} \;\Rightarrow\; \text{the point does not matter} \]

Judge a point's importance by its distance from the line

Why: That is what the outlier rule measures.

\[ (75, 198): \; \varepsilon = 9.5 \text{ and yet } b: 4.83 \to 3.46 \]

A point can sit close to the line and still determine where the line goes, if its x is far from the others.

The fix

\[ \text{check residuals for outliers AND } x \text{ positions for influence} \]

Treat the two as separate checks

Why: The book defines them separately for this reason.

The reason a small residual can accompany large influence is circular in a useful way: an influential point pulls the line toward itself, so the fitted line ends up close to it — and its residual is small BECAUSE it had so much influence. The residual measures the outcome of the fit, not the point's importance to it.

12. One of these is false

Two truths and a lie

All three concern the two kinds.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. An influential point is far horizontally from the others
  • C. A point can be influential without being an outlier
  • B. Every influential point has a large residual

Survives elimination: B

Why: The survivor is false, and backwards in an instructive way. An influential point pulls the line toward itself, so its residual is often SMALL precisely because of the influence it exerted.

13. How is influence tested?

Prediction

Commit before reasoning.

Predict first

The book's suggested test for an influential point is what?

  • Remove it and see whether the slope changes significantly
  • Compute its residual
  • Check whether it exceeds two standard deviations
  • Test its correlation

Correct: Remove it and refit.

Why: The book says to begin identifying an influential point by removing it from the data set and seeing if the slope of the regression line is changed significantly. There is no formula given; the test is direct and requires refitting.

14. How much did removal move the slope?

Estimation

The full data gives a slope of 4.83; removing (75, 198) gives 3.46.

Predict first

Roughly what proportional change is that?

  • About 28 percent
  • About 5 percent
  • About 50 percent
  • About 3 percent

Correct: About 28 percent.

Why: The drop of 1.37 against 4.83 is 28 percent, from a point whose residual was under a third of the outlier threshold. That is a substantial change from one observation in eleven, and no residual-based check would have surfaced it.

15. The residual standard deviation

Section

Section 2

16. The root of SSE over n minus two

Concept

The standard deviation of the residuals is calculated from the SSE as the square root of the SSE divided by n minus two. We divide by n minus two because the regression model involves two estimates.

s, the residual standard deviation — A measure of the typical vertical miss. The calculator reports it as s in the LinRegTTest output, and it is the quantity the outlier rule is stated in terms of.

\[ s = \sqrt{\frac{\text{SSE}}{n-2}} \]

The divisor is the same accounting section 12.4 used for its degrees of freedom. A regression estimates two quantities from the data, an intercept and a slope, so two degrees of freedom are spent and n minus two remain. That is the same principle behind chapter 8's n minus one for a single mean, applied to a model with one more parameter.

Figure (svg): A card reconciling the three values the book gives for the residual standard deviation

Where a book's own numbers disagree, the honest response is to show all of them and check whether anything turns on the difference.

OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 638-639 — the formula for s and the note on dividing by n minus two

17. Three values in the book

Picture it

For the same quantity, in the same example.

Figure (svg): A card reconciling the three values the book gives for the residual standard deviation

Where a book's own numbers disagree, the honest response is to show all of them and check whether anything turns on the difference.

The first is the calculator's exact value and the second comes from the book's hand computation on rounded residuals. The third appears only as a threshold in one sentence and matches neither, but since the largest residual is 34.7 and the next is 19.1, all three flag the same single point.

18. Worked example: computing s from the residuals

Worked example

The book squares the residuals 35, 17, 16, 6, 19, 9, 3, 1, 10, 9 and 1, and adds them.

\[ \text{SSE} = 2440, \; n = 11 \]

The divisor

Why: n minus two.

\[ 9 \]

Divide

Why: 2440 over 9.

\[ 271.1 \]

Take the root

Why: Of 271.1.

\[ 16.47 \]

Double it

Why: The threshold.

\[ 32.94 \]

Figure (svg): The solution to Worked example computing s from the residuals shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ s = \sqrt{\frac{2440}{9}} = 16.47 \]

Verify: confirm against the calculator's value and account for the gap

Why: LinRegTTest reports s = 16.412, from an exact SSE of 2424.30 rather than 2440. The difference comes entirely from the book's hand computation using residuals rounded to whole numbers, which inflates the SSE by about 16. Both give thresholds near 33, and since the second-largest residual is 19.1, no point sits between them — the conclusion is identical either way.

OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, p. 639

19. Compute the threshold

Faded example

A regression on 12 points has SSE = 900.

Fill in the blanks

s = \sqrt1018.97}} = 9.49, \qquad 2s = ___

Why: Any point whose residual exceeds about 19 in magnitude would be flagged as a potential outlier under the book's rough rule.

20. Worked example: why n minus two

Worked example

Comparing with the divisors used earlier in the book.

\[ n-1 \text{ against } n-2 \]

A single mean

Why: One estimate.

\[ \text{divide by } n - 1 \]

A regression line

Why: Intercept and slope.

So the divisor

Why: Two fewer.

\[ n - 2 \]

Check against section 12.4

Why: Same count.

\[ d f = n - 2 \]

Figure (svg): The solution to Worked example why n minus two shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{df} = n - 2 = 9 \]

Verify: confirm the two uses of n minus two are the same quantity

Why: Section 12.4's significance test uses n minus two degrees of freedom for exactly this reason, and this section uses n minus two in the divisor for the same one. They are not two coincidental rules but one accounting applied twice: a fitted line costs two degrees of freedom, wherever that cost shows up.

OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, p. 638

21. Error analysis: four calculations of s

Error analysis

With SSE = 2424.30 and n = 11, which are correct?

Annotate

On: \( \begin{aligned} &(1)\; s = \sqrt{2424.30/9} = 16.41 \\ &(2)\; s = \sqrt{2424.30/11} = 14.85 \\ &(3)\; s = \sqrt{2424.30/10} = 15.57 \\ &(4)\; s = 2424.30/9 = 269.4 \end{aligned} \)

  • (1) is correct: the divisor is n minus two, and the calculator reports 16.412.
  • (2) divides by n, which no version of this formula uses.
  • (3) divides by n minus one, the rule for a single mean rather than for a fitted line.
  • (4) omits the square root, giving a variance rather than a standard deviation.

Error (3) is the most tempting because n minus one is so familiar, and it understates s by about 5 percent — which lowers the outlier threshold and would flag points the correct rule does not.

22. One of these is false

Two truths and a lie

All three concern s.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The divisor is n minus two
  • C. It measures the typical vertical miss
  • B. It is the standard deviation of the y values

Survives elimination: B

Why: The survivor is false. The standard deviation of the y values here is 20.80, describing how much final scores vary overall; s is 16.41, describing how much they vary about the LINE. The second is smaller precisely because the line explains some of the variation.

23. Why is s smaller than the spread of y?

Prediction

Commit before reasoning.

Predict first

For the exam data, sy is 20.80 and s is 16.41. Why is s smaller?

  • The line explains part of y's variation, leaving less as scatter about it
  • Because of the division by n minus two
  • It is a coincidence
  • Because the residuals are squared

Correct: The line explains part of the variation.

Why: This is r-squared in a different guise. The line accounts for about 44 percent of y's variation, so what remains scattered around it is smaller than the total spread. For a perfect fit s would be zero, and for a useless one it would approach the spread of y itself.

24. How much did rounding inflate SSE?

Estimation

The exact SSE is 2424.30; the book's hand computation gives 2440.

Predict first

By roughly what percentage?

  • Under 1 percent
  • About 10 percent
  • About 25 percent
  • About 50 percent

Correct: Under 1 percent.

Why: The gap of 15.7 against 2424.3 is about 0.6 percent, which moves s from 16.412 to 16.47 — a difference of 0.06. Rounding eleven residuals to whole numbers costs remarkably little here, which is why the book can compute by hand without changing any conclusion.

25. Flagging outliers

Section

Section 3

26. Two standard deviations, graphically or numerically

Concept

As a rough rule of thumb, flag any point located further than two standard deviations above or below the best-fit line. This can be done visually by drawing an extra pair of lines two standard deviations above and below the best-fit line, or numerically by calculating each residual and comparing it to twice the standard deviation.

the rule of thumb — The book's own description. It is a screening device rather than a test, and it produces candidates for examination rather than verdicts.

\[ |y - \hat{y}| \ge 2s \;\Longrightarrow\; \text{flag} \]

The two methods are the same rule seen two ways. The extra lines are the fitted line shifted vertically by two standard deviations, so a point outside them is exactly a point whose residual exceeds that amount. The book notes that on a calculator screen the graphical version can be hard to read when a point is close to a band — which is why the numerical version exists as a check.

Figure (svg): A scatter plot with the fitted line and two dashed parallel lines two residual standard deviations above and below, one point falling outside them

The book's graphical method: any point outside the two extra lines is flagged as a potential outlier.

OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 636-639 — the graphical and numerical identification methods

27. The bands drawn

Picture it

Two standard deviations above and below the fitted line.

Figure (svg): A scatter plot with the fitted line and two dashed parallel lines two residual standard deviations above and below, one point falling outside them

The book's graphical method: any point outside the two extra lines is flagged as a potential outlier.

The flagged point is only just outside, which is the book's own observation: on the calculator screen it is just barely outside these lines. When a point is that close, the numerical comparison settles it.

28. Worked example: the graphical method

Worked example

The book draws Y2 and Y3 alongside the best-fit line.

\[ s = 16.4 \]

The fitted line

Why: As Y1.

\[ -173.5 + 4.83 x \]

Two standard deviations

Why: Twice 16.4.

\[ 32.8 \]

The lower band

Why: Subtract it.

\[ Y 2 = -173.5 + 4.83 x - 32.8 \]

The upper band

Why: Add it.

\[ Y 3 = -173.5 + 4.83 x + 32.8 \]

Figure (svg): The solution to Worked example the graphical method shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ Y_2, Y_3 = -173.5 + 4.83x \mp 2(16.4) \]

Verify: confirm the bands are parallel and why that matters

Why: Both have the same slope as the best-fit line, differing only in intercept — so the band has constant width across the whole range. That is consistent with section 12.4's third assumption, that the standard deviations of the y values about the line are equal for each value of x. A residuals plot that fanned outward would violate it, and the constant-width band would then be the wrong screening tool.

OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, p. 637

29. Apply the rule

Faded example

A regression has s = 12, and a point has a residual of -26.

Fill in the blanks

2s = 24, \textexceeds |___| = 26 ___ \text___

Why: The point is flagged. The sign is irrelevant to the rule — the comparison uses the magnitude, so points below the line are flagged on the same terms as points above it.

30. Worked example: the numerical method

Worked example

Comparing every residual against twice the standard deviation.

\[ 2s = 32.8 \]

The largest residual

Why: At (65, 175).

\[ 34.73 \]

Compare

Why: Exceeds 32.8.

The next largest

Why: At (66, 126).

\[ 19.09 \]

Compare

Why: Well below.

Figure (svg): The solution to Worked example the numerical method shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 34.73 > 32.8 > 19.09 \]

Verify: confirm the verdict is robust to which threshold is used

Why: The book's three candidate thresholds are 31.29, 32.8 and 32.94, and the two largest residuals are 34.7 and 19.1. Every threshold falls in the wide gap between them, so all three identify exactly the same point. When a screening rule's exact value is uncertain, checking whether the answer is sensitive to it is more useful than settling which value is right.

OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 639-640

31. Trap: treating the rule as a test

Trap

The trap

\[ |\varepsilon| > 2s \;\Rightarrow\; \text{the point is wrong; delete it} \]

Read a flag as a verdict

Why: The rule produced a definite answer.

\[ \text{but the book calls it a rough rule of thumb} \]

It flags candidates for examination, not points for removal.

The fix

\[ |\varepsilon| > 2s \;\Rightarrow\; \text{examine this observation} \]

Investigate the flagged point before doing anything with it

Why: The key is to examine what causes it to be an outlier.

The book is explicit that an outlier may hold valuable information about the population and should remain included. In a sample of eleven, roughly one point would be expected beyond two standard deviations by chance alone under a normal model — so a single flag is close to what ordinary variation produces, and is not on its own evidence of anything wrong.

32. One of these is false

Two truths and a lie

All three concern flagging.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The graphical and numerical methods are the same rule
  • C. The rule uses the magnitude, so sign does not matter
  • B. A flagged point should be deleted from the analysis

Survives elimination: B

Why: The survivor is false and reverses the book's guidance. A flagged point should be examined; it may be erroneous, or it may hold valuable information about the population and should remain included. The book deletes one only as a learning exercise, after supposing it was found to be an error.

33. Why draw the bands parallel?

Prediction

Commit before reasoning.

Predict first

Why do Y2 and Y3 have the same slope as the fitted line?

  • The threshold is a constant vertical distance, matching the equal-spread assumption
  • To make them easier to graph
  • Because the slope is positive
  • So they meet at the ends

Correct: The threshold is a constant vertical distance.

Why: Two standard deviations is the same distance at every x, which is exactly section 12.4's third assumption — that the spread of y about the line is equal for each x. Bands of constant width encode that assumption, and if it failed the screening rule would be too strict at some x values and too lenient at others.

34. How many flags should be expected?

Estimation

Under the test's assumptions, residuals are roughly normal about the line.

Predict first

In a sample of 11, how many points would be expected beyond two standard deviations by chance?

  • About half a point, so one is unremarkable
  • About three
  • None ever
  • About five

Correct: About half a point.

Why: A normal distribution puts about 5 percent of its values beyond two standard deviations, and 5 percent of eleven is 0.55. So finding one flagged point in eleven is entirely ordinary and is not by itself evidence that anything is wrong — which is precisely why the book calls this a rough rule of thumb and asks for examination rather than action.

35. What to do with a flagged point

Section

Section 4

36. Examine, correct, keep or record

Concept

Outliers need to be examined closely. Sometimes they should not be included in the analysis of the data — it is possible that an outlier is a result of erroneous data. Other times an outlier may hold valuable information about the population under study and should remain included. The key is to examine carefully what causes a data point to be an outlier.

the reporting requirement — When outliers are deleted, the researcher should either record that data was deleted and why, or should provide results both with and without the deleted data.

\[ \text{examine} \to \text{correct, keep, or delete with a record} \]

The book adds a third option that is easy to overlook: if the data are erroneous and the correct values are known, the correction can simply be made — a student who actually scored 70 rather than 65 should have their value fixed rather than removed. Deletion is the last resort among three responses, not the default one.

Figure (svg): Handling a potential outlier

The book's own guidance, and its final note on reporting is the part most often skipped.

OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 636-640 — what outliers need, and the note on recording deletions

37. The procedure

Picture it

Six steps, ending in a reporting obligation.

Figure (svg): Handling a potential outlier

The book's own guidance, and its final note on reporting is the part most often skipped.

The final step is the one most often skipped and the one that makes an analysis checkable. A reader who is told which point was removed and why can judge the decision; a reader given only the final line cannot.

38. Worked example: the book's own decision

Worked example

Having flagged (65, 175), what does the book do?

\[ (65, 175) \text{ flagged} \]

First response

Why: Re-examine the data.

If an error

Why: Fix or delete.

If correct

Why: Leave it in.

The book's course

Why: Supposes an error.

Figure (svg): The solution to Worked example the book's own decision shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{delete, having recorded why} \]

Verify: confirm the book flags its own artificiality

Why: It writes that for this problem we will suppose that we examined the data and found that this outlier data was an error, and adds a parenthetical reminder that we do not always delete an outlier. The supposition is stated rather than assumed, which is exactly the transparency its own reporting note asks for.

OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 639-640

39. Consistent with the book's guidance?

Sorting

Each is a response to a flagged point.

Sort into buckets

Sort by whether the book supports it.

Supported
re-examine the observation for an error; correct it if the true value is known; keep it, since the data are correct; delete it and report both lines
Not supported
delete it because it lowers r
yes
The book names this response explicitly.
no
Improving the fit is not a reason to remove an observation.

Item (e) is supported because of the second half: the book permits deletion provided the results are given both with and without the removed data, or the deletion is recorded with its reason.

40. Worked example: applying the rule to Try It 12.13

Worked example

The points (1,5), (2,7), (2,6), (3,9), (4,12), (4,13), (5,18), (6,19), (7,12) and (7,21).

\[ n = 10 \]

Fit the line

Why: Least squares.

\[ 2.897 + 2.269 x \]

The residual standard deviation

Why: Root of SSE over 8.

\[ 3.063 \]

The threshold

Why: Twice it.

\[ 6.13 \]

The largest residual

Why: At (7, 12).

\[ -6.78:\text{ flagged} \]

Refit without it

Why: Nine points.

\[ 1.035 + 2.961 x \]

Figure (svg): The solution to Worked example applying the rule to Try It 12.13 shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \hat{y}(10) = 1.035 + 2.961(10) = 30.65 \]

Verify: confirm the flagged point is the one the picture suggests

Why: At x = 7 the other observation is 21 while this one is 12, so two points at the same x differ by nine — and the line, drawn through a generally rising pattern, passes well above 12. The residual of -6.78 exceeds the threshold of 6.13, confirming numerically what the plot shows. Note that the prediction at x = 10 is extrapolation, since the observed x values stop at 7.

OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, p. 640

41. Error analysis: four responses to a flagged point

Error analysis

Which are consistent with the book's guidance?

Annotate

On: \( \begin{aligned} &(1)\; \text{examine the observation for a recording error} \\ &(2)\; \text{delete it, since it worsens the fit} \\ &(3)\; \text{correct it, if the true value is known} \\ &(4)\; \text{delete it and report only the new line} \end{aligned} \)

  • (1) is the book's first instruction: outliers need to be examined closely.
  • (2) is exactly the wrong reason. Improving the fit is not evidence that a point is wrong, and deleting on that basis guarantees a better-looking result from any data at all.
  • (3) is the book's own suggestion, using the example of a student who actually scored 70 instead of 65.
  • (4) violates the reporting note, which requires recording that data was deleted and why, or giving results both ways.

Errors (2) and (4) compound each other. Deleting whatever fits poorly and reporting only the survivor produces a correlation that describes the deletion process rather than the population — and leaves a reader no way to detect it.

42. One of these is false

Two truths and a lie

All three concern handling.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. An outlier may hold valuable information and should remain
  • C. A deletion should be recorded with its reason
  • B. Outliers should be removed before fitting a line

Survives elimination: B

Why: The survivor is false on two counts. Outliers cannot be identified before fitting, since the rule uses residuals from the fitted line — and even once identified, the book's first instruction is to examine rather than to remove.

43. The reporting rule

Faded example

A researcher removes a flagged point from an analysis.

Fill in the blanks

\textwhy both ways, \text___ ___

Why: Either satisfies the requirement, and both leave a reader able to judge the decision rather than having to accept it.

44. Explain the danger

Explain it

A classmate removes every point with a large residual, one at a time, until the correlation exceeds 0.95.

Discussion prompt

In two sentences or fewer, say what is wrong.

Hint: Ask what that procedure would do to unrelated data.

Answer:

Removing whatever fits poorly will raise the correlation of ANY dataset, including one with no relationship at all, so the resulting number measures the deletion process rather than the variables.

The book's rule is to examine a flagged point for a reason it might be wrong, not to delete points until the fit looks good.

45. What deletion changes

Section

Section 5

46. A new line, a stronger correlation, a different prediction

Concept

Computing a new best-fit line using the ten remaining points gives a slope of 7.39 and an intercept of -355.19, with r equal to 0.9121. The new line is a stronger correlation than the original because 0.9121 is closer to one, so it fits the ten remaining data values better.

what moves — Everything: the equation, the correlation, the coefficient of determination, the SSE and every prediction made from the line.

\[ \hat{y} = -355.19 + 7.39x, \qquad r = 0.9121 \]

The correlation rising is not evidence that removing the point was right. Any least-squares fit improves when its worst-fitting observation is discarded, so a higher r is guaranteed by the procedure rather than earned by it. The justification for a deletion has to come from examining the observation, never from what happens to the fit afterwards.

Figure (svg): A table comparing the regression before and after the outlier is removed

One point in eleven, and every number moves substantially — which is why outliers have to be examined rather than ignored.

OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 639-640 — the new line and correlation after deletion

47. Before and after

Picture it

The same data, one point apart.

Figure (svg): A table comparing the regression before and after the outlier is removed

One point in eleven, and every number moves substantially — which is why outliers have to be examined rather than ignored.

The SSE falls by more than two thirds and r-squared nearly doubles, from one observation in eleven. That magnitude is the argument for the reporting requirement: a reader shown only the second column would form a very different impression of the relationship.

48. Worked example: Example 12.13, the new prediction

Worked example

Using the new line based on the remaining ten points, what would a student who scores 73 on the third exam expect on the final?

\[ \hat{y} = -355.19 + 7.39x, \; x = 73 \]

Multiply

Why: 7.39 times 73.

\[ 539.47 \]

Add the intercept

Why: Subtract 355.19.

\[ 184.28 \]

The original prediction

Why: From section 12.5.

\[ 179.08 \]

The difference

Why: Subtract.

\[ 5.2\text{ points} \]

Figure (svg): The solution to Worked example Example 12.13, the new prediction shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \hat{y}(73) = 184.28 \quad\text{against}\quad 179.08 \]

Verify: confirm the difference is meaningful relative to the model's precision

Why: Five points on a 200-point exam is small, but the residual standard deviation has also fallen — from 16.4 to about 9.3 on the ten remaining points — so the new line is both different and more precise. Both changes come from one deleted observation, which is why the book's instruction to report results with and without it is not a formality.

OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, p. 640

49. Predict from the new line

Faded example

The refitted line is ŷ = -355.19 + 7.39x.

Fill in the blanks

\hat517.3(70) = -355.19 + 7.39(70) = -355.19 + 162.11 = ___

Why: The original line predicted 164.59 at the same x, so the two differ by only 2.5 points near the centre of the data — much less than the 5.2-point gap at x = 73, since the lines cross near the middle and diverge toward the edges.

50. Worked example: Example 12.14, when nothing is flagged

Worked example

The Consumer Price Index data, with fourteen points and a fitted line ŷ = -3204 + 1.662x.

\[ s = 25.4 \]

The threshold

Why: Twice 25.4.

\[ 50.8 \]

Draw the bands

Why: Above and below.

\[ Y 2\text{ and } Y 3 \]

Check every point

Why: None outside.

The closest

Why: 1999.

Figure (svg): The solution to Worked example Example 12.14, when nothing is flagged shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{all } |\varepsilon| < 50.8 \]

Verify: confirm the more important caution in this example

Why: The book adds a note that matters more than the outlier check: although the correlation coefficient is significant, the pattern in the scatterplot indicates that a curve would be a more appropriate model than a line. So the data pass the outlier screen while failing the linearity condition — a reminder that the checks in this chapter are independent, and passing one says nothing about the others.

OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 641-642

51. Trap: justifying a deletion by the improvement it produces

Trap

The trap

\[ r: 0.6631 \to 0.9121 \;\Rightarrow\; \text{removing the point was right} \]

Take the better fit as evidence for the decision

Why: The improvement is dramatic.

\[ \text{but removing the worst point always improves the fit} \]

The same procedure applied to entirely unrelated data would also raise r, so the improvement proves nothing.

The fix

\[ \text{delete only on evidence about the OBSERVATION} \]

Justify from examination of the data point, then report both

Why: The book supposes an error was found, and says so.

The asymmetry is worth stating plainly: a deletion can be justified by finding a recording error, a measurement fault or a subject who does not belong to the population — all facts about the observation. It can never be justified by the effect on r, because that effect is guaranteed in advance.

52. One of these is false

Two truths and a lie

All three concern the effect of deletion.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The correlation rose from 0.6631 to 0.9121
  • C. Every prediction from the line changed
  • B. The rise in r shows the deletion was justified

Survives elimination: B

Why: The survivor is false. Removing the worst-fitting point raises r in any dataset whatsoever, so the improvement carries no information about whether the point deserved removal. That justification has to come from examining the observation itself.

53. How much did the SSE fall?

Estimation

From 2424.3 on eleven points to 688.2 on ten.

Predict first

Roughly what fraction of the SSE did that one point contribute?

  • About 70 percent
  • About 10 percent
  • About 30 percent
  • About 95 percent

Correct: About 70 percent.

Why: The drop of 1736 against 2424 is 72 percent, from one observation in eleven. Its own squared residual was 1206, and removing it also let the line shift to fit the rest better — so a single point accounted for the majority of the total squared miss.

54. What does Example 12.14 warn about?

Prediction

Commit before reasoning.

Predict first

The CPI data has a significant correlation and no outliers. What does the book still object to?

  • The pattern bends, so a curve would be more appropriate than a line
  • The sample size is too small
  • The correlation is negative
  • There are too many outliers

Correct: The pattern bends.

Why: Its note says a statistician should prefer to use other methods to fit a curve to this data, and adds the general lesson that in addition to doing the calculations, it is always important to look at the scatterplot when deciding whether a linear model is appropriate. Passing the outlier screen and the significance test does not make a line the right model.

55. Outliers against influential points

Comparison

Fill the blanks. Two different properties, and two different checks.

Comparison matrix

OutlierInfluential point
Far in which directionvertically, from the linehorizontally, from the other points
How it is detecteda residual beyond two standard deviationsremove it and refit
What it threatensthe accuracy of one observationthe slope of the whole line
Its own residuallarge, by definitionoften small, because the line follows it

The last row is why a residual check alone is not enough. An influential point pulls the line toward itself, so the very leverage that makes it matter also hides it from the outlier rule.

56. Checking the points a line was fitted to, in order

Pattern

Six steps, and the last is a reporting obligation rather than a calculation.

  1. Fit the line and compute the residual standard deviation as the root of SSE over n minus two.
  2. Flag any point whose residual exceeds two of those in magnitude, graphically or numerically.
  3. Separately, look for points whose x values are far from the rest, and remove each to see whether the slope shifts.
  4. Examine every flagged or influential observation for a recording or measurement error.
  5. Correct it if the true value is known; otherwise keep it unless there is a reason to remove it.
  6. If a point is removed, record which and why, or report the results both with and without it.

A better fit after deletion is guaranteed by the procedure and is never a justification for it.

OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 636-643

57. Check yourself 1 of 3

Check

The two kinds.

Check your understanding

A point has a small residual but its x value is far from every other. What is it?

  • A. Possibly influential, but not an outlier (correct)
  • B. An outlier
  • C. Both an outlier and influential
  • D. Neither

Answer: A

Why: An outlier is far from the line vertically, which a small residual rules out. Being far horizontally is what makes a point influential, and the test is to remove it and see whether the slope changes.

Why B tempts people
An outlier by definition has a large residual.
Why C tempts people
The small residual rules out the outlier half.
Why D tempts people
A distant x value is exactly the influential-point condition.

58. Check yourself 2 of 3

Check

The standard deviation.

Check your understanding

A regression on 10 points has SSE = 512. What is the residual standard deviation?

  • A. 8 (correct)
  • B. 7.16
  • C. 64
  • D. 51.2

Answer: A

Why: The divisor is n minus two, so 512 over 8 is 64, whose square root is 8.

Why B tempts people
That divides by n rather than n minus two.
Why C tempts people
That is the variance, before taking the square root.
Why D tempts people
That divides by n and omits the square root.

59. Check yourself 3 of 3

Check

Handling.

Check your understanding

A flagged point is examined and the data are found to be correct. What should be done?

  • A. Leave it in: it may hold valuable information about the population (correct)
  • B. Delete it, since it was flagged
  • C. Delete it and report only the new line
  • D. Replace it with the predicted value

Answer: A

Why: The book says an outlier may hold valuable information about the population under study and should remain included. Being unusual is not a reason for removal.

Why B tempts people
The flag identifies a candidate for examination, not a point for deletion.
Why C tempts people
That compounds an unjustified deletion with a reporting failure.
Why D tempts people
Replacing an observation with the model's own prediction invents data and guarantees a better fit.

60. Where this shows up outside the textbook

Real world

A lab technician measures reaction rate against temperature at 20, 25, 30, 35, 40 and 95 degrees, fits a line, and reports a correlation of 0.97. A colleague notes that the 95-degree run had a small residual and was clearly consistent with the others, so nothing needs checking.

Discussion prompt

Assess the colleague's reasoning and say what should be done.

Hint: Ask where the 95-degree point sits relative to the other temperatures.

Answer:

The colleague has applied the outlier check and missed the influential-point one. A small residual rules out the point being an outlier and says nothing about influence — and this point's x value is 55 degrees beyond the next highest, while the other five span only 20 degrees in total.

The small residual is itself a symptom rather than reassurance. A point that far from the others exerts enormous leverage: the line is pulled toward it, so it ends up fitting closely BECAUSE it dominated the fit. The correlation of 0.97 is largely a statement that a five-point cluster and one distant point can be joined by a line, which two well-separated groups almost always can.

\[ \text{other } x: [20, 40]; \quad \text{this } x = 95, \; 55 \text{ beyond the next} \]

The book's own test applies directly: remove the 95-degree run and refit. If the slope changes substantially, the point is influential and the reported line describes it more than it describes the other five. That is a two-minute check and the colleague's argument gives no reason to skip it.

What should be done: refit without the 95-degree point and report both lines; collect runs at intermediate temperatures — 50, 65, 80 — so the gap is filled and the fit no longer rests on one observation; and check whether the reaction's behaviour is even linear across so wide a range, since many rates rise exponentially with temperature. The correlation is high; what it is high ABOUT is the open question.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

Why can an influential point have a small residual?

  • Because it was measured accurately
  • Because it pulls the line toward itself, so the fitted line ends up close to it
  • Because small residuals cause influence
  • It cannot: influential points always have large residuals

Correct: It pulls the line toward itself.

\[ (75, 198): \; \varepsilon = 9.46 \text{ and yet } b: 4.83 \to 3.46 \text{ on removal} \]

Why: A least-squares fit minimises squared vertical misses, and a miss at an x far from the others is most cheaply reduced by tilting the whole line. So the point gets fitted closely, and its residual comes out small precisely because of the influence it exerted. The residual measures the outcome of the fit rather than the point's importance to it, which is why influence needs its own check.

62. Explain it to someone a year behind you

Explain it

They deleted three points with the largest residuals, refitted, and report a correlation of 0.96 as their result.

Discussion prompt

In two sentences or fewer, say what is wrong.

Hint: Ask what that procedure does to data with no relationship.

Answer:

Removing the worst-fitting points raises the correlation of any dataset at all, so 0.96 describes the deletion procedure rather than the two variables.

A point may be removed only after examining it and finding a reason it is wrong, and the deletion has to be recorded with its reason or the results given both ways.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Distinguishing an outlier from an influential point
  • Computing the residual standard deviation and the two-standard-deviation threshold
  • Saying what should and should not be done with a flagged point
  • Explaining why a better fit after deletion proves nothing

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the first, far vertically against far horizontally. For the second, root of SSE over n minus two, then double it. For the third, examine, then correct or keep; delete only with a recorded reason. For the fourth, removing the worst point improves any fit. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Fifteen minutes.

Draw it

At the top, write the two definitions side by side and box the words VERTICALLY and HORIZONTALLY: an outlier is far from the line vertically, an influential point is far from the other points horizontally. Beneath each, name the exam data's example — (65, 175) with residual 34.7 for the first, and (75, 198) with residual only 9.5 for the second — and note that removing the second changes the slope from 4.83 to 3.46. In the middle, plot the eleven points with the fitted line and two dashed parallel bands 32.8 above and below it, circling the single point that falls outside. Beside the plot, write s equals the root of SSE over n minus two, work it as the root of 2440 over 9 giving 16.47, and note the calculator's 16.412 from the exact SSE of 2424.3. At the bottom left, write the six-step procedure ending with the reporting requirement. At the bottom right, make a two-column before-and-after table: the equation, r, r-squared and the prediction at 73, showing 0.6631 becoming 0.9121 and 179.08 becoming 184.28 — and write one sentence saying that the rise in r is guaranteed by removing the worst point and therefore justifies nothing.

Check your bands by confirming they are parallel to the fitted line, since the threshold is the same vertical distance at every x. Check your before-and-after table by confirming every single row changed, which is the argument for reporting both.

65. What you can do now

Recap

Six things, and the distinction in the first governs the rest.

If you seeThen
A residual beyond two standard deviationsFlag it and examine the observation
An x far from all the othersRemove it and refit to test for influence
A small residual at a distant xSuspect influence, not accuracy
SSE and ns is the root of SSE over n minus two
A flagged point with correct dataKeep it: it may be informative
A flagged point with a known correct valueCorrect it rather than deleting it
A deletionRecord it with its reason, or report both ways
A higher r after deleting a pointExpected, and no justification at all

That completes chapter 12, and with it the chapter's full sequence: look at the data, fit a line, test whether it means anything, predict within its range, and check which points it was answering to. Chapter 13 turns to comparing several population means at once, which needs a distribution the course has not yet met.

OpenStax Introductory Statistics 2e, §12.6 Outliers §12.6, pp. 636-643 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §12.6 Outliers — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 636-643

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108