The section that comes before any arithmetic, and says so: before taking up linear regression and correlation, the relation between two variables has to be looked at. A scatter plot is the most common and easiest way to display it, and three things are read off one — the direction of the relationship, its strength, and the overall pattern together with any deviations from it. A clear direction means either that high values of one variable occur with high values of the other, or that high values of one occur with low values of the other. Strength is judged by how close the points fall to a line or some other curve, with one exception the book flags: points lying exactly on a horizontal line look like a perfect fit and in fact show no relationship, since y is the same whatever x does. The section closes with the condition for going further, which is that one variable must actually help explain or predict the other.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 12 — Linear Regression and Correlation
Scatter Plots
Objectives
Five outcomes, and the last is a judgement rather than a reading.
OpenStax Introductory Statistics 2e, §12.2 Scatter Plots §12.2, pp. 620-623 — the section these objectives are drawn from
Warm-up
Chapter 2 displayed one variable at a time; section 12.1 wrote an exact line.
Discussion prompt
You have four children's ages and vocabulary sizes: (3, 655), (4, 1098), (6, 2463) and (7, 3195). What is the first thing to do with them?
Hint: Section 12.1's methods assumed an exact rule. Do these follow one?
Answer:
Plot them. Section 12.1 could write an equation straight from a price list because the rule was known in advance; here there is no rule, only four measured pairs, and nothing can be assumed about how they relate until they have been looked at.
The book puts it directly: before we take up the discussion of linear regression and correlation, we need to examine a way to display the relation between two variables x and y. The most common and easiest way is a scatter plot.
Plotted, these four rise steadily and lie close to a straight line — so a line is a reasonable thing to fit, which is what section 12.3 will do. Had they bent sharply or scattered without pattern, that decision would have gone the other way, and no amount of arithmetic would have revealed it.
Concept
A scatter plot displays the relation between two variables by plotting each observation as a point. It shows the direction of a relationship, allows the strength to be judged by how close the points lie to a line or curve, and reveals the overall pattern along with any deviations from it.
a scatter plot — One point per observation, with the independent variable x on the horizontal axis and the dependent variable y on the vertical. The most common and easiest way to display a relation between two variables.
\[ \text{each observation } (x_i, y_i) \text{ as one point} \]
The book's instruction is worth quoting in full: when you look at a scatterplot, you want to notice the overall pattern and any deviations from the pattern. Both halves matter. The pattern decides whether a line is the right model; the deviations are the points section 12.6 will return to as outliers and influential points.
Figure (svg): Four scatter plots showing a positive trend, a negative trend, no pattern, and points lying on a horizontal line
OpenStax Introductory Statistics 2e, §12.2 Scatter Plots §12.2, pp. 620-622
Section
Section 1
Concept
A scatter plot shows the direction of a relationship between the variables. A clear direction happens when there is either high values of one variable occurring with high values of the other and low with low, or high values of one variable occurring with low values of the other.
direction — Positive when the two variables move together, negative when one rises as the other falls. A plot may show neither, in which case there is no clear direction.
\[ \text{high-high} \;\nearrow\; \qquad \text{high-low} \;\searrow\; \]
Direction is the easiest thing to read and the first thing to record, because it fixes the sign of everything that follows. Section 12.3 will note that the sign of the correlation coefficient r is the same as the sign of the slope b, so a plot sloping downward guarantees both will be negative — which makes the picture a check on the arithmetic rather than a decoration.
Figure (svg): Four scatter plots showing a positive trend, a negative trend, no pattern, and points lying on a horizontal line
OpenStax Introductory Statistics 2e, §12.2 Scatter Plots §12.2, p. 622 — the two kinds of clear direction
Picture it
Two directions, one with no direction, and the exception.
Figure (svg): Four scatter plots showing a positive trend, a negative trend, no pattern, and points lying on a horizontal line
The first two are the book's two cases. The third has no clear direction at all, and the fourth is the horizontal case the book singles out as a warning.
Worked example
An educational researcher records four children's ages and vocabulary sizes: (3, 655), (4, 1098), (6, 2463) and (7, 3195). Is there a relationship?
\[ x = \text{age}, \; y = \text{words} \]
Plot the four points
Why: Age across, words up.
Compare the smallest x
Why: Age 3.
Compare the largest x
Why: Age 7.
Name the direction
Why: High with high.
Figure (svg): The solution to Worked example Example 12.5, age and vocabulary shown as a ladder of expressions, one row per legal move
\[ \text{high } x \text{ with high } y \]
Verify: confirm the direction holds across every pair, not just the extremes
Why: Reading the four y-values in order of x gives 655, 1098, 2463, 3195 — strictly increasing, so every pair of points agrees on the direction. A single reversal would not overturn a positive direction, but no reversals at all is stronger evidence than comparing only the first and last points, which two unusual observations could easily mislead.
OpenStax Introductory Statistics 2e, §12.2 Scatter Plots §12.2, pp. 620-621
Sorting
Each describes a pair of variables.
Sort into buckets
Sort by the direction you would expect.
Predicting the direction before plotting is worth doing every time: when the plot disagrees, either the expectation was wrong or the data were entered wrongly, and both are worth knowing.
Worked example
SCUBA divers have maximum dive times: (50, 80), (60, 55), (70, 45), (80, 35), (90, 25) and (100, 22), with depth in feet.
\[ x = \text{depth}, \; y = \text{minutes} \]
Read the smallest depth
Why: 50 feet.
\[ 80\text{ minutes} \]
Read the largest depth
Why: 100 feet.
\[ 22\text{ minutes} \]
Check the middle
Why: Steadily falling.
Name the direction
Why: High with low.
Figure (svg): The solution to Worked example Try It 12.6, depth and dive time shown as a ladder of expressions, one row per legal move
\[ \text{high } x \text{ with low } y \]
Verify: confirm the direction is what physical sense predicts
Why: Deeper water means faster gas consumption and greater decompression obligation, so shorter permitted times are exactly what should be expected. When a plot's direction contradicts what the situation predicts, the first suspicion should be that the two columns have been swapped or entered wrongly — the picture is a check on the data as well as on the arithmetic.
OpenStax Introductory Statistics 2e, §12.2 Scatter Plots §12.2, p. 622
Trap
\[ \text{first point low, last point high} \;\Rightarrow\; \text{positive} \]
Compare only the extremes
Why: They are the easiest to see.
\[ \text{but the middle may do anything} \]
Points that rise, fall and rise again can share their endpoints with a straight positive trend and mean something entirely different.
\[ \text{look at the whole set, then name the direction} \]
Read the overall pattern, which is what the book asks for
Why: Notice the pattern and any deviations from it.
This is why the section exists at all. A correlation coefficient computed from a set that rises then falls can come out near zero or moderately positive depending on the balance, and neither number describes the shape. Only the plot does — which is why the book puts it before any calculation.
Two truths and a lie
All three concern direction.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The book describes when a clear direction HAPPENS, which implies it may not. A cloud of points with no tilt has no direction, and neither does a set following a symmetric curve.
Prediction
Commit before reasoning.
Predict first
A plot slopes clearly downward. What can be said before computing anything?
Correct: Both the slope and r will be negative.
Why: Section 12.3 notes that the sign of r is the same as the sign of the slope b, and a downward-sloping plot means both are below zero. The magnitude is not predictable by eye, and the intercept is unconstrained — it is often positive for a downward-sloping line, as the dive-time data show.
Faded example
Depth and dive time: as depth rises from 50 to 100 feet, time falls from 80 to 22 minutes.
Fill in the blanks
\textlow x \textnegative ___ \; y \;\Rightarrow\; \text___ ___ \text___
Why: That is the second of the book's two cases, and it guarantees a negative slope when a line is fitted in section 12.3.
Section
Section 2
Concept
You can determine the strength of the relationship by looking at the scatter plot and seeing how close the points are to a line, a power function, an exponential function, or to some other type of function. For a linear relationship there is an exception: a scatter plot where all the points fall on a horizontal line provides a perfect fit, yet the horizontal line would in fact show no relationship.
the horizontal exception — Points on a horizontal line lie perfectly on a line and yet y does not depend on x at all. Closeness to a line therefore indicates strength everywhere except in this one case.
\[ \text{all } y_i \text{ equal} \;\Rightarrow\; \text{perfect fit, no relationship} \]
The exception is sharper than the book states it. If every y value is identical, the standard deviation of the y values is zero — and since r divides by that standard deviation, the correlation is not merely zero but undefined. The picture and the arithmetic agree that nothing is there, which is the right conclusion arrived at two different ways.
Figure (svg): Four scatter plots showing a positive trend, a negative trend, no pattern, and points lying on a horizontal line
OpenStax Introductory Statistics 2e, §12.2 Scatter Plots §12.2, p. 622 — strength, and the horizontal-line exception
Picture it
Three ordinary plots and the one that misleads.
Figure (svg): Four scatter plots showing a positive trend, a negative trend, no pattern, and points lying on a horizontal line
The fourth panel would score perfectly on a rule that says closer to a line means stronger. It is the one plot in the set where that rule must not be applied, which is why the book states it as an explicit exception.
Worked example
The vocabulary data and the dive-time data, both plotted.
\[ \text{which is tighter?} \]
Vocabulary
Why: Four points, near-straight.
Its correlation
Why: Computed later.
\[ 0.9971 \]
Dive times
Why: Six points, slight bend.
Its correlation
Why: Also strong.
\[ -0.9629 \]
Figure (svg): The solution to Worked example judging strength in two datasets shown as a ladder of expressions, one row per legal move
\[ r = 0.9971 \quad\text{against}\quad r = -0.9629 \]
Verify: confirm that the weaker one is weaker for a reason worth seeing
Why: The dive-time points do not just scatter around a line; they bend, falling steeply from 50 to 70 feet and then flattening. So the shortfall from a perfect correlation reflects a genuine curve rather than noise, and a straight line will systematically over- and under-predict in different parts of the range. That is exactly the kind of thing the book means by noticing the overall pattern.
OpenStax Introductory Statistics 2e, §12.2 Scatter Plots §12.2, pp. 621-623
Two truths and a lie
All three concern strength.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false because of the horizontal case. Points lying exactly on a horizontal line fit perfectly while showing that y does not depend on x at all — the one situation where a perfect fit carries no information.
Worked example
Suppose four measurements give (1, 7), (2, 7), (3, 7) and (4, 7).
\[ \text{all } y = 7 \]
Fit a line by eye
Why: Every point on it.
Ask what it predicts
Why: Seven, always.
So does x matter?
Why: No.
Try to compute r
Why: Divide by sd of y.
Figure (svg): The solution to Worked example why the horizontal case breaks the rule shown as a ladder of expressions, one row per legal move
\[ s_y = 0 \;\Longrightarrow\; r \text{ undefined} \]
Verify: confirm the two readings agree rather than conflict
Why: The picture says y never changes, so nothing about x can explain it; the arithmetic refuses to produce a correlation because there is no variation in y to explain. Both are saying the same thing. The apparent paradox — a perfect fit meaning nothing — dissolves once strength is understood as how much of y's variation the line accounts for, since here there is none to account for.
OpenStax Introductory Statistics 2e, §12.2 Scatter Plots §12.2, p. 622
Error analysis
Which are sound?
Annotate
On: \( \begin{aligned} &(1)\; \text{points close to a line indicate a strong relationship} \\ &(2)\; \text{points on a horizontal line indicate the strongest relationship} \\ &(3)\; \text{points close to a curve may indicate a strong non-linear relationship} \\ &(4)\; \text{a wide scatter indicates a weak linear relationship} \end{aligned} \)
Statement (3) matters for the rest of the chapter. A relationship can be strong and still be the wrong shape for a straight line, which is precisely the situation Example 12.14's CPI data presents.
Prediction
Commit before reasoning.
Predict first
For points all sharing the same y value, why can r not be computed?
Correct: The standard deviation of y is zero.
Why: Section 12.3 gives the slope as r times the standard deviation of y over the standard deviation of x, and r itself is built by dividing by both standard deviations. With no variation in y that division is undefined — the arithmetic declines to answer, which is the honest response to a question about explaining variation that is not there.
Sorting
Each describes a plotted pattern.
Sort into buckets
Sort by strength as a linear relationship.
Item (b) is the interesting one: the relationship is genuinely strong, but not LINEARLY strong, so a straight line is the wrong model even though something clearly connects the variables.
Estimation
Four points rising almost exactly in a line: (3, 655), (4, 1098), (6, 2463), (7, 3195).
Predict first
Roughly what correlation would you expect?
Correct: About 0.99.
Why: The actual value is 0.9971. Points that lie nearly on a rising straight line give a correlation just below one, and the direction being positive fixes the sign. Estimating r from a plot before computing it is a habit worth keeping: a computed value far from the estimate usually means an arithmetic slip.
Section
Section 3
Concept
When you look at a scatterplot, you want to notice the overall pattern and any deviations from the pattern. The pattern decides whether a straight line is the right model; the deviations are individual points that do not follow it.
the overall pattern — The shape the bulk of the points follow — straight, bending, or none. Chapter 12 fits straight lines, so a bending pattern calls for a different method rather than a worse line.
\[ \text{pattern} \;\to\; \text{which model}; \qquad \text{deviations} \;\to\; \text{which points} \]
The book returns to this with unusual force in Example 12.14, where a fitted line to the Consumer Price Index has a significant correlation and the note observes that the pattern in the scatterplot indicates that a curve would be a more appropriate model than a line. A significant r does not certify that a line is the right shape — only that the linear component of the relationship is real.
Figure (svg): Two scatter plots side by side, one following a tight upward curve and one following a tight straight line
OpenStax Introductory Statistics 2e, §12.2 Scatter Plots §12.2, pp. 622-623 — notice the overall pattern and any deviations
Picture it
One curved, one straight.
Figure (svg): Two scatter plots side by side, one following a tight upward curve and one following a tight straight line
Judging strength alone would rank these as equals. Judging the pattern separates them, and only the right-hand one is a candidate for the straight line chapter 12 knows how to fit.
Worked example
The dive-time data: (50, 80), (60, 55), (70, 45), (80, 35), (90, 25), (100, 22).
\[ \text{does it bend?} \]
Drop from 50 to 60
Why: 80 to 55.
\[ -25\text{ minutes} \]
Drop from 60 to 70
Why: 55 to 45.
\[ -10\text{ minutes} \]
Drop from 90 to 100
Why: 25 to 22.
\[ -3\text{ minutes} \]
Compare the drops
Why: Shrinking.
Figure (svg): The solution to Worked example a bending pattern shown as a ladder of expressions, one row per legal move
\[ -25, -10, -10, -10, -3 \text{ per 10 feet} \]
Verify: confirm what a straight line would do to this data
Why: A line has one constant drop per ten feet, and the fitted slope is about -1.11 minutes per foot, so -11 per ten feet. That over-predicts the drop at the deep end, where the real drop is 3, and under-predicts it at the shallow end, where it is 25. The residuals would therefore run positive, negative, then positive — a pattern rather than noise, which is exactly what section 12.3's residuals plot is designed to reveal.
OpenStax Introductory Statistics 2e, §12.2 Scatter Plots §12.2, pp. 622-623
Matching
Match each reading of a scatter plot to what it calls for.
Match the pairs
Why: Four readings and four different responses. Only the first leads directly into section 12.3, which is why looking comes before fitting.
Worked example
Suppose the vocabulary data gained a fifth child: (5, 400).
\[ (5, 400) \text{ added} \]
The pattern
Why: Rising steadily.
\[ \text{roughly } 640\text{ words } a\text{ year} \]
Predict at age 5
Why: Between 1098 and 2463.
\[ \text{about } 1750 \]
The observed value
Why: Four hundred.
Classify it
Why: One point, not the shape.
Figure (svg): The solution to Worked example a deviation from a pattern shown as a ladder of expressions, one row per legal move
\[ \text{observed } 400 \text{ against about } 1750 \]
Verify: confirm why the distinction between the two matters
Why: A bending PATTERN means the model is wrong and a curve is needed; a single DEVIATION means the model may be right and one observation needs examining — a recording error, an unusual child, or a genuine but rare case. The responses are completely different, which is why the book asks for both readings rather than one overall impression. Section 12.6 takes up the second case in detail.
OpenStax Introductory Statistics 2e, §12.2 Scatter Plots §12.2, p. 622
Trap
\[ \text{the ends miss the line, so remove those points} \]
Read a systematic curve as a handful of deviations
Why: Both show up as large residuals.
\[ \text{but removing them leaves the bend} \]
The remaining points still curve, so new extremes will miss the new line in the same way.
\[ \text{a bending pattern needs a curve, not fewer points} \]
Distinguish a wrong model from wrong observations
Why: Residuals with a PATTERN mean the first.
The test is whether the misses are systematic. Deviations scatter on both sides at random; a bend produces residuals that run positive in one region and negative in another. Section 12.3 introduces the residuals plot precisely to make that distinction visible, since a residuals plot should appear random with no pattern.
Two truths and a lie
All three concern pattern and deviations.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and the book says so explicitly at Example 12.14: the correlation coefficient there is significant while the scatterplot's pattern indicates a curve would be more appropriate. Significance concerns whether the linear component is real, not whether the shape is right.
Faded example
A fitted line's residuals run positive at both ends and negative in the middle.
Fill in the blanks
\textpattern bends, \text___ ___
Why: Residuals that change sign in a regular way indicate a curve. Random residuals, changing sign unpredictably, indicate a line is adequate — which is what section 12.3's residuals plot is checked for.
Prediction
Commit before reasoning.
Predict first
Why does the book put scatter plots before regression rather than after?
Correct: The plot decides whether fitting a line makes sense.
Why: A least-squares line can be computed from any set of paired numbers, including ones that bend, scatter without pattern, or have no variation in y. Nothing in the arithmetic objects. Only the picture reveals that the answer would be meaningless, which is why looking has to come first.
Section
Section 4
Concept
Each observation contributes one point, with the independent variable on the horizontal axis and the dependent variable on the vertical. The book's Example 12.5 says explicitly: let x be the child's age and y be the vocabulary size.
which variable goes where — x on the horizontal axis is the explaining variable; y on the vertical is the one being explained. Swapping them changes the picture and, in section 12.3, the fitted line.
\[ (x_i, y_i) \text{ plotted at } x_i \text{ across}, \; y_i \text{ up} \]
The axis assignment repeats section 12.1's point about independent and dependent variables, and it has real consequences here. A regression line predicts y from x by minimising vertical distances, so swapping the axes minimises a different set of distances and produces a different line. The correlation r, by contrast, is unaffected by the swap — which is one reason it measures association rather than prediction.
Figure (svg): A scatter plot of four points showing vocabulary size rising steadily with a child's age
OpenStax Introductory Statistics 2e, §12.2 Scatter Plots §12.2, pp. 620-621 — Example 12.5, and the axis assignment
Picture it
Four children, age across and vocabulary up.
Figure (svg): A scatter plot of four points showing vocabulary size rising steadily with a child's age
Four points is a small sample, and the book uses it because the direction is unmistakable. Section 12.4 will make the sample size matter formally, since the reliability of a linear model depends on how many observed data points there are.
Worked example
Amelia records hours practising against points scored: (5, 15), (7, 22), (9, 28), (10, 31), (11, 33) and (12, 36). She believes more practice raises her scoring.
\[ x = \text{hours}, \; y = \text{points} \]
Assign the axes
Why: Practice explains scoring.
Plot six points
Why: One per game.
Read the direction
Why: High with high.
Read the strength
Why: Very close to a line.
Figure (svg): The solution to Worked example Try It 12.5, practice and points shown as a ladder of expressions, one row per legal move
\[ r = 0.9976 \]
Verify: confirm what the plot does and does not establish
Why: It establishes that in these six games more practice went with more points, strongly and consistently. It does not establish that practising causes the improvement — Amelia might practise more in weeks when she is playing well anyway, or against weaker opponents. Section 12.3 will state the rule plainly: correlation does not imply causation, however tight the plot.
OpenStax Introductory Statistics 2e, §12.2 Scatter Plots §12.2, pp. 621-622
Faded example
Depth of a dive and maximum dive time.
Fill in the blanks
x = depth, \qquad y = time
Why: A diver chooses a depth and the permitted time follows, so depth explains time. Reversing them would ask what depth a given time implies, which is a different and less natural question.
Worked example
The vocabulary data, plotted both ways.
\[ \text{age against words, or words against age} \]
Age explains vocabulary
Why: The sensible reading.
\[ x = a g e \]
The reverse
Why: Vocabulary explains age?
Effect on r
Why: Symmetric measure.
\[ \text{unchanged at } 0.9971 \]
Effect on the line
Why: Minimises other distances.
Figure (svg): The solution to Worked example why the axes are not interchangeable shown as a ladder of expressions, one row per legal move
\[ r \text{ symmetric}; \quad \text{the line is not} \]
Verify: confirm the asymmetry with the numbers
Why: Regressing words on age gives a slope of 644.5 words per year. Regressing age on words gives a slope of about 0.0015 years per word, and one over 644.5 is 0.00155 — close but not equal, and they would be equal only if the correlation were exactly one. The two lines coincide only for a perfect fit, which is why choosing the right dependent variable matters.
OpenStax Introductory Statistics 2e, §12.2 Scatter Plots §12.2, pp. 620-623
Error analysis
Which are correct?
Annotate
On: \( \begin{aligned} &(1)\; \text{the independent variable goes on the horizontal axis} \\ &(2)\; \text{the points should be joined by line segments} \\ &(3)\; \text{each observation contributes exactly one point} \\ &(4)\; \text{repeated pairs should be entered once} \end{aligned} \)
Error (4) is easy to make and quietly damaging. The third exam data contains 71 paired with three different final scores and 69 with two — entering each pair once would change every subsequent calculation in the chapter.
Two truths and a lie
All three concern construction.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. A regression line minimises VERTICAL distances, so swapping which variable is vertical minimises a different quantity and gives a different line. The two coincide only when the correlation is exactly one or minus one.
Prediction
Commit before reasoning.
Predict first
Why does a scatter plot leave its points unconnected?
Correct: It would impose an unjustified shape.
Why: The whole purpose of the plot is to reveal what shape the points suggest. Connecting them draws a jagged path through every observation, which hides the pattern behind the noise and makes any bend look like a series of straight segments. The points are left alone so the eye can judge the pattern itself.
Estimation
Example 12.5 uses four children; the third-exam example uses eleven students.
Predict first
What does a small sample cost when reading a scatter plot?
Correct: A convincing pattern can arise by chance.
Why: Four points can lie nearly on a line even when the variables are unrelated, simply because there are few ways for four points to disagree. This is exactly why section 12.4 exists: it tests whether an observed correlation is strong enough, given the sample size, to conclude that a relationship holds in the population.
Section
Section 5
Concept
We only calculate a regression line if one of the variables helps to explain or predict the other. If x is the independent variable and y the dependent variable, then we can use a regression line to predict y for a given value of x.
the purpose condition — A judgement about the situation, not the plot. Two variables can correlate strongly without either explaining the other, in which case a regression line answers no question worth asking.
\[ x \text{ explains } y \;\Longrightarrow\; \text{fit } \hat{y} = a + bx \]
This condition is easy to skip because nothing enforces it. A least-squares line can be computed from any two columns of numbers, and it will come out looking authoritative. The classic failure is two variables that both track a third: shoe size and reading ability correlate strongly among schoolchildren, and neither explains the other — both follow age.
Figure (svg): A card giving the condition for calculating a regression line
OpenStax Introductory Statistics 2e, §12.2 Scatter Plots §12.2, p. 623 — we only calculate a regression line if one variable helps explain the other
Picture it
One case that meets it and one that does not.
Figure (svg): A card giving the condition for calculating a regression line
The right-hand case would produce a perfectly ordinary regression line with a respectable correlation. Nothing in the output would reveal that the question it answers is not one anybody asked.
Worked example
Four pairs of variables, each strongly correlated.
\[ \text{fit a line or not?} \]
Age and vocabulary
Why: Age explains words.
Third and final exam
Why: The third comes first.
Shoe size and reading score
Why: Both follow age.
Depth and dive time
Why: Depth sets the limit.
Figure (svg): The solution to Worked example applying the condition shown as a ladder of expressions, one row per legal move
\[ \text{correlation} \ne \text{explanation} \]
Verify: confirm the third case would look fine numerically
Why: Among children aged five to twelve, shoe size and reading score would produce a strong positive correlation, a significant test in section 12.4, and a line that predicts reading scores from shoe sizes with real accuracy. Every number would pass. What fails is the premise, and only knowing that both variables track age reveals it — which is why the book makes this a condition rather than a calculation.
OpenStax Introductory Statistics 2e, §12.2 Scatter Plots §12.2, p. 623
Sorting
Each pair is strongly correlated.
Sort into buckets
Sort by whether a regression line is worth fitting.
Items (b) and (d) both follow a hidden third variable — age in one case and the season in the other. Nothing in either dataset would reveal it, which is why the condition has to be applied by judgement.
Worked example
The third-exam and final-exam scores from section 12.3.
\[ \text{which is } x? \]
Which comes first in time
Why: The third exam.
Which could predict which
Why: Earlier predicts later.
Would the reverse be useful
Why: Predicting a past score.
Assign
Why: x explains y.
\[ x =\text{ third exam} \]
Figure (svg): The solution to Worked example which variable explains which shown as a ladder of expressions, one row per legal move
\[ x = \text{third exam}, \quad y = \text{final exam} \]
Verify: confirm the reasoning is about usefulness, not about arithmetic
Why: Nothing prevents regressing the third exam on the final; the numbers would work and the correlation would be the same 0.6631. But the resulting line would predict a score already known from one not yet taken, which answers no question a student or instructor has. Time order is a strong guide here, though not a universal rule — what matters is which prediction is worth making.
OpenStax Introductory Statistics 2e, §12.2 Scatter Plots §12.2, p. 623
Trap
\[ r = 0.94 \;\Rightarrow\; \text{fit the line and use it} \]
Take a high correlation as licence to model
Why: It certainly indicates a strong association.
\[ \text{but association is not explanation} \]
Two variables driven by a common third will correlate strongly while neither explains the other.
\[ \text{ask what explains what, then fit} \]
Settle the purpose from the situation before computing
Why: The condition is about the variables, not the numbers.
The practical consequence is about what the line may be used for. A line fitted to variables that merely travel together can describe and even predict within the observed range, but it will not survive an intervention: buying a child larger shoes does not improve their reading. Prediction and explanation come apart exactly here.
Two truths and a lie
All three concern the condition.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and section 12.3 states the rule as correlation does not imply causation. A strong association is compatible with one variable explaining the other, with the reverse, or with both following something else entirely.
Prediction
Commit before reasoning.
Predict first
A line is fitted to two variables that both track a hidden third. What still works, and what fails?
Correct: Prediction can work; intervention will not.
Why: If shoe size and reading score both follow age, then shoe size really does predict reading score among children — the association is real. What fails is any claim that changing x would change y, since the connection runs through age rather than between them. That is the practical content of correlation not implying causation.
Explain it
A classmate has computed a correlation of 0.94 between two variables and is about to fit a line, saying the number justifies it.
Discussion prompt
In two sentences or fewer, say what still needs checking.
Hint: Ask what question the line would answer.
Answer:
A high correlation shows the two variables travel together, but a regression line is only worth calculating if one of them helps to explain or predict the other.
If both simply follow some third variable, the line will predict accurately within the data and still support no claim about what would happen if x were changed.
Comparison
Fill the blanks. The left column is read off the picture; the right is not.
Comparison matrix
| The plot shows | The plot cannot show | |
|---|---|---|
| Direction | positive, negative or none | which variable causes which |
| Strength | how close the points lie to a curve | whether the strength is significant |
| Shape | whether the pattern bends | what a fair sample size would be |
| Purpose | nothing | whether one variable explains the other |
The bottom-left cell is the point of the last idea: purpose is not visible in a picture at all. The middle-right cell is section 12.4's subject, and the top-right is a limit no method in this chapter removes.
Pattern
Five steps, and only after all five does fitting a line make sense.
Predict the direction from the situation before plotting; a disagreement means either the expectation or the data entry is wrong.
OpenStax Introductory Business Statistics 2e, §13.1 The Correlation Coefficient r §13.1 The Correlation Coefficient r
Check
The exception.
Check your understanding
Every point in a scatter plot lies exactly on a horizontal line. What does that show?
Answer: A
Why: The book states this exception directly. A horizontal line fits perfectly and yet shows that y does not depend on x at all.
Check
Pattern.
Check your understanding
Points follow a tight upward curve. What does that indicate?
Answer: A
Why: The book lists power and exponential functions alongside lines: closeness to any curve indicates strength of that kind, but chapter 12 fits straight lines only.
Check
Purpose.
Check your understanding
Two variables correlate at 0.9, but both are driven by a third. Should a regression line be fitted?
Answer: A
Why: The book's condition is that one variable must help explain or predict the other. Here the association is real and usable for prediction, but neither variable explains the other.
Real world
A city analyst plots monthly ice cream sales against monthly drowning deaths across five years, finds the points hugging a rising line with a correlation of 0.91, and recommends restricting ice cream sales at public pools.
Discussion prompt
Assess the plot's reading and the recommendation separately.
Hint: Ask what else varies month to month.
Answer:
The plot has been read correctly. The direction is positive, the points lie close to a line, and a correlation of 0.91 is a fair summary of a tight linear pattern. Nothing about the display or the number is wrong, and that is what makes the case instructive.
The recommendation fails the section's own condition. A regression line is only worth calculating if one variable helps explain or predict the other, and here neither does: both ice cream sales and drownings rise in hot months and fall in cold ones. Temperature drives both, so the association between them is real and entirely indirect.
\[ \text{temperature} \;\nearrow\; \;\Rightarrow\; \text{sales} \;\nearrow\; \text{ and drownings} \;\nearrow\; \]
Prediction and intervention come apart here, and only one of them survives. Knowing this month's ice cream sales genuinely does help predict this month's drownings, because both encode the season. But removing the ice cream would not lower the drownings by a single case, since nothing flows from one to the other. A line that predicts well can still support no claim about what an intervention would do.
Two things would have caught it. Plotting each variable against the month, or against temperature, shows both tracking the same seasonal shape — the standard check for a common cause. And asking the section's question first, whether one variable explains the other, settles it before any correlation is computed. The scatter plot did its job; the judgement it cannot make was skipped.
Commit first
Answer, then rate your confidence honestly.
Predict first
Points in a scatter plot fall exactly on a horizontal line. How strong is the relationship?
Correct: There is no relationship.
\[ \text{all } y_i \text{ equal} \;\Rightarrow\; s_y = 0 \;\Rightarrow\; r \text{ undefined} \]
Why: This is the exception the book attaches to judging strength by closeness to a line. Everywhere else, points hugging a line means a strong relationship; here the line is horizontal, so it predicts the same y for every x, which is precisely what no relationship means. The arithmetic agrees in a stronger way: with no variation in y, the correlation is not zero but undefined.
Explain it
They have plotted six points that follow a tight upward curve and concluded a straight line will fit well because the points are so close together.
Discussion prompt
In two sentences or fewer, correct them.
Hint: Ask close to WHAT.
Answer:
Closeness indicates strength only relative to a particular shape, and these points are close to a curve rather than to a straight line.
A line fitted here would over-predict in some parts of the range and under-predict in others, which shows up as residuals with a pattern rather than random scatter.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, high-with-high or high-with-low, then closeness to a curve. For the second, the same y for every x is what no relationship means. For the third, systematic misses indicate a curve while scattered ones indicate points. For the fourth, ask whether one variable explains the other. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Twelve minutes, and most of it is drawing.
Draw it
Across the top, draw four small scatter plots side by side: one with points rising along a line, one with points falling along a line, one with a shapeless cloud, and one with every point on a horizontal line. Label the first two positive and negative, label the third no direction, and box the fourth with the words PERFECT FIT, NO RELATIONSHIP. Underneath the fourth, write that the standard deviation of y is zero there, so r is undefined rather than zero. In the middle of the page, plot Example 12.5's four points — (3, 655), (4, 1098), (6, 2463), (7, 3195) — with age across and words up, and write beside them that the direction is positive and the pattern close to linear. To their right, draw two more small plots: one following a tight curve and one following a tight line, and write beneath them that only the second is a candidate for chapter 12's straight line. At the bottom, write the five reading steps in order — axes, points, direction, strength, pattern and deviations — and then the closing condition in a box: only fit a line if one variable helps explain or predict the other.
Check your four top plots by covering the labels and asking whether each could be identified from its shape alone. Check your bottom box by naming one pair of variables that would pass the condition and one that would correlate strongly and fail it.
Recap
Five things, and the last one no picture can settle.
| If you see | Then |
|---|---|
| High x with high y throughout | A positive direction |
| High x with low y throughout | A negative direction |
| Points hugging a straight tilted line | A strong linear relationship |
| Points hugging a curve | Strong, but a line is the wrong model |
| Every point on a horizontal line | No relationship, and r is undefined |
| A shapeless cloud | No linear relationship to model |
| One point far from an otherwise clear pattern | A deviation to examine, not a change of model |
| A strong correlation with no explanatory link | Do not fit a line for explanation |
Section 12.3 fits the line. Given a plot that shows a linear pattern, it chooses the one line that makes the total squared vertical miss as small as possible, and defines the correlation coefficient that measures how well it does.
OpenStax Introductory Statistics 2e, §12.2 Scatter Plots §12.2, pp. 620-623 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.