SAT Math Type 14: Two-Variable Data and Scatterplots

The first of the long-tail Math types (3.8% of the bank), and an interpretive one: a scatterplot with a fitted model, and questions about what that model says. Covers reading the slope and intercept of a line of best fit in the story's own units, choosing between linear, quadratic and exponential models by asking whether the data bends, the difference between interpolation and extrapolation, and why correlation does not license a causal claim — with six worked examples, four traps and three checks.

Subject: SAT Prep · 61 slides · applied lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Two-Variable Data and Scatterplots

Title

SAT Math · Type 14 of 19

3.8% of the question bank — 63 of 1675 questions

2. By the end of this deck you can

Objectives

A scatterplot with a line or curve drawn through it, and questions about what the model means. The arithmetic is almost nothing. What is being tested is whether you can say what a slope represents in the units of the story, whether you know how far a model can be trusted, and whether you can resist a choice that says one thing causes another.

  1. State what the slope and the intercept of a fitted line mean in the story's units.
  2. Predict a value from a model, and say how far that prediction can be trusted.
  3. Distinguish interpolation from extrapolation and explain why one is safer.
  4. Choose between linear, quadratic and exponential models by the shape of the data.
  5. Explain why an association in observational data does not establish a cause.

One reading habit does most of the work: scan the answer choices for causal language and treat it as a red flag.

Naruhodo Tutoring — SAT Math question bank export (sat-question-index.json) — 63 tagged questions of this type in the site's bank

3. Recognise It

Section

Section 1

4. What this question actually asks

Concept

A scatterplot with a line or curve drawn through it, and questions about what the model means. The arithmetic is almost nothing. What is being tested is whether you can say what a slope represents in the units of the story, whether you know how far a model can be trusted, and whether you can resist a choice that says one thing causes another.

You will see it phrased in these ways:

The third phrasing is where the marks are. Model choice is decided by whether the data bends, and conclusion questions are decided by whether the word cause appears.

5. The card for this type

Picture it

The same three-panel card as the survey deck, so the shorthand carries over: what identifies it, what you write first, and what is built to catch you.

Figure (svg): Two-variable data: models and scatterplots: the tell, the move, and the trap

Two-variable data: models and scatterplots — 63 of 1675 bank questions (3.8%)

Both halves of the red panel are about the LIMITS of a model rather than its content. That is unusual on the Math section, and it is what makes this type feel different.

6. Which conclusion does a scatterplot license?

Prediction

The most-tested idea in the type, up front.

Predict first

A scatterplot shows that towns with more libraries have higher reading scores. Which conclusion is supported?

  • There is a positive association between library count and reading scores
  • Building more libraries would raise reading scores
  • Reading more causes towns to build libraries
  • Libraries have no effect on reading scores

Correct: There is a positive association between library count and reading scores

Why: Observational data can establish that two things vary together, and nothing more. The second and third choices both assert a direction of causation, which the data cannot distinguish — and a third factor such as town funding could produce the pattern with neither causing the other. The fourth overstates in the opposite direction: the data does not establish an absence of effect either. Association is exactly what a scatterplot shows.

7. The faces this type wears

Concept

Four shapes of question, and only one of them involves computation.

variantwhat it wantsthe move
Interpret a coefficientwhat the slope or intercept meansstate the quantity with its units and a per
Predict a valuey at a given x, or x at a given ysubstitute into the model
Choose a modellinear, quadratic or exponentialask whether the data bend, and how
Judge a conclusionwhich statement the data supportsreject causal language; accept association

Interpretation and conclusion questions together are the majority. Both are graded on wording rather than arithmetic.

8. What shape is the data?

Definition probe

Model choice comes from the shape, not from the context.

Sort into buckets

What kind of model does each pattern suggest?

Linear
The points rise along a straight path
Quadratic
The points rise to a peak and then fall
Exponential
The points rise slowly, then much more steeply; The points fall steeply, then level off toward a floor
lin
A constant rate of change plots as a straight path. Equal steps in x give equal steps in y anywhere along the data.
quad
A single turning point — rising then falling, or falling then rising — is the signature of a parabola.
exp
Exponential growth accelerates without turning, and exponential decay falls toward a floor without ever crossing it. Neither has a turning point.

9. Inside the data, or beyond it?

Discrimination

The data runs from x equals 5 to x equals 40.

Sort into buckets

Is each prediction interpolation or extrapolation?

Interpolation — inside the data
Predicting y at x equals 20; Predicting y at x equals 38; Predicting y at x equals 12
Extrapolation — beyond the data
Predicting y at x equals 100; Predicting y at x equals 0; Predicting y at x equals negative 5
in
These inputs fall between 5 and 40, where the model was fitted and where data actually exists. Predictions here are supported by observations on both sides.
out
These fall outside the observed range. Note that (d), x equals 0, is extrapolation even though it seems modest — which matters, because the y-intercept is often meaningless for exactly this reason.

10. Interpret a slope

Warm-up

Answer in words, with units.

Discussion prompt

A scatterplot of car age against value has a line of best fit with equation V equals 24000 minus 1800a, where a is age in years and V is value in dollars. What does the 1800 represent?

Hint: The units of a slope are always output units per input unit.

Answer:

The car loses about 1,800 dollars of value per year. The units are dollars per year, which is output units over input units.

The minus sign is part of the meaning: the value decreases, so the model predicts a fall of 1,800 dollars for each additional year of age.

A weak answer that scores nothing: the 1800 is the slope. Interpretation questions want the quantity, its units, and its direction.

And the 24,000 is the predicted value at age zero — the value of a new car, according to the model.

11. The routine, every time

Pattern

Three steps, and none of them is a calculation.

Read the axes: what quantity is on each, and in what units.

Why: Every interpretation depends on this, and the axes are frequently in thousands, percentages or other scaled units that change the answer.

Name the slope and the intercept in the story's own words, with units and direction.

Why: The slope is the change in the vertical quantity per one unit of the horizontal one; the intercept is the vertical value when the horizontal one is zero.

Before selecting, scan the choices for causal language and for predictions outside the data range.

Why: Those are the two designed traps, and both are visible in the wording without any further reasoning.

Step 3 is the one that converts a reasonable answer into the right one. Both traps are caught by reading rather than by thinking harder.

12. The Rules

Section

Section 2

13. Rule 1 · The slope is a rate, in the story's units

Concept

The slope of a fitted line is the predicted change in the vertical quantity for each one-unit increase in the horizontal quantity.

A complete interpretation names the quantity, the units, the direction and the per. Saying it is the slope earns nothing.

14. Rule 2 · The intercept is the value at zero, and may be meaningless

Concept

The y-intercept is the model's prediction when the horizontal quantity is zero.

Interpret the intercept as what the model says at zero, and be ready to observe that the model is not trustworthy there.

15. Interpret with units

Prediction

Name the quantity, not the word slope.

Predict first

A model gives P equals 240 plus 15t, the number of plants after t weeks. What does the 15 mean?

  • the number of plants added each week
  • the number of plants at the start
  • the number of weeks to reach 240 plants
  • the total number of plants

Correct: the number of plants added each week

Why: The 15 multiplies t, so its units are plants per week — the increase for each additional week. The 240 is the value at t equals 0, which is the starting number of plants. The other two choices describe quantities that do not correspond to either coefficient. A good interpretation always carries the units and the per.

16. Rule 3 · Interpolation is safe, extrapolation is not

Concept

A model is trustworthy inside the range of the data and speculative outside it.

predictionnamehow much to trust it
inside the observed rangeinterpolationsupported by data on both sides
just outside the rangeextrapolationplausible, with caution
far outside the rangeextrapolationunreliable — the pattern may not continue

If a question asks about a value far beyond the plotted data, expect the correct answer to include a caution rather than a confident number.

17. Interpolation or extrapolation

Prediction

Compare the input to the data range.

Predict first

Data covers ages 20 to 65. Predicting a value at age 80 is what?

  • extrapolation, and should be treated cautiously
  • interpolation, and is reliable
  • impossible to compute
  • the same as predicting at age 65

Correct: extrapolation, and should be treated cautiously

Why: Age 80 lies outside the observed range of 20 to 65, so the model has no data supporting its behaviour there. The prediction can certainly be computed by substituting into the equation — that is not the issue. The issue is whether the pattern continues, and nothing in the data speaks to that.

18. Rule 4 · Choose a model by asking whether the data bends

Concept

Straight means linear, one turn means quadratic, and accelerating without turning means exponential.

The multiply-versus-add test settles linear against exponential: linear data adds a constant each step, exponential data multiplies by one.

19. Rule 5 · Association is not causation

Concept

Observational data can show that two quantities vary together and cannot establish that one produces the other.

Treat cause, causes, leads to, results in and would raise as red-flag words in the answer choices. They are almost always wrong on this type.

20. Which model?

Prediction

Ask whether the data bends, and whether it turns.

Predict first

A scatterplot's points rise gently at first and then increasingly steeply, never turning downward. Which model fits best?

  • exponential
  • linear
  • quadratic
  • no model fits

Correct: exponential

Why: Accelerating growth without a turning point is the signature of an exponential model. A linear model would rise at a constant rate throughout, and a quadratic would eventually turn — a parabola opening upward would have fallen before rising. The absence of a turn is what rules out the quadratic.

21. Rule 6 · A line of best fit describes a trend, not every point

Concept

Individual points may sit well away from the line, and the model predicts the average behaviour.

That last mapping is worth fixing in memory: overestimate means the line sits above the point, so the point is below the line.

22. Rule 7 · Read the axis scale before anything else

Concept

Axes on a scatterplot are frequently in thousands, or start away from zero.

Every interpretation in this type is stated in units, so getting the units wrong makes the whole answer wrong regardless of the reasoning.

23. The red-flag word

Prediction

Scan the wording before the content.

Predict first

An observational study finds a strong positive association between hours of sleep and test scores. Which conclusion is supported?

  • Students who sleep more tend to score higher
  • Sleeping more causes higher test scores
  • Higher test scores cause students to sleep more
  • There is no relationship between the two

Correct: Students who sleep more tend to score higher

Why: The first choice states an association and nothing more, which is exactly what observational data supports. Both causal choices go beyond the evidence, and the data cannot distinguish between them or rule out a third factor such as household stability affecting both. The fourth contradicts the stated association.

24. Three of these are true

Two truths and a lie

Three of these statements about fitted models are correct. The one left standing is false.

Eliminate the wrong options

Which statement is FALSE?

  • a. A point above the line of best fit was underpredicted by the model
  • b. The intercept is an extrapolation when the data does not reach zero
  • c. A strong association in observational data shows one variable causes the other
  • d. Exponential data multiplies by roughly a constant factor at each equal step

Survives elimination: c

Why: Observational data establishes only that two quantities vary together. The association may run in either direction, or arise from a third factor affecting both. Only a randomised experiment, where the researcher assigns the treatment, can license a causal claim — and a scatterplot of observed data never is one. This distinction is the single most tested idea in the type.

25. Check 1 · Interpretation

Check

Answer with units.

Check your understanding

A model gives W equals 3.1 plus 0.42m, an infant's weight in kilograms after m months. What does 0.42 represent?

  • A. the predicted weight gain in kilograms each month (correct)
  • B. the birth weight in kilograms
  • C. the number of months to gain one kilogram
  • D. the total weight gained

Answer: A

Why: The 0.42 multiplies m, so its units are kilograms per month — the predicted increase for each additional month. The 3.1 is the value at m equals 0, which is the birth weight. Checking: after 10 months the model predicts 3.1 plus 4.2, which is 7.3 kg, a plausible figure.

Why B tempts people
That is the 3.1, the value at zero months.
Why C tempts people
That would be 1 divided by 0.42, about 2.4 months — a genuine quantity derived from the model, but not what the coefficient itself represents.
Why D tempts people
The total gained depends on how many months have passed, so it is not a fixed number at all.

Choice C is worth noticing: it is a real quantity computed from the slope, offered because reciprocals of rates feel like interpretations of them.

26. Worked Examples

Section

Section 3

27. Example 1 · Interpreting slope and intercept

Worked example

A study of house size and heating cost gives C equals 180 plus 0.09s, where s is floor area in square feet and C is annual cost in dollars. Interpret both numbers.

Figure (svg): A scatterplot of floor area against heating cost with a rising line of best fit

The slope is dollars per square foot; the intercept is the cost at zero area.

The slope 0.09 multiplies s, so its units are dollars per square foot.

Why: Slope units are always output units over input units.

So each additional square foot is associated with about 9 cents more annual heating cost.

Why: A complete interpretation names the quantity, the units and the direction.

The intercept 180 is the predicted cost when floor area is zero — a fixed cost of 180 dollars.

Why: The intercept is the model's value at an input of zero.

The intercept is meaningful here — a standing charge — but a house of zero area does not exist, so it is still an extrapolation and worth describing as a fixed component rather than a real house.

Verify: check the units combine correctly: square feet times dollars per square foot gives dollars.

Why: Adding that to the 180 dollars gives a cost, which is what C measures.

Answer: the slope is about 9 cents per square foot; the intercept is a 180 dollar fixed cost

28. Example 2 · Predicting from a model

Worked example

Using C equals 180 plus 0.09s, predict the heating cost for a 2,200 square foot house, and say whether the prediction is trustworthy given data from 800 to 2,800 square feet.

Figure (svg): The same scatterplot with the prediction point at 2200 square feet highlighted

2,200 lies inside the observed range, so this is interpolation.

Substitute s equals 2200: C equals 180 plus 0.09 times 2200.

Why: Prediction is substitution into the fitted model.

Compute: 0.09 times 2200 is 198, so C equals 378 dollars.

Why: The arithmetic is the easy part of this variant.

Note that 2,200 lies between 800 and 2,800, so this is interpolation and is well supported.

Why: Predictions inside the observed range have data on both sides.

Had the question asked about a 10,000 square foot house, the arithmetic would be just as easy and the answer would need a caution attached.

Verify: check the answer sits between the model's values at 2,000 and 2,800, which are 360 and 432.

Why: 378 lies between them, as it must for an increasing model.

Answer: 378 dollars, and the prediction is trustworthy because it interpolates

29. Example 3 · Choosing between models

Worked example

A data set gives y values of 3, 6, 12, 24, 48 at x equals 1, 2, 3, 4, 5. Which model fits, and why?

Figure (svg): A table showing the differences growing while the ratios stay constant at two

Constant ratio, not constant difference: exponential.

Compute the differences: 3, 6, 12, 24. They are not constant, so the data is not linear.

Why: Linear data has constant first differences over equal steps in x.

Compute the ratios: 6 over 3 is 2, 12 over 6 is 2, and so on. They are constant.

Why: A constant ratio at equal steps is the signature of an exponential model.

So the model is exponential, with a base of 2 — the values double at each step.

Why: The common ratio is the base of the exponential.

The differences-versus-ratios test is the reliable way to separate linear from exponential when you have a table rather than a picture.

Verify: check the model: 1.5 times 2 to the power x gives 3, 6, 12, 24, 48.

Why: The model reproduces every data point exactly, confirming the choice.

Answer: exponential, with the values doubling each step

30. What do you read first?

Step zero

Before interpreting anything.

Discussion prompt

A scatterplot question asks you to interpret the slope of a line of best fit. What do you read before looking at the line at all?

Hint: Two labels, and they may not say what you expect.

Answer:

Read both axis labels, including their units. The vertical axis may be in thousands, percentages, or minutes rather than hours.

Check whether either axis starts away from zero, which makes the visible intercept misleading.

Only then read the slope, and state it as output units per input unit.

Why it matters here more than elsewhere: every answer on this type is a sentence containing units, so getting the units wrong makes the whole interpretation wrong even when the reasoning is perfect.

This is the same axis-scale discipline as type 4, applied to a question whose answer is words rather than a number.

31. Example 4 · Reading residuals

Worked example

On a scatterplot with a line of best fit, four points lie above the line and six below. For how many points did the model OVERESTIMATE the actual value?

Figure (svg): A scatterplot with points scattered above and below the fitted line

Overestimated means the line is above the point.

Overestimating means the model predicted MORE than the actual value.

Why: The prediction is the height of the line; the actual value is the height of the point.

So the line sits above the point, which means the point sits BELOW the line.

Why: Reversing the comparison is where this question is lost.

Six points lie below the line, so the model overestimated six of them.

Why: Counting the points below the line answers the question directly.

The mapping is worth memorising because it inverts: overestimate means point below the line, underestimate means point above it.

Verify: check the complement: four points above the line were underestimated, and 4 plus 6 is 10.

Why: Every point is either above or below, so the two counts must sum to the total.

Answer: six points

32. Example 5 · Rejecting a causal claim

Worked example

A scatterplot shows that cities with more coffee shops have lower unemployment. Which conclusion does the data support?

Figure (svg): Bars showing that only the association claim is supported by the data

One claim survives; two directions of causation do not.

The data shows the two quantities vary together, which is an association.

Why: A scatterplot displays how two observed quantities move relative to each other.

It cannot establish direction: more shops might follow prosperity rather than create it.

Why: Observational data does not distinguish which variable moved first.

And a third factor — a growing local economy — could raise both, with neither causing the other.

Why: A confounding variable produces the same pattern without any causal link between the two plotted quantities.

Every question of this shape has one choice that merely describes the pattern. That choice is nearly always the answer.

Verify: check what would be needed for a causal claim: random assignment of coffee shops to cities.

Why: That experiment is not what happened, so the causal conclusion is unavailable.

Answer: only that there is a negative association between coffee shop count and unemployment

33. Complete the interpretation rules

Faded example

From memory. Four lines that carry the type.

Fill in the blanks

The slope of a fitted line is the change in the vertical quantity for each one-unit increase in the horizontal quantity. Predicting inside the range of the data is called interpolation. And observational data can establish an association but never a cause.

Why: The last blank is the most-tested idea in the type. Association means the two quantities move together; causation means one produces the other, and distinguishing them requires an experiment with random assignment rather than a scatterplot of observed data.

34. Example 6 · A scaled axis

Worked example

A scatterplot's vertical axis is labelled in thousands of dollars. The line of best fit rises 2 units vertically for every 5 units horizontally, where the horizontal axis is years. Interpret the slope.

Figure (svg): A line rising two vertical units over five horizontal units on an axis marked in thousands

The vertical units are thousands, so the slope carries a factor of 1,000.

The slope as read off the grid is 2 over 5, which is 0.4.

Why: Rise over run, in the units of the axes as drawn.

But the vertical axis is in THOUSANDS of dollars, so 0.4 means 0.4 thousand dollars.

Why: The axis label rescales every vertical reading.

So the slope is 400 dollars per year.

Why: Converting to plain dollars gives the interpretation in the story's real units.

Answering 0.4 dollars per year would be off by a factor of a thousand, and nothing in the reasoning would look wrong. Read the axis labels first.

Verify: check across five years: 5 times 400 is 2,000 dollars, which is the 2 thousand read off the graph.

Why: The conversion reproduces the original reading, confirming the factor.

Answer: the value rises about 400 dollars per year

35. Fill in the model shapes

Fill the middle

Model choice comes from the shape of the data.

Fill in the blanks

Data that rises at a constant rate is linear. Data with exactly one turning point is quadratic. Data that curves without turning, multiplying by a constant factor at each step, is exponential. And a point lying below the line of best fit was overestimated by the model.

Why: The last blank inverts, which is why it is worth practising. A point below the line means the line is above the point, so the model predicted a larger value than actually occurred — an overestimate.

36. Estimate a prediction

Estimation

A rough value confirms you substituted correctly.

Predict first

A model gives y equals 45 plus 3.2x. Roughly what is y when x equals 30?

  • about 140
  • about 100
  • about 1,400
  • about 75

Correct: about 140

Why: Round 3.2 to 3: 3 times 30 is 90, plus 45 gives about 135, so the answer is a little above that. The exact value is 141. The answer 75 comes from adding 45 and 30 without applying the coefficient, and 1,400 multiplies everything by ten somewhere. Having an expectation first makes both errors visible immediately.

37. Check 2 · Model choice

Check

Differences or ratios?

x12345
y100806451.240.96

Check your understanding

Which model best fits the data in the table?

  • A. exponential decay with a factor of 0.8 (correct)
  • B. linear with a slope of negative 20
  • C. quadratic
  • D. exponential growth

Answer: A

Why: The differences are negative 20, negative 16, negative 12.8, negative 10.24 — not constant, so the data is not linear. The ratios are 80 over 100 equals 0.8, 64 over 80 equals 0.8, and so on — constant. A constant ratio below 1 is exponential decay with that factor.

Why B tempts people
The first difference is indeed negative 20, which is why this tempts, but the later differences shrink. Linear data has constant differences throughout.
Why C tempts people
A quadratic has a turning point; this data falls throughout and levels toward a floor without turning.
Why D tempts people
The values are decreasing, so any exponential model must have a factor below 1, making it decay rather than growth.

Choice B shows why checking only the first difference is not enough. Compute at least two differences and two ratios before deciding.

38. The Traps

Section

Section 4

39. Reading a correlation as a cause

Trap

The trap

The trap. A scatterplot shows two quantities rising together, and a choice reads: increasing the first would raise the second.

It matches the data, it sounds sensible, and it is the answer most students pick.

Observational data cannot separate the two quantities from everything that travels with them, and it cannot tell which one moved first.

The fix

The fix. Scan the choices for causal language before evaluating any of them on content.

Prefer the choice that describes an association, a relationship, or a tendency.

  1. Look for cause, causes, leads to, results in, would raise, would reduce.
  2. Ask whether subjects were randomly ASSIGNED. If they were merely observed, no causal claim survives.
  3. Choose the choice that only describes the pattern.
  4. The exception: if the stem explicitly describes a randomised experiment, causal language may be legitimate.

40. Trusting an extrapolation

Trap

The trap

The trap. Data covers ages 5 to 15, and a question asks what the model predicts at age 40. You substitute and report a confident number.

The arithmetic is right and the prediction is unsupported. A model of childhood growth extended to age 40 predicts an impossible height.

Nothing in the data speaks to whether the pattern continues that far.

The fix

The fix. Compare the requested input with the range of the plotted data before answering.

Far outside means the answer should carry a caution, and the correct choice usually says so.

  1. Note the smallest and largest x-values in the data.
  2. Ask whether the requested input falls inside that range.
  3. If it does not, expect the right answer to question the reliability rather than state a confident figure.

41. Annotate a causal conclusion

Error analysis

A student's conclusion from a scatterplot. The reading is right and the claim is not.

Annotate

On: \( \text{more hours studied} \;\leftrightarrow\; \text{higher scores} \;\Rightarrow\; \text{studying causes higher scores} \)

  • The first part is a correct reading of the data. The two quantities do rise together, and describing that association is entirely legitimate.
  • The arrow to a causal claim is where it fails. The data shows co-variation and nothing about mechanism.
  • Consider the alternatives the data cannot rule out. The direction might be reversed: students who find the material easy may score well AND enjoy studying it more.
  • Or a third factor may drive both: a stable home environment could increase study time and raise scores independently.
  • What would license the causal claim is random ASSIGNMENT — allocating students to different study durations and comparing outcomes. That is an experiment, not an observation.
  • Note that the causal claim may well be TRUE. The point is not that it is false, but that this data does not establish it.

That last note matters. The SAT is not asking whether the claim is plausible; it is asking what the evidence supports. Those are different questions.

42. Interpreting a slope without units

Trap

The trap

The trap. Asked what the 15 in P equals 240 plus 15t represents, you answer that it is the slope.

That is true and it earns nothing. The question wants the meaning in the story: 15 more plants for each additional week.

The wrong answers are usually correct sentences about a different number, so a vague answer cannot distinguish them.

The fix

The fix. Answer with a quantity, its units, and a per.

State the direction too: rising or falling.

  1. Identify which coefficient the question names.
  2. Write its units as output units per input unit.
  3. Say the sentence: for each additional [input unit], the [output] changes by this much.

43. Eliminate three by wording

Elimination

An observational study finds that students who eat breakfast score higher on tests.

Eliminate the wrong options

Which conclusion is supported? Three can be eliminated on their wording alone.

  • a. Eating breakfast causes higher test scores
  • b. There is a positive association between eating breakfast and test scores
  • c. Requiring all students to eat breakfast would raise average scores
  • d. Higher test scores lead students to eat breakfast

Survives elimination: b

Why: Only one choice describes the pattern without claiming a mechanism, and it is the one the data supports. Notice that all three eliminations were made on wording rather than content: cause, would raise and lead to are each causal, and observational data licenses none of them. Scanning for those words is faster than reasoning about each choice.

44. Reversing over and under estimates

Trap

The trap

The trap. Asked how many points the model overestimated, you count the points above the line.

Overestimating means the model predicted too HIGH, so the line sits above the point and the point is BELOW the line.

The two counts are usually both offered, so the reversal is a whole question.

The fix

The fix. Compare the line's height to the point's height explicitly.

Overestimate means line above point; underestimate means line below point.

  1. Say aloud: the model predicted the height of the LINE.
  2. Overestimate means that prediction was too large, so the actual point is lower.
  3. Count points below the line for overestimates and above for underestimates.

45. Find the counterexample

Counterexample

A claim that a fitted model tempts you into.

Discussion prompt

A student says: a line of best fit lets you predict the value at any x. Give a counterexample and explain the limit.

Hint: Think about what happens far outside the data.

Answer:

Counterexample: a model of a child's height against age, fitted to ages 2 to 12, giving height equals 80 plus 6a centimetres.

At age 12 it predicts 152 cm, which is reasonable. At age 40 it predicts 320 cm, which is impossible.

The limit: a fitted line describes the pattern WITHIN the range of the data. Outside it, the relationship may change shape entirely — and human growth stops.

A second counterexample in the other direction: at age 0 the model predicts 80 cm, far taller than any newborn.

So both ends fail, which is why the intercept of a fitted model is often meaningless and why far extrapolation is unreliable.

The correct statement: a model can be evaluated at any x, but it can only be TRUSTED near the data it was fitted to.

46. Push the causation rule to its edge

Edge cases

Association is not causation. Is causal language ever correct?

Discussion prompt

When would a causal conclusion actually be supported, and how would the question signal it?

Hint: The signal is a single feature of the study design.

Answer:

Causal language is supported when the study used random ASSIGNMENT — the researcher allocated subjects to treatment and control groups.

Random assignment balances every other factor, known and unknown, between the groups. So a difference in outcome can be attributed to the treatment.

How a question signals it: phrases such as subjects were randomly assigned, participants were randomly allocated, or a randomised controlled trial.

Random SELECTION is a different thing. It licenses generalising to the population sampled, not a causal claim. The two words look alike and do different jobs.

So the full rule: random assignment licenses cause; random selection licenses generalisation. A scatterplot of observed data has neither.

This is the same distinction that type 19 tests directly, which is why the two types reward being studied together.

47. Check 3 · What the data supports

Check

Scan for causal language first.

Check your understanding

An observational study of 500 adults finds that those who walk more have lower blood pressure. Which conclusion is best supported?

  • A. Among these adults, walking more is associated with lower blood pressure (correct)
  • B. Walking more reduces blood pressure
  • C. Lower blood pressure causes people to walk more
  • D. Walking has no effect on blood pressure

Answer: A

Why: The study observed rather than assigned, so it can establish only that the two vary together. The phrase among these adults is also appropriately cautious about generalising beyond the sample studied.

Why B tempts people
This asserts causation, which requires random assignment of walking amounts rather than observation of existing habits.
Why C tempts people
This asserts causation in the reverse direction, equally unsupported by observational data.
Why D tempts people
This claims an absence of effect, which the data does not establish either — a positive association is evidence against it, if anything.

Note that choice A restricts itself to the adults studied. Both the causal caution and the generalisation caution appear in correct answers on this type.

48. Drill and Plan

Section

Section 5

49. Match each feature to what it tells you

Matching

Six features of a fitted model, six meanings.

Match the pairs

  • slope. The slope of the line
  • int. The y-intercept
  • above. A point above the line
  • inside. Predicting inside the data range
  • far. Predicting far outside the range
  • ratio. Constant ratios at equal steps
  • rate. Change in output per one unit of input
  • zero. The predicted output when the input is zero
  • under. The model underestimated that value
  • interp. Interpolation, well supported
  • extrap. Extrapolation, treat with caution
  • exp. An exponential model fits

Why: The third row is the one that inverts. A point above the line means the actual value exceeded the prediction, so the model predicted too little — an underestimate. Getting that mapping the wrong way round is a reliable way to lose a residual question.

50. Sort six conclusions

Sorting

Each follows a study. Which are supported?

Sort into buckets

Does the study design support the conclusion?

Supported
Observational data; conclusion states an association; Randomised experiment; conclusion states a cause; Randomised experiment; conclusion states an association
Not supported
Observational data; conclusion states a cause; Observational data; conclusion predicts the effect of an intervention; Observational data; conclusion says one variable has no effect
yes
Observational data supports association claims, and a randomised experiment supports causal claims — and also association claims, since causation implies the two vary together.
no
Causal claims and predictions about interventions both require random assignment. And a claim of NO effect overstates in the opposite direction: failing to establish a cause is not the same as establishing its absence.

Item (f) is the over-correction worth noticing. Students who learn that correlation is not causation sometimes conclude that the data shows no relationship, which is a different error in the opposite direction.

51. Three models, side by side

Comparison

Fill the blanks from memory.

Comparison matrix

modelwhat stays constantshape of the plot
Linearthe difference between successive y-valuesa straight path
Quadraticthe second differenceone turning point
Exponentialthe ratio between successive y-valuescurves without turning

The differences-versus-ratios test is what separates linear from exponential in a table, and the presence or absence of a turning point separates quadratic from both. Two questions decide all three.

52. How far can you trust a model?

Trade off

Fill in the reliability of each prediction.

Comparison matrix

where you predictreliabilitywhy
In the middle of the datahighobservations on both sides support it
At the edge of the datamoderatesupported on one side only
Just beyond the datalow, use cautionthe pattern may continue, but nothing confirms it
Far beyond the dataunreliablerelationships change shape outside the observed range

The reliability falls off with distance from the data rather than dropping suddenly at the edge. That is why correct answers on extrapolation questions usually express caution rather than refusal.

53. Where this shows up outside the test

Real world

One minute on why the causation rule matters.

Discussion prompt

Newspaper headlines routinely report observational studies with causal language. Why does the distinction matter, and what should a reader look for?

Answer:

Because the intervention implied by a causal headline may not work. If a third factor drives both quantities, changing one will not move the other.

A classic example: studies once found that people taking a particular supplement had better health outcomes. Randomised trials later showed no benefit — the supplement-takers were simply healthier people to begin with.

What a reader should look for: were participants randomly ASSIGNED to groups, or merely observed and compared?

The vocabulary is a reliable tell. Observational studies report that something is associated with or linked to an outcome; experiments report that it reduces or causes one.

And the same reasoning is the SAT question: the test is asking whether you can tell what the evidence supports from how the study was run, which is a genuinely useful skill.

54. Order these predictions by reliability

Ranking

Data covers x from 10 to 50. Order from most reliable to least.

Put in order

  1. Predicting at x equals 30
  2. Predicting at x equals 48
  3. Predicting at x equals 55
  4. Predicting at x equals 200

Why: x equals 30 sits in the middle of the data with observations on both sides, so it is the most reliable. x equals 48 is inside the range but near its edge. x equals 55 is just outside, so the pattern probably continues but nothing confirms it. x equals 200 is four times beyond the observed range, where the relationship may have changed shape entirely. Reliability falls off with distance from the data rather than dropping sharply at the boundary.

55. How to practise this type

Concept

This type is 3.8 per cent of the section and is graded almost entirely on wording, so the practice looks different from the rest of the Math section.

sessionwhat you dowhy
1Fifteen interpretation questions, answering aloud with the quantity, units and direction.Interpretation is the most common variant and is graded on units.
2Ten model-choice questions using the differences-and-ratios test on tables.Separates linear from exponential reliably, which a picture alone sometimes does not.
3Fifteen conclusion questions, scanning for causal language before reading for content.The single highest-value habit in the type.
4Ten prediction questions, stating whether each is interpolation or extrapolation first.Builds the range check that decides how confidently to answer.

Naruhodo Tutoring — SAT Math question bank export (sat-question-index.json) — filter the bank to this skill tag — 63 questions, roughly a third of each difficulty

56. Explain the causation rule from memory

Explain it to yourself

Close the deck.

Discussion prompt

Without looking, explain why a scatterplot cannot establish causation, and say what would be needed instead.

Hint: There are two distinct reasons the data falls short.

Answer:

First, direction is unidentified. The data shows the two quantities move together but not which one moved first, so the causal arrow could point either way.

Second, a confounding factor may drive both. A third variable can produce the same pattern with no causal link between the two plotted quantities at all.

What would be needed: random ASSIGNMENT. Allocating subjects to treatment and control balances every other factor between the groups, so a difference in outcome can be attributed to the treatment.

Random selection is a different thing — it licenses generalising to the population sampled, not a causal claim.

If you gave both reasons rather than just one, you have the version that also answers the type 19 questions.

57. Teach the association distinction

Explain it

Two minutes, out loud.

Discussion prompt

A friend says the data clearly shows that studying more causes better grades, so why is that answer wrong? How do you explain it without seeming pedantic?

Answer:

Grant the point first: the claim is probably true in real life. That is not what is being tested.

Then separate the two questions: is the claim plausible, and does THIS data establish it? The test asks the second.

Give the two gaps concretely. Maybe students who find the material easy both study more and score higher — so the arrow runs the other way, or from a third cause.

Then give the test that would settle it: randomly assign students to different study durations. That is an experiment, and nobody ran one here.

Finish with the practical rule: on this type, scan for cause and would raise in the choices. The choice that only describes the pattern is nearly always the answer.

58. How confident are you on residuals?

Commit first

Commit before you check.

Predict first

A data point lies below the line of best fit. What does that mean?

  • The model overestimated that value
  • The model underestimated that value
  • The model predicted it exactly
  • The point is an outlier

Correct: The model overestimated that value

Why: The line's height is the prediction and the point's height is the actual value. If the point is below the line, the line is above the point, so the prediction exceeded reality — an overestimate. The mapping inverts, which is why it is worth stating explicitly rather than trusting instinct. Being below the line says nothing about whether the point is an outlier, which is about distance rather than direction.

59. Draw the whole type on one page

Connect it up

Blank paper.

Draw it

Draw a scatterplot with a line through it. Label the slope with an arrow and write beside it: change in output per one unit of input, with units. Label the intercept and write: the value at zero, and note that it is an extrapolation if the data does not reach zero. Mark one point above the line and one below, labelling them underestimated and overestimated. Along the horizontal axis, shade the data range and mark regions outside it as EXTRAPOLATION. In a box at the side, write the three model shapes with their tests: constant differences means linear, one turning point means quadratic, constant ratios means exponential. At the bottom, in large letters: ASSOCIATION IS NOT CAUSATION, and beneath it, random assignment is what licenses cause.

60. Exit ticket

Exit ticket

One question before you close the deck.

Predict first

You reach a scatterplot question asking which conclusion the data supports. What do you do first?

  • Scan the choices for causal language
  • Compute the slope of the line of best fit
  • Find the point furthest from the line
  • Predict the value at the largest x

Correct: Scan the choices for causal language

Why: Conclusion questions on this type are decided by wording, and the causal choices can usually be eliminated before you evaluate any of them on content. Computing the slope answers an interpretation question rather than a conclusion one, finding the furthest point identifies an outlier that was not asked about, and predicting at the largest x is an extrapolation nobody requested. The scan takes about five seconds and frequently leaves one choice standing.

61. What to take away

Recap

One type, one discipline: say what the model means, and refuse to say what it cannot.

never do thisdo this instead
Say a scatterplot shows one thing causes anotherSay the two are associated
Report a confident prediction far outside the dataNote that it extrapolates and may not hold
Answer that a number is the slopeGive the quantity, its units and its direction
Count points above the line as overestimatesOverestimate means the point is below the line
Read a slope off the grid without checking the axis unitsConvert using the axis labels first
Conclude there is no relationship because there is no proven causeAbsence of a causal claim is not evidence of no effect

Naruhodo Tutoring — SAT Math question bank export (sat-question-index.json) — and the site's SAT pages to drill this type in isolation, then mixed

Sources

  1. College Board — Digital SAT Suite: test description and format
  2. College Board — Digital SAT Suite Assessment Specifications, Math section: domain weightings and skill definitions — College Board, 2023
  3. Naruhodo Tutoring — SAT Math question bank export (sat-question-index.json) — 1675 questions carrying official domain, skill and difficulty tags; the frequency figures in this deck are counted from this file
  4. Khan Academy — Official Digital SAT Prep, Math

Want this taught 1-on-1? Alexander tutors SAT Prep — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108