Exam 1 Review: Description, Distributions, Inference and Regression

A full sweep of the first-exam material: summarising a sample, the normal and binomial models, standard error and confidence intervals, hypothesis testing, and simple linear regression — each built on an economic scenario before any formula appears.

Subject: Economic Statistics · 71 slides · applied lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Economic Statistics Exam Review

Title

Session 1

Describe it, model it, infer from it, test it, fit a line through it

2. What you will be able to do by the end

Objectives

Your professor said the exam follows the shape of the practice papers. Those papers walk one path five times, so this deck walks it once, carefully.

OpenStax, Introductory Business Statistics 2e chapters 2, 6, 7, 9, 13 — the same order this deck follows

3. Describing a Sample

Section

Section 1

4. Before anything new: what do you already have?

Warm-up

Every exam question in this course starts by handing you a sample and asking a question about the population behind it.

Discussion prompt

Write down, from memory, the difference between a parameter and a statistic. Which one do you ever actually get to see?

Hint: Which of the two would change if you collected the data again tomorrow?

Answer:

parameter — A number describing the whole population. Fixed, unknown, and you almost never observe it.

statistic — A number computed from your sample. You observe it, and it changes if you draw a different sample.

Every technique in this course exists because the first one is hidden and the second one is not.

OpenStax, Introductory Statistics 2e §1.2

5. Seven workers, one payroll

Concept

A small firm reports monthly wages for its seven employees. This is the data set the next six slides live on.

workermonthly wage, dollars
A2400
B2600
C2600
D2800
E3000
F3200
G9800

Worker G is the owner. Nothing about that value is a data-entry error, and that is exactly what makes it interesting.

U.S. Bureau of Labor Statistics, Current Employment Statistics — average hourly earnings — wage distributions in real firms are right-skewed for this reason

6. Guess before you compute

Estimation

Do not calculate. Estimate.

Predict first

Where will the mean wage fall relative to the median wage?

  • Well below the median
  • About equal to the median
  • Somewhat above the median
  • Far above the median

Correct: Far above the median.

Why: The mean is a balance point, so a single value at 9800 pulls it a long way to the right. The median only cares about position in the ordered list, so it barely notices. Mean 3771.43 versus median 2800 — a gap of nearly a thousand dollars caused by one person.

7. Worked example: mean, median, mode

Worked example

Order the data first. It is already ordered, which is the only reason the median is easy here.

Add all seven wages

Why: The sum is 2400 plus 2600 plus 2600 plus 2800 plus 3000 plus 3200 plus 9800, which is 26400.

\[ \bar{x} = \frac{26400}{7} = 3771.43 \]

Take the middle value of the ordered list

Why: Seven values, so the median is the fourth one: 2800. No arithmetic at all.

Find the value that repeats

Why: 2600 appears twice and nothing else repeats, so the mode is 2600.

Verify: that five of the seven workers earn less than the mean

Why: If a measure of centre sits above five of seven observations, it is not describing a typical worker. That is the whole argument for reporting the median here.

8. What the outlier actually does

Picture it

The green line is the median. The yellow line is the mean. One data point is responsible for the whole gap.

Figure (svg): Dot plot of seven wages from 2400 to 9800 with the median marked at 2800 and the mean marked at 3771.43, and the value 9800 labelled as an outlier

The mean is the balance point of the dots; the median is the middle dot.

This is why the government reports median household income, not mean household income.

9. Trap: reporting the mean because it uses all the data

Trap

The trap

"The mean uses every observation, so it must be the better summary."

\[ \bar{x} = 3771.43 \]

Report the mean as the typical wage

Why: It is true that the mean uses every value. It is also true that this is the problem, not the virtue.

The fix

Ask what the number is for. If it is 'what does a typical employee earn', the answer must sit among typical employees.

\[ \text{median} = 2800 \]

Report the median, and mention the mean separately

Why: Both are correct numbers. Only one answers the question that was asked. On an exam, say which you chose and why.

10. Which measure of centre survives the outlier?

Sorting

Sort each quantity by whether one extreme value can move it a lot.

Sort into buckets

Resistant, or not resistant, to a single extreme observation?

resistant
the median; the interquartile range; the mode
not resistant
the mean; the range; the standard deviation
res
These depend on position or frequency, not on magnitude. Push the largest value to a million and the median, the IQR and the mode do not move at all.
not
These arithmetic on the actual values. One huge number enters the sum (mean), the subtraction (range), or the squared deviations (standard deviation), and drags the answer with it.

An exam question that mentions skew or an outlier is telling you which column to pick from.

11. Spread: the second half of any description

Concept

Two towns can report the same average wage and be nothing alike. Centre without spread is half a description.

variance — The average squared distance from the mean. Squared, so that distances above and below do not cancel.

standard deviation — The square root of the variance, which puts the number back into the units of the data.

OpenStax, Introductory Business Statistics 2e §2.7

12. Why squares, and why the awkward divisor

Intuition

Deviations from the mean always sum to exactly zero — that is what being the balance point means. So the plain average deviation is useless: it is zero for every data set on earth.

Squaring removes the cancellation. Taking the square root at the end undoes the squaring so the answer is in dollars rather than dollars squared.

The divisor for a sample is one less than the sample size. The intuition: once you know the mean and all but one deviation, the last deviation is forced. Only that many of them are free to vary.

NIST/SEMATECH e-Handbook of Statistical Methods §1.3.5.6 — the standard treatment of degrees of freedom

13. Deviations cancel — that is the whole problem

Picture it

Five values with mean 8. Green bars are above the mean, red bars below.

Figure (svg): Bar chart of deviations from the mean for the values 4, 6, 8, 10 and 12, showing bars of minus 4, minus 2, 0, plus 2 and plus 4 that cancel to zero

Sum of deviations equals zero, always, for every data set.

14. Worked example: sample standard deviation

Worked example

The five values are 4, 6, 8, 10 and 12.

Find the mean

Why: The sum is 40 and there are five values, so the mean is 8.

valuedeviationsquared deviation
4-416
6-24
800
1024
12416

Add the squared deviations

Why: 16 plus 4 plus 0 plus 4 plus 16 gives 40.

\[ s^2 = \frac{40}{5-1} = 10 \]

\[ s = \sqrt{10} \approx 3.16 \]

Verify: that the deviations sum to zero

Why: Minus 4, minus 2, 0, plus 2, plus 4 add to 0. If they do not, the mean was computed wrongly and everything after it is wrong too.

15. Same mean, different worlds

Picture it

Two towns both report a mean weekly wage of 800 dollars. Here is their standard deviation, to scale.

Figure (svg): Two horizontal bars comparing a standard deviation of 40 for Town A against 260 for Town B

Identical centres. Nobody would call these the same labour market.

This is why a summary that reports only a mean is an incomplete summary, and why exam questions almost always ask for both.

16. Find the error in this standard deviation

Error analysis

A classmate hands you this. One line is wrong, and it is a line that produces a plausible-looking answer.

Annotate

On: \( s^2 = \frac{40}{5} = 8, \qquad s = \sqrt{8} \approx 2.83 \)

  • The divisor is the sample size, not the sample size minus one. This is the population variance formula applied to a sample.
  • It always understates the spread, and it understates it most when the sample is small — exactly when you can least afford it.
  • Correct: 40 divided by 4 equals 10, so s is about 3.16.

On an exam, the word 'sample' in the question is the instruction to use the smaller divisor.

17. Two towns, same mean

Comparison

Both towns report a mean weekly wage of 800 dollars. Fill in what is missing.

Comparison matrix

Town ATown B
mean800800
standard deviation40260
coefficient of variation5%32.5%
which town is riskier to move to—Town B

The coefficient of variation is the standard deviation as a percentage of the mean. It is how you compare spread across series measured in different units, which is most of macroeconomics.

18. Modelling with the Normal Curve

Section

Section 2

19. A distribution, not a data set

Concept

Monthly household electricity spending in a metro area is roughly normal with mean 120 dollars and standard deviation 18 dollars.

Notice what changed. You are no longer given seven numbers; you are given a model of the whole population, and asked what fraction of it lies somewhere.

FRED, Federal Reserve Bank of St. Louis — economic time series used for the scenarios — utility expenditure series behave this way at metro scale

20. The empirical rule, before any table

Picture it

Three bands you should be able to draw from memory.

Figure (svg): Normal curve with nested shaded bands at one, two and three standard deviations, labelled 68, 95 and 99.7 percent

68 within one sd, 95 within two, 99.7 within three.

For this model that means about 95 percent of households spend between 84 and 156 dollars.

21. Predict before you standardize

Prediction

A household spends 150 dollars this month.

Predict first

Roughly what fraction of households spend more than that?

  • About 1 in 3
  • About 1 in 10
  • About 1 in 20
  • About 1 in 200

Correct: About 1 in 20.

Why: 150 sits 30 dollars above the mean, and 30 is about 1.67 standard deviations. Two standard deviations would leave 2.5 percent in the upper tail; 1.67 leaves a bit more, 4.75 percent, which is roughly one household in twenty.

22. Worked example: from a dollar amount to a probability

Worked example

What proportion of households spend more than 150 dollars a month?

Write the standardizing formula

Why: A z-score is the distance from the mean measured in standard deviations, which is the only currency the standard normal table accepts.

\[ z = \frac{x - \mu}{\sigma} = \frac{150 - 120}{18} \]

\[ z = \frac{30}{18} = 1.67 \]

Look up the area to the left, then subtract

Why: The table gives the area below a z-score. You want the area above, so subtract from one.

\[ P(X > 150) = 1 - 0.9525 = 0.0475 \]

Verify: that the answer is smaller than 0.05

Why: z of 1.67 is inside two standard deviations, so the tail must be larger than 2.5 percent and smaller than 16 percent. 4.75 percent passes both checks.

23. The tail you just computed

Picture it

The shaded region is the 4.75 percent of households spending over 150 dollars.

Figure (svg): Normal curve centred at 120 with the region to the right of 150 shaded red, representing 4.75 percent of the distribution

Area under the curve is probability. There is nothing else to it.

24. Trap: subtracting in the wrong order

Trap

The trap

Computing the z-score for a value below the mean.

\[ z = \frac{\mu - x}{\sigma} = \frac{120 - 96}{18} = 1.33 \]

Report a positive z-score for a below-average value

Why: The arithmetic is fine. The sign is a lie, and the sign is what tells you which tail you are in.

The fix

The value always comes first. The z-score inherits the sign of the deviation, and that sign is the information.

\[ z = \frac{x - \mu}{\sigma} = \frac{96 - 120}{18} = -1.33 \]

Report the negative z-score

Why: Now the table lookup lands in the left tail, where a below-average household belongs.

25. Reading the notation on the exam

Notation

Exam questions hide the method inside the notation. Decode it before you compute anything.

Annotate

On: \( X \sim N(\mu = 120,\; \sigma = 18) \)

  • The tilde reads 'is distributed as'. It names the model, not a data set.
  • Greek letters mean population parameters. If the question had given you a sample, they would be x-bar and s instead.
  • The second argument is the standard deviation here. Some textbooks put the variance in that slot, so check which convention your paper uses.

26. When the outcome is a count, not a measurement

Concept

A regional bank finds that 30 percent of loan applications are incomplete. Ten applications arrive today.

Nothing here is measured on a continuous scale. You are counting successes in a fixed number of independent trials, each with the same probability. That is the binomial setting.

OpenStax, Introductory Business Statistics 2e §4.3

27. Pick the model, do not solve

Discrimination

The commonest exam mistake is not bad arithmetic. It is reaching for the wrong distribution.

Sort into buckets

Which model does each scenario call for?

binomial
Out of 10 applications, how many are incomplete?; In 200 calls, how many are abandoned?
normal
What share of households spend over 150 dollars?; What is the 90th percentile of delivery times?
bin
A fixed number of independent trials, each either a success or a failure, and you are counting the successes. The answer is a whole number.
norm
A continuous measurement, and you are asked about a proportion of the distribution or a percentile. The answer can be any real value.

28. Worked example: exactly three incomplete applications

Worked example

Ten applications, each incomplete with probability 0.3, independently.

Write the binomial probability formula

Why: The combination counts which three of the ten are the incomplete ones; the powers give the probability of any one such arrangement.

\[ P(X = 3) = \binom{10}{3}(0.3)^3(0.7)^7 \]

Evaluate the three pieces separately

Why: The combination is 120. The cube of 0.3 is 0.027. The seventh power of 0.7 is 0.0823543.

\[ P(X = 3) = 120 \times 0.027 \times 0.0823543 = 0.2668 \]

State the mean and variance while you are here

Why: Exam papers ask for these in the same question about half the time.

\[ \mu = np = 3, \qquad \sigma^2 = np(1-p) = 2.1 \]

Verify: that three is the most likely single count

Why: The mean is exactly 3, so the probability at 3 should be the largest of the eleven possible counts. About 27 percent for the single most likely value out of eleven is the right order of magnitude.

29. All eleven outcomes, not just the one you were asked for

Picture it

The bar at 3 is the 0.2668 you computed. Seeing the rest is how you sanity-check it.

Figure (svg): Bar chart of binomial probabilities for zero through ten incomplete applications, peaking at three

Ten trials with probability 0.3 each: the distribution peaks at the mean, which is 3.

The bars must add to one. If a single probability you computed is bigger than about 0.4 in a distribution this wide, something went wrong.

30. Pattern: turning a question into a distribution

Pattern

Before the first check, lock in the routing. Almost every marked-down answer in this course is a right method applied to the wrong question.

  1. Read the outcome. A count of successes in a fixed number of trials is binomial. A measurement on a continuous scale is normal.
  2. Read what is asked. 'Exactly k' means one binomial term. 'At least k' means a sum of them, or the complement. 'What proportion' means an area under the normal curve.
  3. Standardize, if normal. Subtract the mean, divide by the standard deviation, and keep the sign.
  4. Look up the area to the left, then adjust. Tables give left areas; subtract from one for an upper tail, subtract two left areas for a middle band.
  5. Sanity-check against the empirical rule. Anything past two standard deviations must come out under 5 percent.

OpenStax, Introductory Business Statistics 2e §4.3 and §6.2

31. Check: one binomial, one normal

Check

Solve it on paper before you click.

Check your understanding

Thirty percent of loan applications are incomplete. In a batch of 10, what is the probability that none is incomplete?

  • A. 0.0282 (correct)
  • B. 0.3000
  • C. 0.7000
  • D. 0.0000

Answer: A

Why: Zero successes means all ten trials are failures, so the probability is 0.7 raised to the tenth power, which is 0.0282. The combination term is 1 and the 0.3 term is raised to the power zero, which is also 1.

Why B tempts people
That is the probability that a single application is incomplete, not the probability that none of ten is.
Why C tempts people
That is the probability that one particular application is complete. Ten of them must all be complete, so the 0.7 has to be raised to the tenth power.
Why D tempts people
The probability is small but far from impossible — roughly one batch in thirty-five.

32. From Sample to Population

Section

Section 3

33. The statistic is random; the parameter is not

Concept

Draw 36 households, average their spending, write the number down. Throw them back and draw 36 more. You get a different number.

That variation is not error in your arithmetic. It is the sampling distribution, and quantifying it is the entire content of inference.

OpenStax, Introductory Statistics 2e §7.1

34. The one question the whole section answers

Socratic

Take a minute on this before the formula appears.

Discussion prompt

Why should the average of 36 households be closer to the truth than a single household, when both are drawn from exactly the same population?

Hint: What happens to one unusually large value when you average it with 35 ordinary ones?

Answer:

Because unusually high and unusually low households partly cancel inside the average. One extreme household is the whole estimate; one extreme household among 36 is a thirty-sixth of it.

The cancellation is only partial, which is why the spread shrinks by the square root of the sample size rather than by the sample size itself.

\[ SE = \frac{\sigma}{\sqrt{n}} \]

NIST/SEMATECH e-Handbook of Statistical Methods §1.3.5

35. Same centre, tighter spread

Picture it

Three distributions: one household, the mean of nine, the mean of thirty-six.

Figure (svg): Three normal curves with the same centre and progressively narrower spread, labelled one household, mean of nine and mean of thirty-six

Quadrupling the sample halves the standard error.

Note what does not change: the centre. A larger sample makes the estimate more precise, never less biased.

36. The standard error, from memory

Fill the middle

You are given a population standard deviation of 18 and a sample of 36.

Fill in the blanks

SE = \frac1836} = \frac3}___}}} = ___

Why: The square root of 36 is 6, and 18 divided by 6 is 3. The standard error is in the same units as the data — dollars — and it is always smaller than the population standard deviation, which is the point of averaging.

37. Worked example: a 95 percent confidence interval

Worked example

A sample of 36 households has a mean spend of 126.40 dollars. Population standard deviation is known to be 18.

Compute the standard error

Why: Eighteen divided by the square root of thirty-six, which is 3.

\[ SE = \frac{18}{6} = 3 \]

Multiply by the critical value for 95 percent

Why: For a known population standard deviation the critical value is 1.96, the z-score cutting 2.5 percent into each tail.

\[ E = 1.96 \times 3 = 5.88 \]

\[ 126.40 \pm 5.88 \Rightarrow (120.52,\; 132.28) \]

Verify: that the sample mean sits exactly at the centre

Why: Half the width is 5.88, and 126.40 minus 120.52 is 5.88. If the estimate is not centred, the margin was applied to the wrong number.

38. What the interval is, drawn

Picture it

The yellow dot is your estimate. The band is the interval.

Figure (svg): A number line from 118 to 134 with a shaded confidence band from 120.52 to 132.28 and the sample mean marked at 126.4

The interval is a property of the procedure, not of the population.

39. Trap: the sentence that loses the mark

Trap

The trap

"There is a 95 percent probability that the true mean lies between 120.52 and 132.28."

Assign a probability to the parameter

Why: The true mean is a fixed number. It is either inside this interval or it is not; there is no randomness left in it to carry a probability.

Lose the mark even though the interval is right

Why: Graders mark this sentence wrong every single time, and it is the most common way to lose points on an otherwise perfect answer.

The fix

"We are 95 percent confident that the true mean lies between 120.52 and 132.28."

Assign the confidence to the method

Why: Ninety-five percent of intervals built this way, over repeated samples, contain the true mean. This one either does or does not.

Say what varies

Why: The interval moved because the sample moved. The parameter never moved at all.

40. Three statements about the interval — two hold up

Two truths and a lie

The interval is (120.52, 132.28) from a sample of 36.

Eliminate the wrong options

Which statement survives scrutiny?

  • A. Repeating this procedure on many samples, about 95 percent of the intervals produced would contain the true mean.
  • B. Ninety-five percent of households spend between 120.52 and 132.28 dollars.
  • C. If we increased the sample to 144, the interval would be about twice as wide.
  • D. The true mean has a 95 percent chance of being in this particular interval.

Survives elimination: A

Why: Confidence is a long-run property of the recipe. The only statement that describes the recipe rather than this one interval or this one population is the first, which is why it is the one that survives.

41. When the population standard deviation is unknown

Concept

In practice you almost never know the population standard deviation. You estimate it from the sample, and that extra estimation costs you something.

The cost is paid by switching from the normal critical value to a t critical value, which is slightly larger, giving a slightly wider interval. The degrees of freedom are one less than the sample size.

OpenStax, Introductory Business Statistics 2e §8.2

42. What happens as the sample grows without limit?

Edge cases

The t distribution has one parameter: degrees of freedom.

Discussion prompt

As the degrees of freedom grow larger and larger, what does the t distribution turn into, and why does that make sense?

Hint: What is the t distribution correcting for, and does that correction still matter when n is 10000?

Answer:

It converges to the standard normal. With a huge sample, the sample standard deviation is essentially the population standard deviation, so the extra uncertainty the t distribution exists to price disappears.

Practically: past about 30 degrees of freedom the two critical values differ in the second decimal place. That is why older textbooks use 30 as a rule of thumb.

NIST/SEMATECH e-Handbook of Statistical Methods §1.3.6.6.4

43. Worked example: a t interval

Worked example

A sample of 25 firms reports a mean delivery cost of 48.20 dollars with sample standard deviation 6.50. Build a 95 percent interval.

Compute the standard error using the sample standard deviation

Why: Six point five divided by the square root of twenty-five, which is 5.

\[ SE = \frac{6.50}{5} = 1.30 \]

Find the critical value with 24 degrees of freedom

Why: One less than the sample size. For 95 percent and 24 degrees of freedom the t critical value is 2.064.

\[ E = 2.064 \times 1.30 = 2.68 \]

\[ 48.20 \pm 2.68 \Rightarrow (45.52,\; 50.88) \]

Verify: that this interval is wider than the z version would be

Why: Using 1.96 instead of 2.064 would give a margin of 2.55. The t interval is wider, which is the price of not knowing the population standard deviation.

44. What the t distribution costs you

Picture it

Same sample, same confidence level. The only difference is whether the population standard deviation was known.

Figure (svg): Two bars comparing a margin of error of 2.55 using the z critical value against 2.68 using the t critical value

Estimating the spread from the sample widens the interval by about five percent here.

The gap shrinks as the sample grows, which is the whole reason the t distribution converges to the normal.

45. Testing a Claim

Section

Section 4

46. A courier promises thirty minutes

Concept

A courier advertises a mean delivery time of 30 minutes. You sample 40 deliveries and find a mean of 32.1 minutes with a sample standard deviation of 5.4.

The question is not whether 32.1 differs from 30. It obviously does. The question is whether a gap that size is surprising when the claim is true.

OpenStax, Introductory Business Statistics 2e §9.1

47. Plan it in English first

Step zero

Before any symbol goes on the page.

Discussion prompt

State, in one plain sentence each: what the null hypothesis says, what the alternative says, and what would have to be true of the data for you to reject the null.

Hint: Which of the two hypotheses gets the benefit of the doubt?

Answer:

Null: the true mean delivery time is 30 minutes, and the 2.1 minute gap is ordinary sampling noise.

Alternative: the true mean is not 30 minutes.

You reject when the observed gap is too large to be comfortably explained by noise — where 'too large' was fixed in advance by choosing a significance level.

Note the asymmetry: you never prove the null. You either find enough evidence against it or you do not.

48. Put the test in order

Ranking

These are the steps of every hypothesis test in the course, shuffled.

Put in order

  1. State the null and alternative hypotheses
  2. Choose the significance level
  3. Compute the test statistic from the sample
  4. Find the p-value or the critical value
  5. Compare, and decide whether to reject
  6. State the conclusion in the words of the original problem

Why: The significance level is chosen before the statistic is computed. Choosing it afterwards, once you have seen the p-value, is the definitional form of cheating in this course — and the last step is the one students skip and graders always look for.

49. Worked example: a two-tailed t-test

Worked example

Null: the mean is 30. Alternative: the mean is not 30. Significance level 5 percent, so 2.5 percent in each tail.

Compute the standard error

Why: The sample standard deviation is 5.4 and the square root of 40 is about 6.3246.

\[ SE = \frac{5.4}{\sqrt{40}} = 0.8538 \]

Compute the test statistic

Why: How many standard errors the observed mean sits from the hypothesized mean.

\[ t = \frac{32.1 - 30}{0.8538} = 2.46 \]

Find the p-value with 39 degrees of freedom

Why: Two-tailed, so double the upper-tail area. The result is about 0.018.

\[ p \approx 0.018 < 0.05 \]

State the conclusion in the language of the problem

Why: Not 'reject the null' on its own — say what that means about couriers.

At the 5 percent level there is sufficient evidence that the courier's mean delivery time differs from the advertised 30 minutes.

Verify: that the statistic is beyond the critical value

Why: For 39 degrees of freedom the two-tailed critical value is about 2.02. The observed 2.46 is further out, which must agree with a p-value below 0.05 — and it does.

50. The decision, drawn

Picture it

Red tails are the 5 percent you agreed in advance to risk. The yellow mark is what you observed.

Figure (svg): A t distribution with both tails beyond plus and minus 2.02 shaded red and the observed statistic 2.46 marked in the right tail

The statistic landed in the rejection region, so the p-value is below the significance level.

The p-value and the critical value are two readings of the same picture. They can never disagree.

51. Explain the p-value to someone a year behind you

Explain it

Say it out loud. If it takes more than two sentences, it is not clear yet.

Discussion prompt

What does a p-value of 0.018 actually mean here? Deliberately avoid the words 'probability the null is true'.

Hint: The p-value conditions on the null being true. Start your sentence with 'if'.

Answer:

If the courier's true mean really were 30 minutes, then samples of 40 would produce a mean at least 2.1 minutes away from 30 about 1.8 percent of the time.

It is the probability of data this extreme given the null, not the probability of the null given the data. Those are different quantities and reversing them is the single most common misreading in applied statistics.

Wasserstein & Lazar, "The ASA Statement on p-Values: Context, Process, and Purpose", The American Statistician 70(2), 129-133 (2016) principles 2 and 3 — the ASA wrote a whole statement because of how often this is reversed

52. The two ways to be wrong

Picture it

Every test can fail in exactly two directions, and lowering the risk of one raises the risk of the other.

Figure (svg): A two by two table showing correct decisions on the diagonal, a Type I error when rejecting a true null, and a Type II error when failing to reject a false null

The significance level is the Type I error rate you chose to accept.

Increasing the sample size is the only move that reduces both at once. That is why it is the answer to so many exam questions.

53. Trap: accepting the null hypothesis

Trap

The trap

A test returns a p-value of 0.32.

Conclude that the mean is 30 minutes

Why: This treats absence of evidence as evidence of absence.

Write 'we accept the null hypothesis'

Why: The test never had the power to establish the null. A tiny sample fails to reject almost any null, including badly wrong ones.

The fix

A test returns a p-value of 0.32.

Conclude that there is insufficient evidence that the mean differs from 30 minutes

Why: This says exactly what happened: the data did not clear the bar, and the null survives by default rather than by proof.

Write 'we fail to reject the null hypothesis'

Why: Clumsy English, deliberately. The clumsiness is carrying real logical content.

54. One conclusion is defensible

Elimination

The test gave t equal to 2.46 with a p-value of 0.018 at the 5 percent level.

Eliminate the wrong options

Which conclusion would you write on the exam?

  • A. There is sufficient evidence at the 5 percent level that the mean delivery time differs from 30 minutes.
  • B. The probability that the courier's claim is false is 98.2 percent.
  • C. The mean delivery time is 32.1 minutes.
  • D. The courier is deliberately misleading customers.

Survives elimination: A

Why: A defensible conclusion names the significance level, refers to evidence rather than proof, and talks about the parameter in the units of the original problem. Only the first choice does all three.

55. Fitting a Line

Section

Section 5

56. Does advertising move sales?

Concept

Five months of data from one regional store. Advertising spend and sales are both in thousands of dollars.

advertisingsales
13
27
35
411
514

The relationship is clearly positive and clearly not perfect. Regression is the tool for describing exactly how positive and exactly how imperfect.

Wooldridge, Introductory Econometrics: A Modern Approach, 7th ed., Ch. 2 (The Simple Regression Model) §2.1

57. Look at the scatter before fitting anything

Picture it

Five points, one fitted line. Draw the scatter first, every time.

Figure (svg): Scatter plot of advertising spend against sales with five points and an upward sloping fitted line

The line minimises the total squared vertical distance to the points.

If the scatter shows a curve, or one point far from the rest, no amount of correct arithmetic afterwards will save the answer.

58. Guess the slope before computing it

Hypothesis

Sales rise from about 3 to about 14 as advertising rises from 1 to 5.

Predict first

Roughly what slope do you expect?

  • About 0.4
  • About 1.3
  • About 2.6
  • About 5.5

Correct: About 2.6.

Why: The total rise is roughly 11 over a run of 4, giving about 2.75 as a crude eyeball estimate. The least-squares slope comes out at 2.6, close to the eyeball value because the relationship really is close to linear.

59. Worked example: the least-squares line

Worked example

Compute the two means first. The sum of advertising is 15 and the sum of sales is 40, over five months.

\[ \bar{x} = 3, \qquad \bar{y} = 8 \]

xyx - xbary - ybarproductsquared x deviation
13-2-5104
27-1-111
350-300
4111331
51426124

Add the last two columns

Why: The products sum to 26 and the squared deviations sum to 10.

\[ b_1 = \frac{26}{10} = 2.6 \]

Put the line through the point of averages

Why: Every least-squares line passes through the pair of means. That fact is what pins down the intercept.

\[ b_0 = 8 - 2.6 \times 3 = 0.2 \]

\[ \hat{y} = 0.2 + 2.6x \]

Verify: that the line passes through the point of averages

Why: Substituting x equal to 3 gives 0.2 plus 7.8, which is 8, exactly the mean of the sales column. If it does not, the intercept is wrong.

60. Residuals: what the line could not explain

Picture it

Hollow circles are the fitted values, filled dots the actual sales. The dashed gaps are the residuals.

Figure (svg): The same scatter with vertical dashed segments joining each observed point to the fitted line, marking the residuals

Residuals from a least-squares fit always sum to zero.

The largest miss is at three thousand of advertising, where the line predicts 8 and sales were only 5.

61. Watch one column as the fit proceeds

Invariant

Step through the five residuals and keep a running total.

Step through it

What is true of the running total at the end, and would it still be true for a different data set?

  1. First residual: sales were 0.2 above the line.
  2. Second: another 1.6 above.
  3. Third: 3.0 below — the worst miss.
  4. Fourth: back up by 0.4.
  5. Fifth: 0.8 above, and the total lands exactly on zero.

It is zero for every least-squares fit with an intercept. It is not a coincidence and it is not evidence that the fit is good — it is forced by the arithmetic that produced the intercept.

62. How much of the variation did the line explain?

Concept

Total variation in sales is the sum of squared deviations from the mean sales, which is 80. The line leaves 12.4 of that unexplained.

R-squared — The share of the total variation in the outcome that the fitted line accounts for. Always between zero and one for a simple regression with an intercept.

\[ R^2 = \frac{67.6}{80} = 0.845 \]

Wooldridge, Introductory Econometrics: A Modern Approach, 7th ed., Ch. 2 (The Simple Regression Model) §2.3

63. Eighty units of variation, split in two

Picture it

The green bar is what the line explains; the red bar is what it misses.

Figure (svg): Two horizontal bars comparing explained variation of 67.6 against unexplained variation of 12.4 out of a total of 80

R-squared is simply the green bar as a share of the whole.

64. Trap: reading R-squared as proof of causation

Trap

The trap

R-squared is 0.845.

Conclude that advertising causes 84.5 percent of sales

Why: Two errors in one sentence: R-squared is about variation, not about levels of sales, and it is a measure of fit, not of cause.

Recommend tripling the advertising budget

Why: Nothing in the data rules out a third variable — a seasonal demand cycle, say — driving both advertising and sales upward together.

The fix

R-squared is 0.845.

Conclude that the line accounts for 84.5 percent of the variation in sales across these five months

Why: Correct scope: variation, these months, this sample.

Say what would be needed for a causal claim

Why: An experiment, or a design that rules out confounders. Regression on observational data describes association and stops there.

65. The five-step pattern behind every question on this exam

Pattern

Every practice-paper question you showed me is an instance of this. Name the step you are on and the formula follows.

  1. Describe — centre and spread, and say which measure of centre the shape of the data justifies
  2. Model — decide whether the outcome is a count or a measurement, then standardize or use the binomial formula
  3. Quantify the noise — compute the standard error, which is the standard deviation divided by the root of the sample size
  4. Estimate or test — a confidence interval when the question asks 'how big', a hypothesis test when it asks 'is it different'
  5. Fit and interpret — slope in the units of the problem, R-squared as a share of variation, and no causal language
the question saysthe tool
'typical', 'average', 'skewed'mean versus median, and the coefficient of variation
'what proportion', 'what percentage of'z-score, then the normal table
'exactly k out of n'binomial formula
'estimate the mean', 'margin of error'confidence interval
'test the claim', 'is there evidence'hypothesis test, then a sentence in context
'predict', 'for each additional'regression slope

OpenStax, Introductory Business Statistics 2e chapters 2, 4, 6, 8, 9, 13

66. Match the phrase to the tool

Matching

Exam questions signal the method in their wording. Learn the signals and half the work is done before you compute anything.

Match the pairs

  • l1. "for each additional thousand dollars spent"
  • l2. "is there sufficient evidence at the 1 percent level"
  • l3. "what percentage of households"
  • l4. "estimate the true mean with 95 percent confidence"
  • l5. "the data are strongly right-skewed"
  • r1. the regression slope
  • r2. a hypothesis test
  • r3. a z-score and the normal table
  • r4. a confidence interval
  • r5. report the median, not the mean

Why: Slope answers 'per unit change'. A stated significance level always means a test. 'What percentage' means an area under a curve. 'Estimate with confidence' means an interval. And any mention of skew is a hint about which measure of centre to report.

67. Check: interpret the slope

Check

Solve it on paper before you click.

Check your understanding

The fitted line is sales equals 0.2 plus 2.6 times advertising, both measured in thousands of dollars. Which interpretation is correct?

  • A. Each additional thousand dollars of advertising is associated with 2,600 dollars more in sales, on average. (correct)
  • B. Each additional thousand dollars of advertising causes sales to rise by 2,600 dollars.
  • C. Sales are 2.6 times advertising spend.
  • D. When advertising is zero, sales are 2,600 dollars.

Answer: A

Why: The slope is a rate of change per unit of the predictor, stated as an association because the data are observational. Both variables are in thousands, so 2.6 means 2,600 dollars.

Why B tempts people
Correct arithmetic, wrong verb. Observational data support association; only a designed experiment supports 'causes'.
Why C tempts people
That would be true only if the intercept were zero and every point sat exactly on the line. It confuses a slope with a ratio.
Why D tempts people
That is the intercept, which is 0.2 thousand, or 200 dollars — and even that is an extrapolation outside the observed range of advertising.

68. Check: which interval is wider

Check

Solve it on paper before you click.

Check your understanding

You build a 95 percent confidence interval for a mean from a sample of 25, then rebuild it from a sample of 100 with the same sample standard deviation. What happens to the width?

  • A. It shrinks to about half its original width. (correct)
  • B. It shrinks to about a quarter of its original width.
  • C. It stays the same, because the confidence level did not change.
  • D. It grows, because more data means more variability.

Answer: A

Why: Width depends on the standard error, which divides by the square root of the sample size. Going from 25 to 100 quadruples n, and the square root of 4 is 2, so the standard error halves and so does the width.

Why B tempts people
That would happen if the standard error divided by n rather than by the root of n. The square root is why precision is expensive.
Why C tempts people
The confidence level sets the critical value, but the standard error is the other factor in the width, and it fell.
Why D tempts people
More data reduces the variability of the estimate. The variability of individual observations is unchanged, but the interval is about the mean.

69. Exit ticket: name your weakest spot

Exit ticket

Be honest — this is what we spend the next session on.

Predict first

Which of the five steps would you least want to see first on Monday's paper?

  • Describing a sample and choosing between mean and median
  • Standardizing and reading normal probabilities
  • Confidence intervals and the standard error
  • Hypothesis tests and stating the conclusion
  • Regression: slope, residuals and R-squared

Correct: Whichever one you picked is the one we open with next session.

Why: There is no wrong answer here. Naming the weak step is worth more than another pass over the strong ones, and four days is enough time to fix exactly one thing properly.

70. Draw the whole course on one page

Connect it up

One page, no notes. This is the single best use of twenty minutes before the exam.

Draw it

Start with a box labelled 'sample' and a box labelled 'population'. Draw every arrow between them that this course has given you, and label each arrow with the formula that travels along it.

If an arrow has no formula on it, that is the gap. Go find it.

71. What you can do now

Recap

Four days out, this is the checklist. Everything on the exam is one of these five moves, or two of them stacked.

quantityformula in wordsvalue in this deck
standard errorsample standard deviation over the root of n3.00
z-scorevalue minus mean, over standard deviation1.67
margin of errorcritical value times standard error5.88
test statisticestimate minus hypothesized value, over standard error2.46
slopesum of cross products over sum of squared x deviations2.60
R-squaredexplained variation over total variation0.845

OpenStax, Introductory Business Statistics 2e — every formula above appears there with a worked example

Sources

  1. OpenStax, Introductory Business Statistics 2e
  2. OpenStax, Introductory Statistics 2e
  3. NIST/SEMATECH e-Handbook of Statistical Methods
  4. Wasserstein & Lazar, "The ASA Statement on p-Values: Context, Process, and Purpose", The American Statistician 70(2), 129-133 (2016)
  5. Wooldridge, Introductory Econometrics: A Modern Approach, 7th ed., Ch. 2 (The Simple Regression Model) — Cengage, 2020
  6. FRED, Federal Reserve Bank of St. Louis — economic time series used for the scenarios
  7. U.S. Bureau of Labor Statistics, Current Employment Statistics — average hourly earnings

Want this taught 1-on-1? Alexander tutors Economic Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108