Every lesson in the Statistics slide course, in full text: 58 decks, 3890 slides.
1.1 Definitions of Statistics, Probability, and Key TermsSeparate a population from a sample, a parameter from a statistic, and numerical from categorical data using a college whose every first-year is visible; then see probability as the proportion long runs of coin tosses settle on.
1.2 Data, Sampling, and Variation in Data and SamplingThe taxonomy of data and the methods for collecting it: qualitative data that name a category against quantitative data that are always numbers, and within quantitative the split between discrete data that are counted and continuous data that are measured; pie charts and bar graphs for qualitative data and the two situations where a pie chart is not merely ugly but wrong; the four random sampling methods — simple random, stratified, cluster and systematic — with convenience sampling as the non-random alternative; and the difference between sampling error, which every sample has and a larger sample reduces, and sampling bias, which no sample size can cure.
1.3 Frequency, Frequency Tables, and Levels of MeasurementThe four levels of measurement as a ladder — nominal names with no order, ordinal adding order but no measurable differences, interval adding meaningful differences but no true zero, and ratio adding a true zero so that ratios become meaningful — and which operations each level permits. Then the frequency table and its two derived columns: relative frequency as each count divided by the sample size, and cumulative relative frequency as the running total that must finish at one. Closes with grouped frequency tables for continuous data, and the book's rounding convention of carrying one more decimal place than the original data.
1.4 Experimental Design and EthicsWhat a study must do before its results can support a claim about cause: the anatomy of a randomized experiment with its explanatory and response variables, treatments and experimental units; lurking variables and why random assignment rather than random selection is what neutralises them; the control group and the placebo that make the power of suggestion measurable; blinding and double-blinding and what each one protects against; and the studies that cannot be randomized at all, where an observational design shows association and cannot establish cause. Closes with the ethical obligations a researcher carries toward the people being studied.
2.1 Stem-and-Leaf Graphs (Stemplots), Line Graphs, and Bar GraphsCut every value into a stem and a leaf, read counts and far values off the finished plot, put two data sets on one stem column, join counted points across equal steps, and turn head counts into shares of a named whole.
2.2 Histograms, Frequency Polygons, and Time Series GraphsThe display the rest of the book uses. What a histogram is and why its boxes are contiguous rather than separated; how to choose a starting point one decimal place finer than the data so that no value lands on a boundary, and how to compute the class width once the number of bars is chosen; the discrete case, where a width of one centres each value in its own bar; the frequency polygon, which plots each class midpoint at its frequency and joins them, anchored at empty classes so the shape closes, and which can be overlaid to compare two distributions; and the time series graph, which plots a measurement against its date and so preserves the chronological order every other display discards.
2.3 Measures of the Location of the DataWhere a value sits inside its data set. Quartiles cut the ordered data into four parts, with the first quartile the median of the lower half and the third the median of the upper half; the interquartile range measures the spread of the middle fifty percent and supplies the 1.5 IQR rule that turns a visual judgement about outliers into arithmetic anyone can repeat; percentiles cut the data into hundredths and are what a score in the ninetieth percentile actually means; and two formulas run in opposite directions, one taking a percentile and returning a value by the index rule, the other taking a value and returning its percentile by counting what lies below it.
2.4 Box PlotsFive numbers drawn as a picture. A box plot is built from the minimum, the first quartile, the median, the third quartile and the maximum, with a box spanning the middle fifty percent and whiskers reaching the extremes; because each of the four parts holds about a quarter of the data, the width of a part reports how tightly that quarter is packed rather than how many values it contains. Then the modified box plot, whose whiskers stop at the most extreme non-outliers and whose outliers are marked as separate points, and parallel box plots on one shared number line, which are the cheapest honest comparison of several distributions. Closes with what a box plot cannot show.
2.5 Measures of the Center of the DataBuild the mean from equal sharing and balance, locate the median and see why extreme values cannot move it, count modes, pool frequency tables, estimate grouped means, and meet the law of large numbers.
2.6 Skewness and the Mean, Median, and ModeFold a dot plot to test it for symmetry, watch the three centres split when the fold fails, measure how far one value moves the mean and why it leaves the median alone, and read a data set's stretch back out of its mean and median.
2.7 Measures of the Spread of the DataBuild variance and standard deviation from squares you can see, prove why a sample divides by n − 1, derive Chebyshev's inequality, and compare data sets with z-scores.
3.1 TerminologyCount probabilities from pictures of equally likely outcomes, read probability as a long-run relative frequency, combine events with OR, AND and complements, and shrink the sample space to find conditional probabilities.
3.2 Independent and Mutually Exclusive EventsBuild independence from a grid whose columns are one event and whose rows are the other, derive all three tests from it, separate mutually exclusive events from independent ones, and work draws with and without replacement.
3.3 Two Basic Rules of ProbabilityThe two rules that do the arithmetic of the chapter, and the two special cases that simplify them. The multiplication rule computes an AND probability as P(B) times the conditional P(A given B), and collapses to a plain product exactly when the events are independent. The addition rule computes an OR probability as the sum of the two probabilities minus their intersection, and loses the subtraction exactly when the events are mutually exclusive. Both general forms are always valid, so section 3.2's two properties turn out to be permissions to drop a term rather than separate techniques — which is why identifying them was worth the effort.
3.4 Contingency TablesThe display that turns the arithmetic of this chapter into reading. A contingency table shows sample values in relation to two variables that may be dependent on one another, with a cell for every combination, row and column totals in the margins, and a grand total in the corner. A joint probability is one cell over the grand total; a marginal probability is a margin over the grand total; and a conditional probability is a cell over its own row or column total, which is section 3.1's reduced sample space made literal. Section 3.2's independence test becomes a comparison of two fractions read from the same table.
3.5 Tree and Venn DiagramsThe last two displays of the chapter, for when a probability problem is complex enough to be worth drawing. A tree diagram lays out a sequence of stages with a probability or a frequency on every branch, so a path's probability is the product along it and a conditional probability is read by keeping one subtree and discarding the rest — which makes it the multiplication rule of section 3.3 drawn, with the conditional impossible to overlook. A Venn diagram shows two events as overlapping regions with a count in each, so unions, intersections and complements all become sums of regions. Closes by choosing among the tree, the Venn diagram and the contingency table.
4.1 Probability Distribution Function (PDF) for a Discrete Random VariableThe chapter's opening move: instead of asking for the probability of one event, ask how probability is distributed across every value a quantity can take. A random variable attaches a number to each outcome of an experiment, and a discrete probability distribution function lists each value the variable can take beside the probability of taking it. Such a function has exactly two characteristics — each probability lies between zero and one inclusive, and the probabilities sum to one — and a table failing either is not a distribution at all. Because distinct values are mutually exclusive, every range question is a plain sum, and because the variable takes isolated values the distribution is drawn as separated stems rather than a curve.
4.2 Mean or Expected Value and Standard DeviationThe centre and spread of a distribution rather than of a data set. The expected value is the long-term average of a random variable, found by multiplying each value by its probability and adding the products, and it is denoted by the Greek letter mu because it is a parameter of a process rather than a statistic computed from data. It need not be one of the values the variable can take. The standard deviation is the square root of the sum of squared deviations weighted by their probabilities, with no divisor because the probabilities already sum to one. Both are computed from a single expected value table, and the Law of Large Numbers is what connects them to what actually happens over many trials.
4.3 Binomial DistributionThe first named family of distributions. A binomial experiment has three characteristics: a fixed number of trials n, exactly two outcomes on each trial with success probability p and failure probability q summing to one, and independent trials repeated under identical conditions so that p never changes. The random variable X counting the successes then has a binomial distribution, written X follows B of n and p, and its probabilities come from a formula whose three factors count the orderings and give the probability of one of them. Its mean is np and its standard deviation is the square root of npq, so two numbers determine the whole distribution — which is what makes a named family worth having.
4.4 Geometric DistributionThe binomial's first characteristic relaxed. A geometric experiment repeats a trial until the first success occurs and then stops, so the number of trials is the random variable rather than the number of successes, and it takes the values one, two, three and upward with no bound. The probability of a first success on trial x is q to the power x minus one times p, with no combination factor because exactly one sequence produces it. The mean is one over p, matching the intuition that a one-in-three chance takes about three tries. The distribution is also the only discrete one that is memoryless: previous failures carry no information about what comes next, which is a genuine property of independent trials and a poor model for anything that learns.
4.5 Hypergeometric DistributionThe binomial's independence characteristic relaxed. A hypergeometric experiment samples without replacement from two groups, so each pick changes what remains, the picks are not independent, and there is no single success probability to raise to a power. The count of items drawn from the group of interest has a hypergeometric distribution, written X follows H of r, b and n, and its probabilities are computed by counting combinations: the ways of choosing x from the group of interest, times the ways of choosing the rest from the second group, over the ways of choosing the sample at all. The possible values of X are limited by the group sizes as well as by the sample size, and the mean is the sample size times the proportion in the group of interest.
4.6 Poisson DistributionThe last discrete family in the chapter and the only one that does not count successes among trials. A Poisson variable counts occurrences in a fixed interval of time or space, when those events happen with a known average rate and independently of the time since the last one. It is written X follows P of mu, where mu is the mean number of occurrences per interval, and mu scales in proportion when the interval is lengthened or shortened. Its single parameter is both its mean and its variance, so the standard deviation is the square root of mu, and the family doubles as an approximation to the binomial when the number of trials is large and the success probability small.
5.1 Continuous Probability FunctionsThe chapter's hinge, and a genuine change of machinery. A continuous random variable can take any value in an interval rather than isolated whole numbers, and its distribution is described by a probability density function whose defining property is that the area between it and the horizontal axis equals a probability. Since the largest probability is one, the largest area is one. Three consequences follow immediately: probabilities belong to intervals rather than to single values, the probability of any one exact value is zero because a vertical line has no width and therefore no area, and strict and non-strict inequalities give the same answer because they enclose the same region. The cumulative distribution function gives the area to the left of a value, and the area to the right is one minus it.
5.2 The Uniform DistributionSection 5.1's flat density, named and given parameters. The uniform distribution is a continuous distribution concerned with events that are equally likely to occur across an interval, written X follows U of a and b, where a is the lowest value the variable can take and b the highest. Its density is one over b minus a, a height forced by the requirement that a rectangle of that width enclose an area of one. Because the region is always a rectangle, every question reduces to arithmetic on a base and a height: a probability is base times height, the mean is the midpoint of a and b, the standard deviation is the width divided by the square root of twelve, a percentile is found by setting the area equal to the given proportion and solving for the boundary, and a conditional probability is found by redrawing the density on the reduced interval.
5.3 The Exponential DistributionThe chapter's second named density and the first that is not flat. The exponential distribution is concerned with the amount of time until some specific event occurs, and its shape is a declining curve: there are fewer large values and more small ones. It has a single parameter, the decay parameter m, which is the reciprocal of the mean, and unusually its standard deviation equals its mean. Probabilities come from a closed form rather than a table, with the area to the left given by one minus e to the minus m x and the area to the right by e to the minus m x. Its defining property is memorylessness: the probability of waiting a further t, given that you have already waited r, equals the unconditional probability of waiting t. The lesson closes on the equivalence between exponential gaps between events and Poisson counts of events.
6.1 The Standard Normal DistributionThe most important distribution in the book, and the foundation of everything from chapter 7 onward. A normal distribution is bell shaped and symmetric about its mean, and it is fixed by exactly two numbers: the mean, which slides the curve left or right, and the standard deviation, which makes it fatter or skinnier. Since those two may be anything, there are infinitely many normal distributions, and one of special interest is the standard normal, whose mean is zero and standard deviation one. The z-score converts any normal value into a standard one by measuring its distance from the mean in units of the standard deviation, which both reduces every normal distribution to a single one and makes values on completely different scales directly comparable. The Empirical Rule then reads off the areas that matter most: about 68, 95 and 99.7 percent of values lie within one, two and three standard deviations of the mean.
6.2 Using the Normal DistributionSection 6.1's machinery put to work. With a calculator or computer returning areas under the normal curve, three questions can be answered about any normally distributed quantity: the probability of a tail, the probability of an interval, and the value at a given percentile. The technology reports only the area to the left, so a right tail is found by subtracting from one and an interval by subtracting one left area from another, while a percentile reverses the process by taking an area and returning the boundary. The genuine difficulty in this section is neither the arithmetic nor the technology but the translation of English phrases into areas to the left: at most, at least, the bottom quartile, the top ten percent and the middle twenty percent all specify an area, and two of them specify the complement of the number they mention.
7.1 The Central Limit Theorem for Sample Means (Averages)The result that makes the rest of the course possible. Chapter 6's normal machinery applies only to quantities that happen to be bell shaped, but the central limit theorem says that if random samples of size n are drawn from any population at all, then as n increases the distribution of the sample means tends toward a normal distribution. That distribution has the same mean as the original, and a standard deviation equal to the original standard deviation divided by the square root of the sample size, a quantity called the standard error of the mean. Two consequences follow. Sample means are far less variable than individual values, and they become less variable in proportion to the square root of the sample size rather than to the sample size itself. The sample size n counts the values averaged together to form one mean, not the number of times the sampling is repeated.
7.2 The Central Limit Theorem for SumsThe same theorem stated for totals rather than averages. If random samples of size n are drawn from any population with mean mu and standard deviation sigma, then as n increases the sum of the sample tends to be normally distributed, with a mean equal to the population mean multiplied by the sample size and a standard deviation equal to the population standard deviation multiplied by the square root of the sample size. The asymmetry between those two is the section's whole content: the centre of a total grows in proportion to n while its spread grows only in proportion to the square root of n, so a total becomes proportionally more predictable as more values are added. Probabilities and percentiles for a sum are then ordinary normal questions asked on that distribution, and because a sum is simply a mean multiplied by n, either form can be converted into the other as a check.
7.3 Using the Central Limit TheoremThe chapter's decision section. Sections 7.1 and 7.2 supplied the distributions of a sample mean and of a sum; this one supplies the question that must be answered before either can be used, namely whether the problem concerns an individual value, a mean, or a total. The book is explicit that a question about an individual value must be answered from that variable's own distribution and not from the central limit theorem at all. The section also states the law of large numbers, which says that as samples get larger their means get closer to the population mean, and observes that the central limit theorem illustrates it because the standard error shrinks toward zero. Two worked examples take populations that are conspicuously not normal, a uniform and an exponential, which is where the theorem does real work rather than restating what chapter 6 already covered.
8.1 A Single Population Mean using the Normal DistributionThe first act of statistical inference in the course. Chapter 7 asked how far a sample mean is likely to fall from a known population mean; this section reverses the question and asks which values of an unknown population mean are consistent with an observed sample mean. The answer is not a single number but an interval, formed as the point estimate plus and minus a margin of error called the error bound, which equals a z-score chosen by the confidence level multiplied by the standard error of the mean. The confidence level and its complement alpha determine that z-score, since alpha is split equally between the two tails. Raising the confidence level or shrinking the sample widens the interval, and the error bound formula can be solved backwards to find the sample size a study needs. The interpretation is the section's real difficulty: the interval is the random thing and the parameter is fixed, so the confidence level describes the procedure across repeated samples rather than any one interval.
8.2 A Single Population Mean using the Student t DistributionSection 8.1 assumed the population standard deviation was known, which in practice it almost never is. Replacing sigma with the sample standard deviation adds a second source of uncertainty, and for small samples the normal distribution no longer describes the resulting statistic. William Gosset, working at the Guinness brewery where experiments yielded very few samples, found the distribution that does, and published it under the pen name Student. The t distribution is symmetric about zero like the normal but has more probability in its tails and less in its centre, and its exact shape depends on the degrees of freedom, which equal the sample size minus one because the deviations used to compute the sample standard deviation must sum to zero. As the degrees of freedom grow the t curve approaches the normal, so the price of not knowing sigma is large for a small sample and negligible for a large one.
8.3 A Population ProportionSections 8.1 and 8.2 estimated a population mean. This one changes the parameter rather than the multiplier: when the data are categorical, the quantity of interest is the proportion of the population in one category, and the underlying distribution is the binomial of section 4.3. The point estimate is the sample proportion, written p prime, the number of successes divided by the number of trials, and its error bound is a z-score times the square root of p prime times q prime over n. No population standard deviation appears anywhere, because a binomial's variability is determined once its probability is. The section also introduces the plus-four adjustment, which adds two successes and two failures to the observed counts before proceeding, and gives more accurate intervals when the confidence level is at least ninety percent and the sample at least ten. Solving the error bound for n gives a planning formula, and using one half for the unknown proportion makes it as large as it can be.
9.1 Null and Alternative HypothesesChapter 8 asked which values of a parameter the data supports. Chapter 9 starts from a specific claim about the parameter and asks whether the data are consistent with it, and the whole apparatus rests on stating two contradictory hypotheses correctly before any arithmetic begins. The null hypothesis is a statement of no difference, often the status quo, and always contains a symbol with an equal in it. The alternative hypothesis contradicts it, never contains an equal sign, and is usually what the researcher is trying to show. Because the two are contradictory, evidence in the form of sample data can support one over the other, and the decision is always either to reject the null hypothesis or to decline to reject it. It is never to accept the null, because failing to find evidence against a claim is not evidence for it.
9.2 Outcomes and the Type I and Type II ErrorsA hypothesis test decides on incomplete evidence, so it can be wrong in two distinct ways. Crossing the two possible decisions with the two possible truths gives four outcomes: two of them correct, and two of them errors. A Type I error is rejecting the null hypothesis when it is true, and its probability is called alpha. A Type II error is failing to reject the null when it is false, and its probability is called beta. Both should be as small as possible and neither is ever zero, and they trade against each other — moving the decision boundary to reduce one increases the other. Only a larger sample reduces both at once. The probability of correctly rejecting a false null is one minus beta, and is called the power of the test. Which of the two errors matters more is a judgement about consequences that the statistics cannot supply.
9.3 Probability Distribution Needed for Hypothesis TestingA short section that does one job: it says which probability distribution each of the three one-sample tests uses, and what has to be true for that choice to be legitimate. A test of a mean when the population standard deviation is known uses a normal distribution; a test of a mean when it is unknown and estimated by the sample standard deviation uses a Student t; and a test of a proportion uses a normal distribution built on the binomial. The three rows are chapter 8's three confidence intervals in the same order and for the same reasons. What the section adds is an explicit statement of the assumptions behind each: a simple random sample in every case, approximate normality of the population for the two mean tests, and for a proportion the binomial conditions together with the requirement that n times p and n times q both exceed five, so that the binomial is close enough in shape to the normal for the approximation to hold.
9.4 Rare Events, the Sample, Decision and ConclusionThe section that supplies the logic behind a hypothesis test. The reasoning is the rare-event argument: make an assumption about the population, gather data, and if the sample has properties that would be very unlikely to occur when the assumption is true, conclude that the assumption is probably incorrect. The p-value measures that unlikeliness — it is the probability that, if the null hypothesis is true, another randomly selected sample would give results as extreme or more extreme than the ones obtained. A large p-value indicates the null should not be rejected, and the smaller the p-value the stronger the evidence against it. Comparing the p-value against a significance level chosen before the data were collected turns the measurement into a decision, and the conclusion must then be written in the words of the problem. The p-value is conditional on the null being true, which is what makes almost every common misreading of it wrong.
9.5 Additional Information and Full Hypothesis Test ExamplesThe chapter assembled. Sections 9.1 through 9.4 supplied the hypotheses, the two error types, the distributions and the decision rule, and this section adds the one remaining piece and then works complete tests. The remaining piece is the tail: when the p-value is drawn it occupies the left tail, the right tail, or is split evenly between both, and the alternative hypothesis is what decides which. A less-than alternative gives a left-tailed test, a greater-than gives a right-tailed one, and a not-equal gives a two-tailed test whose p-value is twice the area beyond the observed statistic. The significance level is chosen before the data are collected, and a common standard when none is given is five percent. The section then works four full tests, one for each of the three distributions, ending each with a conclusion in the words of the problem and an identification of the two error types in context.
10.1 Two Population Means with Unknown Standard DeviationsChapter 9 tested a parameter against a fixed number; this chapter tests two populations against each other, and the six steps of a hypothesis test carry over unchanged. What changes is the parameter, which is now the difference of two population means, its point estimate, which is the difference of the two sample means, and the standard error, which combines the variability of both samples by adding their variances. Because the population standard deviations are unknown and estimated by the two sample standard deviations, the test statistic follows a Student t distribution, and its degrees of freedom come from the Aspin-Welch formula — a calculation the book calls complicated, which is rarely a whole number and is done by machine. The book is emphatic that the sample variances are not pooled, since pooling would assume the two populations share a standard deviation that this test deliberately does not assume.
10.2 Two Population Means with Known Standard DeviationsThe shortest section in the chapter, and the book concedes at the outset that the situation it describes is not likely, since knowing the population standard deviations rarely happens. Everything from section 10.1 carries over unchanged except two things. The two population standard deviations replace the two sample standard deviations in the standard error, so the variances added are the true ones rather than estimates, and the sampling distribution of the difference becomes normal rather than a Student t. That second change removes the Aspin-Welch degrees of freedom entirely, which is the whole practical simplification. Both populations must be normal, the samples must be independent, and the test statistic is a z-score formed exactly as before: the difference of the sample means divided by the standard error of that difference.
10.3 Comparing Two Independent Population ProportionsThe same comparison as the previous two sections, with categorical data. The parameter is the difference of two population proportions, the point estimate is the difference of the two sample proportions, and because the difference of two proportions follows an approximate normal distribution, the test statistic is a z-score with no degrees of freedom. One feature is genuinely new and is the section's central idea: because the null hypothesis generally states that the two proportions are the same, the two samples are estimating one common value, so their successes and their trials are combined into a single pooled proportion before the standard error is formed. That is the opposite of the instruction in section 10.1, and the reason is that a binomial's spread follows from its probability — so a null of equal proportions is also a null of equal spreads, while a null of equal means says nothing about the two standard deviations.
10.4 Matched or Paired SamplesThe chapter's last section changes the design rather than the parameter. The three previous tests all required the two samples to be independent, and here they are not: two measurements are drawn from the same pair of individuals or objects, so a subject who scores high on the first will tend to score high on the second. The remedy is to calculate a difference for each pair, and those differences become the data. The population mean of the differences is then tested using a Student t test for a single population mean with n minus one degrees of freedom, where n is the number of differences — which is chapter 9's one-sample test applied to a single column of numbers. The design has a real advantage beyond correctness: because each subject serves as their own control, the variation between subjects is cancelled and only each subject's own change remains to be analysed.
11.1 Facts About the Chi-Square DistributionA short section introducing the distribution the whole chapter runs on. A chi-square random variable with k degrees of freedom is the sum of k independent squared standard normal variables, and every property the book lists follows from that definition. Because the terms are squares the statistic is never negative; because there are k of them the mean is k and the standard deviation is the square root of twice k; and because it is a sum of many independent quantities, the central limit theorem makes the curve approximately normal once the degrees of freedom exceed about ninety. The curve is nonsymmetrical and skewed to the right, there is a different curve for every degrees of freedom, and the mean sits just to the right of the peak. How the degrees of freedom are counted depends on which of the chapter's three tests is being run.
11.2 Goodness-of-Fit TestThe chi-square distribution's first use, and the chapter's first genuinely new question: instead of testing a single mean or proportion, a goodness-of-fit test asks whether a whole set of observed counts is consistent with a claimed distribution. The statistic sums observed minus expected, squared, over expected across the cells, so a departure in any cell and in either direction makes it larger — which is why the test is almost always right-tailed and why small statistics mean good agreement. The degrees of freedom are the number of categories minus one, never the number of observations minus one, and every expected count must be at least five, a condition that in the book's first example forces two categories to be combined before the test can be run at all. Worked through the book's four examples: absenteeism, absences by weekday, streaming services and a pair of coins.
11.3 Test of IndependenceThe same chi-square statistic as a goodness-of-fit test, with the expected counts built from the data's own margins rather than supplied by an outside claim. A test of independence uses a contingency table and asks whether two factors are related. The expected count in each cell is its row total times its column total divided by the grand total — a formula the book derives directly from chapter 3's definition of independence, since assuming the two factors are unrelated means each cell's probability is the product of its marginal probabilities. The degrees of freedom are rows minus one times columns minus one, which counts the cells still free once every margin is fixed, and the test is always right-tailed. Worked through the book's cell-phone and volunteer examples.
11.4 Test for HomogeneityA new question answered by an existing procedure. A goodness-of-fit test can decide whether a population fits a given distribution, but it cannot decide whether two populations follow the same unknown distribution — and that second question is often the one that matters, since a researcher comparing men with women or before with after usually has no external distribution to test against. The test for homogeneity fills the gap, and the book is explicit that its statistic is computed exactly as the test of independence's: expected counts from the margins, the same sum of squared departures, the same right tail, the same requirement of at least five per cell. The hypotheses are worded about two populations having the same distribution, the degrees of freedom are the number of columns minus one for the usual two-population case, and the conclusion is only that the distributions differ — never how.
11.5 Comparison of the Chi-Square TestsA summary section, and a necessary one. The chi-square statistic has now been used in three different circumstances, and since all three share the statistic, the right tail and the at-least-five condition, nothing in the arithmetic distinguishes them. The book's list separates them by the study instead, and two counts settle nearly every case: how many populations were sampled, and how many questions were asked. A goodness-of-fit test has one population, one question, and a known distribution to test against; a test of independence has one population and two questions, arranged in a contingency table; a test for homogeneity has two populations, one question, and no known distribution at all. The surest signal of all is how the hypotheses are worded, and this lesson works through the book's three pairs.
11.6 Test of a Single VarianceThe chapter's outlier. Every pattern the previous four sections established breaks here: the data are quantitative rather than categorical, the null names a parameter and is written as an equation, the degrees of freedom are the familiar n minus one, and the test may be right-tailed, left-tailed or two-tailed. That last point is the genuinely new one — sections 11.2 to 11.4 had no left-tailed version because their statistics measured disagreement, which cannot be negative, whereas this statistic compares a sample variance against a claimed one and a variance can be too small as well as too large. The test assumes the underlying distribution is normal, and that assumption does not soften with sample size the way the t procedures' did. Worked through the book's exam-score and post-office examples.
12.1 Linear EquationsThe chapter opens by fixing a convention. Linear regression for two variables is based on a linear equation with one independent variable, written y = a + bx — with the intercept first and the slope second, the reverse of the algebra form y = mx + b. Every formula in the rest of the chapter follows that convention, so reading a as the intercept and b as the slope from their positions rather than from their letters is what this section exists to establish. The graph of such an equation is a straight line, and any line that is not vertical can be described this way: b greater than zero slopes upward, b equal to zero is horizontal, and b less than zero slopes downward. Beyond the algebra, the section asks for something new — an interpretation of both numbers in the units of the problem, in complete sentences, which is exactly what sections 12.3 and 12.5 will demand of a fitted regression line.
12.2 Scatter PlotsThe section that comes before any arithmetic, and says so: before taking up linear regression and correlation, the relation between two variables has to be looked at. A scatter plot is the most common and easiest way to display it, and three things are read off one — the direction of the relationship, its strength, and the overall pattern together with any deviations from it. A clear direction means either that high values of one variable occur with high values of the other, or that high values of one occur with low values of the other. Strength is judged by how close the points fall to a line or some other curve, with one exception the book flags: points lying exactly on a horizontal line look like a perfect fit and in fact show no relationship, since y is the same whatever x does. The section closes with the condition for going further, which is that one variable must actually help explain or predict the other.
12.3 The Regression EquationThe chapter's central section. Data rarely fit a straight line exactly, so the question is which of the many lines that miss should be used — and the answer is the one minimising the sum of the squared vertical distances from the points to the line. Each such distance is a residual, positive where the line underestimates and negative where it overestimates, and the sum of their squares is the SSE. Two properties of the resulting line are worth more than its formulas: it always passes through the point of the two means, and its slope is the correlation coefficient times the ratio of the two standard deviations — which is why r and the slope always share a sign. The section closes with r itself, measuring the strength and direction of the linear association, and with r-squared, the share of the variation in y that the line explains. For the running example that share is 44 percent, which is a useful corrective to reading a correlation of 0.66 as a strong result.
12.4 Testing the Significance of the Correlation CoefficientSection 12.3 computed a correlation of 0.6631 from eleven students and left open whether that is enough to conclude anything. The correlation coefficient tells us about the strength and direction of the linear relationship, but the reliability of the linear model also depends on how many observed data points are in the sample — so r and n have to be looked at together. This section performs a hypothesis test of the significance of the correlation coefficient, with a null that the population correlation rho equals zero and a two-tailed alternative that it does not. The book gives two methods, a p-value from a t statistic on n minus two degrees of freedom and a table of critical values at a fixed 5 percent level, and calls them equivalent. They are in fact the same test written twice: every critical value in the table equals the two-tailed t critical value divided by the square root of df plus that value squared.
12.5 PredictionA short section that finally uses the line for what it was fitted for. Everything needed was established earlier — section 12.2 confirmed a linear pattern, section 12.3 fitted the least-squares line, section 12.4 established that its correlation is significant — so predicting is a substitution. The section's real content is the boundary around where that substitution may be trusted. Predicting for an x inside the range of observed x values is interpolation and is licensed; predicting outside that range is extrapolation and is not, however significant the correlation. The book demonstrates rather than asserts this: substituting a third-exam score of 90, well above the observed maximum of 75, gives a predicted final exam score of 261.19 when the largest score the final exam can carry is 200. The prediction is also a prediction of a mean, not of what any individual will score.
12.6 OutliersThe chapter's last section asks which points the line should have been fitted to. Two different kinds of point matter and the book keeps them apart: an outlier is far from the least-squares line vertically, meaning it has a large residual, while an influential point is far from the other observations horizontally and may have a big effect on the slope. The rough rule for flagging an outlier is any point further than two standard deviations from the line, where the standard deviation is that of the residuals, computed as the square root of the SSE divided by n minus two — dividing by n minus two because the regression model involves two estimates. For the running example that flags one point, the student who scored 65 on the third exam and 175 on the final. Deleting it changes the fitted line from a slope of 4.83 to 7.39 and the correlation from 0.6631 to 0.9121, which is why the book insists a deleted point be recorded and explained, or the results reported both ways.
13.1 One-Way ANOVAThe setup section for the chapter, containing one genuinely surprising sentence: the purpose of a one-way ANOVA test is to determine the existence of a statistically significant difference among several group means, and the test actually uses variances to help determine whether those means are equal. The reason is visible in a picture. If every group has the same population mean, pooling the groups changes nothing and the combined data has about the variance each group has; if the means differ, pooling spreads the data out and the combined variance is larger than any single group's. So a variance computed between groups, compared against one computed within groups, detects a difference in means. The null hypothesis is that all k group means are equal and the alternative is that at least two are not, which means a rejection identifies no particular pair. Five assumptions have to hold, and the third — equal population variances — is the one quietly carrying the argument.
13.2 The F Distribution and the F-RatioSection 13.1's box-plot picture turned into arithmetic. Two estimates of the population variance are computed: the variance between samples, which is the variance of the sample means multiplied by n and is called the explained variation, and the variance within samples, which is the average of the sample variances and is called the unexplained variation. The F statistic is their ratio. Its logic is stated plainly by the book: MSwithin is an estimate of the population variance, while MSbetween consists of the population variance plus a variance produced from the differences between the samples — so under the null both estimate the same value and F should be approximately one, while differing means add a term to the numerator and push the ratio above it. The distribution is named after Sir Ronald Fisher and has two sets of degrees of freedom, k minus one for the numerator and n minus k for the denominator.
13.3 Facts About the F DistributionThe chapter's counterpart to section 11.1: a short list of properties, followed by the section where the machinery is actually used. The five facts are that the curve is skewed right rather than symmetrical, that there is a different curve for each set of degrees of freedom, that the F statistic is always at least zero, that the curve approximates the normal as both degrees of freedom grow, and that the distribution has other uses including comparing two variances. Those first four parallel chi-square's facts almost word for word, which is no coincidence — an F statistic is a ratio of quantities built from sums of squares, so it inherits non-negativity and right skew from the same source. The section then works the chapter's three complete tests: tomato yields under five mulches, which reject at 5 percent; sorority grade means, which do not reject at 1 percent; and bean plant heights, which do not reject at 3 percent and are computed by the balanced shortcut.
13.4 Test of Two VariancesThe chapter's second use of the F distribution, and the course's last section. It is often desirable to compare two variances rather than two averages — college administrators would like two professors to have the same variation in their grading, a lid and a container must vary alike to fit, a supermarket may care about the variability of two checkers' times. The statistic is the ratio of the two sample variances, since the population variances cancel under a null of equality, and its degrees of freedom are one less than each sample size. Because a variance can be larger or smaller than another, the test may be left-tailed, right-tailed or two-tailed. The section also carries a warning unlike anything else in the book: this test is very sensitive to deviations from normality, can give p-values too high or too low in unpredictable ways, and many texts suggest students not use it at all — a caveat this lesson treats as the section's most important content.
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.