The display the rest of the book uses. What a histogram is and why its boxes are contiguous rather than separated; how to choose a starting point one decimal place finer than the data so that no value lands on a boundary, and how to compute the class width once the number of bars is chosen; the discrete case, where a width of one centres each value in its own bar; the frequency polygon, which plots each class midpoint at its frequency and joins them, anchored at empty classes so the shape closes, and which can be overlaid to compare two distributions; and the time series graph, which plots a measurement against its date and so preserves the chronological order every other display discards.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 2 — Descriptive Statistics
Histograms, Frequency Polygons, and Time Series Graphs
Objectives
Five outcomes. The second is the one that goes wrong most often and the one a calculator will not do for you.
OpenStax Introductory Statistics 2e, §2.2 Histograms, Frequency Polygons, and Time Series Graphs §2.2, pp. 75-85 — the section these objectives are drawn from
Warm-up
Section 1.3 grouped 100 soccer players' heights into eight classes: 5, 3, 15, 40, 17, 12, 7 and 1. Section 2.1 drew bar graphs with separated bars.
Discussion prompt
If you drew those eight frequencies as a bar graph with gaps between the bars, what would the gaps be claiming? Is the claim true here?
Hint: Ask what lies between the class 63.95 to 65.95 and the class 65.95 to 67.95.
Answer:
The gaps would claim that nothing lies between one class and the next. That was true for regions of a country and for streaming services, where there really is nothing in between.
It is false here. The classes are adjoining intervals on a number line: 65.95 is where one ends and the next begins, with no space at all. A player 65.95 inches tall would be exactly at the join, which is why the boundaries were chosen so that no recorded height can land there.
So this data needs a display whose boxes touch, and that display is the histogram. The single visual difference from a bar graph — closing the gaps — is a statement that the horizontal axis is a continuous number line rather than a list.
Concept
A histogram consists of contiguous, adjoining boxes. The horizontal axis is labelled with what the data represent and the vertical axis with frequency or relative frequency; the graph has the same shape either way. One advantage of a histogram is that it can readily display large data sets, and a rule of thumb is to use one when the data set has 100 values or more.
histogram — A display of quantitative data consisting of contiguous boxes over a continuous horizontal axis. The horizontal axis carries the variable, divided into classes; the vertical axis carries frequency, relative frequency, percent frequency or probability. Like the stemplot, it shows the shape, the centre and the spread.
\[ RF = \frac{f}{n}, \qquad \text{where } f = \text{frequency and } n = \text{total number of values} \]
Labelling the vertical axis with frequency or with relative frequency changes only the numbers on it, never the shape, because every frequency is divided by the same n. That is worth knowing because it means the choice is free: use frequency when the counts matter and relative frequency when the shape is to be compared with another data set of a different size. The book's own example is three students out of forty scoring 90 to 100, giving a relative frequency of 0.075, or 7.5 percent.
Figure (svg): Two columns contrasting a bar graph with separated bars over categories and a histogram with touching bars over a continuous axis
OpenStax Introductory Statistics 2e, §2.2 Histograms, Frequency Polygons, and Time Series Graphs §2.2, p. 75
Section
Section 1
Concept
A histogram consists of contiguous boxes over a horizontal axis labelled with what the data represent. Because the classes are adjoining intervals of a number line, the boxes touch — the absence of a gap is the statement that the axis runs continuously from one class into the next.
contiguous boxes — Adjoining boxes with no gaps between them. In a histogram the boxes touch because each class ends exactly where the next begins, so there is no region of the horizontal axis that belongs to no class.
\[ \text{bar graph: gaps} \qquad \text{histogram: no gaps} \]
The rule of thumb is to use a histogram when the data set has 100 values or more, which is exactly where the stemplot of section 2.1 becomes unwieldy. The two displays answer the same questions — shape, centre, spread — and differ in what they cost: the stemplot writes every value and the histogram writes none of them, keeping only how many fell in each class.
Figure (svg): A histogram of one hundred soccer players' heights in eight touching classes, with a tall bar of forty between 65.95 and 67.95 inches
OpenStax Introductory Statistics 2e, §2.2 Histograms, Frequency Polygons, and Time Series Graphs §2.2, p. 75 — the histogram defined, and the 100-value rule of thumb
Picture it
The book's Example 2.7, built from the grouped table of section 1.3.
Figure (svg): A histogram of one hundred soccer players' heights in eight touching classes, with a tall bar of forty between 65.95 and 67.95 inches
Forty of the hundred players fall in the single class from 65.95 to 67.95 inches, and the distribution falls away on both sides — more steeply below than above. That asymmetry is what section 2.6 will call skewness. Note also what has been lost: the histogram cannot tell you that eleven players were exactly 66.5 inches, which the stemplot could.
Worked example
The book's own small example, and why the axis label is a free choice.
\[ \text{3 students of 40 scored between } 90\% \text{ and } 100\% \]
Identify f and n
Why: The frequency of the class, and the sample size.
\[ f = 3, n = 40 \]
Divide
Why: The frequency over the total.
\[ \frac{3}{40} = 0.075 \]
Express as a percent
Why: Multiply by one hundred.
\[ 7.5 \% \]
Note what changes on the graph
Why: Every bar is divided by the same 40.
Figure (svg): The solution to Worked example frequency against relative frequency shown as a ladder of expressions, one row per legal move
\[ RF = \frac{f}{n} = \frac{3}{40} = 0.075 \]
Verify: confirm the shape really cannot change
Why: Dividing every bar height by the same number 40 rescales the vertical axis uniformly, which stretches or shrinks the whole picture without altering any bar's height relative to another. So the tallest bar is still tallest and every ratio between bars is preserved. That is why the book says the graph has the same shape with either label, and why relative frequency is the right choice when comparing two data sets of different sizes — the same argument section 1.2 made for percent columns.
OpenStax Introductory Statistics 2e, §2.2 Histograms, Frequency Polygons, and Time Series Graphs §2.2, p. 75
Discrimination
Ask whether the horizontal axis is a number line or a list.
Sort into buckets
Sort each variable.
Worked example
Every question a histogram answers is a question about classes, not about values.
\[ \text{How many players are shorter than } 65.95 \text{ inches?} \]
Identify the classes below the cut
Why: The first three: 59.95-61.95, 61.95-63.95, 63.95-65.95.
Read their frequencies
Why: Five, three and fifteen.
\[ 5, 3, 15 \]
Add them
Why: The total of the qualifying classes.
\[ 5 + 3 + 15 = 23 \]
Express as a share
Why: Twenty-three of one hundred.
\[ 0.23 \]
Figure (svg): The solution to Worked example reading the heights histogram shown as a ladder of expressions, one row per legal move
\[ 5 + 3 + 15 = 23 \quad \text{so} \quad RF = 0.23 \]
Verify: confirm against the cumulative column of section 1.3
Why: The grouped table in section 1.3 gives a cumulative relative frequency of 0.23 at the class ending 65.95, which is the same answer arrived at without any addition. A histogram and the cumulative column of its frequency table are two views of the same object, and either can check the other. Note the question had to be asked at a class BOUNDARY: the histogram cannot say how many players are under 65 inches, because that cut falls inside a class and the individual values are gone.
OpenStax Introductory Statistics 2e, §2.2 Histograms, Frequency Polygons, and Time Series Graphs §2.2, pp. 75-76
Trap
\[ \text{the class } 65.95\text{-}67.95 \text{ has frequency } 40 \]
Conclude that forty players are 67 inches tall
Why: The student picks a representative value from the class and attaches the whole frequency to it.
\[ 40 \text{ players at } 67 \text{ inches} \quad \text{(wrong)} \]
The forty players are spread across a two-inch interval, and the histogram records nothing about how they are spread within it.
\[ 40 \text{ players between } 65.95 \text{ and } 67.95 \text{ inches} \]
Report a class frequency as a statement about the interval
Why: The bar's height belongs to the whole width beneath it.
This is the cost the histogram pays for handling large data sets: it keeps counts per class and destroys the individual values. The stemplot of section 2.1 would have shown all forty heights individually. When a question needs a value rather than a class — the exact median, say — the histogram cannot answer it and the raw data must be consulted.
Fill the middle
Three students of forty in the top class.
Fill in the blanks
RF = \frac0.075___ = \frac______ = ___
Why: Three divided by forty is 0.075, or 7.5 percent. Relabelling the vertical axis from frequency to relative frequency divides every bar by the same n, so the shape of the histogram is completely unchanged — only the numbers on the axis differ.
Two truths and a lie
All three concern what a histogram shows.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and it is the exact difference from the stemplot. A histogram records only how many values fell in each class, so the forty players in one class could be anywhere within their two-inch interval. That loss is what allows the histogram to handle a hundred or a hundred thousand values equally well.
Estimation
The book gives a rule of thumb for choosing between the displays.
Predict first
At roughly what data set size does a histogram become the right choice over a stemplot?
Correct: About 100 values or more.
Why: The book's rule of thumb is to use a histogram when the data set consists of 100 values or more, which is just past the point where a stemplot's rows become too long to compare by eye. Below that the stemplot is often better precisely because it keeps every value. The last option overstates: for twenty values a stemplot shows everything a histogram would and more.
Section
Section 2
Concept
To construct a histogram, first decide how many bars or classes represent the data; many histograms consist of five to fifteen for clarity. Choose a starting point for the first interval that is less than the smallest data value, carried out to one more decimal place than the value with the most decimal places. Then the width is the ending value minus the starting point, divided by the number of bars.
starting point — A value less than the smallest data value, carried to one more decimal place than the data. Because it and the boundaries derived from it have more precision than any recorded value, no data value can fall on a boundary.
\[ \text{width} = \frac{\text{ending value} - \text{starting point}}{\text{number of bars}} \]
The book's examples of the rule: if the most precise value has one decimal and the smallest is 60, use 60 minus 0.05, giving 59.95. If the most precise has two decimals and the smallest is 1.5, use 1.5 minus 0.005, giving 1.495. If the values are all integers and the smallest is 2, use 2 minus 0.5, giving 1.5. There is also a guideline for the number of bars: take the square root of the number of data values and round, so 150 values suggests about twelve classes.
Figure (svg): A number line showing the smallest data value of 60 and the first class boundary placed at 59.95, one decimal place finer
OpenStax Introductory Statistics 2e, §2.2 Histograms, Frequency Polygons, and Time Series Graphs §2.2, pp. 75-76 — the starting point rule and the width formula
Picture it
The data carry one decimal place; the boundary carries two.
Figure (svg): A number line showing the smallest data value of 60 and the first class boundary placed at 59.95, one decimal place finer
The extra decimal place is doing one job: making it impossible for a recorded value to equal a boundary. Had the classes started at 60 with width 2, a player recorded at exactly 62.0 would satisfy both the class ending at 62 and the class beginning at 62, and the two class frequencies would depend on an unstated convention. Section 1.3 solved the same problem the same way for its grouped table.
Worked example
Example 2.7 in full. The smallest height is 60 and the largest is 74.
\[ \text{100 heights, smallest } 60, \text{ largest } 74, \text{ most precise value has one decimal} \]
Choose the starting point
Why: One decimal finer than the data, below the smallest value.
\[ 60 - 0.05 = 59.95 \]
Choose the ending value
Why: The same distance above the largest.
\[ 74 + 0.05 = 74.05 \]
Choose the number of bars and divide
Why: Eight bars across a span of 14.1.
\[ \frac{14.1}{8} = 1.7625 \]
Round UP to a convenient width
Why: Two units wide, which keeps values off the boundaries.
\[ \text{width } 2 \]
Figure (svg): The solution to Worked example building the eight classes shown as a ladder of expressions, one row per legal move
\[ \frac{74.05 - 59.95}{8} = 1.7625 \;\longrightarrow\; \text{width } 2 \]
Verify: confirm rounding up rather than to nearest was correct
Why: Standard rounding would give 1.8, and the book notes that rounding up is often necessary even when it goes against the usual rule. The reason is coverage: a width of 1.76 times eight is exactly 14.1 and just reaches the ending value, but any rounding DOWN would leave the largest values outside the last class. Rounding up to 2 gives eight classes spanning 16 units, comfortably covering the data, with the last class running to 75.95 and containing the single height of 74. The book notes 1.76 would also work; what would not work is 1.7.
OpenStax Introductory Statistics 2e, §2.2 Histograms, Frequency Polygons, and Time Series Graphs §2.2, p. 76
Matching
The smallest value, and how much to subtract from it.
Match the pairs
Why: Each subtracts half a unit of the finest precision present in the data: 0.5 for integers, 0.05 for one decimal, 0.005 for two, 0.0005 for three. The result always has one more decimal place than any recorded value, which is exactly what makes a collision with a boundary impossible.
Worked example
The rule in three more cases, all from the book.
\[ \text{smallest } 6.1 \text{ (one decimal)}; \quad \text{smallest } 1.5, \text{ most precise } 2.23; \quad \text{integers, smallest } 2 \]
One decimal place in the data
Why: Subtract 0.05 from the smallest value.
\[ 6.1 - 0.05 = 6.05 \]
Two decimal places in the data
Why: Subtract 0.005 from the smallest value.
\[ 1.5 - 0.005 = 1.495 \]
Three decimal places in the data
Why: Subtract 0.0005.
\[ 1.0 - 0.0005 = 0.9995 \]
All integers
Why: Subtract 0.5.
\[ 2 - 0.5 = 1.5 \]
Figure (svg): The solution to Worked example the starting point for other data shown as a ladder of expressions, one row per legal move
\[ \text{subtract } 0.5 \text{, } 0.05 \text{, } 0.005 \text{, } 0.0005 \text{ as the data get finer} \]
Verify: confirm the rule keys off the MOST precise value, not the smallest
Why: In the second case the smallest value is 1.5, which has one decimal place, but the most precise value in the data set is 2.23 with two. The starting point subtracts 0.005 rather than 0.05, because a boundary at 1.45 could be equalled by a value like 1.45 elsewhere in the data. It is the precision of the finest value anywhere in the set that determines how fine the boundaries must be, and reading the rule as being about the smallest value is the commonest way to misapply it.
OpenStax Introductory Statistics 2e, §2.2 Histograms, Frequency Polygons, and Time Series Graphs §2.2, p. 75
Error analysis
The 100 heights run from 60 to 74, recorded to the nearest half inch. Four students propose a scheme. Only one is right.
Annotate
On: \( \begin{aligned} &(1)\; \text{start } 60, \text{ width } 2 \\ &(2)\; \text{start } 59.95, \text{ width } 1.7 \\ &(3)\; \text{start } 59.95, \text{ width } 2, \text{ but only } 6 \text{ classes} \\ &(4)\; \text{start } 59.95, \text{ width } 2, \; 8 \text{ classes} \end{aligned} \)
Errors (2) and (3) are the same failure — classes that do not reach the largest value — and both are caught by one check: multiply the width by the number of classes, add the starting point, and confirm the result exceeds the largest data value. Error (1) is caught by asking whether any boundary is a number the data could actually take.
Faded example
Starting point 59.95, ending value 74.05, eight bars.
Fill in the blanks
\text14.1 = \frac1.7625___ = \frac___}___ = ___
Why: The computed width is 1.7625, which the book rounds UP to 2 for convenience and to keep values off boundaries. Rounding up rather than to nearest is deliberate: a width rounded down would leave the largest values outside the final class.
Estimation
A data set has 150 values, and the book gives a guideline some people follow.
Predict first
About how many classes does that guideline suggest?
Correct: About 12.
Why: The guideline is to take the square root of the number of data values and round to a whole number, and the square root of 150 is about 12.2. It sits comfortably inside the five-to-fifteen range the book recommends for clarity. Like all such rules it is a starting suggestion rather than a requirement — the right number is whichever makes the shape clearest, and it is worth trying two.
Edge cases
Suppose the classes had started at 60 with width 2, and a player is recorded at exactly 64 inches.
Discussion prompt
Which class does that player belong to, and what does the ambiguity cost?
Hint: Ask whether 64 satisfies '62 to 64' or '64 to 66', or both.
Answer:
It satisfies both, and nothing in the scheme decides between them. Whoever tallies the data must adopt a convention — usually that a class includes its lower boundary and excludes its upper — and that convention is invisible to anyone reading the finished histogram.
The cost is that two people can tally the same data into the same classes and get different frequencies, and a reader cannot tell which convention produced the picture. With seven players at exactly 64 inches, two class heights change by seven.
The starting-point rule removes the question rather than answering it. A boundary of 63.95 cannot be equalled by any height recorded to the nearest half inch, so every value belongs to exactly one class and any two people tallying will agree.
Section
Section 3
Concept
Discrete data can be displayed in a histogram too. When the data are integers and there are not too many different values, the most convenient width is one that places the data values in the middle of the bars, so each possible value gets a bar of its own.
centring a discrete value — Choosing class boundaries at the half-integers so that each whole-number value sits at the centre of its own class. With a starting point of 0.5 and a width of 1, the value 1 is in the middle of the interval from 0.5 to 1.5, the value 2 in the middle of 1.5 to 2.5, and so on.
\[ \text{bars} = \frac{6.5 - 0.5}{1} = 6 \]
The starting-point rule already gives the right answer here: the data are integers, so subtract 0.5 from the smallest value of 1 to get 0.5, and add 0.5 to the largest value of 6 to get 6.5. A width of 1 then produces six bars, each centred on one of the six possible values. The book's data are the numbers of books bought by fifty part-time students, with frequencies 11, 10, 16, 6, 5 and 2.
Figure (svg): A histogram of the number of books bought by fifty students, six bars of width one, each centred on a whole number
OpenStax Introductory Statistics 2e, §2.2 Histograms, Frequency Polygons, and Time Series Graphs §2.2, pp. 77-78 — the discrete case, with the books example
Picture it
Example 2.8, with each whole number at the centre of its bar.
Figure (svg): A histogram of the number of books bought by fifty students, six bars of width one, each centred on a whole number
The bars still touch, and they should: the horizontal axis is still a number line, and the interval from 2.5 to 3.5 really does adjoin the interval from 1.5 to 2.5. What the boxes now assert is only that everything in the interval from 2.5 to 3.5 was recorded as 3, which for a count of books is exactly true. Compare section 2.1's bar graph of regions, where nothing at all lies between one bar and the next.
Worked example
Example 2.8. The starting-point rule does all the work.
\[ \text{books bought: } 1\text{ through }6, \text{ frequencies } 11, 10, 16, 6, 5, 2 \]
Apply the integer starting-point rule
Why: Subtract 0.5 from the smallest value of 1.
\[ \text{starting point } 0.5 \]
Find the ending value
Why: Add 0.5 to the largest value of 6.
\[ \text{ending value } 6.5 \]
Choose a width that centres the values
Why: A width of 1 puts 1 in the interval 0.5 to 1.5.
\[ \text{width } 1 \]
Compute the number of bars
Why: The span divided by the width.
\[ \frac{6.5 - 0.5}{1} = 6 \]
Figure (svg): The solution to Worked example the books histogram shown as a ladder of expressions, one row per legal move
\[ \text{bars} = \frac{6.5 - 0.5}{1} = 6 \]
Verify: confirm the frequencies account for all fifty students
Why: Adding 11, 10, 16, 6, 5 and 2 gives 50, which is the stated number of part-time students. The frequency column of any histogram totals the sample size for the same reason section 1.3's did: every student contributed exactly one observation to exactly one class. Running that check before drawing catches a mis-tallied class while it is still cheap to fix.
OpenStax Introductory Statistics 2e, §2.2 Histograms, Frequency Polygons, and Time Series Graphs §2.2, pp. 77-78
Sorting
Count the distinct values the variable can take in this data set.
Sort into buckets
Sort each discrete data set.
Every item here is discrete, which is the point: being discrete does not decide the question. Salaries in whole pounds are as discrete as children per household and need grouping just as badly as a continuous measurement would.
Worked example
The centring trick only works while the number of distinct values is small.
\[ \text{the number of words in each of } 500 \text{ essays, ranging from } 380 \text{ to } 1240 \]
Note the data are discrete
Why: Words are counted, so every value is a whole number.
Count the distinct values
Why: Anything from 380 to 1240 could occur.
\[ \text{up to } 861\text{ possible values} \]
Reject the centring approach
Why: 861 bars of width one is not a display.
Group as though continuous
Why: Choose ten or so classes and apply the usual rule.
\[ \text{width about } 90 \]
Figure (svg): The solution to Worked example when discrete data need grouping anyway shown as a ladder of expressions, one row per legal move
\[ \text{width} \approx \frac{1240.5 - 379.5}{10} \approx 87 \]
Verify: confirm the deciding question is the number of distinct values
Why: It is not whether the data are discrete or continuous. The books example is discrete and takes six values, so each gets a bar; the word counts are equally discrete and take hundreds, so they must be grouped. Section 1.3 made the same point about frequency tables, where twenty answers to a work-hours question needed no grouping and a hundred measured heights did. The test in both places is how many distinct values there are, and 'discrete' is only a strong hint that the answer may be small.
OpenStax Introductory Statistics 2e, §2.2 Histograms, Frequency Polygons, and Time Series Graphs §2.2, pp. 77-78
Trap
\[ \text{starting point } 1, \text{ width } 1 \]
Use boundaries 1, 2, 3, 4, 5, 6, 7
Why: The values are whole numbers, so the student uses them as boundaries.
\[ \text{every recorded value sits on a boundary} \quad \text{(wrong)} \]
Every one of the fifty observations is an integer, so every single one is ambiguous between two classes.
\[ \text{starting point } 0.5, \text{ width } 1, \text{ boundaries } 0.5, 1.5, 2.5, \ldots \]
Put the boundaries at the half-integers
Why: Then each whole number sits at the centre of a class rather than at its edge.
The general rule already covers this: integers are data with zero decimal places, so subtract 0.5 to get a starting point one place finer. The half-integer boundaries are not a special trick for discrete data but the same precision rule applied to the coarsest kind of data there is, and the pleasant side effect is that each value ends up centred under its own bar.
Fill the middle
Starting point 0.5, ending value 6.5, width 1.
Fill in the blanks
\text6 = \frac______ = ___
Why: Six bars, one for each of the values 1 through 6. When the centring approach is used the number of bars always equals the number of distinct values, which is a useful check that the width and boundaries were chosen consistently.
Prediction
Commit before reasoning.
Predict first
The books data take six whole-number values. Why draw a histogram with touching bars rather than a bar graph with gaps?
Correct: Because the number of books is quantitative.
Why: The axis is a number line: 3 is genuinely between 2 and 4, the bars could not be reordered, and the distance from 1 to 2 is the same as from 5 to 6. A bar graph's gaps would deny all of that. Neither the number of categories nor the size of the frequencies has any bearing on the choice — it is decided entirely by whether the horizontal variable is quantitative.
Two truths and a lie
All three concern discrete histograms.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The bars still touch, because the horizontal axis is still a continuous number line even though only certain points on it can occur. The gap in a bar graph asserts that the axis is categorical, and the number of books is not a category — it is a quantity with order and meaningful distances.
Section
Section 4
Concept
Frequency polygons are analogous to line graphs, and just as line graphs make continuous data visually easy to interpret, so do frequency polygons. To construct one, decide on the class intervals, plot each class at its midpoint against its frequency, and draw line segments connecting the points.
frequency polygon — A display in which each class is plotted as a single point at the midpoint of the class, at a height equal to the class frequency, and the points are joined by line segments. Empty classes are added at each end so the polygon comes down to the horizontal axis.
\[ \text{midpoint} = \frac{\text{lower bound} + \text{upper bound}}{2} \]
The anchors are the detail worth attending to. In the book's example the first label on the horizontal axis is 44.5, representing the interval from 39.5 to 49.5, and it is used only to allow the graph to touch the axis — the lowest test score is above it and the interval contains no data. The same is done at the top with 104.5. Without those two empty classes the polygon would begin and end in mid-air, which would suggest the distribution continues off the edge of the picture.
Figure (svg): A frequency polygon over a faint histogram of calculus test scores, with points at class midpoints joined by lines and anchored at zero on both ends
OpenStax Introductory Statistics 2e, §2.2 Histograms, Frequency Polygons, and Time Series Graphs §2.2, pp. 80-81 — constructing a frequency polygon, with the anchors explained
Picture it
The frequency polygon drawn over the histogram it comes from.
Figure (svg): A frequency polygon over a faint histogram of calculus test scores, with points at class midpoints joined by lines and anchored at zero on both ends
Each yellow point sits at the middle of its class at the height of that class's frequency, so the polygon is the histogram's tops connected. The book reads one further thing off it: the distribution is skewed, because one side of the graph does not mirror the other — it rises gently to a peak at 84.5 and falls away sharply. Section 2.6 gives that observation a name and connects it to the mean and the median.
Worked example
Example 2.10. Five real classes and two anchors.
\[ \text{classes } 49.5\text{-}59.5, \; 59.5\text{-}69.5, \; 69.5\text{-}79.5, \; 79.5\text{-}89.5, \; 89.5\text{-}99.5 \text{ with frequencies } 5, 10, 30, 40, 15 \]
Find each class midpoint
Why: The average of the two bounds.
\[ 54.5, 64.5, 74.5, 84.5, 94.5 \]
Plot each midpoint at its frequency
Why: One point per class.
Add an empty class at each end
Why: One class below the first and one above the last.
\[ 44.5\text{ and } 104.5\text{ at zero} \]
Join the points in order
Why: Line segments left to right.
Figure (svg): The solution to Worked example plotting the points shown as a ladder of expressions, one row per legal move
\[ \text{midpoint of } 49.5\text{-}59.5 = \frac{49.5 + 59.5}{2} = 54.5 \]
Verify: confirm the frequencies total the sample size
Why: Adding 5, 10, 30, 40 and 15 gives 100, matching the cumulative frequency column's final entry of 100 in the book's table. The two anchor classes contribute nothing to that total, which is exactly right — they are drawing devices holding no data, and including them in the count would be an error. Their frequencies are zero precisely because no score falls in them.
OpenStax Introductory Statistics 2e, §2.2 Histograms, Frequency Polygons, and Time Series Graphs §2.2, pp. 80-81
Faded example
The class running from 79.5 to 89.5.
Fill in the blanks
\text169 = \frac84.5___ = \frac___}___ = ___
Why: The midpoint 84.5 is where this class's point is plotted, at a height of 40. Notice the midpoints are spaced exactly one class width apart — 54.5, 64.5, 74.5, 84.5, 94.5 — which is a quick check that they were computed correctly.
Worked example
Example 2.11. The comparison a histogram cannot make.
\[ \text{test scores } 5, 10, 30, 40, 15 \quad \text{against final grades } 10, 10, 30, 45, 5 \]
Check both use the same classes
Why: Both tables run 49.5 to 99.5 in steps of ten.
Check both have the same total
Why: Each column sums to 100 students.
Plot both polygons on one pair of axes
Why: Two lines, distinguished by colour.
Read the difference
Why: Grades sit lower at the top and higher at the bottom.
Figure (svg): The solution to Worked example overlaying two polygons shown as a ladder of expressions, one row per legal move
\[ \text{top class: } 15 \text{ scores against } 5 \text{ grades} \]
Verify: confirm why two histograms could not do this
Why: Two histograms drawn on the same axes would have overlapping opaque boxes, so whichever was drawn second would hide the other wherever it was shorter. A polygon is a single line and takes up almost no area, so several can share one pair of axes and all remain readable. That is the whole reason the book introduces the polygon immediately after the histogram — it is the same information in a form that superimposes. Comparing the two totals first matters too: both are 100 here, so raw frequencies are comparable; had they differed, relative frequencies would have been needed.
OpenStax Introductory Statistics 2e, §2.2 Histograms, Frequency Polygons, and Time Series Graphs §2.2, pp. 81-82
Trap
\[ \text{plot the class } 49.5\text{-}59.5 \text{ at } x = 49.5 \]
Use the lower bound of each class as the horizontal coordinate
Why: The bound is the number written in the table, so the student reaches for it.
\[ \text{every point shifted half a class to the left} \quad \text{(wrong)} \]
The whole polygon slides five units left, so the peak is reported in the wrong class and comparisons with another polygon are shifted too.
\[ \text{plot at the MIDPOINT: } \frac{49.5 + 59.5}{2} = 54.5 \]
Compute each class's midpoint before plotting anything
Why: The point represents the whole class, so it belongs at its centre.
The midpoint is the right choice because the class frequency describes the entire interval, and the centre is the only point that does not favour one end. It is also what makes the polygon sit on top of the histogram's bars: a bar spans its class and its top is level, so the point that best represents that top is the one above its middle. Plotting at a boundary would place the point above the join between two bars, where the height is ambiguous.
Socratic
The anchors at 44.5 and 104.5 contain no data at all.
Discussion prompt
What would the polygon look like without them, and why does the book add them?
Hint: Ask where the line would start and stop.
Answer:
Without them the line would begin at the point (54.5, 5) and end at (94.5, 15), both well above the horizontal axis. The polygon would appear to be cut off, as though the distribution continued beyond the edges of the picture.
The book says the interval is used only to allow the graph to touch the axis. Bringing the line down to zero at both ends states that there are no data below 49.5 and none above 99.5 — the distribution is complete, not truncated.
The anchors carry no data and must not be counted in the total, which stays at 100. They are a drawing convention, and the honest reading of them is 'zero students here', which is true.
Two truths and a lie
All three concern frequency polygons.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The anchors are empty classes plotted at height zero purely so the polygon meets the axis. They contain no observations, contribute nothing to the total, and the book's example still totals 100 students across the five real classes.
Prediction
Commit before reasoning.
Predict first
Two frequency polygons are overlaid. What must be true for the comparison to be fair?
Correct: Same classes, and comparable totals.
Why: Different classes would put the two lines on different horizontal scales, so the shapes could not be compared at all. And if one data set has 100 values and the other 400, the larger will sit above the smaller everywhere simply for being larger — which is section 1.2's point about comparing groups with different totals, and the fix is the same: switch both to relative frequency.
Section
Section 5
Concept
A histogram of temperatures over a month ignores part of the data collected: each date is paired with its reading, and that pairing imposes a chronological order. A graph that recognises this ordering and displays the changing value as time passes is called a time series graph.
time series graph — A graph of paired data in which the horizontal axis carries the date or time increment and the vertical axis carries the measured value, so each point corresponds to a date and a quantity. The points are typically connected by straight lines in the order in which they occur.
\[ \text{each point is a pair } (\text{date}, \text{measurement}) \]
The book's argument for it is worth quoting in substance: you could compute the mean temperature for the month, or build a histogram of how many days fell in each range, and all of those methods ignore a portion of the data you collected. A histogram of a month's temperatures is the same picture whether the month warmed steadily or alternated hot and cold days, because it counts days into classes and forgets which day was which.
Figure (svg): A time series graph of the Consumer Price Index over seven months, rising to a peak in March and then drifting slightly downward
OpenStax Introductory Statistics 2e, §2.2 Histograms, Frequency Polygons, and Time Series Graphs §2.2, pp. 83-85 — the time series graph and the Consumer Price Index example
Picture it
Example 2.12, the Annual Consumer Price Index for the first seven months of year 1.
Figure (svg): A time series graph of the Consumer Price Index over seven months, rising to a peak in March and then drifting slightly downward
The rise to a March peak of 184.2 and the slight drift down afterwards is a feature of the ORDER, and no histogram of these seven numbers could show it. Note the vertical axis begins at 181 rather than zero. For an index that never approaches zero this is a defensible choice, but it magnifies a total range of 2.5 points into the full height of the picture, and a reader should register that before reading drama into the shape.
Worked example
Example 2.12. Both halves of each pair go on the graph.
\[ \text{Jan } 181.7; \;\text{Feb } 183.1; \;\text{Mar } 184.2; \;\text{Apr } 183.8; \;\text{May } 183.5; \;\text{Jun } 183.7; \;\text{Jul } 183.9 \]
Put time on the horizontal axis
Why: The date or time increment.
Put the measurement on the vertical axis
Why: The index value.
Plot one point per pair
Why: Each point is a date and a quantity.
Join them in chronological order
Why: Straight lines in the order they occur.
Figure (svg): The solution to Worked example building the time series graph shown as a ladder of expressions, one row per legal move
\[ \text{peak} = 184.2 \text{ in March} \]
Verify: confirm the order may not be rearranged
Why: Sorting these seven values from smallest to largest would produce a smoothly rising line and destroy the entire content of the display: the graph would then say the index rose steadily, which it did not. Every other display in this chapter is indifferent to the order of the data — a histogram of these seven values is identical however they are shuffled — and this one is defined by it. That is exactly what the book means when it says the other methods ignore a portion of the data collected.
OpenStax Introductory Statistics 2e, §2.2 Histograms, Frequency Polygons, and Time Series Graphs §2.2, pp. 83-84
Discrimination
Ask whether the question being asked involves when the observations occurred.
Sort into buckets
Sort each question by the display that answers it.
Worked example
Two months with identical histograms and opposite stories.
\[ \text{Month A: temperatures rise steadily. Month B: hot and cold days alternate.} \]
Suppose both months contain the same thirty readings
Why: Only the order in which they occurred differs.
Build a histogram of each
Why: Counting days into temperature classes.
Build a time series of each
Why: Plotting each day's reading against its date.
Say what distinguishes them
Why: Only the chronological pairing.
Figure (svg): The solution to Worked example what a histogram of the same data loses shown as a ladder of expressions, one row per legal move
\[ \text{same values, different order} \;\Longrightarrow\; \text{same histogram, different time series} \]
Verify: confirm which display is right for which question
Why: If the question is how often the temperature exceeded 30 degrees, the histogram answers it and the time series is unnecessary. If the question is whether the month warmed, only the time series can answer, and the histogram is not merely unhelpful but actively misleading if read as though it could. Neither display is better; they answer different questions, and choosing between them means deciding whether the order of the observations is part of what you are asking about.
OpenStax Introductory Statistics 2e, §2.2 Histograms, Frequency Polygons, and Time Series Graphs §2.2, p. 83
Trap
\[ \text{sort the seven index values: } 181.7, 183.1, 183.5, 183.7, 183.8, 183.9, 184.2 \]
Order the data before plotting, as one would for a stemplot
Why: Every other display in this chapter begins by sorting.
\[ \text{a smoothly rising line} \quad \text{(a graph of nothing)} \]
The sorted graph rises steadily whatever the data did, because sorting guarantees it. The picture is now a property of the sorting, not of the index.
\[ \text{plot in chronological order: } 181.7, 183.1, 184.2, 183.8, \ldots \]
Never reorder a time series; the horizontal axis is fixed by the dates
Why: The pairing of value with date is the data.
This is the clearest case of a rule stated in section 2.1: any feature of a display that can be changed by an arbitrary rearrangement is a feature of the rearrangement. For a bar graph, reordering is harmless because the bars carry their labels with them. For a time series it is fatal, because the horizontal position IS the date and moving a point changes what it says.
Fill the middle
The construction of a time series graph.
Fill in the blanks
\textdate ___ \qquad \text___
Why: Time always goes on the horizontal axis, so that reading the graph left to right is reading forward in time. That convention is universal and worth relying on: any graph with dates along the bottom is asserting a chronological story, and one without them is not.
Elimination
A shop records its daily takings for a year and wants to know whether takings grew.
Eliminate the wrong options
Eliminate the three that cannot answer it and keep the one that can.
Survives elimination: B
Why: Only the time series preserves the pairing of each amount with its date, and growth is a statement about that pairing. The three rejected options are all legitimate displays that answer other questions well — how variable takings were, what a typical day brought — which is why the choice must be driven by the question rather than by which display is most familiar.
Prediction
Commit before reasoning.
Predict first
A time series graph of an index shows a dramatic-looking climb. What should you check first?
Correct: Where the vertical axis starts.
Why: An axis beginning near the data rather than at zero stretches a small range across the full height of the picture, turning a two-point move into a mountain. The seven index values here span 2.5 points on an axis 4 points tall, which looks dramatic and is not. Even spacing of dates is worth checking too, but the truncated axis is by far the commonest way a time series exaggerates — the misleading display section 1.2 warned about.
Comparison
Fill the blanks. The last column is what each one throws away.
Comparison matrix
| Display | Best for | What it discards |
|---|---|---|
| Stemplot | under about 100 values | nothing: every value is kept |
| Histogram | 100 values or more | the individual values within each class |
| Frequency polygon | comparing two or more distributions by overlay | the same as a histogram, plus the class widths |
| Time series graph | data paired with dates, when change over time is the question | nothing about order, but it shows no distribution |
The first three all treat the data as an unordered collection and differ only in how much detail they keep. The fourth is a different kind of object altogether: it keeps the order and shows no distribution at all, which is why a data set collected over time usually deserves both a histogram and a time series.
Pattern
Six steps. Steps two and three are where the errors are.
For discrete data with few distinct values, steps two and four take care of themselves: integers give a starting point at a half-integer, and a width of one then centres each value under its own bar. The check at step five is the one that catches a width rounded the wrong way.
OpenStax Introductory Business Statistics 2e, §2.1 Display Data §2.1 Display Data
Check
The starting point rule.
Check your understanding
A data set is recorded to two decimal places, and its smallest value is 3.40. What starting point does the book's rule give?
Answer: A
Why: Two decimal places in the data means subtracting 0.005, giving 3.395 — a number with three decimal places, so no recorded value can equal it.
Check
Compute the width.
Check your understanding
A starting point of 0.5 and an ending value of 6.5 are to be divided into 6 classes. What is the class width?
Answer: A
Why: The span is 6.5 minus 0.5, which is 6, and dividing by 6 classes gives a width of 1 — which then centres each whole number in its own bar.
Check
Where the polygon's points go.
Check your understanding
A frequency polygon is built from classes 69.5 to 79.5 and so on. At what horizontal position is the class 69.5 to 79.5 plotted?
Answer: A
Why: Each class is plotted at its midpoint, and the midpoint of 69.5 and 79.5 is 74.5. The frequency determines the height of the point, never its horizontal position.
Real world
A company reports customer wait times for a year in two ways. The annual report shows a histogram: most calls answered within four minutes, a thin tail out to twenty. The operations team keeps a time series of the daily mean wait, which shows a flat line near three minutes for ten months and then a steady climb to nine minutes over the final two.
Discussion prompt
Are the two displays consistent? Which one should drive a decision, and what would you have missed with only the other?
Hint: Ask what the histogram does with the dates.
Answer:
They are entirely consistent. The histogram counts every call into a wait-time class and forgets when it happened, so two months of long waits appear as the thin upper tail. Nothing in it is wrong; it simply cannot distinguish a year of occasional long waits from a good year followed by a deteriorating one.
The time series should drive the decision, because the question a manager is asking is whether service is getting worse, and that is a question about order. It shows a clear and recent deterioration that the histogram averages away across twelve months.
With only the time series you would miss the distribution. It plots daily means, so a day with half the calls answered instantly and half after half an hour looks identical to a day where every call took three minutes. The histogram is what shows the spread within days and the length of the tail.
\[ \text{histogram: what the values were} \qquad \text{time series: when they were} \]
The general lesson is the one the book makes when introducing the time series: a display that ignores the dates is ignoring part of the data you collected. Data gathered over time almost always deserve both displays, and reporting only the histogram is the standard way a recent deterioration stays invisible for a year.
Commit first
Answer, then rate your confidence honestly.
Predict first
Why are a histogram's boundaries carried to one more decimal place than the data?
Correct: So that no data value can fall on a boundary.
\[ \text{data to 1 decimal} \;\Longrightarrow\; \text{boundaries to 2} \;\Longrightarrow\; \text{no collision possible} \]
Why: A value equal to a boundary satisfies the class below and the class above, and nothing in the scheme decides between them, so two people tallying the same data would get different frequencies. Carrying the boundaries one decimal finer than any recorded value makes the collision arithmetically impossible. The other options confuse this with unrelated concerns: the width rarely comes out evenly and is rounded up, and equal widths are a separate choice entirely.
Explain it
They have drawn a histogram of a month of daily temperatures and concluded the month did not warm up.
Discussion prompt
In three sentences or fewer, explain what their histogram cannot tell them and what to draw instead.
Hint: Ask them to shuffle the days and redraw.
Answer:
Tell them to shuffle the thirty readings into a random order and draw the histogram again: it will be exactly the same picture, because a histogram only counts how many days fell in each temperature range.
Any display that is unchanged by shuffling cannot possibly answer a question about the order, and 'did the month warm up' is entirely a question about order.
What they need is a time series graph, with the date along the bottom and the temperature up the side, which is the only display in this chapter that keeps the pairing of each reading with its day.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the starting point, subtract half a unit of the FINEST precision anywhere in the data, then check the last boundary exceeds the largest value. For discrete data, put the boundaries at the half-integers and each value lands centred in its own bar. For the polygon, compute every class midpoint first and add one empty class at each end. For the choice of display, ask whether shuffling the observations would change the answer to your question. Do five examples of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top left, work the class scheme for the book's 100 heights: smallest 60, largest 74, data to one decimal place, eight bars. Write the starting point, the ending value, the computed width, the rounded width, and the full list of boundaries, then check that the last boundary exceeds 74. Beside it, sketch the histogram with frequencies 5, 3, 15, 40, 17, 12, 7, 1 and confirm they total 100. Below, draw the books histogram for values 1 to 6 with frequencies 11, 10, 16, 6, 5, 2, using boundaries at the half-integers, and write one sentence saying why the bars still touch even though the data are counts. In the middle right, take the classes 49.5-59.5, 59.5-69.5, 69.5-79.5, 79.5-89.5, 89.5-99.5 with frequencies 5, 10, 30, 40, 15: compute all five midpoints, plot the frequency polygon including both anchor classes, and mark which two points carry no data. At the bottom, draw a time series of these seven index values in order — 181.7, 183.1, 184.2, 183.8, 183.5, 183.7, 183.9 — and write two sentences: one saying what a histogram of the same seven numbers would fail to show, and one noting where your vertical axis starts and what that does to the apparent size of the change.
Check your boundary list two ways: no boundary should be a number the data could actually take, and the last boundary must be above the largest value. If either fails, the width was rounded the wrong way or the starting point used too coarse a precision.
Recap
Five things, and the second is the one a calculator will not do for you.
| If you see | Then |
|---|---|
| Quantitative data, 100 values or more | Histogram: contiguous boxes |
| Data recorded to one decimal place | Boundaries to two decimals: subtract 0.05 from the smallest |
| A computed width like 1.7625 | Round UP, then check the last boundary clears the largest value |
| Discrete data with few distinct values | Boundaries at the half-integers, width one |
| Discrete data with hundreds of values | Group as though continuous |
| Two distributions to compare | Overlay frequency polygons, matching classes and totals |
| Each value paired with a date | Time series: never reorder it |
Section 2.3 stops drawing and starts measuring. Percentiles and quartiles locate a value's position within the data — the same accumulation the cumulative relative frequency column performed in section 1.3, now made precise enough to answer questions like which score beats three quarters of the class.
OpenStax Introductory Statistics 2e, §2.2 Histograms, Frequency Polygons, and Time Series Graphs §2.2, pp. 75-85 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.