1.3 Frequency, Frequency Tables, and Levels of Measurement

The four levels of measurement as a ladder — nominal names with no order, ordinal adding order but no measurable differences, interval adding meaningful differences but no true zero, and ratio adding a true zero so that ratios become meaningful — and which operations each level permits. Then the frequency table and its two derived columns: relative frequency as each count divided by the sample size, and cumulative relative frequency as the running total that must finish at one. Closes with grouped frequency tables for continuous data, and the book's rounding convention of carrying one more decimal place than the original data.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 1.3 Frequency, Frequency Tables, and Levels of Measurement

Title

Statistics · Chapter 1 — Sampling and Data

Frequency, Frequency Tables, and Levels of Measurement

2. By the end of this lesson you can

Objectives

Five outcomes. The first decides what arithmetic is legal; the rest are the arithmetic.

OpenStax Introductory Statistics 2e, §1.3 Frequency, Frequency Tables, and Levels of Measurement §1.3, pp. 26-33 — the section these objectives are drawn from

3. What you already have

Warm-up

Section 1.2 split data into qualitative and quantitative, and warned that averaging a categorical variable is meaningless.

Discussion prompt

Rank these by how much they let you do: eye colour, a cruise rating of excellent or good or satisfactory, temperature in Celsius, and an exam score out of 100. What can you do with each that you cannot do with the one before?

Hint: Try ordering the values, then subtracting two of them, then dividing one by another.

Answer:

Eye colour can only be counted into groups: there is no first or last. A cruise rating adds order — excellent is above good — but the gap between excellent and good is not a measurable quantity.

Celsius adds real differences: 40 degrees is exactly 100 minus 60, and that subtraction means something. But it has no true zero, since zero Celsius is not the absence of temperature and negative values exist.

An exam score adds the last thing: a genuine zero meaning none. That is what makes 80 four times 20 a sensible statement, which is not true of temperatures — 40 degrees is not twice as hot as 20.

Those four steps are exactly the four levels of measurement, and this section names them. The level a variable sits at decides which statistical procedures are legitimate, which is why it comes before any calculation.

4. How data are measured decides what you may do with them

Concept

The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on knowing it, because not every operation can be used with every set of data. There are four levels, and each permits everything the one below it does and one thing more.

level of measurement — The way a set of data is measured, classified from lowest to highest as nominal, ordinal, interval and ratio. The level determines which statistical operations are valid: nominal and ordinal data cannot be used in calculations, while interval and ratio data can.

\[ \text{nominal} \;\subset\; \text{ordinal} \;\subset\; \text{interval} \;\subset\; \text{ratio} \]

Section 1.2's two-way split still holds and this refines it: nominal and ordinal data are the qualitative ones, interval and ratio the quantitative. The refinement earns its keep in the middle, where ordinal data look calculable and are not. A survey coding excellent as 4 and good as 3 invites an average, and that average assumes the step from good to excellent is the same size as the step from satisfactory to good — an assumption the ordinal scale explicitly does not make.

Figure (svg): A four-step ladder of measurement levels, from nominal at the bottom to ratio at the top, each step adding one permitted operation

A ladder, not a list: ratio data can do everything interval data can, and one thing more.

OpenStax Introductory Statistics 2e, §1.3 Frequency, Frequency Tables, and Levels of Measurement §1.3, pp. 26-27

5. The four levels of measurement

Section

Section 1

6. Names, then order, then differences, then a true zero

Concept

Nominal data are categories, colours, names and labels, and they are not ordered. Ordinal data can be ordered but differences between values cannot be measured. Interval data are numerical and have meaningful differences, but no true zero. Ratio data have a true zero, so ratios between values are meaningful.

the four levels — Nominal: categories with no order, and no calculations. Ordinal: ordered categories, but differences are not measurable, and still no calculations. Interval: numerical with meaningful differences but no minimum or zero value. Ratio: like interval, but with a true zero, so ratios can be calculated.

\[ \text{Celsius: } 100^\circ - 60^\circ = 40^\circ \;\checkmark \qquad \frac{40^\circ}{20^\circ} = 2 \;\text{ meaningless} \]

The book's own examples: smartphone brands are nominal, since there is no agreed order however strong personal preferences may be. The top five national parks are ordinal — they can be ranked one to five, but the gap between first and second is not a measurable amount. Celsius and Fahrenheit are interval, because 40 degrees really is 100 minus 60 while zero is not the lowest possible temperature. Exam scores out of 100 are ratio, because zero means none and 80 is genuinely four times 20.

Figure (svg): A four-step ladder of measurement levels, from nominal at the bottom to ratio at the top, each step adding one permitted operation

A ladder, not a list: ratio data can do everything interval data can, and one thing more.

OpenStax Introductory Statistics 2e, §1.3 Frequency, Frequency Tables, and Levels of Measurement §1.3, pp. 26-27 — the four levels, with the book's examples

7. A ladder, not a list

Picture it

Each rung keeps everything below it and adds exactly one capability.

Figure (svg): A four-step ladder of measurement levels, from nominal at the bottom to ratio at the top, each step adding one permitted operation

A ladder, not a list: ratio data can do everything interval data can, and one thing more.

Drawing the levels as nested steps matters because the containment is the content. A ratio variable can be ordered and differenced as well as divided, so any procedure valid for a lower level is valid for it. The practical consequence is that the level tells you the CEILING on what you may do, and working below the ceiling is always safe — treating exam scores as merely ordinal loses information but tells no lies.

8. Worked example: why temperature is interval and not ratio

Worked example

The test is the zero, and it is worth doing carefully because temperature looks so numerical.

\[ \text{Is } 40^\circ\text{C twice as hot as } 20^\circ\text{C?} \]

Check that differences work

Why: The gap is 20 degrees, and that is a real quantity.

Convert both to Fahrenheit

Why: Forty Celsius is 104 Fahrenheit; twenty Celsius is 68 Fahrenheit.

\[ 104\text{ and } 68 \]

Take the ratio in each scale

Why: In Celsius the ratio is 2; in Fahrenheit it is about 1.53.

Draw the conclusion

Why: A physical fact cannot depend on which scale you wrote it in.

Figure (svg): The solution to Worked example why temperature is interval and not ratio shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \frac{40}{20} = 2 \quad \text{but} \quad \frac{104}{68} \approx 1.53 \]

Verify: confirm the differences survive the same test

Why: Convert the difference instead of the ratio: 20 Celsius degrees of separation is 36 Fahrenheit degrees, and while the number changes, the statement 'the gap between them is the same as the gap between 60 and 80 Celsius' stays true in both scales. Differences are preserved by the conversion and ratios are not, which is exactly the signature of interval data. The reason is that the conversion adds a constant, and adding a constant leaves differences alone while destroying ratios.

OpenStax Introductory Statistics 2e, §1.3 Frequency, Frequency Tables, and Levels of Measurement §1.3, p. 26

9. Place each on the ladder

Sorting

Ask in turn: ordered? differences meaningful? true zero?

Sort into buckets

Sort each variable.

Nominal
blood type; favourite food
Ordinal
military rank; finishing position in a race
Interval
temperature in Fahrenheit
Ratio
weight in kilograms; distance travelled in metres
nom
The values are names with no agreed order, so they can only be counted into groups.
ord
The values have a genuine order, but the gap between consecutive values is not a measurable amount — the gap between first and second place is not the same quantity as between second and third.
int
Numerical with meaningful differences, but zero does not mean absence, so ratios do not survive a change of scale.
rat
There is a true zero meaning none, so ratios are meaningful: 80 kilograms really is twice 40.

Finishing position is the one worth pausing on. The values are 1, 2, 3 and look thoroughly numerical, but the winner may have finished a second ahead or an hour ahead, so the difference between positions carries no fixed amount. Times are ratio; positions computed from them are only ordinal.

10. Worked example: classifying five variables

Worked example

Work down the ladder for each: can it be ordered, differenced, divided?

\[ \text{smartphone brand; cruise rating; year of birth; exam score; national park ranking} \]

Smartphone brand

Why: No agreed order among the brands.

Cruise rating and park ranking

Why: Ordered, but the gaps are not measurable amounts.

Year of birth

Why: Differences work; the year zero is a convention, not an absence of time.

Exam score out of 100

Why: Zero means no marks, so ratios are meaningful.

Figure (svg): The solution to Worked example classifying five variables shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{brand} \to \text{nominal}, \; \text{rating} \to \text{ordinal}, \; \text{year} \to \text{interval}, \; \text{score} \to \text{ratio} \]

Verify: confirm the year of birth really is interval

Why: Someone born in 2000 is not twice as old as someone born in 1000, which settles that the ratio is meaningless. But the difference is perfectly meaningful: those births are 1,000 years apart, and that statement survives moving to any other calendar. Year of birth is the cleanest everyday example of interval data after temperature, and it is a useful one to keep because it shows the issue is a missing true zero rather than anything to do with negative numbers.

OpenStax Introductory Statistics 2e, §1.3 Frequency, Frequency Tables, and Levels of Measurement §1.3, pp. 26-27

11. Trap: averaging an ordinal scale

Trap

The trap

\[ \text{excellent} = 4, \; \text{good} = 3, \; \text{satisfactory} = 2, \; \text{unsatisfactory} = 1 \]

Average the coded responses to get a satisfaction score of 3.4

Why: The codes are numbers, so the arithmetic runs without complaint.

\[ \bar{x} = 3.4 \quad \text{(what is 0.4 of the way from good to excellent?)} \]

The average assumes the step from good to excellent is the same size as the step from satisfactory to good, which the ordinal scale explicitly refuses to claim.

The fix

\[ \text{report the proportion in each category, or the median rating} \]

Use only operations the level supports

Why: Ordinal data can be ordered, so a median is legitimate; they cannot be added, so a mean is not.

The median works because finding the middle response only requires putting the responses in order, which ordinal data support. Reporting proportions works for the same reason a proportion works for any categorical variable. This is a genuine and common dispute in survey research rather than a textbook nicety — averaged Likert scales appear in published work constantly — and the honest position is that the mean requires an assumption about equal spacing that the scale does not supply.

12. What ratio data have that interval data lack

Fill the middle

The single feature that separates the top two rungs.

Fill in the blanks

\textzero ___ \text___

Why: A true zero means the absence of the quantity. Weight has one, so 80 kilograms is genuinely twice 40. Celsius does not, so 40 degrees is not twice 20 — and the proof is that the ratio changes when the same two temperatures are written in Fahrenheit. Everything else about interval and ratio data is identical.

13. One of these is false

Two truths and a lie

All three are about the levels.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Ordinal data can be ordered but not meaningfully subtracted
  • C. Nominal data cannot be used in calculations
  • B. Data recorded as numbers are at least interval level

Survives elimination: B

Why: The survivor is false, and finishing positions are the counterexample: 1, 2 and 3 are numbers, and the differences between them mean nothing fixed. Coding ordinal categories as 1, 2, 3 and 4 does the same thing deliberately. The level is a fact about what the values represent, never about how they are written — the same lesson section 1.2 taught with postal codes.

14. Level to its legal summary

Matching

The highest summary each level supports.

Match the pairs

  • l1. nominal
  • l2. ordinal
  • l3. interval
  • l4. ratio
  • r1. counts and proportions; the mode
  • r2. everything above, plus the median
  • r3. everything above, plus the mean and standard deviation
  • r4. everything above, plus ratios and percent change

Why: The right column is cumulative, matching the ladder. The median appears at ordinal because finding the middle value needs only an ordering; the mean appears at interval because adding values needs differences to be meaningful; and ratios appear only at the top, where a true zero exists. This table is the practical payoff of the whole idea — it says which chapter 2 summary you are entitled to use.

15. Frequency tables

Section

Section 2

16. How often each value occurs

Concept

Once you have a set of data you need to organise it so you can analyse how frequently each datum occurs. A frequency is the number of times a value of the data occurs, and a frequency table lists the distinct data values in ascending order beside their frequencies.

frequency — The number of times a value of the data occurs. The sum of the values in the frequency column equals the total number of members in the sample.

\[ \sum \text{frequencies} = n \]

Twenty students were asked how many hours they work per day, and three answered two hours, five answered three, three answered four, six answered five, two answered six and one answered seven. Those six frequencies sum to twenty, which is the sample size — and that check is the reason to write the total row at all. A frequency table that does not sum to n has lost or duplicated an observation, and finding out now is far cheaper than finding out after three more columns have been computed from it.

Figure (svg): A frequency table of twenty students' daily work hours with frequency, relative frequency and cumulative relative frequency columns

The book's Table 1.12. Every value here was recomputed from the twenty raw responses rather than transcribed.

OpenStax Introductory Statistics 2e, §1.3 Frequency, Frequency Tables, and Levels of Measurement §1.3, p. 27 — frequency and the work-hours table

17. The same table, drawn

Picture it

A bar for each distinct value, its height the frequency.

Figure (svg): A bar chart of how many students work each number of hours per day, from two to seven hours

A frequency table and a bar chart carry exactly the same information; the chart makes the shape visible at a glance.

The table and the chart carry identical information, and the chart is worth drawing because it makes the shape available at a glance: a peak at five hours, a lighter tail out to seven, nothing below two. Chapter 2 will formalise this picture as a histogram, whose only real difference is that the bars touch, because there the horizontal axis is continuous.

18. Worked example: building the table from raw data

Worked example

The book's twenty responses, in the order they were collected.

\[ 5; \;6; \;3; \;3; \;2; \;4; \;7; \;5; \;2; \;3; \;5; \;6; \;5; \;4; \;4; \;3; \;5; \;2; \;5; \;3 \]

List the distinct values in ascending order

Why: Every value that occurs at least once.

\[ 2, 3, 4, 5, 6, 7 \]

Tally each one

Why: Count occurrences of each value across the raw list.

\[ 3, 5, 3, 6, 2, 1 \]

Add the frequency column

Why: The total must be the number of students.

\[ 3 + 5 + 3 + 6 + 2 + 1 = 20 \]

Compare with the sample size

Why: Twenty students were asked.

Figure (svg): The solution to Worked example building the table from raw data shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \sum f = 20 = n \]

Verify: confirm no value was skipped

Why: The distinct values run 2, 3, 4, 5, 6, 7 with no gaps, so nothing between the minimum and maximum was overlooked. A gap would not be an error on its own — a value that genuinely never occurred belongs in the table with a frequency of zero when the values are consecutive counts — but a missing ROW is a real error, and reading the value column for gaps catches it. Here the total of 20 confirms the tally independently.

OpenStax Introductory Statistics 2e, §1.3 Frequency, Frequency Tables, and Levels of Measurement §1.3, p. 27

19. Read the row

Notation

One row of the work-hours frequency table, and four questions it answers differently.

Annotate

On: \( \text{value } = 5, \quad f = 6, \quad \text{rel. } f = 0.30, \quad \text{cum. rel. } f = 0.85 \)

  • The value 5 is the answer some students gave: five hours of work per day.
  • The frequency 6 is how many students gave that answer, out of the twenty asked.
  • The relative frequency 0.30 is that count divided by twenty: 30 percent of the sample work exactly five hours.
  • The cumulative relative frequency 0.85 covers everything up to AND INCLUDING five hours: 85 percent of the sample work five hours or fewer.

Four numbers in one row, and every one answers a different question. The last is the one students misread most often, because 'cumulative' is easy to skim past — it is a statement about a range of values, while the other three are all about the single value five.

20. Worked example: reading a question off the table

Worked example

Frequency tables answer counting questions without returning to the raw data.

\[ \text{How many of the twenty students work more than four hours per day?} \]

Identify which rows qualify

Why: More than four means five, six and seven.

\[ \text{rows } 5, 6\text{ and } 7 \]

Read their frequencies

Why: Six, two and one.

\[ 6, 2, 1 \]

Add them

Why: The total of the qualifying rows.

\[ 6 + 2 + 1 = 9 \]

Express as a share if wanted

Why: Nine of the twenty students.

\[ \frac{9}{20} = 0.45 \]

Figure (svg): The solution to Worked example reading a question off the table shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 6 + 2 + 1 = 9 \quad \text{so} \quad \frac{9}{20} = 0.45 \]

Verify: confirm by counting the complement

Why: Four hours or fewer covers rows 2, 3 and 4 with frequencies 3, 5 and 3, totalling 11. Since 11 plus 9 is 20, the two counts account for every student exactly once and the answer checks. Working the complement is the standard verification for any counting question on a frequency table, and it is worth the extra line because it catches the commonest error here — including the boundary value when the question said 'more than' rather than 'at least'.

OpenStax Introductory Statistics 2e, §1.3 Frequency, Frequency Tables, and Levels of Measurement §1.3, p. 27

21. Trap: confusing the value with its frequency

Trap

The trap

\[ \text{row: } 5 \text{ hours}, \; \text{frequency } 6 \]

Answer 'six hours' when asked for the commonest number of hours worked

Why: Two numbers sit in the row and the student reads the wrong one.

\[ \text{most common} = 6 \text{ hours} \quad \text{(wrong)} \]

Six is how many students gave the answer. The answer they gave was five hours.

The fix

\[ \text{most common} = 5 \text{ hours}, \; \text{given by } 6 \text{ students} \]

Read the value column for what, and the frequency column for how many

Why: Label the two columns in words before answering anything.

The confusion is worth naming because it recurs in every display built on a frequency table. On the bar chart, the horizontal axis carries values and the vertical axis carries frequencies; a question about which value is commonest is answered by pointing at the tallest bar's position, not its height. Chapter 2's mode is defined as the value with the highest frequency, and the definition is stated that carefully for exactly this reason.

22. Complete the check

Faded example

The frequencies in the work-hours table are 3, 5, 3, 6, 2 and 1.

Fill in the blanks

3 + 5 + 3 + 6 + 2 + 1 = 20 = n

Why: The frequency column must sum to the sample size, because every member of the sample contributed exactly one observation to exactly one row. This is the cheapest error check available on any frequency table and it catches both a dropped observation and a double-counted one.

23. How many rows will the table have?

Estimation

A frequency table is built from 100 measured heights, recorded to one decimal place.

Predict first

Roughly how many rows will an ungrouped frequency table have?

  • About 6
  • About 20
  • Close to 100
  • Exactly 10

Correct: Close to 100.

Why: Continuous measurements recorded to a decimal place are nearly all distinct, so almost every observation gets its own row with a frequency of one. Such a table is as long as the raw data and summarises nothing, which is precisely why continuous data are grouped into class intervals instead. Discrete data with few possible values, like the work-hours example, need no grouping — six rows covered twenty students there.

24. What does the total row buy you?

Prediction

Commit before reasoning.

Predict first

Why write the sum of the frequency column at the bottom of the table?

  • Because the sum is needed for the relative frequency column, and because it checks the tally against the sample size
  • Because tables conventionally have a total row
  • Because the sum is the mean of the data
  • Because it is needed to find the mode

Correct: It is the denominator for relative frequencies, and it checks the tally.

Why: The total does double duty: it is exactly the number every frequency will be divided by in the next column, and comparing it with the known sample size verifies that no observation was lost or counted twice. The third option confuses a sum with a mean, and the fourth is wrong because the mode is read off the frequency column directly without needing any total.

25. Relative frequency

Section

Section 3

26. Each count as a share of the whole

Concept

A relative frequency is the ratio of the number of times a value occurs to the total number of outcomes. To find them, divide each frequency by the total number of members in the sample. Relative frequencies can be written as fractions, percents or decimals.

relative frequency — The ratio of the number of times a value of the data occurs to the total number of outcomes. Computed by dividing each frequency by the sample size, and the column totals one.

\[ \text{relative frequency} = \frac{f}{n}, \qquad \sum \frac{f}{n} = 1 \]

For the work-hours data the sample size is twenty, so the frequency 3 becomes 0.15, the frequency 5 becomes 0.25, and so on down to 1 becoming 0.05. The column sums to one, which is the second free check in the table. The book adds an important caveat: because of rounding, the relative frequency column may not always sum exactly to one, and the last cumulative entry may not be exactly one — but each should be close.

Figure (svg): A frequency table of twenty students' daily work hours with frequency, relative frequency and cumulative relative frequency columns

The book's Table 1.12. Every value here was recomputed from the twenty raw responses rather than transcribed.

OpenStax Introductory Statistics 2e, §1.3 Frequency, Frequency Tables, and Levels of Measurement §1.3, pp. 27-28 — relative frequency and the total of one

27. Frequencies and their shares, in one table

Picture it

Three columns computed from one, with two totals that check them.

Figure (svg): A frequency table of twenty students' daily work hours with frequency, relative frequency and cumulative relative frequency columns

The book's Table 1.12. Every value here was recomputed from the twenty raw responses rather than transcribed.

Relative frequency is what makes two samples comparable when they are different sizes — the same argument section 1.2 made for percent columns beside enrolment counts. A frequency of six means nothing until you know whether the sample was twenty or two thousand; a relative frequency of 0.30 means the same thing either way.

28. Worked example: computing the relative frequency column

Worked example

Divide each frequency by the sample size, then check the total.

\[ \text{frequencies } 3, 5, 3, 6, 2, 1 \text{ with } n = 20 \]

Divide the first frequency by n

Why: Three students out of twenty.

\[ \frac{3}{20} = 0.15 \]

Continue down the column

Why: Five, three, six, two and one over twenty.

\[ 0.25, 0.15, 0.30, 0.10, 0.05 \]

Add the new column

Why: The six relative frequencies.

\[ 0.15 + 0.25 + 0.15 + 0.30 + 0.10 + 0.05 \]

Check the total

Why: It must come to one.

\[ = 1.00 \]

Figure (svg): The solution to Worked example computing the relative frequency column shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \frac{3}{20}, \frac{5}{20}, \frac{3}{20}, \frac{6}{20}, \frac{2}{20}, \frac{1}{20} \;\longrightarrow\; \sum = 1 \]

Verify: confirm the total had to be one, whatever the data

Why: The relative frequencies are the frequencies divided by n, so their sum is the sum of the frequencies divided by n, which is n over n. The total is forced to be one by the arithmetic itself, independently of what the data were — which is what makes it a genuine check rather than a coincidence worth noticing. If your column does not total one, either a frequency is wrong or you divided by something other than the sample size.

OpenStax Introductory Statistics 2e, §1.3 Frequency, Frequency Tables, and Levels of Measurement §1.3, p. 27

29. The denominator

Fill the middle

What every frequency is divided by.

Fill in the blanks

\textn = \frac______}

Why: The denominator is always the total number of outcomes, never the value in the row and never the count of everything else. Getting this wrong is the single most common error in building the table, and the column total catches it immediately: dividing by anything but n gives a column that does not sum to one.

30. Worked example: comparing two unequal samples

Worked example

Why the relative column exists at all.

\[ \text{Sample A: } 6 \text{ of } 20 \text{ work 5 hours. Sample B: } 24 \text{ of } 120 \text{ do.} \]

Compare the raw frequencies

Why: Twenty-four is four times six.

Compute A's relative frequency

Why: Six out of twenty.

\[ 0.30 \]

Compute B's relative frequency

Why: Twenty-four out of one hundred and twenty.

\[ 0.20 \]

Compare the shares

Why: Thirty percent against twenty percent.

Figure (svg): The solution to Worked example comparing two unequal samples shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \frac{6}{20} = 0.30 > 0.20 = \frac{24}{120} \]

Verify: confirm the reversal is real and not an arithmetic slip

Why: Scale A up to B's size for a direct comparison: 0.30 of 120 would be 36 students, against B's actual 24. So on equal footing A really does have more five-hour workers, and the raw counts pointed the opposite way purely because B's sample was six times larger. This reversal is the whole reason the relative frequency column exists, and it is the same point the two colleges made in section 1.2 with enrolment percentages.

OpenStax Introductory Statistics 2e, §1.3 Frequency, Frequency Tables, and Levels of Measurement §1.3, pp. 27-28

31. Error analysis: four attempts at a relative frequency

Error analysis

The work-hours table has frequency 6 at the value 5 hours, with a sample of 20. Four students compute the relative frequency.

Annotate

On: \( \begin{aligned} &(1)\; \frac{6}{5} = 1.2 \\ &(2)\; \frac{6}{14} \approx 0.43 \\ &(3)\; \frac{5}{20} = 0.25 \\ &(4)\; \frac{6}{20} = 0.30 \end{aligned} \)

  • (1) divides the frequency by the VALUE. Nothing about five hours is a total, and the giveaway is the answer exceeding one, which no relative frequency can do.
  • (2) divides by the number of students NOT in this row. That produces a ratio of this row to the rest, not a share of the whole, and such a ratio can exceed one.
  • (3) divides the value by the sample size, swapping the two numbers in the row. It gives a legal-looking answer of 0.25, which is why this error survives checking.
  • (4) is correct: the frequency divided by the sample size, giving the share of students who work five hours.

Errors (1) and (2) are caught free by the range check, since a relative frequency must lie between 0 and 1. Error (3) is the dangerous one because 0.25 passes that check — the only defence is the column total, which comes to something other than one as soon as a value is used where a frequency belongs.

32. One of these is false

Two truths and a lie

All three concern relative frequency.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. A relative frequency always lies between 0 and 1
  • C. Rounding can make the column total slightly different from one
  • B. Relative frequencies must be written as decimals

Survives elimination: B

Why: The survivor is false. The book states that relative frequencies can be written as fractions, percents or decimals, and in the probability chapter it actively recommends leaving answers as unreduced fractions. The form is a presentation choice; what matters is that the quantity is a count divided by the sample size.

33. Spot the impossible column

Estimation

Four proposed relative frequency columns for a six-row table.

Predict first

Which could NOT be a relative frequency column?

  • 0.15, 0.25, 0.15, 0.30, 0.10, 0.05
  • 0.20, 0.20, 0.20, 0.20, 0.10, 0.10
  • 0.30, 0.30, 0.30, 0.20, 0.10, 0.05
  • 0.50, 0.20, 0.10, 0.10, 0.05, 0.05

Correct: The third: it totals 1.25.

Why: Adding the third column gives 1.25, which is impossible because the shares of a whole cannot exceed the whole. The other three total exactly one. This check takes a few seconds and catches errors that would otherwise propagate into the cumulative column and then into every conclusion drawn from it, so it is worth running every time a relative frequency column is built.

34. Statement to column

Translation

Each sentence, and the column it is read from.

Match the pairs

  • l1. Six students work exactly five hours
  • l2. Thirty percent of students work exactly five hours
  • l3. Eighty-five percent work five hours or fewer
  • l4. Twenty students were surveyed
  • r1. the frequency column
  • r2. the relative frequency column
  • r3. the cumulative relative frequency column
  • r4. the total of the frequency column

Why: Each column answers a distinct kind of question: how many, what share, what share up to here, and how many altogether. Being able to route a question to the right column before computing anything is most of the skill, and it is what makes a frequency table faster than returning to the raw data.

35. Cumulative relative frequency

Section

Section 4

36. The running total, which must finish at one

Concept

Cumulative relative frequency is the accumulation of the previous relative frequencies. To find them, add all the previous relative frequencies to the relative frequency for the current row. The last entry is one, indicating that one hundred percent of the data has been accumulated.

cumulative relative frequency — The accumulation of the previous relative frequencies: each row's relative frequency added to the running total of all rows above it. The final entry is one, because by then every observation has been counted.

\[ \text{cum. rel. } f_k = \sum_{i \le k} \frac{f_i}{n}, \qquad \text{cum. rel. } f_{\text{last}} = 1 \]

For the work-hours data the column runs 0.15, then 0.15 plus 0.25 giving 0.40, then 0.55, 0.85, 0.95 and finally 1.00. Read it as 'this value or fewer': the 0.85 at five hours says 85 percent of the sample work five hours or less. That reading is what makes the column the ancestor of the percentiles in section 2.3, where the same accumulation is used to locate a value's position in the data.

Figure (svg): A rising staircase of six bars showing cumulative relative frequency climbing from 0.15 to 1.00 across work hours two to seven

Cumulative relative frequency only ever rises, and it must finish at one. Both facts are checks you can run without the data.

OpenStax Introductory Statistics 2e, §1.3 Frequency, Frequency Tables, and Levels of Measurement §1.3, p. 28 — cumulative relative frequency and the final entry of one

37. A staircase that must reach the top

Picture it

Each bar is everything counted so far, and the last one is everybody.

Figure (svg): A rising staircase of six bars showing cumulative relative frequency climbing from 0.15 to 1.00 across work hours two to seven

Cumulative relative frequency only ever rises, and it must finish at one. Both facts are checks you can run without the data.

Two properties are visible and both are checks. The staircase never descends, because relative frequencies are never negative, so a cumulative column that drops has an arithmetic error in it. And the final step must land exactly on one, because by the last row every observation in the sample has been accumulated.

38. Worked example: building the cumulative column

Worked example

Each entry is the one above it plus this row's relative frequency.

\[ \text{relative frequencies } 0.15, \;0.25, \;0.15, \;0.30, \;0.10, \;0.05 \]

The first row is its own relative frequency

Why: Nothing has been accumulated before it.

\[ 0.15 \]

Add the second

Why: 0.15 plus 0.25.

\[ 0.40 \]

Continue down

Why: 0.40 plus 0.15, then plus 0.30, then plus 0.10.

\[ 0.55, 0.85, 0.95 \]

Add the last

Why: 0.95 plus 0.05.

\[ 1.00 \]

Figure (svg): The solution to Worked example building the cumulative column shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 0.15 \to 0.40 \to 0.55 \to 0.85 \to 0.95 \to 1.00 \]

Verify: confirm the column is non-decreasing and ends at one

Why: Every entry is at least as large as the one above it, because each is formed by adding a non-negative relative frequency, and the final entry is exactly 1.00. Both checks pass. A useful third check is that the difference between consecutive cumulative entries recovers the relative frequency of the lower row: 0.85 minus 0.55 gives 0.30, which is the relative frequency at five hours, so the two columns are consistent.

OpenStax Introductory Statistics 2e, §1.3 Frequency, Frequency Tables, and Levels of Measurement §1.3, p. 28

39. Which column answers it?

Sorting

Route each question before computing anything.

Sort into buckets

Sort each question by the column that answers it directly.

Frequency column
How many students work exactly three hours?
Relative frequency column
What share work exactly three hours?
Cumulative relative frequency column
What share work three hours or fewer?; What share work more than five hours?
The total row
How many students were surveyed?
f
The question asks for a count of a single value, which is exactly what a frequency is.
rel
The question asks for a share of the whole at a single value, which is the frequency divided by the sample size.
cum
The question asks about a RANGE of values, which the running total answers directly or by one subtraction.
tot
The question asks for the size of the sample, which is the sum of the frequency column.

Item (d) needs the cumulative column and one subtraction from one: more than five hours is everything not covered by the 0.85 accumulated at five hours, so the answer is 0.15. Recognising 'more than' as a complement of a cumulative entry saves adding the tail rows individually.

40. Worked example: reading a range off the column

Worked example

Cumulative entries answer 'how many at most' directly, and 'how many between' by subtraction.

\[ \text{What share of students work more than three and at most six hours?} \]

Find the cumulative entry at six hours

Why: Everything up to and including six.

\[ 0.95 \]

Find the cumulative entry at three hours

Why: Everything up to and including three.

\[ 0.40 \]

Subtract

Why: Removing everything at or below three from everything at or below six.

\[ 0.95 - 0.40 = 0.55 \]

State the range covered

Why: The rows for four, five and six hours.

\[ 0.15 + 0.30 + 0.10 \]

Figure (svg): The solution to Worked example reading a range off the column shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 0.95 - 0.40 = 0.55 \quad\text{, and }\quad 0.55 \times 20 = 11 \]

Verify: confirm against the relative frequency column directly

Why: Adding the individual relative frequencies for four, five and six hours gives 0.15 plus 0.30 plus 0.10, which is 0.55 — the same answer by a different route. The subtraction method is faster for wide ranges and the direct sum is faster for narrow ones, and agreeing with each other confirms both columns. Note which endpoint got subtracted: taking the cumulative entry at three rather than at four is what makes the answer exclude the three-hour students, matching 'more than three'.

OpenStax Introductory Statistics 2e, §1.3 Frequency, Frequency Tables, and Levels of Measurement §1.3, p. 28

41. Trap: subtracting at the wrong endpoint

Trap

The trap

\[ \text{share working more than 3 and at most 6} \]

Subtract the cumulative entry at four hours from the one at six

Why: The student sees 'more than three' and reaches for the next row up.

\[ 0.95 - 0.55 = 0.40 \quad \text{(wrong: the four-hour students were removed)} \]

Subtracting at four removes everyone up to and including four hours, so the 15 percent who work exactly four hours have been discarded, even though they satisfy the condition.

The fix

\[ 0.95 - 0.40 = 0.55 \]

Subtract the cumulative entry at the value BELOW the range's lower end

Why: Cumulative entries include their own row, so subtracting at three removes exactly the students who fail the condition.

A reliable way to avoid the slip: write the condition as a list of rows first — more than three and at most six means rows 4, 5 and 6 — then subtract the cumulative entry at the row immediately below the first one you want. The check afterwards is to add those rows' relative frequencies directly, which costs one extra line and settles it.

42. Fill the running total

Faded example

Relative frequencies 0.15, 0.25, 0.15, 0.30, 0.10, 0.05.

Fill in the blanks

0.15, \;\; 0.40, \;\; 0.55, \;\; 0.85, \;\; 0.95, \;\; 1.00

Why: Each entry is the previous cumulative value plus the current row's relative frequency. The column is forced to be non-decreasing and to end at one, so those two properties can be used to check any entry you are unsure of — and the gap between consecutive entries must reproduce a relative frequency from the other column.

43. Why must the last entry be one?

Prediction

Commit before reasoning.

Predict first

What guarantees the final cumulative relative frequency is one?

  • A convention adopted to make tables tidy
  • That every observation has been counted by the last row, so the accumulated share is the whole sample
  • That the data happen to be evenly spread
  • Nothing guarantees it; it varies with the data

Correct: Every observation has been counted by the last row.

Why: The cumulative entry at the final row is the sum of all relative frequencies, which is n divided by n. It is forced by the arithmetic and does not depend on the data at all. The only caveat is the book's own: rounding may leave the last entry very slightly off one, and it should still be close. The last option is the misconception worth removing, because it would make the check useless.

44. What happens with an open-ended top class?

Edge cases

A table's last row reads 'seven hours or more' rather than exactly seven.

Discussion prompt

Can the cumulative column still be built, and what is lost?

Hint: Ask which of the four columns still has a definite value for that row.

Answer:

The cumulative column is fine. The last row still accumulates every remaining observation, so the final entry is still one, and every entry above it is unaffected.

What is lost is any information about the values inside that class. You know how many students work seven hours or more, and nothing about whether they work seven or seventy — so no mean, no ratio and no upper endpoint can be computed from the table.

This is why open-ended top classes are common in income tables and awkward in analysis: they are honest about a long tail and they make the mean of the grouped data impossible to compute without an extra assumption about where the tail's values sit.

45. Grouping continuous data, and rounding

Section

Section 5

46. Class intervals, and how many decimals to carry

Concept

Continuous measurements are nearly all distinct, so an ungrouped frequency table has one row per observation and summarises nothing. The remedy is to group the values into class intervals. Separately, the book's rounding convention is to carry the final answer one more decimal place than was present in the original data, and to round only the final answer.

class interval — A range of values that a frequency table treats as one row. Grouping continuous data into class intervals trades the exact values for a visible shape, and is what makes a frequency table useful for measurements.

\[ \text{round the final answer to one more decimal place than the data} \]

The book's table of 100 semiprofessional soccer players' heights uses intervals of two inches, from 59.95 up to 71.95, and the endpoints end in .95 deliberately — heights recorded to one decimal place can never land exactly on a boundary, so no observation is ambiguous. On rounding: the mean of the quiz scores four, six and nine is reported as 6.3, one decimal place because the data are whole numbers, and intermediate results should not be rounded at all where possible.

Figure (svg): A comparison showing an ungrouped frequency table with one row per distinct height against a grouped table with eight class intervals

Continuous measurements are almost all distinct, so an ungrouped table has one row each and summarises nothing.

OpenStax Introductory Statistics 2e, §1.3 Frequency, Frequency Tables, and Levels of Measurement §1.3, pp. 26-33 — rounding, and the grouped table of soccer players' heights

47. One hundred heights, two ways

Picture it

The same data ungrouped and grouped into eight classes.

Figure (svg): A comparison showing an ungrouped frequency table with one row per distinct height against a grouped table with eight class intervals

Continuous measurements are almost all distinct, so an ungrouped table has one row each and summarises nothing.

The ungrouped table is not wrong, merely useless: a hundred rows almost all reading one. Grouping discards the exact heights and buys the thing the table was built for, which is the shape — a heavy concentration between 66 and 68 inches, thinning in both directions. Chapter 2's histogram is this grouped table drawn, with the bars touching to show the axis is continuous.

48. Worked example: why the boundaries end in .95

Worked example

The book's intervals look odd until you ask what could land on a boundary.

\[ \text{intervals } 59.95\text{-}61.95, \; 61.95\text{-}63.95, \; \ldots \text{ for heights recorded to one decimal} \]

Ask what values the data can take

Why: Heights to one decimal place: 61.9, 62.0, 62.1.

Ask whether any can equal a boundary

Why: A boundary of 61.95 has two decimals.

Contrast with a boundary of 62.0

Why: That is a possible recorded height.

State the design rule

Why: Put boundaries one decimal finer than the data.

Figure (svg): The solution to Worked example why the boundaries end in .95 shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{data: } 1 \text{ decimal} \;\Longrightarrow\; \text{boundaries: } 2 \text{ decimals} \]

Verify: confirm the alternative really is a problem

Why: With boundaries at 60, 62, 64 and a recorded height of exactly 62.0, the value satisfies both 60 to 62 and 62 to 64, and the table's total depends on an arbitrary tie-breaking rule that the reader cannot see. Some texts solve this with a stated convention such as 'include the lower boundary, exclude the upper'. The book's approach removes the question rather than answering it, which is why the boundaries look strange and are in fact the tidier choice.

OpenStax Introductory Statistics 2e, §1.3 Frequency, Frequency Tables, and Levels of Measurement §1.3, p. 28

49. Group it, or not?

Discrimination

Ask how many distinct values there are, and whether they are counts or measurements.

Sort into buckets

Sort each data set.

Group into class intervals
100 heights measured to one decimal place; 500 salaries in pounds; 200 reaction times in milliseconds
One row per value is fine
20 answers to 'how many hours do you work?'; 40 answers to 'how many siblings do you have?'
group
The values are measurements or span a wide range, so almost every observation is distinct and an ungrouped table would be as long as the raw data.
raw
The values are counts with only a handful of possibilities, so a table with one row per value is already short and shows the shape.

50. Worked example: applying the rounding convention

Worked example

The book's own example, and the rule for intermediate values.

\[ \text{the mean of the quiz scores } 4, \; 6, \; 9 \]

Compute without rounding

Why: Nineteen divided by three.

\[ 6.3333... \]

Count the decimals in the data

Why: Whole numbers: zero decimal places.

\[ 0\text{ decimals} \]

Carry one more than the data

Why: Zero plus one.

\[ 1\text{ decimal place} \]

Round only at the end

Why: Six point three three rounds to 6.3.

\[ 6.3 \]

Figure (svg): The solution to Worked example applying the rounding convention shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \bar{x} = \frac{4+6+9}{3} = 6.333\ldots \approx 6.3 \]

Verify: confirm why intermediate rounding is forbidden

Why: Had the division been rounded to 6.3 and then used in a further calculation — say multiplying by three to recover the total — it would give 18.9 rather than 19, and the error grows with every further step. The book's instruction is to round off only the final answer, and if an intermediate must be rounded, to carry at least twice as many decimal places as the final answer will have. This matters most in chapter 2's standard deviation, where a squared and re-rooted quantity magnifies any early rounding.

OpenStax Introductory Statistics 2e, §1.3 Frequency, Frequency Tables, and Levels of Measurement §1.3, p. 26

51. Trap: grouping data that did not need grouping

Trap

The trap

\[ \text{work hours: } 2, 3, 4, 5, 6, 7 \;\longrightarrow\; \text{classes } 2\text{-}3, 4\text{-}5, 6\text{-}7 \]

Group the six distinct values into three classes for tidiness

Why: The student groups by habit rather than by need.

\[ \text{three rows instead of six} \quad \text{(information thrown away for nothing)} \]

Six rows already fitted on a page and already showed the shape. The grouping has hidden the peak at five hours inside a class that also contains four.

The fix

\[ \text{keep one row per value when the values are few} \]

Group only when the values are too many or too nearly distinct

Why: Discrete data with a handful of possible values need no grouping.

Grouping is a trade: exact values are given up to make a shape visible. When the shape is already visible the trade is all cost. The rough guide is that a table wants somewhere between five and about fifteen rows — fewer than five hides the shape, many more stops being a summary — so twenty distinct values want grouping and six do not.

52. The rounding rule

Fill the middle

How many decimal places the final answer carries.

Fill in the blanks

\textoriginal data \text___ ___

Why: Whole-number data give an answer to one decimal place, which is why the mean of 4, 6 and 9 is reported as 6.3. Data recorded to one decimal place give an answer to two. The rule keeps the reported precision honest: an answer with six decimal places computed from whole numbers claims an accuracy the data never had.

53. One of these is false

Two truths and a lie

All three concern grouping and rounding.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Grouping discards information about exact values
  • C. Intermediate results should not be rounded where possible
  • B. Class intervals must all be the same width

Survives elimination: B

Why: The survivor is false, though it is usually good practice. Equal widths make a histogram's bar heights directly comparable, which is why the book uses two-inch intervals throughout its heights table. But unequal intervals are legitimate and sometimes necessary, as with income data where a long upper tail is often collected into one wide top class — the cost being that the histogram must then plot density rather than frequency for the bars to be comparable.

54. How many classes?

Estimation

Two hundred reaction times are to be grouped.

Predict first

Roughly how many class intervals should the table have?

  • About 2 or 3
  • Somewhere between about 5 and 15
  • About 50
  • One per observation, so 200

Correct: Somewhere between about 5 and 15.

Why: Too few classes flatten the distribution into a couple of bars that hide the shape entirely; too many recreate the ungrouped table and summarise nothing. The working range for most data sets is five to fifteen, and the book's heights table sits at six. This is a guideline rather than a rule — the right number is whichever makes the shape clearest — but answers outside the range almost always indicate a mistake in the plan.

55. The four columns of a frequency table

Comparison

Fill the blanks. Each column answers a different kind of question.

Comparison matrix

ColumnHow it is computedThe question it answers
Valuethe distinct data values, ascendingwhat answers were given
Frequencycount the occurrences of each valuehow many gave that answer
Relative frequencyeach frequency divided by nwhat share gave that answer
Cumulative relative freq.add each relative frequency to the running totalwhat share gave that answer or a lower one

Two checks come free and should be run every time: the frequency column totals the sample size, and the relative frequency column totals one. A third follows from them — the last cumulative entry is one — and together they catch nearly every arithmetic slip a table can contain.

56. Building a frequency table, in order

Pattern

Six steps from raw data to a table you can answer questions from.

  1. Determine the level of measurement, since it decides which summaries the table may later support.
  2. Decide whether to group: count the distinct values, and group if they are many or if the data are measurements.
  3. List the values or class intervals in ascending order, choosing boundaries one decimal finer than the data so nothing is ambiguous.
  4. Tally the frequencies, then check that they total the sample size.
  5. Divide each frequency by the sample size for the relative frequencies, then check that they total one.
  6. Accumulate the relative frequencies down the column, and check that the last entry is one.

Three checks are built into the procedure, at steps four, five and six, and each costs a few seconds. Skipping them means an error found in step four is not discovered until a conclusion drawn in chapter 2 turns out to be nonsense.

OpenStax Introductory Business Statistics 2e, §1.3 Levels of Measurement §1.3 Levels of Measurement

57. Check yourself 1 of 3

Check

Place it on the ladder.

Check your understanding

A survey records each respondent's highest qualification as school, college, or university. What level of measurement is this?

  • A. Ordinal (correct)
  • B. Nominal
  • C. Interval
  • D. Ratio

Answer: A

Why: The three categories have a genuine order — university is above college, which is above school — but the gap between them is not a measurable amount, which is exactly the ordinal level.

Why B tempts people
Nominal data have no order at all. These three categories are clearly ranked, so something more than nominal is available.
Why C tempts people
Interval data are numerical with measurable differences. There is no number of units between school and college.
Why D tempts people
Ratio requires a true zero and meaningful ratios. University is not twice school in any sense.

58. Check yourself 2 of 3

Check

Read the right column.

Check your understanding

In the work-hours table, the cumulative relative frequency at four hours is 0.55. What does that mean?

  • A. 55 percent of students work four hours or fewer (correct)
  • B. 55 percent of students work exactly four hours
  • C. 55 percent of students work more than four hours
  • D. 55 students work four hours

Answer: A

Why: A cumulative relative frequency accumulates every row up to and including the current one, so it reports the share at or below that value.

Why B tempts people
That is the relative frequency, which is 0.15 for four hours. The cumulative column includes the rows above it as well.
Why C tempts people
That would be the complement, 0.45, found by subtracting the cumulative entry from one.
Why D tempts people
The sample is only 20 students, so 55 of them is impossible. The 0.55 is a share, not a count.

59. Check yourself 3 of 3

Check

Apply the rounding convention.

Check your understanding

Three whole-number quiz scores have a mean of 7.16666. How should it be reported?

  • A. 7.2 (correct)
  • B. 7.17
  • C. 7.167
  • D. 7

Answer: A

Why: The data are whole numbers, carrying zero decimal places, so the final answer carries one more than that: a single decimal place, giving 7.2.

Why B tempts people
Two decimal places claims more precision than whole-number data support.
Why C tempts people
Three decimal places claims far more precision than the data carry.
Why D tempts people
Rounding to a whole number carries the same precision as the data rather than one more, and discards real information.

60. Where this shows up outside the textbook

Real world

A hospital rates each patient's pain on a scale from 0, no pain, to 10, the worst imaginable. An analyst reports that the mean pain score fell from 6.2 to 5.4 after a new protocol, and concludes that pain fell by 13 percent.

Discussion prompt

Identify the level of measurement, say which of the two claims is defensible and which is not, and give a summary the analyst could report instead.

Hint: Ask whether the step from 5 to 6 is the same size as the step from 9 to 10, and whether 0 really means none.

Answer:

The scale is ordinal, not ratio. Patients are ordered by how much pain they report, but nothing guarantees that the step from 5 to 6 is the same amount of pain as the step from 9 to 10 — and reports differ systematically between people, so the numbers are not comparable across patients in the way an interval scale requires.

The 13 percent claim is not defensible. Percent change is a ratio operation and needs a true zero with equal spacing above it. Even granting that 0 means no pain, unequal spacing alone destroys the calculation.

The direction of the change is defensible. Ordinal data support order, so 'scores went down' is a legitimate reading, and a median or a comparison of proportions supports it properly: report the median score before and after, or the share of patients reporting 7 or above.

\[ \text{ordinal} \;\Longrightarrow\; \text{median and proportions} \;\checkmark \qquad \text{mean and percent change} \;\times \]

This is not a contrived example. Pain scales, satisfaction surveys and severity ratings are averaged in published clinical and business work constantly, and the practice is genuinely contested. The defensible position is the one this section supplies: report what the level of measurement supports, and if a mean is reported anyway, state the equal-spacing assumption it rests on rather than leaving it hidden.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

Which check would catch a frequency mistakenly entered where a data value belongs?

  • The relative frequency column totalling one
  • The cumulative column ending at one
  • The frequency column totalling the sample size
  • None of these; only re-tallying the raw data would catch it

Correct: The frequency column totalling the sample size.

\[ \sum f = n \;\text{ checks the tally}; \quad \sum \frac{f}{n} = 1 \;\text{ checks only the division} \]

Why: Swapping a value into the frequency column changes the frequency total, so it no longer matches the known sample size. The relative frequency and cumulative checks would not catch it independently, because both are computed from the frequency column and would simply inherit the error — dividing the wrong frequencies by the wrong total still gives a column summing to one. This is why the sample size must be known from outside the table for the first check to have any force.

62. Explain it to someone a year behind you

Explain it

They cannot see why temperature is not ratio data when it is obviously numbers.

Discussion prompt

In three sentences or fewer, give them a test they can run themselves that settles it.

Hint: Have them convert to another scale and take the ratio again.

Answer:

Tell them to write down that 40 degrees Celsius is twice 20, then convert both to Fahrenheit: 104 and 68. Now the ratio is about 1.53, not 2, and no physical fact can change depending on which scale it was written in.

The reason is that the conversion adds a constant as well as scaling, and adding a constant leaves differences alone while destroying ratios — which is exactly why interval data support subtraction and not division.

\[ \frac{40}{20} = 2 \quad \text{but} \quad \frac{104}{68} \approx 1.53 \]

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Telling ordinal from interval data
  • Building a frequency table and checking it
  • Reading a range off the cumulative column
  • Choosing class intervals for continuous data

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For ordinal against interval, ask whether the gap between consecutive values is a measurable amount. For the table, always write the total row and check it against the sample size before computing anything else. For the cumulative column, write the condition as a list of rows first, then subtract at the row below the range's lower end. For class intervals, aim for five to fifteen classes with boundaries one decimal finer than the data. Do five examples of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Fifteen minutes.

Draw it

At the top, draw the four levels of measurement as a staircase with nominal at the bottom and ratio at the top. Beside each step write the one capability it adds, two examples, and the highest summary it permits — mode, median, mean, and ratios respectively. Include one example whose values are digits but whose level is nominal. In the middle of the page, build the complete frequency table for these twenty responses: 5, 6, 3, 3, 2, 4, 7, 5, 2, 3, 5, 6, 5, 4, 4, 3, 5, 2, 5, 3. Use four columns — value, frequency, relative frequency, cumulative relative frequency — and write the two totals underneath with a tick beside each. Then answer three questions from your table without looking at the raw data: how many work exactly three hours, what share work five or fewer, and what share work more than three and at most six. Show the subtraction for the last one and state which cumulative entry you subtracted and why. At the bottom, write the rounding rule and apply it to the mean of 4, 6 and 9, and write one sentence saying when continuous data must be grouped and why the class boundaries carry an extra decimal place.

Check your cumulative column two ways: it must never decrease, and the difference between consecutive entries must reproduce the relative frequency of the lower row. If either fails, the error is in the relative frequency column rather than in the accumulation.

65. What you can do now

Recap

Five things, and the first one governs everything you are allowed to do with the other four.

If you seeThen
Categories with no orderNominal: counts, proportions and the mode only
Ordered categories with unmeasurable gapsOrdinal: add the median, but not the mean
Numbers with meaningful differences, no true zeroInterval: add the mean, but not ratios
Numbers with a true zeroRatio: everything, including percent change
A frequency columnCheck it totals the sample size
A relative frequency columnCheck it totals one
A question about a range of valuesUse the cumulative column, subtracting below the range
Continuous data with many distinct valuesGroup into five to fifteen classes

Section 1.4 closes the chapter by asking how a study must be DESIGNED for any of this to mean anything: what a treatment and a control group are, why blinding exists, and what obligations a researcher has to the people being studied.

OpenStax Introductory Statistics 2e, §1.3 Frequency, Frequency Tables, and Levels of Measurement §1.3, pp. 26-33 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §1.3 Frequency, Frequency Tables, and Levels of Measurement — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 26-33
  2. OpenStax Introductory Business Statistics 2e, §1.3 Levels of Measurement — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0
  3. OpenStax Introductory Business Statistics 2e, §2.1 Display Data — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108