Where a value sits inside its data set. Quartiles cut the ordered data into four parts, with the first quartile the median of the lower half and the third the median of the upper half; the interquartile range measures the spread of the middle fifty percent and supplies the 1.5 IQR rule that turns a visual judgement about outliers into arithmetic anyone can repeat; percentiles cut the data into hundredths and are what a score in the ninetieth percentile actually means; and two formulas run in opposite directions, one taking a percentile and returning a value by the index rule, the other taking a value and returning its percentile by counting what lies below it.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 2 — Descriptive Statistics
Measures of the Location of the Data
Objectives
Five outcomes. The last two are the same idea run in opposite directions, and telling them apart is the whole trick.
OpenStax Introductory Statistics 2e, §2.3 Measures of the Location of the Data §2.3, pp. 86-93 — the section these objectives are drawn from
Warm-up
Section 1.3's cumulative relative frequency column answered questions like what share of students work five hours or fewer.
Discussion prompt
That column runs from 0 to 1 up the data. If you wanted to run it backwards — to ask which value has three quarters of the data at or below it — what would you need?
Hint: You have the share and you want the value, rather than the other way round.
Answer:
The cumulative column takes a value and returns a share. The question asked here takes a share and returns a value, which is the same relationship read in the opposite direction.
Both directions are useful and both get names in this section. The value with three quarters of the data at or below it is the third quartile, and the general version — the value with any given share below it — is a percentile.
Section 2.1 also left something unfinished: it called a value an outlier when it did not fit the pattern, and admitted that was a judgement rather than a rule. The interquartile range, built from two of these quartiles, is what turns that judgement into arithmetic.
Concept
The common measures of location are quartiles and percentiles. Quartiles divide an ordered data set into four equal parts and percentiles divide it into a hundred. To calculate either, the data must first be ordered from smallest to largest.
quartiles — Three numbers that separate ordered data into quarters. About one fourth of the data falls on or below the first quartile Q1, about one half on or below the second quartile Q2, and about three fourths on or below the third quartile Q3. A quartile may or may not be one of the data values.
\[ Q_1: \; 25\% \text{ at or below} \qquad Q_2: \; 50\% \qquad Q_3: \; 75\% \]
A measure of location says nothing about how big a value is in absolute terms and everything about where it sits relative to the rest. That is what makes percentiles useful for comparing values across different tests and different populations, and it is why universities use them so heavily — the book's example is a university accepting scores at or above the 75th percentile, which translates into a particular score only once you know the distribution behind it.
Figure (svg): Fourteen ordered data values on a number line with three dashed cuts marking the first quartile, median and third quartile, dividing the values into four groups
OpenStax Introductory Statistics 2e, §2.3 Measures of the Location of the Data §2.3, p. 86
Section
Section 1
Concept
To find the quartiles, first find the median, which is the second quartile. The first quartile is the middle value of the lower half of the data and the third quartile is the middle value, or median, of the upper half. Quartiles may or may not be part of the data.
first and third quartiles — Q1 is the median of the lower half of the ordered data and Q3 is the median of the upper half. When the number of values is odd, the overall median belongs to neither half and is excluded from both. One fourth of the values are at or below Q1 and three fourths are at or below Q3.
\[ Q_1 = \text{median of the lower half}, \qquad Q_3 = \text{median of the upper half} \]
The book's own worked case: for the ordered values 1, 1, 2, 2, 4, 6, 6.8, 7.2, 8, 8.3, 9, 10, 10, 11.5 the median lies between the seventh and eighth values, giving 7. The lower half is 1, 1, 2, 2, 4, 6, 6.8 with middle value 2, and the upper half is 7.2, 8, 8.3, 9, 10, 10, 11.5 with middle value 9. So Q1 is 2 and Q3 is 9, both of which happen to be data values, while the median 7 is not.
Figure (svg): Fourteen ordered data values on a number line with three dashed cuts marking the first quartile, median and third quartile, dividing the values into four groups
OpenStax Introductory Statistics 2e, §2.3 Measures of the Location of the Data §2.3, pp. 86-87 — the median and the quartiles, with the fourteen-value example
Picture it
The book's fourteen values, ordered, with the cuts marked.
Figure (svg): Fourteen ordered data values on a number line with three dashed cuts marking the first quartile, median and third quartile, dividing the values into four groups
Note that the median here is 7 and no observation equals 7 — a quartile is a position in the data rather than necessarily a member of it. Note also that the four parts hold three, four, three and four values rather than three and a half each: with fourteen values an exact quarter is impossible, which is why the book says ABOUT one fourth falls at or below Q1.
Worked example
The book's introductory set, worked in the order the definition prescribes.
\[ 1; \;11.5; \;6; \;7.2; \;4; \;8; \;9; \;10; \;6.8; \;8.3; \;2; \;2; \;10; \;1 \]
Order the data
Why: Smallest to largest, keeping repeats.
\[ 1, 1, 2, 2, 4, 6, 6.8, 7.2, 8, 8.3, 9, 10, 10, 11.5 \]
Find the median
Why: Fourteen values, so average the seventh and eighth.
\[ \frac{6.8 + 7.2}{2} = 7 \]
Take the median of the lower half
Why: The first seven values, whose middle is the fourth.
\[ Q 1 = 2 \]
Take the median of the upper half
Why: The last seven values, whose middle is the fourth of those.
\[ Q 3 = 9 \]
Figure (svg): The solution to Worked example the quartiles of fourteen values shown as a ladder of expressions, one row per legal move
\[ Q_1 = 2, \quad Q_2 = 7, \quad Q_3 = 9 \]
Verify: confirm the counts either side of each cut
Why: Two of the fourteen values are below 2 and four are at or below it, so about a quarter sit at or below Q1 as promised. Ten of the fourteen are at or below 9, which is about three quarters. And exactly seven lie below 7 and seven above, which is what the median guarantees. The word ABOUT in the definition is doing real work: with fourteen values no cut can leave exactly 3.5 on a side, so the quartiles land as close as the data allow.
OpenStax Introductory Statistics 2e, §2.3 Measures of the Location of the Data §2.3, pp. 86-87
Fill the middle
Complete the definition.
Fill in the blanks
Q_1 = \textlower ___ \text___
Why: Q1 is the median of the lower half and Q3 the median of the upper half. Splitting at the overall median and taking the middle of each piece is the whole procedure, and when the count is odd the median itself goes into neither piece.
Worked example
Example 2.13's thirteen prices. With an odd count the median is excluded from both halves.
\[ \text{13 home prices, from } 114\,950 \text{ to } 5\,500\,000 \]
Find the median
Why: Thirteen values, so the seventh is the middle one.
\[ \text{median } = 488 800 \]
Form the lower half WITHOUT the median
Why: The six values below it.
Take their median
Why: Average the third and fourth of those six.
\[ Q 1 = 308 750 \]
Do the same above
Why: The six values above the median, averaged at the middle.
\[ Q 3 = 649 000 \]
Figure (svg): The solution to Worked example an odd number of values shown as a ladder of expressions, one row per legal move
\[ Q_1 = 308\,750, \quad Q_2 = 488\,800, \quad Q_3 = 649\,000 \]
Verify: confirm why the median is excluded from both halves
Why: Including it would put thirteen values into two halves of seven each, counting it twice, and each half's median would shift toward it. The book's rule is that Q1 is the median of the LOWER HALF, and with an odd count the middle value is in neither half — it is the divider. This convention matters because it is not universal: a spreadsheet's QUARTILE function interpolates instead and returns a different Q1 on the same thirteen prices, so a student checking by computer may see a number the book does not.
OpenStax Introductory Statistics 2e, §2.3 Measures of the Location of the Data §2.3, p. 87
Trap
\[ 1; \;11.5; \;6; \;7.2; \;4; \;8; \;9; \;10; \;6.8; \;8.3; \;2; \;2; \;10; \;1 \]
Average the seventh and eighth values as they were written down
Why: The rule says the seventh and eighth, so the student counts along the list given.
\[ \frac{9 + 10}{2} = 9.5 \quad \text{(wrong: that is not the median)} \]
The seventh and eighth values of the ORDERED set are 6.8 and 7.2, giving 7. The unordered list's seventh and eighth are two arbitrary observations.
\[ 1, 1, 2, 2, 4, 6, 6.8, 7.2, 8, 8.3, 9, 10, 10, 11.5 \;\Longrightarrow\; \text{median} = 7 \]
Order the data before any measure of location is computed
Why: Every rule in this section counts positions, and positions only mean something in a sorted list.
The book states this as the first step for both quartiles and percentiles, and it is the single most common source of a wrong answer here. It is worth writing the ordered list out once and working from it for the rest of the question, because the median, both quartiles and every percentile all read positions off the same sorted sequence.
Sorting
A quartile may or may not be one of the observations. Decide for each.
Sort into buckets
Sort each quartile from the book's fourteen-value example.
The pattern is exactly the parity of the half being taken. It matters because students sometimes reject a computed quartile on the grounds that no such observation exists, when a quartile is a position rather than a member.
Prediction
Commit before reasoning.
Predict first
A data set has 13 values. When forming the lower half to find Q1, what happens to the seventh value?
Correct: It is excluded from both halves.
Why: The seventh value is the median, and with an odd count it is the divider rather than a member of either half. That leaves six values below and six above, and the median of each six gives Q1 and Q3. Putting it into one half would make the halves unequal, and putting it into both would count it twice — and both choices shift the quartiles toward the centre.
Faded example
The six values above the median are 529,000; 575,000; 639,000; 659,000; 1,095,000; 5,500,000.
Fill in the blanks
Q_3 = \frac659000}}649000 = ___
Why: Six values have no single middle, so Q3 is the average of the third and fourth of them, which are 639,000 and 659,000. Notice that the enormous value of 5,500,000 has no effect on Q3 at all — it is simply the last of the six, and only the middle two are used. That insensitivity is exactly what makes the quartiles usable for detecting outliers.
Section
Section 2
Concept
The interquartile range is a number that indicates the spread of the middle half, or middle fifty percent, of the data. It is the difference between the third quartile and the first. It can help determine potential outliers: a value is suspected to be one if it is more than 1.5 times the IQR below Q1 or above Q3.
interquartile range — The difference Q3 minus Q1, measuring the spread of the middle fifty percent of the data. A value more than 1.5 times the IQR below Q1, or more than 1.5 times the IQR above Q3, is suspected to be a potential outlier and always requires further investigation.
\[ IQR = Q_3 - Q_1, \qquad \text{fences at } Q_1 - 1.5(IQR) \text{ and } Q_3 + 1.5(IQR) \]
The rule works because the quartiles ignore the extremes. A single enormous value can drag the largest observation anywhere it likes without moving Q1 or Q3 at all, so the fences are computed from a part of the data the outlier cannot influence. Building the detector out of quantities the suspect cannot affect is exactly why the rule is trustworthy, and it is the same reason section 2.6 will prefer the median to the mean when a distribution has a long tail.
Figure (svg): Thirteen home prices on a number line with the interquartile box, the two 1.5 IQR fences, and one price marked as an outlier beyond the upper fence
OpenStax Introductory Statistics 2e, §2.3 Measures of the Location of the Data §2.3, p. 87 — the IQR and the 1.5 IQR rule
Picture it
Example 2.13, with the middle half boxed and the fences drawn.
Figure (svg): Thirteen home prices on a number line with the interquartile box, the two 1.5 IQR fences, and one price marked as an outlier beyond the upper fence
The lower fence comes out negative, which is not a problem — it simply means no price could possibly be low enough to be a low outlier, since prices cannot be negative. The upper fence at 1,159,375 is exceeded by exactly one price. Note that 1,095,000 sits just inside it and is therefore NOT flagged, even though it is more than twice the median: the rule is about distance from the middle half, not about looking large.
Worked example
Example 2.13. Four lines of arithmetic, in a fixed order.
\[ Q_1 = 308\,750, \quad Q_3 = 649\,000 \]
Compute the IQR
Why: The third quartile minus the first.
\[ 649 000 - 308 750 = 340 250 \]
Multiply by 1.5
Why: The distance each fence sits beyond its quartile.
\[ 1.5(340 250) = 510 375 \]
Place the lower fence
Why: Q1 minus that distance.
\[ 308 750 - 510 375 = -201 625 \]
Place the upper fence
Why: Q3 plus that distance.
\[ 649 000 + 510 375 = 1 159 375 \]
Figure (svg): The solution to Worked example the full outlier calculation shown as a ladder of expressions, one row per legal move
\[ 5\,500\,000 > 1\,159\,375 \;\Longrightarrow\; \text{potential outlier} \]
Verify: confirm nothing else is flagged
Why: The next largest price is 1,095,000, which is below the upper fence of 1,159,375 and so is not flagged, and no price is below the lower fence because none is negative. Exactly one of the thirteen is identified. This matches what section 2.1's eye test would have said about a value nearly five times the next largest, and the agreement is the point: the rule is a way of making a judgement reproducible, not of overriding it.
OpenStax Introductory Statistics 2e, §2.3 Measures of the Location of the Data §2.3, p. 87
Faded example
The evening class has Q1 = 78 and Q3 = 89.
Fill in the blanks
IQR = 89 - 78 = 11, \qquad \text61.5 = 78 - 1.5(___) = ___
Why: One and a half times 11 is 16.5, and 78 minus 16.5 is 61.5. Any evening score below 61.5 is a potential outlier, which catches both the 45 and the 25.5. The tightness of the evening class's middle half is exactly what makes its fences close in and its low scores conspicuous.
Worked example
Example 2.14. The day and evening statistics classes.
\[ \text{day: } Q_1 = 56, \; Q_3 = 82.5 \qquad \text{evening: } Q_1 = 78, \; Q_3 = 89 \]
Compute the day class IQR
Why: The difference of its quartiles.
\[ 82.5 - 56 = 26.5 \]
Compute the evening class IQR
Why: The same for the evening scores.
\[ 89 - 78 = 11 \]
Compare them
Why: The day IQR is more than twice the evening IQR.
Apply the fences to the evening class
Why: 78 minus 1.5 times 11.
\[ \text{lower fence } 61.5 \]
Figure (svg): The solution to Worked example comparing two IQRs shown as a ladder of expressions, one row per legal move
\[ IQR_{\text{day}} = 26.5 \quad \text{against} \quad IQR_{\text{evening}} = 11 \]
Verify: confirm the day class has no outliers despite its lower minimum
Why: The day class contains scores of 32, lower than the evening class's 45, and yet 45 is an outlier and 32 is not. The reason is that each class is judged against its own spread: the day fences are 16.25 and 122.25, so 32 sits comfortably inside, while the evening class's tight middle half puts its lower fence at 61.5. Being an outlier is a statement about a value's position within ITS OWN data set, never an absolute judgement about the number.
OpenStax Introductory Statistics 2e, §2.3 Measures of the Location of the Data §2.3, p. 88
Error analysis
For the home prices, Q1 is 308,750 and Q3 is 649,000. Four students look for outliers.
Annotate
On: \( \begin{aligned} &(1)\; Q_3 + 1.5 = 649\,001.5 \\ &(2)\; 1.5 \times 649\,000 = 973\,500 \\ &(3)\; Q_3 + IQR = 989\,250 \\ &(4)\; Q_3 + 1.5(IQR) = 1\,159\,375 \end{aligned} \)
Errors (2) and (3) both produce fences that look reasonable and both wrongly flag the second largest price, which is why the arithmetic is worth writing out in the four separate steps rather than combining them. The quantity multiplied by 1.5 is always the IQR, never a quartile.
Discrimination
A data set has Q1 = 20, Q3 = 40, so the IQR is 20 and the fences are at -10 and 70.
Sort into buckets
Sort each value.
Two truths and a lie
All three concern the IQR and the rule.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The book says potential outliers always require further investigation, and calls them a possible key to understanding the data as much as a possible error. Deleting a value because a rule flagged it is how the most interesting observation in a data set gets thrown away — section 2.1 made the same point when it said explaining an outlier takes background information.
Socratic
The rule could have used the mean and the standard deviation instead, or a different multiplier.
Discussion prompt
Why is the rule built from the quartiles rather than from the mean, and what would go wrong otherwise?
Hint: Ask what an extreme value does to each ingredient.
Answer:
An outlier inflates the mean and inflates the standard deviation even more, because the deviation is squared. So a detector built from those two would have its own threshold pushed outward by the very value it is meant to catch — the outlier partly hides itself.
The quartiles are computed from the middle of each half, where a single extreme value has no influence at all. Adding a price of fifty million to the home data would not move Q1 or Q3 by a penny, so the fences stay exactly where they were and the detector still works.
The multiplier 1.5 is a convention rather than a theorem. It is chosen to flag few values in a well-behaved data set while still catching genuinely distant ones, and some texts use 3 as a second, stricter fence for extreme outliers.
Section
Section 3
Concept
Percentiles divide ordered data into hundredths. To score in the 90th percentile of an exam does not necessarily mean that you received 90 percent on the test. It means that 90 percent of test scores are the same as or less than your score, and 10 percent are the same as or greater.
percentile — A measure of location dividing ordered data into hundredths. A value at the kth percentile has about k percent of the data at or below it. Percentiles are mostly used with very large populations, where saying that k percent are LESS than the value is also acceptable.
\[ \text{the } k\text{th percentile: about } k\% \text{ of the data are at or below it} \]
The distinction the book draws is the one that gets misread outside the classroom. A percentile is a position, not a score. A student in the 90th percentile might have answered 62 percent of the questions correctly on a hard exam, or 98 percent on an easy one — what the percentile reports is that they did as well as or better than nine tenths of the people who sat it. That is why universities use percentiles: the book's example is a university accepting scores at or above the 75th percentile, which corresponds to a particular score only once the distribution is known.
Figure (svg): Two columns contrasting finding the value at a given percentile with finding the percentile of a given value
OpenStax Introductory Statistics 2e, §2.3 Measures of the Location of the Data §2.3, p. 86 — percentiles and what a 90th percentile score means
Picture it
Percentile questions come in two shapes, and the formulas are not interchangeable.
Figure (svg): Two columns contrasting finding the value at a given percentile with finding the percentile of a given value
The reliable first move on any percentile question is to ask what you were GIVEN. Given a percentile and asked for a value, use the index formula and read a number out of the data. Given a value and asked for its percentile, count how many observations lie below it. Choosing the wrong direction produces an answer in the wrong units, which is at least a visible error — an answer of 64 years to a question that wanted a percentage.
Worked example
Three statements about one student, only some of which follow.
\[ \text{A student's exam score is at the } 90\text{th percentile.} \]
What does follow
Why: Ninety percent of scores are at or below this one.
What does not follow
Why: Nothing at all about the number of marks obtained.
\[ \text{not } 90 \%\text{ correct} \]
What also follows
Why: Ten percent of scores are the same or greater.
What the score itself is
Why: It depends entirely on the distribution.
Figure (svg): The solution to Worked example interpreting a reported percentile shown as a ladder of expressions, one row per legal move
\[ 90\text{th percentile} \;\ne\; 90\% \text{ correct} \]
Verify: confirm with two different exams
Why: On an exam where the scores cluster between 30 and 50 percent, the 90th percentile might be a mark of 48 percent. On an easy exam where nearly everyone scores above 85, the 90th percentile might be 97 percent. The same percentile corresponds to entirely different marks, which is precisely what makes it useful for comparing performance across different tests — and precisely why it says nothing on its own about how much a student knows.
OpenStax Introductory Statistics 2e, §2.3 Measures of the Location of the Data §2.3, p. 86
Matching
The same three cuts under two names.
Match the pairs
Why: Quartiles are the percentiles at multiples of 25, and having both vocabularies is a convenience rather than a distinction. The median carries three names in this book — the median, the second quartile and the 50th percentile — and section 2.5 will add a fourth role for it as a measure of centre.
Worked example
The two vocabularies name the same cuts.
\[ \text{Which percentile is the first quartile?} \]
Recall what Q1 marks
Why: One fourth of the data at or below it.
\[ 25 \% \]
Express that as a percentile
Why: The value with 25 percent at or below it.
Do the same for the median
Why: One half at or below it.
And for Q3
Why: Three fourths at or below it.
Figure (svg): The solution to Worked example quartiles are percentiles shown as a ladder of expressions, one row per legal move
\[ Q_1 = P_{25}, \quad Q_2 = P_{50}, \quad Q_3 = P_{75} \]
Verify: confirm the two methods agree on the book's own data
Why: Example 2.17 computes the first quartile of a fifty-value data set by the percentile index formula rather than by the median-of-halves method, and gets 6 — treating the two as interchangeable, which they are in intent. In practice the two routes can differ by a little on small data sets, because the index formula interpolates between neighbours while the median-of-halves rule does not. On large data sets the difference vanishes, which is why the book is content to move between them.
OpenStax Introductory Statistics 2e, §2.3 Measures of the Location of the Data §2.3, p. 91
Trap
\[ \text{a student scores at the } 30\text{th percentile} \]
Report that the student got 30 percent of the questions right
Why: The number is followed by a percent-like word, so it reads as a mark.
\[ 30\text{th percentile} \;\longrightarrow\; 30\% \text{ correct} \quad \text{(wrong)} \]
The student may have answered 70 percent of the questions correctly on a difficult exam and still sat below seven tenths of the candidates.
\[ 30\% \text{ of candidates scored at or below this student} \]
Read a percentile as a position in a group, never as a quantity achieved
Why: It is a statement about other people's scores as much as about this one.
A useful habit is to say the sentence out loud in full: 'thirty percent of the scores were the same or lower.' That phrasing makes it impossible to confuse with a mark, and it also makes clear that the percentile depends on who else sat the exam. A student whose performance never changes can move between percentiles simply by joining a stronger or weaker group.
Two truths and a lie
All three concern what a percentile reports.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and it is the misreading the book warns about explicitly. The 80th percentile means eighty percent of the scores were the same or lower. What mark that corresponds to depends entirely on the exam and on who sat it, and could be anything from 45 percent to 99 percent.
Prediction
Commit before reasoning.
Predict first
A runner's time is unchanged but she moves from a club race to a national championship. What happens to her percentile?
Correct: It falls, because the field is stronger.
Why: A percentile is computed from the other values in the data set, so changing the group changes the percentile even when the individual's value is fixed. In a stronger field fewer runners are at or below her time, so her percentile drops. This is the practical content of a measure of LOCATION: it locates a value within a particular set, and moving the set moves the answer.
Socratic
The book gives the careful definition and then relaxes it.
Discussion prompt
The book says 90 percent of scores are 'the same or less' but then allows saying 'less'. When is that relaxation safe, and when is it not?
Hint: Ask how much difference one observation makes.
Answer:
The book says percentiles are mostly used with very large populations, so removing one particular data value is not significant. In a group of fifty thousand candidates, whether one score is counted as at-or-below or strictly-below changes the percentile by two thousandths of a percent.
It stops being safe when the data set is small or when there are many ties. In a set of twenty values with eight of them equal to the value in question, the two readings differ by twenty percentage points, and which one is meant genuinely matters.
This is why the book's second formula counts x values strictly below and adds HALF the y values that tie — it splits the difference rather than choosing a side, which keeps the answer symmetric and keeps percentiles of tied values sensible.
Section
Section 4
Concept
There are several formulas for calculating the kth percentile in circulation; the book gives one. Order the data from smallest to largest, compute the index as k over 100 times the quantity n plus 1, and then either take the value at that position or average the two values around it.
the index formula — For the kth percentile of n ordered values, the index is i equal to k over 100 times (n + 1). If i is an integer, the kth percentile is the data value in the ith position. If i is not an integer, round it up and down to the nearest integers and average the two data values in those positions.
\[ i = \left(\frac{k}{100}\right)(n + 1) \]
The n plus 1 rather than n is worth a moment. It places the percentiles at evenly spaced positions between the observations rather than on them: with 29 values the 50th percentile lands at position 15, exactly the middle, whereas using n would give 14.5. The formula also explains why the book says there are several formulas in circulation — different choices here give slightly different answers on small data sets, and a calculator may not use this one.
Figure (svg): A diagram of the percentile index formula showing an integer index taking a single value and a non-integer index averaging two neighbouring values
OpenStax Introductory Statistics 2e, §2.3 Measures of the Location of the Data §2.3, p. 90 — the index formula, with the actors' ages examples
Picture it
Whether the index comes out whole decides everything that follows.
Figure (svg): A diagram of the percentile index formula showing an integer index taking a single value and a non-integer index averaging two neighbouring values
The rule for a non-integer index is to round UP and DOWN and average the two values — never to round the index to the nearest whole number. Rounding 24.9 to 25 would give 72 rather than 71.5, and while the difference is small here it is systematic: it would push every non-integer percentile toward one side. Averaging is what keeps the answer between the two neighbouring observations, where it belongs.
Worked example
Example 2.16a. Twenty-nine ages of Academy Award winning best actors.
\[ \text{Find the } 70\text{th percentile of the } 29 \text{ ordered ages.} \]
Write down k and n
Why: The percentile wanted and the number of values.
\[ k = 70, n = 29 \]
Compute the index
Why: k over 100, times n plus 1.
\[ i = (0.70) (30) = 21 \]
Note that i is an integer
Why: So no averaging is needed.
\[ \text{take the } 21 s t\text{ value} \]
Read the 21st ordered value
Why: Counting along the ordered list.
\[ 64 \]
Figure (svg): The solution to Worked example an index that is a whole number shown as a ladder of expressions, one row per legal move
\[ i = \left(\frac{70}{100}\right)(29+1) = 21 \;\longrightarrow\; 64 \text{ years} \]
Verify: confirm the answer sits sensibly in the data
Why: Twenty of the twenty-nine ages lie below 64 and eight lie above it, so about seventy percent of the data sit at or below the answer, which is what the seventieth percentile promises. Checking that roughly the right fraction falls below is the fastest way to catch an index computed with n instead of n plus 1, or a miscount along the ordered list — both of which would land the answer a position or two away.
OpenStax Introductory Statistics 2e, §2.3 Measures of the Location of the Data §2.3, p. 90
Faded example
The 20th percentile of the 29 actors' ages.
Fill in the blanks
i = \left(\frac306\right)(29 + 1) = (0.20)(___) = ___
Why: The index is 6, a whole number, so the 20th percentile is the sixth ordered age, which is 27. When the index comes out whole the work is finished in one step — no averaging, no rounding.
Worked example
Example 2.16b. The same ages, a different percentile.
\[ \text{Find the } 83\text{rd percentile of the same } 29 \text{ ages.} \]
Compute the index
Why: Eighty-three hundredths of thirty.
\[ i = (0.83) (30) = 24.9 \]
Note that i is not an integer
Why: So two values will be averaged.
Identify the two positions
Why: The 24th and the 25th.
\[ 24\text{ and } 25 \]
Average the values there
Why: The ages 71 and 72.
\[ \frac{71 + 72}{2} = 71.5 \]
Figure (svg): The solution to Worked example an index that is not a whole number shown as a ladder of expressions, one row per legal move
\[ i = 24.9 \;\longrightarrow\; \frac{71 + 72}{2} = 71.5 \]
Verify: confirm the answer lies between the two neighbours
Why: The result 71.5 sits between the 24th value of 71 and the 25th of 72, which it must: averaging two numbers always lands between them. An answer outside that interval would mean the wrong positions were read, and an answer equal to one of them would mean the index was rounded rather than averaged. Note also that the percentile is not one of the observed ages, which is entirely normal — the book says at the outset that a percentile may or may not be part of the data.
OpenStax Introductory Statistics 2e, §2.3 Measures of the Location of the Data §2.3, p. 90
Trap
\[ i = 24.9 \]
Round the index to the nearest whole number, 25
Why: Rounding is what one usually does with a decimal that has to become a position.
\[ \text{the } 25\text{th value} = 72 \quad \text{(wrong)} \]
The index landed between two positions precisely because the percentile falls between two observations, and rounding discards that information.
\[ \text{round both ways: } 24 \text{ and } 25, \text{ then average } 71 \text{ and } 72 \]
Take both neighbouring positions and average their values
Why: A non-integer index is an instruction to interpolate, not to choose.
The book's wording is explicit: round i up AND round i down, then average the two data values in those positions. Rounding to nearest would also be systematically biased — an index of 24.1 and one of 24.9 would give different answers under averaging but the same answer under rounding, even though they represent quite different positions between the observations.
Prediction
Commit before reasoning.
Predict first
The index formula uses n plus 1 rather than n. What does that accomplish?
Correct: It spaces the percentiles evenly between the observations.
Why: With 29 values, using n plus 1 puts the 50th percentile at position 15, which is exactly the middle observation. Using n would give 14.5 and place the median between two values when a genuine middle one exists. The formula does not make indices whole — the 83rd percentile gave 24.9 — and the choice is a real one rather than arbitrary, which is why different textbooks and calculators disagree slightly on small data sets.
Two truths and a lie
All three concern the index formula.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The book's instruction is to round up AND down and average the two values at those positions. Rounding to nearest would throw away the information that the percentile lies between two observations, and would treat an index of 24.1 exactly like one of 24.9.
Estimation
A data set has 49 values and you want the 50th percentile.
Predict first
What index does the formula give?
Correct: 25.
Why: The index is 0.50 times 50, which is 25 — the middle position of 49 ordered values, with 24 below it and 24 above. That is exactly the median, as it should be, and it is a clean illustration of why the formula uses n plus 1: with n it would have given 24.5 and required averaging two values when a genuine middle observation exists.
Section
Section 5
Concept
The formula runs the other way as well. Order the data, let x be the number of values below the one whose percentile you want, let y be the number of values equal to it, and compute x plus half of y, divided by n, times a hundred. Round to the nearest integer.
the percentile of a value — Order the data. Let x be the number of values counting from the bottom up to but not including the value in question, y the number of values equal to it, and n the total. The percentile is (x + 0.5y)/n times 100, rounded to the nearest integer.
\[ \text{percentile} = \left(\frac{x + 0.5y}{n}\right)(100) \]
The half in the formula is doing real work. A value ties with itself, so counting all the ties as below would overstate its position and counting none of them would understate it; taking half splits the difference. On a data set with no ties the y is always 1 and the formula adds half a position, which is what makes the median of an odd-sized set come out at the 50th percentile exactly.
Figure (svg): Two columns contrasting finding the value at a given percentile with finding the percentile of a given value
OpenStax Introductory Statistics 2e, §2.3 Measures of the Location of the Data §2.3, p. 92 — the formula for the percentile of a value
Picture it
The two formulas, and the question that chooses between them.
Figure (svg): Two columns contrasting finding the value at a given percentile with finding the percentile of a given value
The units of the answer are the safety net. A percentile must be a number between 0 and 100 with no units; a value at a percentile carries the data's units — years, dollars, inches. If a question about a student's standing produces an answer of 64 years, the wrong formula was used, and the mistake is visible without checking any arithmetic.
Worked example
Example 2.18a. The same twenty-nine ages.
\[ \text{Find the percentile for the age } 58. \]
Count the values strictly below
Why: Counting from the bottom up to but not including 58.
\[ x = 18 \]
Count the values equal to it
Why: There is one 58 in the list.
\[ y = 1 \]
Apply the formula
Why: x plus half of y, over n, times a hundred.
\[ (18 + 0.5) / 29 x 100 \]
Round to the nearest integer
Why: The formula gives 63.79.
\[ 64 \]
Figure (svg): The solution to Worked example the percentile of an age shown as a ladder of expressions, one row per legal move
\[ \left(\frac{18 + 0.5(1)}{29}\right)(100) = 63.79 \;\longrightarrow\; 64\text{th percentile} \]
Verify: confirm the two formulas are consistent
Why: Running the index formula in the other direction as a check: the 64th percentile has index 0.64 times 30, which is 19.2, giving the average of the 19th and 20th ages — 58 and 62 — or 60. That is not exactly 58, and the discrepancy is expected: the two formulas are inverse in intent but not exact inverses, because one interpolates between positions and the other counts ties. On a large data set they converge; on twenty-nine values they can differ by a position, which is worth knowing before treating one as a check on the other.
OpenStax Introductory Statistics 2e, §2.3 Measures of the Location of the Data §2.3, p. 92
Faded example
In a set of 40 values, 12 lie strictly below the value 19 and 4 values equal 19.
Fill in the blanks
\text4 = \left(\frac35})}___\right)(100) = ___
Why: Twelve plus half of four is fourteen, and fourteen over forty is 0.35, so the value is at the 35th percentile. With four ties the half-counting matters: treating all four as below would give 40, and none of them would give 30, so the convention moves the answer by five percentage points.
Worked example
Example 2.18b. The same method, a low value.
\[ \text{Find the percentile for the age } 25. \]
Count the values strictly below
Why: The ages 18, 21 and 22.
\[ x = 3 \]
Count the values equal to it
Why: One age of 25.
\[ y = 1 \]
Apply the formula
Why: Three and a half over twenty-nine.
\[ (3 + 0.5) / 29 x 100 \]
Round
Why: The formula gives 12.07.
\[ 12 \]
Figure (svg): The solution to Worked example a value near the bottom shown as a ladder of expressions, one row per legal move
\[ \left(\frac{3 + 0.5}{29}\right)(100) = 12.07 \;\longrightarrow\; 12\text{th percentile} \]
Verify: confirm the answer is a sensible position
Why: The age 25 is the fourth smallest of twenty-nine, and four out of twenty-nine is about fourteen percent, so a percentile of 12 is in the right region. A useful sanity check on any percentile-of-a-value answer is to compute the value's rank as a rough fraction of n: the formula should land near it, a little below because half a tie is subtracted. An answer far from that fraction means x was counted wrongly, usually by including the value itself.
OpenStax Introductory Statistics 2e, §2.3 Measures of the Location of the Data §2.3, p. 92
Trap
\[ \text{18 ages are below 58, and 58 itself is in the list, so } x = 19 \]
Include the value in the count of what lies below it
Why: It is in the data, so the student counts it.
\[ \left(\frac{19 + 0.5}{29}\right)(100) = 67.2 \quad \text{(wrong)} \]
The value has now been counted twice: once in x and again in the half of y that the formula adds.
\[ x = 18 \text{ (strictly below)}, \quad y = 1 \text{ (equal)} \]
Count up to but NOT including the value, then handle ties separately with y
Why: The book's wording is explicit about the boundary.
The two counts have disjoint jobs: x is everything strictly less and y is everything exactly equal, so no observation belongs to both. The formula then credits the value with half of its own tie group. Getting this wrong inflates every percentile, and it inflates them most for values with many ties — which are exactly the cases where the careful treatment matters.
Sorting
Ask what you were given and what is wanted.
Sort into buckets
Sort each question.
Item (e) is the one that hides its shape: 'the top 10 percent' means the 90th percentile, so a percentile is being given and a score is wanted. Translating the everyday phrasing into a percentile first is what makes the direction obvious.
Prediction
Commit before reasoning.
Predict first
Why does the formula add half of y rather than all of y or none of it?
Correct: To credit the value with half its own tie group.
Why: Counting all the ties as below would put a value above others identical to it; counting none would put it below them. Neither is defensible, so the convention splits the group. The consequence is real rather than arbitrary: on the example with four ties it moves the answer by five percentage points, and with many ties it would move it a great deal more.
Edge cases
Apply the counting formula to the largest observation in a set of 29 with no ties.
Discussion prompt
What percentile does the formula give, and why is it not 100?
Hint: Count how many lie strictly below the maximum.
Answer:
Twenty-eight values lie below it and one equals it, so the formula gives 28.5 over 29 times 100, which is about 98.3, rounding to the 98th percentile.
It is not 100 because the value is credited with only half of its own tie group. Reaching the 100th percentile would require the value to be above everything including itself, which nothing can be.
This is a feature rather than a defect: it means the maximum of a small data set is not reported as beating a hundred percent of the data, which would be false. It also shows the two formulas are not exact inverses, since the index formula's 100th percentile does return the maximum.
Comparison
Fill the blanks. The middle column is what each one is computed from.
Comparison matrix
| Measure | How it is found | What it tells you |
|---|---|---|
| Median (Q2) | the middle of the ordered data | half the values are at or below it |
| Q1 | the median of the lower half | a quarter of the values are at or below it |
| Q3 | the median of the upper half | three quarters are at or below it |
| IQR | Q3 minus Q1 | the spread of the middle fifty percent |
| kth percentile | the value at index (k/100)(n+1) | about k percent of the values are at or below it |
Every row is about POSITION rather than size, which is what makes these measures comparable across data sets in different units. The IQR is the exception that points forward: it is a measure of spread built out of two measures of location, and section 2.7 takes up spread properly.
Pattern
Five steps, and the first is the one people skip.
Check the units of your answer against the question. A percentile is a bare number between 0 and 100; a value at a percentile carries the data's units. Answering a question about standing with a number of years means the formulas were used the wrong way round.
OpenStax Introductory Business Statistics 2e, §2.2 Measures of the Location of the Data §2.2 Measures of the Location of the Data
Check
The quartiles of an odd-sized set.
Check your understanding
A data set has 13 ordered values. Which values form the lower half when finding Q1?
Answer: A
Why: With an odd count the median is the divider and belongs to neither half, so the lower half is the six values below it and Q1 is their median.
Check
The outlier fences.
Check your understanding
A data set has Q1 = 20 and Q3 = 60. Where is the upper fence?
Answer: A
Why: The IQR is 60 minus 20, which is 40; one and a half times 40 is 60; and 60 plus 60 gives an upper fence of 120.
Check
Which direction is the question running?
Check your understanding
In a set of 29 ordered values, you are asked for the 60th percentile. What is the index?
Answer: A
Why: The index is 60 over 100 times 29 plus 1, which is 0.60 times 30, giving exactly 18 — a whole number, so the answer is the 18th ordered value.
Real world
A paediatrician tells two sets of parents that their child is at the 25th percentile for height. The first child is 3 years old and the second is 14. One set of parents is alarmed and asks whether the child is too short; the other asks what height that corresponds to.
Discussion prompt
Answer both questions, and say what a single percentile can and cannot tell these parents.
Hint: Ask what the percentile is computed against, and what it says about the individual child.
Answer:
What it means in both cases is the same: about a quarter of children of that age and sex are the same height or shorter, and about three quarters are taller. It is a position within a reference population, exactly as an exam percentile is a position within the candidates.
Is the child too short? The percentile alone cannot say. A quarter of all healthy children are at or below the 25th percentile by definition — that is what a percentile is — so being there is unremarkable on its own. What paediatricians actually watch is the child's percentile over TIME: a child steady at the 25th is growing normally, and a child who falls from the 60th to the 25th in a year is the one worth investigating.
What height is it? Only the reference distribution can say, and it differs completely between the two children: the 25th percentile might be about 92 centimetres for the three-year-old and about 158 for the fourteen-year-old. The same percentile, wildly different values — which is the point the book makes about the 90th percentile not meaning 90 percent.
\[ \text{same percentile} \;\ne\; \text{same value} \qquad \text{same value} \;\ne\; \text{same percentile} \]
The general lesson is that a measure of location is meaningless without its reference set, and that the interesting quantity is often the CHANGE in percentile rather than its level. This is why growth charts plot a line rather than printing a number, and it is the same reason a time series was worth drawing in section 2.2.
Commit first
Answer, then rate your confidence honestly.
Predict first
Why is the outlier rule built from the quartiles rather than from the mean and standard deviation?
Correct: Because an extreme value cannot move the quartiles.
\[ \text{an outlier moves the mean and the standard deviation, but not } Q_1 \text{ or } Q_3 \]
Why: The quartiles depend on the middle of each half, where a single distant value has no influence, so adding a price of fifty million to the home data would leave Q1, Q3 and both fences exactly where they were. A detector built from the mean and standard deviation would have its threshold inflated by the very observation it is meant to catch, so an extreme value partly conceals itself. Ease of computation is not the reason, and the mean is perfectly well defined here.
Explain it
They think being in the 30th percentile means getting 30 percent on the test.
Discussion prompt
In three sentences or fewer, correct them with an example that makes the difference undeniable.
Hint: Invent an exam where nearly everyone scores badly.
Answer:
Imagine an exam so hard that the top mark in the year is 40 percent and most people score around 20. A student who gets 25 percent has beaten more than half the year and is somewhere near the 60th percentile, while scoring a quarter of the marks.
The percentile counts PEOPLE below you; the mark counts QUESTIONS you answered. They are measured in completely different things and the same student can be low on one and high on the other.
Say the sentence in full every time — 'thirty percent of the scores were the same or lower' — and the confusion becomes impossible.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For odd counts, the median goes into neither half. For the fences, always write the four steps separately and remember the 1.5 multiplies the IQR, never a quartile. For choosing a formula, ask what you were GIVEN and check the units of your answer. For a non-integer index, round up AND down and average the two values rather than rounding to the nearest. Do five examples of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top, write these fourteen values in order: 1, 11.5, 6, 7.2, 4, 8, 9, 10, 6.8, 8.3, 2, 2, 10, 1. Mark the median, Q1 and Q3 on the ordered list with three vertical cuts, write each value beside its cut, and note which of the three are actual data values and which are not. Below, take the thirteen home prices 114950, 158000, 230500, 387000, 389950, 479000, 488800, 529000, 575000, 639000, 659000, 1095000, 5500000 and work the outlier calculation in four separate lines: the IQR, one and a half times it, the lower fence, the upper fence. Circle any price beyond a fence and write one sentence on what you would do about it. To the right, make a two-column table headed 'given k, find a value' and 'given a value, find k', and under each write the formula, one worked example, and the units of the answer. Then work both directions on the 29 actors' ages: find the 70th percentile using the index formula, and find the percentile of the age 58 using the counting formula. At the bottom, write one sentence explaining why the outlier rule uses the quartiles rather than the mean.
Check your fences by asking whether the second largest price is flagged. It should not be — 1,095,000 sits just inside the upper fence of 1,159,375. If your rule flags it, you multiplied 1.5 by a quartile instead of by the IQR.
Recap
Five things, and the last two are one idea run in opposite directions.
| If you see | Then |
|---|---|
| A request for Q1 | Take the median of the lower half |
| An odd number of values | The median belongs to neither half |
| A question about outliers | IQR, then fences at 1.5 IQR beyond each quartile |
| A value beyond a fence | Potential outlier: investigate, do not delete |
| A percentile given, a value wanted | Index formula: (k/100)(n+1) |
| A non-integer index | Round up AND down, then average the two values |
| A value given, a percentile wanted | Count x below and y equal: (x + 0.5y)/n times 100 |
Section 2.4 draws the five numbers this lesson computed. A box plot shows the minimum, Q1, the median, Q3 and the maximum in one picture, with outliers marked separately — which is why the fences had to come first.
OpenStax Introductory Statistics 2e, §2.3 Measures of the Location of the Data §2.3, pp. 86-93 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.