Printable · GCSE Higher · ages 14-16
Averages, spread and comparing distributions worksheet — GCSE Higher
Fifteen questions on "averages, spread and comparing distributions" — DfE statement S4. Print it, or print three versions so neighbours cannot copy by letter; the key gives the letter for each version.
part Higher
Averages, spread and comparing distributions worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- 1.A shop compares customer waiting times, in minutes, at two branches over one week. Branch A had a median of 12 minutes and an interquartile range of 5 minutes. Branch B had a median of 9 minutes and an interquartile range of 11 minutes. Compare the waiting times at the two branches.
- 2.Box plots summarise house prices, in thousands of pounds, in two towns. Town A: minimum 120, lower quartile 180, median 220, upper quartile 310, maximum 520. Town B: minimum 95, lower quartile 175, median 235, upper quartile 260, maximum 340. A buyer wants the town where house prices are typically higher but more consistent. Which statement is correct?
- 3.Priya travels to work by one of two routes and records her journey times, in minutes, over several weeks. Route 1 has a median of 34 minutes and an interquartile range of 22 minutes. Route 2 has a median of 41 minutes and an interquartile range of 6 minutes. Priya wants the more reliable route for getting to an important meeting on time. Which route should she choose, and why?
- 4.In a spelling test the 20 pupils in Group A had a mean mark of 80, and the 30 pupils in Group B had a mean mark of 70. Work out the mean mark of all 50 pupils.
- 5.The mean of the four numbers 10, 15, 20 and x is 18. Work out the value of x.
- 6.Box plots summarise the test scores of two classes. Class X: minimum 20, lower quartile 45, median 60, upper quartile 70, maximum 95. Class Y: minimum 35, lower quartile 50, median 58, upper quartile 65, maximum 80. Which statement about the two classes is correct?
- 7.The mean of three numbers is 50. A fourth number, 100, is added to the set. Work out the mean of the four numbers.
- 8.The mean of 5 numbers is 8. Work out the total of the 5 numbers.
- 9.Grace asked 12 children how many brothers and sisters they have. Her results, in order, were 0, 0, 1, 1, 1, 1, 2, 2, 3, 3, 4, 6. Work out the median number of brothers and sisters.
- 10.A garden centre records the heights, in cm, of eleven seedlings. In order, the heights are 5, 9, x, 17, 20, 24, 28, 31, 35, 40, 44, where x is unknown. The interquartile range of the eleven heights is 19 cm. Work out the value of x.
- 11.The numbers 2, 4, 6, 8, 10 and 12 have a mean of 7 and a median of 7. The value 100 is now added to the list. Which measure is changed more by adding 100, and why?
- 12.Two classes sat the same test. The 30 pupils in Class A had a mean mark of 72. The 20 pupils in Class B had a mean mark of 82. Work out the mean mark of all 50 pupils.
- 13.The mean of the numbers x, 20 and 30 is equal to the mean of the numbers 15 and 25. Work out the value of x.
- 14.Harry counted the coins he found on each of six days: 5, 7, 8, 9, 11, 30. Write down the outlier.
- 15.The mean mass of four parcels is 17 kg. Three of the parcels have masses 12 kg, 16 kg and 18 kg. Work out the mass of the fourth parcel.
Answer key
- (c) Branch A waits longer, and Branch A is more consistent — Method: compare the two branches using a measure of location (the median) for who waits longer, and a measure of spread (the interquartile range) for who is more consistent — a smaller interquartile range means more consistent. Working: Branch A's median, 12 minutes, is higher than Branch B's, 9 minutes, so Branch A's customers wait longer on average. Branch A's interquartile range, 5 minutes, is smaller than Branch B's, 11 minutes, so Branch A's waiting times vary less. Answer: Branch A waits longer, and Branch A is also the more consistent of the two. Watch that each half of the comparison uses the right statistic and reads it correctly: swapping both readings gives Branch B the longer wait and the greater consistency, when neither is true; keeping the median comparison right but reading a larger interquartile range as 'more consistent' has the direction of spread backwards; and swapping only the median comparison keeps the correct branch for consistency but gives the wrong branch the longer wait.
- (b) Town B — higher median and smaller IQR — Method: 'higher and more consistent' needs two comparisons — the median for typical price, and the interquartile range for spread, with a smaller interquartile range meaning more consistent. Working: Town B's median, £235,000, is higher than Town A's, £220,000. Town A's interquartile range is 310 − 180 = 130 and Town B's is 260 − 175 = 85, so Town B's interquartile range is the smaller of the two. Answer: Town B has both the higher median and the smaller interquartile range, so it is the town with higher, more consistent prices. Watch which combination of median and interquartile range each statement claims, and for which town: claiming Town A has the higher median and the smaller interquartile range gets both comparisons wrong, since Town B leads on both; claiming Town A has the higher median (still wrong) but the larger interquartile range at least reads the spread correctly, without it rescuing the false median claim; and claiming Town B has the higher median (correct) but the larger interquartile range misreads the spread — Town B's interquartile range is the smaller of the two, not the larger.
- (c) Route 2, because its interquartile range is smaller — Method: for a journey where turning up on time matters, what matters is not the typical (median) time but how predictable it is — a smaller interquartile range means the middle half of journeys cluster closer together. Working: Route 1's median, 34 minutes, is in fact lower than Route 2's, 41 minutes, so Route 1 is faster on average; but Route 1's interquartile range, 22 minutes, is far larger than Route 2's, 6 minutes, so Route 1's times are much less predictable. Answer: Priya should choose Route 2, because its interquartile range is smaller, even though it is slower on average. Watch which statistic answers the question actually asked: Route 1 does not have the smaller interquartile range, Route 2 does, so picking Route 1 for that reason misreads the table; Route 1's median genuinely is the lower one, but a lower median answers 'which is faster', not 'which is more reliable'; and Route 2's median is not the lower one, so that claim about Route 2 is simply false.
- (b) 74 marks — Method: the two groups are different sizes, so their means cannot simply be averaged — rebuild each group's total mark, add the totals and divide by all 50 pupils. Working: Group A scored 20 × 80 = 1600 marks and Group B scored 30 × 70 = 2100 marks, giving 1600 + 2100 = 3700 marks altogether, so the overall mean is 3700 ÷ 50 = 74 marks. Answer: 74 marks, which sits nearer to 70 than to 80 because the larger group scored 70. The distractors: 75 marks comes from averaging the two group means, (80 + 70) ÷ 2, as though the groups were the same size; 76 marks comes from attaching each mean to the other group's size, (20 × 70 + 30 × 80) ÷ 50; 150 marks comes from adding the two means together and never dividing at all.
- (a) 27 — Method: turn the mean into a total using total = mean × number of values, then subtract the numbers that are already known. Working: four numbers with a mean of 18 have a total of 18 × 4 = 72; the three known numbers give 10 + 15 + 20 = 45; so x = 72 − 45 = 27. Answer: 27, and checking, (10 + 15 + 20 + 27) ÷ 4 = 72 ÷ 4 = 18. The distractors: 72 comes from stopping at the total the four numbers must reach and never subtracting the known three; 18 comes from assuming the missing number must equal the mean; 45 comes from stopping at the total of the three known numbers.
- (c) Class X has the higher median and the wider spread — Method: compare the two box plots statistic by statistic — median for location, and the interquartile range for spread — checking the true value of each rather than assuming a pattern. Working: Class X's median is 60 and Class Y's is 58, so Class X's median is the higher one. Class X's interquartile range is 70 − 45 = 25 and Class Y's is 65 − 50 = 15 (and the ranges follow the same order: 95 − 20 = 75 against 80 − 35 = 45), so Class X also has the wider spread. Answer: Class X has both the higher median and the wider spread. Watch that each half of a compound statement is checked separately: claiming Class Y has the higher median and the wider spread gets both comparisons backwards; claiming Class X has the higher median but the narrower spread keeps the median right while reading the spread the wrong way round; and claiming Class Y has the higher median but the narrower spread swaps the median comparison while getting the spread right.
- (a) 62.5 — Method: a mean cannot be averaged with a new value — rebuild the total, add the new value to it, then divide by the new count. Working: three numbers with a mean of 50 have a total of 50 × 3 = 150; adding 100 makes the total 150 + 100 = 250; there are now 4 numbers, so the new mean is 250 ÷ 4 = 62.5. Answer: 62.5. The distractors: 75 comes from averaging the old mean with the new value, (50 + 100) ÷ 2, which ignores that three numbers pull against one; 50 comes from assuming an extra value leaves the mean unchanged; 37.5 comes from dividing the old total of 150 by the new count of 4, adding the new value to the count but not to the total.
- (a) 40 — Method: the mean is the total divided by how many values there are, so rearranging gives total = mean × number of values. Working: the mean is 8 and there are 5 numbers, so the total is 8 × 5 = 40. Answer: 40, and checking, 40 ÷ 5 = 8, which is the mean given. The distractors: 13 comes from adding the mean and the count, 8 + 5, instead of multiplying them; 1.6 comes from dividing the mean by the count, 8 ÷ 5, which reverses the relationship; 8 comes from quoting the mean itself as the total, which is only true when there is a single number.
- (a) 1.5 — Method: with an even number of values the median is the mean of the two middle values, which for 12 values are the 6th and the 7th once the data are in order. Working: the results are already in order, and 12 ÷ 2 = 6, so the middle pair are the 6th value, 1, and the 7th value, 2; the median is (1 + 2) ÷ 2 = 1.5. Answer: 1.5 brothers and sisters. The distractors: 1 comes from reading the 6th value and stopping there instead of averaging the middle pair; 2 comes from working out the mean, 24 ÷ 12, instead of the median; 6 comes from working out the range, 6 − 0, which measures spread rather than centre.
- (d) 16 — Method: rearrange interquartile range = upper quartile − lower quartile to make the lower quartile the subject: lower quartile = upper quartile − interquartile range, then check the answer sits in the right place in the list. Working: there are 11 values, so 11 + 1 = 12; the upper quartile sits at position 3 × 12 ÷ 4 = 9, which is 35, and x sits at position 12 ÷ 4 = 3, which is the lower quartile. So x = 35 − 19 = 16, and 16 does sit between the 2nd value, 9, and the 4th value, 17, as it should. Answer: x = 16. Watch how you rearrange and where you count to: adding instead of subtracting, 35 + 19 = 54, treats the interquartile range as something added on rather than a gap taken away; subtracting in the wrong order, 19 − 35 = −16, finds the right two numbers but flips the sign; and counting to the 8th value instead of the 9th treats 31 as the upper quartile, giving 31 − 19 = 12, one position short of where the upper quartile actually sits.
- (c) The mean, because every value counts towards it, so 100 pulls it from 7 up to about 20.3. — Method: work each measure out before the extra value is added and again afterwards, then compare the size of the two changes. Working: before, the six values total 42, so the mean is 42 ÷ 6 = 7, and the middle pair 6 and 8 give a median of (6 + 8) ÷ 2 = 7; after, the seven values total 142, so the mean is 142 ÷ 7 = 20.29 to 2 decimal places, while the median is now the 4th of the seven ordered values, which is 8; the mean has moved by about 13.3 and the median by 1. Answer: the mean, because every value counts towards it, so 100 pulls it from 7 up to about 20.3 — this is why the median is often preferred when a data set contains an outlier. The distractors: the reply that the mean rises by 100 adds the extra value to the mean instead of adding it to the total; the reply that the median moves to 12 takes the largest of the original values as the new middle instead of counting to the 4th of the seven values; the reply about even and odd counts quotes a rule that does not exist, since the median moved because a very large value was added, not because the count of values changed.
- (c) 76 marks — Method: a mean of means only works when the groups are the same size, so rebuild each class's total mark, add the totals and divide by the number of pupils altogether. Working: Class A scored 30 × 72 = 2160 marks and Class B scored 20 × 82 = 1640 marks, giving 2160 + 1640 = 3800 marks between 50 pupils, so the overall mean is 3800 ÷ 50 = 76 marks. Answer: 76 marks. The distractors: 77 marks comes from averaging the two class means, (72 + 82) ÷ 2, which ignores the different class sizes; 78 marks comes from attaching each mean to the other class's size, (30 × 82 + 20 × 72) ÷ 50; 3800 marks comes from stopping at the combined total and never dividing by 50.
- (c) 10 — Method: work out the mean that can be found straight away, then use total = mean × number of values on the group of three to find the missing number. Working: the mean of 15 and 25 is (15 + 25) ÷ 2 = 40 ÷ 2 = 20, so the group of three must also have a mean of 20; three numbers with a mean of 20 have a total of 20 × 3 = 60, and 20 + 30 = 50 of that total is already accounted for, so x = 60 − 50 = 10. Answer: 10, and checking, (10 + 20 + 30) ÷ 3 = 20. The distractors: 20 comes from working out the mean the two groups share and writing that down as x; −10 comes from dividing the group of three by 2 instead of by 3, which gives x + 50 = 40; 70 comes from reading the total 15 + 25 = 40 as the mean of the pair, which sets the target total at 120 and leaves x = 70.
- (a) 30 — Method: an outlier is a value that lies far away from the pattern set by the rest of the data, so compare each value with the group the others form. Working: five of the counts, 5, 7, 8, 9 and 11, lie within 6 of one another and the steps between them are 2, 1, 1 and 2; the remaining count of 30 is 19 above the nearest of them, so it is the value that does not belong to the pattern. Answer: 30. The distractors: 5 comes from picking the smallest value, on the idea that the odd one out must be at the bottom of the list; 11 comes from ordering the data and stopping one value short, taking the largest of the counts that sit close together; 8.5 comes from working out the median, (8 + 9) ÷ 2, and giving a measure of centre where a value standing apart was asked for.
- (c) 22 kg — Method: multiply the mean by the number of parcels to rebuild the total mass, then subtract the masses that are known. Working: four parcels with a mean mass of 17 kg have a total mass of 17 × 4 = 68 kg; the three known parcels total 12 + 16 + 18 = 46 kg; so the fourth parcel has mass 68 − 46 = 22 kg. Answer: 22 kg. The distractors: 68 kg comes from stopping at the total mass of all four parcels; 17 kg comes from assuming the missing parcel must have the mean mass; 5 kg comes from multiplying the mean by 3, the number of parcels whose mass is given, leaving 51 − 46 = 5.
Build your own mix at the worksheet builder.