Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Answer key: Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- (d) 1100 — Method: to combine two samples of different sizes, add the faulty counts together and add the sample sizes together before scaling up, rather than treating the two samples separately. Working: the combined sample found 34 + 21 = 55 scratched cases out of 100 + 50 = 150 cases checked, a proportion of 55 ÷ 150. Applying that proportion to the week's production of 3,000 gives an estimate of 55 ÷ 150 × 3000 = 1100 scratched cases. Averaging the two shifts' proportions instead of combining their totals, (34 ÷ 100 + 21 ÷ 50) ÷ 2 = 0.38, gives 0.38 × 3000 = 1140 — this treats the two samples as equally weighted even though Shift A checked twice as many cases as Shift B. Using only Shift A's sample, 34 ÷ 100 × 3000 = 1020, ignores Shift B's cases completely. Using only Shift B's sample, 21 ÷ 50 × 3000 = 1260, ignores Shift A's cases completely. When two samples are different sizes, combine their totals before finding the proportion — do not average the two proportions, and do not use only one shift's sample.
- (d) 16 — Method: rearrange interquartile range = upper quartile − lower quartile to make the lower quartile the subject: lower quartile = upper quartile − interquartile range, then check the answer sits in the right place in the list. Working: there are 11 values, so 11 + 1 = 12; the upper quartile sits at position 3 × 12 ÷ 4 = 9, which is 35, and x sits at position 12 ÷ 4 = 3, which is the lower quartile. So x = 35 − 19 = 16, and 16 does sit between the 2nd value, 9, and the 4th value, 17, as it should. Answer: x = 16. Watch how you rearrange and where you count to: adding instead of subtracting, 35 + 19 = 54, treats the interquartile range as something added on rather than a gap taken away; subtracting in the wrong order, 19 − 35 = −16, finds the right two numbers but flips the sign; and counting to the 8th value instead of the 9th treats 31 as the upper quartile, giving 31 − 19 = 12, one position short of where the upper quartile actually sits.
- (c) 75 — Method: the number in a class is the area of its bar, frequency density × class width, so work out the frequency of each class that lies at or above 10 minutes and add them. Working: the class 10 ≤ t < 25 is 15 minutes wide with a frequency density of 3.2, giving 3.2 × 15 = 48 members; the class 25 ≤ t < 55 is 30 minutes wide with a frequency density of 0.9, giving 0.9 × 30 = 27 members; the total charged is 48 + 27 = 75. Answer: 75 members pay the extra charge. The distractors: 4.1 comes from adding the two frequency densities, 3.2 + 0.9, as though each height were a count; 93 comes from including the class 0 ≤ t < 10 as well, 1.8 × 10 = 18 added to 48 and 27, which charges every member; 27 comes from using only the class 25 ≤ t < 55 and forgetting that 10 ≤ t < 25 is also at or above 10 minutes.
- (c) 72° — The angle is 12 ÷ 60 × 360 = 72°. Choosing 150° divides football's frequency of 25 instead of badminton's 12: 25 ÷ 60 × 360 = 150. Choosing 20° finds badminton as a percentage of the members, 12 ÷ 60 × 100 = 20, rather than an angle in degrees. Choosing 90° uses 48, the total of the other three activities, as the total instead of the full 60 members: 12 ÷ 48 × 360 = 90.
- (c) The class with times from 10 up to 20 — Method: to find the median class from a histogram, first turn each bar's frequency density into a frequency using density × class width, build up the cumulative frequency, and find the first class whose cumulative frequency reaches or passes n ÷ 2. Working: the four classes have widths 10, 10, 20 and 20, so their frequencies are 5 × 10 = 50, 2 × 10 = 20, 1.5 × 20 = 30 and 1 × 20 = 20, which add to the 120 visitors stated. The median sits at position 120 ÷ 2 = 60. The cumulative frequency is 50 after the first class and 50 + 20 = 70 after the second, so the 60th visitor is reached during the second class. Answer: the median lies in the class 10 ≤ t < 20. Watch which class each shortcut lands on: the tallest bar belongs to the first class, with the highest frequency density, 5 — but the tallest bar shows where visitors are packed most densely, not where the middle visitor falls, and picking it lands one class too early, at 0 ≤ t < 10; taking half of the total TIME span instead of half of the total NUMBER of visitors, 60 minutes ÷ 2 = 30 minutes, lands in the class 20 ≤ t < 40, confusing a value on the horizontal axis with a position in the data; and using the full 120 visitors as the target position, rather than 120 ÷ 2 = 60, reaches all the way to the last class, 40 ≤ t < 60, treating the whole data set's size as though it were the position of a single middle value.
- (c) 70 — Method: estimate the median from the cumulative frequency table by interpolation: find its position, n ÷ 2, locate the class it falls in, then add the fraction of the way through that class (adjusted for the cumulative frequency reached before it) to the class's lower boundary. Working: there are 180 riders, so the median is at position 180 ÷ 2 = 90. Before the class 60 ≤ d < 80 the cumulative frequency is 60, and by the end of it, 120, so the 90th rider falls in this class; its frequency is 120 − 60 = 60 and its width is 80 − 60 = 20. The extra distance needed into the class is 90 − 60 = 30, and 30 ÷ 60 × 20 = 10, so the median is 60 + 10 = 70. Answer: the estimated median distance is 70 km. Watch which numbers the interpolation actually uses: reading off just the class's lower boundary, 60, ignores how far into the class the 90th rider falls; using the target position, 90, as the extra distance instead of subtracting the 60 riders already counted before the class gives 90 ÷ 60 × 20 = 30, so 60 + 30 = 90, overshooting by treating the whole position as if none of it had already been counted; and using the total number of riders, 180, instead of half of it as the target position lands in the very last class, giving an estimate of 130 km — further than any rider is known to have ridden by that point in the table.
- (c) No — median £505 at A vs £510 at B. — Branch A's seven wages in order are £480, £495, £500, £505, £510, £515 and £1,200, so the median, the 4th value, is £505. Branch B's in order are £480, £490, £500, £510, £520, £530 and £540, so the median is £510. Since £505 is lower than £510, the median wage is not higher at Branch A, so the claim is not fairly supported. Choosing 'Yes — mean £600.71 at A vs £510 at B' uses the mean: 480 + 495 + 500 + 505 + 510 + 515 + 1200 = 4205, and 4205 ÷ 7 = 600.71, a figure pulled upward by the £1,200 outlier that does not represent a typical wage. Choosing 'Yes — median £515 at A vs £510 at B' miscounts the middle position, taking the 6th wage, £515, instead of the correct 4th value, £505. Choosing 'Yes — highest wage £1,200 at A vs £540 at B' compares the highest wage at each branch rather than a measure of the typical, or average, wage.
- (b) 120, unreliable — x = 25 is outside 1 to 12 — The line of best fit is y = 4x + 20. 4 × 25 = 100, and 100 + 20 = 120, so the estimate is y = 120. But x = 25 lies far outside the plotted range of 1 to 12 years, so this is an extrapolation, and the estimate is not reliable. Reaching 100 instead of 120 comes from 4 × 25 = 100 with the intercept of 20 left out — still correctly flagged as unreliable, but the wrong value. Calling the estimate reliable simply because it was calculated correctly, giving 120, wrongly assumes that a correct calculation is automatically trustworthy, ignoring that x = 25 lies far beyond the data actually collected. Reaching 68, from 4 × 12 = 48 and 48 + 20 = 68, substitutes x = 12, the top of the plotted range, instead of the assistant's actual x = 25, and wrongly calls that reliable because 12 lies inside the range.
- (c) Neither causes the other; sunshine links both. — Both ice cream sales and sunburn cases tend to rise on hot, sunny days, so the amount of sunshine is a third factor linked to both — neither variable causes the other. Saying ice cream sales cause the sunburn assumes a causal link in one direction that the correlation alone cannot establish. Saying sunburn cases cause the ice cream sales assumes the reverse causal link, which is no more justified. Saying a strong correlation always means causation is the general error this question is testing: correlation, however strong, does not by itself prove that one variable causes the other.
- (d) £15.50 — 2 × 6 = 12, and 12 + 3.50 = 15.50, so the estimated fare is £15.50. Choosing £12.00 stops after 2 × 6 = 12 and forgets to add the £3.50. Choosing £19.00 adds the distance and the constant first and then multiplies: 6 + 3.50 = 9.50, and 9.50 × 2 = 19.00, applying the ×2 to the whole sum instead of only to the distance. Choosing £13.00 multiplies only the constant term by 2 instead of the distance: 2 × 3.50 = 7, and 7 + 6 = 13.00.
- (b) All 2,000 members of the sports centre. — Method: in a survey, the population is the whole group the survey is trying to find out about, and the sample is the smaller group actually asked. Working: this survey wants to know what the sports centre's members think, so the population is every one of the 2,000 members — whether or not they were personally asked. Saying the population is the 100 members who were asked names the sample, not the population; the sample is drawn FROM the population, so it is smaller than it, not the same as it. Saying the population is everybody who lives in Ipswich widens the group far beyond who the survey is actually about — plenty of Ipswich residents are not members of the sports centre at all, so they are outside this survey altogether. Saying the population is the members who say they are unhappy confuses the population with a result of the survey: whether a member turns out to be happy or unhappy is something the survey finds out, not part of the definition of who is being studied. The population is always the whole group the question is about, before any sampling or any results come in.
- (b) 29 — There are 40 − 24 = 16 males, and 15 of them prefer cardio, so 16 − 15 = 1 male prefers weights. There are 24 females, and 10 prefer weights, so 24 − 10 = 14 females prefer cardio. Altogether, 15 + 14 = 29 people prefer cardio. Choosing 15 only counts the males who prefer cardio and forgets the females. Choosing 11 adds the two weights figures, 1 + 10 = 11, instead of the two cardio figures. Choosing 30 comes from 40 − 10, subtracting only the number of females who prefer weights from the grand total, rather than finding both cardio sub-totals separately.
- (b) Town B — higher median and smaller IQR — Method: 'higher and more consistent' needs two comparisons — the median for typical price, and the interquartile range for spread, with a smaller interquartile range meaning more consistent. Working: Town B's median, £235,000, is higher than Town A's, £220,000. Town A's interquartile range is 310 − 180 = 130 and Town B's is 260 − 175 = 85, so Town B's interquartile range is the smaller of the two. Answer: Town B has both the higher median and the smaller interquartile range, so it is the town with higher, more consistent prices. Watch which combination of median and interquartile range each statement claims, and for which town: claiming Town A has the higher median and the smaller interquartile range gets both comparisons wrong, since Town B leads on both; claiming Town A has the higher median (still wrong) but the larger interquartile range at least reads the spread correctly, without it rescuing the false median claim; and claiming Town B has the higher median (correct) but the larger interquartile range misreads the spread — Town B's interquartile range is the smaller of the two, not the larger.
- (d) 23.57 °C — Method: add all seven temperatures, divide by the number of readings and round only at the end. Working: 22 + 24 + 23 + 25 + 26 + 21 + 24 = 165, and 165 ÷ 7 = 23.5714…, which rounds to 23.57 to 2 decimal places. Answer: 23.57 °C. The distractors: 24 °C comes from writing down the mode, the only temperature recorded twice, instead of the mean; 27.5 °C comes from dividing the total by 6 instead of by the 7 days recorded; 5 °C comes from working out the range, 26 − 21, which is a measure of spread and not an average.
- (d) 19 kg — Method: estimate the mean of grouped data by multiplying each class's midpoint by its frequency, adding the four totals, then dividing by the total frequency. Working: the midpoints are 5, 15, 25 and 35 kg. The weighted totals are 11 × 5 = 55, 5 × 15 = 75, 5 × 25 = 125 and 9 × 35 = 315, which add to 570. Dividing by the 30 dogs gives an estimate of 570 ÷ 30 = 19 kg. Giving 5 kg reads off the midpoint of the modal class, 0 < m ≤ 10, the class with the most dogs — but the class with the most dogs is not where the mean falls, and neither is a substitute for actually calculating it. Giving 20 kg averages the four midpoints, (5 + 15 + 25 + 35) ÷ 4, treating every class as equally likely and ignoring that far more dogs are in the lightest and heaviest classes than in the middle two. Giving 570 kg stops after finding the correct weighted total and forgets the final division by the 30 dogs. Always weight each midpoint by its own frequency, and always finish by dividing by the total frequency, not the number of classes.
Build your own mix at the worksheet builder.