Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Answer key: Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- (c) The mean, because every value counts towards it, so 100 pulls it from 7 up to about 20.3. — Method: work each measure out before the extra value is added and again afterwards, then compare the size of the two changes. Working: before, the six values total 42, so the mean is 42 ÷ 6 = 7, and the middle pair 6 and 8 give a median of (6 + 8) ÷ 2 = 7; after, the seven values total 142, so the mean is 142 ÷ 7 = 20.29 to 2 decimal places, while the median is now the 4th of the seven ordered values, which is 8; the mean has moved by about 13.3 and the median by 1. Answer: the mean, because every value counts towards it, so 100 pulls it from 7 up to about 20.3 — this is why the median is often preferred when a data set contains an outlier. The distractors: the reply that the mean rises by 100 adds the extra value to the mean instead of adding it to the total; the reply that the median moves to 12 takes the largest of the original values as the new middle instead of counting to the 4th of the seven values; the reply about even and odd counts quotes a rule that does not exist, since the median moved because a very large value was added, not because the count of values changed.
- (d) 25 — Method: the median is estimated at position n ÷ 2 in the cumulative frequency table, then interpolated across the class it falls in: lower boundary, plus the fraction of the way through the class, times the class width. Working: there are 80 sacks, so the median sits at position 80 ÷ 2 = 40. Before the class 20 ≤ m < 30 the cumulative frequency is 22, and by the end of it, it is 58, so this class holds the 40th sack; its frequency is 58 − 22 = 36 and its width is 30 − 20 = 10. The extra distance needed into the class is 40 − 22 = 18, and 18 ÷ 36 × 10 = 5, so the median is 20 + 5 = 25. Answer: the estimated median mass is 25 kg. Watch which numbers the interpolation uses: reading off just the lower boundary of the median class, 20, ignores how far into that class the 40th sack actually falls; treating n ÷ 2 = 40 itself as the median mass mistakes a position in the list for a mass in kilograms; and using the target position, 40, as the extra distance into the class instead of subtracting the sacks already counted changes the calculation to 20 + 40 ÷ 36 × 10. That comes to 20 + 11.1 = 31.1, overshooting the class because it never subtracts the 22 sacks already counted before it.
- (d) 60 — Method: on a histogram the frequency of a class is its frequency density × its class width, and the frequencies of all the classes add up to the total, so turn each labelled bar into a frequency and subtract their total from 250. Working: 10 ≤ age < 20 has width 20 − 10 = 10, so its frequency is 4.5 × 10 = 45; 20 ≤ age < 35 has width 35 − 20 = 15, so its frequency is 6 × 15 = 90; 50 ≤ age < 70 has width 70 − 50 = 20, so its frequency is 2.75 × 20 = 55. Those three come to 45 + 90 + 55 = 190, and the total is 250, so the missing frequency is 250 − 190 = 60. Answer: the class 35 ≤ age < 50 has 60 members. Watch what you do with the total and the three frequencies you have found: giving the total, 250, as the answer forgets that three bars have already accounted for some of the members; giving 190, the total of the other three classes, reports how many members are not in this class rather than how many are; and leaving one of the three out of the subtraction, for example 45 + 90 = 135 and 250 − 135 = 115, still owes the class at 50 ≤ age < 70 its 55 members.
- (c) The class with times from 10 up to 20 — Method: to find the median class from a histogram, first turn each bar's frequency density into a frequency using density × class width, build up the cumulative frequency, and find the first class whose cumulative frequency reaches or passes n ÷ 2. Working: the four classes have widths 10, 10, 20 and 20, so their frequencies are 5 × 10 = 50, 2 × 10 = 20, 1.5 × 20 = 30 and 1 × 20 = 20, which add to the 120 visitors stated. The median sits at position 120 ÷ 2 = 60. The cumulative frequency is 50 after the first class and 50 + 20 = 70 after the second, so the 60th visitor is reached during the second class. Answer: the median lies in the class 10 ≤ t < 20. Watch which class each shortcut lands on: the tallest bar belongs to the first class, with the highest frequency density, 5 — but the tallest bar shows where visitors are packed most densely, not where the middle visitor falls, and picking it lands one class too early, at 0 ≤ t < 10; taking half of the total TIME span instead of half of the total NUMBER of visitors, 60 minutes ÷ 2 = 30 minutes, lands in the class 20 ≤ t < 40, confusing a value on the horizontal axis with a position in the data; and using the full 120 visitors as the target position, rather than 120 ÷ 2 = 60, reaches all the way to the last class, 40 ≤ t < 60, treating the whole data set's size as though it were the position of a single middle value.
- (d) 25% — First find the number of fruit cakes: 80 − 34 − 26 = 20. Then write this as a percentage of the total: 20 ÷ 80 × 100 = 25%. Giving 20% comes from reporting the count of fruit cakes, 20, directly as a percentage, without dividing by the total of 80 first. Giving 32.5% computes the percentage of chocolate cakes instead of fruit cakes: 26 ÷ 80 × 100 = 32.5%. Giving 42.5% computes the percentage of sponge cakes instead of fruit cakes: 34 ÷ 80 × 100 = 42.5%.
- (d) 1100 — Method: to combine two samples of different sizes, add the faulty counts together and add the sample sizes together before scaling up, rather than treating the two samples separately. Working: the combined sample found 34 + 21 = 55 scratched cases out of 100 + 50 = 150 cases checked, a proportion of 55 ÷ 150. Applying that proportion to the week's production of 3,000 gives an estimate of 55 ÷ 150 × 3000 = 1100 scratched cases. Averaging the two shifts' proportions instead of combining their totals, (34 ÷ 100 + 21 ÷ 50) ÷ 2 = 0.38, gives 0.38 × 3000 = 1140 — this treats the two samples as equally weighted even though Shift A checked twice as many cases as Shift B. Using only Shift A's sample, 34 ÷ 100 × 3000 = 1020, ignores Shift B's cases completely. Using only Shift B's sample, 21 ÷ 50 × 3000 = 1260, ignores Shift A's cases completely. When two samples are different sizes, combine their totals before finding the proportion — do not average the two proportions, and do not use only one shift's sample.
- (a) A vertical line chart (discrete numerical data) — The number of pets is discrete numerical data — whole-number values such as 0, 1, 2, 3 or 4 — recorded for one variable, so a vertical line chart is the chart specified for this kind of data. A bar chart is used for categorical data, such as favourite colour, not numerical values counted like this. A pie chart shows proportions of a whole and does not show the frequency of each separate value. A scatter graph compares two different variables against each other, and only one variable, the number of pets, is recorded here.
- (c) Class X has the higher median and the wider spread — Method: compare the two box plots statistic by statistic — median for location, and the interquartile range for spread — checking the true value of each rather than assuming a pattern. Working: Class X's median is 60 and Class Y's is 58, so Class X's median is the higher one. Class X's interquartile range is 70 − 45 = 25 and Class Y's is 65 − 50 = 15 (and the ranges follow the same order: 95 − 20 = 75 against 80 − 35 = 45), so Class X also has the wider spread. Answer: Class X has both the higher median and the wider spread. Watch that each half of a compound statement is checked separately: claiming Class Y has the higher median and the wider spread gets both comparisons backwards; claiming Class X has the higher median but the narrower spread keeps the median right while reading the spread the wrong way round; and claiming Class Y has the higher median but the narrower spread swaps the median comparison while getting the spread right.
- (c) 90 — Method: the height of a bar is its frequency density, so twice as tall means twice the frequency density — not twice the frequency, because the two classes have different widths. Then frequency = frequency density × class width. Working: the first bar has frequency density 3 per cm, so the second has frequency density 2 × 3 = 6 per cm; the class 30 ≤ x < 45 is 45 − 30 = 15 cm wide, so its frequency is 6 × 15 = 90. Answer: 90 rods. The distractors: 120 comes from doubling the first bar's frequency instead of its height — the first class holds 3 × 20 = 60 rods, and doubling that ignores the fact that the second class is narrower; 45 comes from using the first bar's frequency density, 3, for the second bar, 3 × 15, and so never using the information that it is twice as tall; 6 comes from stopping at the frequency density of the taller bar and quoting a height as though it were a count.
- (c) 40 minutes — Method: for grouped data, estimate the mean using the midpoint of each class — multiply each midpoint by its frequency, add the results, then divide by the total frequency. Working: the midpoints are 10, 30, 50 and 70 minutes. 10 × 5 = 50. 30 × 10 = 300. 50 × 10 = 500. 70 × 5 = 350. Σfx = 50 + 300 + 500 + 350 = 1200. Σf = 5 + 10 + 10 + 5 = 30. Estimated mean = 1200 ÷ 30 = 40 minutes. Using the upper boundary of each class instead of the midpoint — 20 × 5 = 100, 40 × 10 = 400, 60 × 10 = 600, 80 × 5 = 400 — gives a total of 1500 and an estimate of 1500 ÷ 30 = 50 minutes, too high because a boundary is not the middle of the class. Averaging the frequencies themselves, 5, 10, 10 and 5, ignores the times altogether and gives 7.5. Stopping after Σfx = 1200 without dividing by the total frequency gives a number far too large to be a time in minutes. Always find the midpoint of each class before multiplying by the frequency, and always divide by Σf at the end.
- (c) 42 — Method: find the target cumulative frequency, 90% of the total, locate the class it falls in from the plotted points, then interpolate: lower boundary, plus the extra distance needed into the class divided by the class's frequency, times its width. Working: 90% of 320 is 0.9 × 320 = 288. The plotted points show a cumulative frequency of 280 at d = 40 and 320 at d = 50, so the class 40 ≤ d < 50 has frequency 320 − 280 = 40 and width 50 − 40 = 10, and 288 falls inside it. The extra distance needed into the class is 288 − 280 = 8, and 8 ÷ 40 × 10 = 2, so the diameter is 40 + 2 = 42. Answer: the estimated diameter is 42 mm. Watch which point and which class the interpolation actually uses: reading off d = 40, the plotted point just below the target, instead of interpolating the extra 8 ball bearings into the next 10 mm, stops one step short of the true answer; finding the diameter below which only 10% lie instead of 90% gives a target of 0.1 × 320 = 32, which falls in the class 10 ≤ d < 20 — the extra distance into that class is 32 − 30 = 2, and 2 ÷ 60 × 10 = 0.3, so this route gives 10 + 0.3 = 10.3, the bottom decile rather than the top 90%; and interpolating within the class 30 ≤ d < 40 instead of 40 ≤ d < 50, as though 288 had not yet reached a cumulative frequency of 280, treats the extra distance as 288 − 190 = 98, and 98 ÷ 90 × 10 = 10.9, giving 30 + 10.9 = 40.9, one class too early.
- (a) No, the size of the fire affects both of the quantities — Method: correlation says that two quantities change together; a claim that one of them produces the other is a further claim, and it needs evidence that a scatter graph on its own cannot give. Working: the graph does show strong positive correlation, so more engines did go with greater damage. But neither quantity was set by the researchers: both were decided by how large the fire was. A large blaze brings many appliances and also destroys a great deal, while a small one brings few and destroys little, so a third quantity is driving both of the recorded ones. Answer: no, because the size of the fire affects both of the quantities. The distractors: saying the correlation is negative contradicts the graph, which shows the two quantities rising together, and reaching the right verdict from a false reading of the data is not the reason the mark is for; saying that strong positive correlation shows one quantity causes the other is the assumption the question exists to test, and no strength of correlation can establish cause; saying the points lie close to the line of best fit describes how strong the correlation is, and strength and cause are different matters entirely.
- (c) Yes — £80 is above the boundary, £78 — Method: a value counts as an outlier when it lies more than 1.5 times the interquartile range beyond the nearer quartile; here that means checking it against upper quartile + 1.5 × interquartile range. Working: the interquartile range is 42 − 18 = 24. 1.5 × 24 = 36, and 42 + 36 = 78, so any saving above £78 is an outlier. Amara saved £80, and 80 is greater than 78. Answer: yes, Amara's saving is an outlier, because £80 is above the outlier boundary, £78. Watch how you build the boundary and what you compare it with: adding the two quartiles instead of subtracting them, 42 + 18 = 60, gives an interquartile range three times too big, and 42 + 1.5 × 60 = 42 + 90 = 132 puts the boundary so far out that £80 wrongly looks ordinary; comparing £80 with the upper quartile alone, £42, checks only that it lies in the top quarter of the data, which every value above £42 does, not that it lies unusually far beyond it; and adding the interquartile range on once instead of one and a half times, 42 + 24 = 66, uses the wrong multiplier, even though £80 still happens to clear that lower boundary too.
- (d) 19 kg — Method: estimate the mean of grouped data by multiplying each class's midpoint by its frequency, adding the four totals, then dividing by the total frequency. Working: the midpoints are 5, 15, 25 and 35 kg. The weighted totals are 11 × 5 = 55, 5 × 15 = 75, 5 × 25 = 125 and 9 × 35 = 315, which add to 570. Dividing by the 30 dogs gives an estimate of 570 ÷ 30 = 19 kg. Giving 5 kg reads off the midpoint of the modal class, 0 < m ≤ 10, the class with the most dogs — but the class with the most dogs is not where the mean falls, and neither is a substitute for actually calculating it. Giving 20 kg averages the four midpoints, (5 + 15 + 25 + 35) ÷ 4, treating every class as equally likely and ignoring that far more dogs are in the lightest and heaviest classes than in the middle two. Giving 570 kg stops after finding the correct weighted total and forgets the final division by the 30 dogs. Always weight each midpoint by its own frequency, and always finish by dividing by the total frequency, not the number of classes.
- (b) £1,000 — Wages take up 150° out of 360°, so the amount spent on wages is 150 ÷ 360 × 2400 = £1,000. Choosing £600 uses the repairs angle, 90°, instead of the wages angle: 90 ÷ 360 × 2400 = 600. Choosing £3,600 treats the angle in degrees as if it were a percentage, 150 ÷ 100 × 2400 = 3600, instead of dividing by 360°. Choosing £800 uses the angle for the 'other costs' sector, 360 − 90 − 150 = 120°, instead of the wages sector: 120 ÷ 360 × 2400 = 800.
Build your own mix at the worksheet builder.