Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Non-calculator
Answer key: Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- (a) No, 150 pupils are a tenth of the school, chosen at random — Method: judge a sample on two things, whether every member of the population had the same chance of being chosen, and whether the sample is large enough to carry a pattern. Working: the 150 pupils were drawn from the register of every pupil in the school, so no year group or set is shut out and no pupil chooses to take part; and 150 ÷ 1,500 = 0.1, so one pupil in ten has been asked. A random sample of that share is ample for an estimate of how long the school's pupils spend on homework. Answer: no, because 150 pupils are a tenth of the school and were chosen at random. The distractors: saying a random sample always gives the exact school figure reaches the same verdict for a reason that is false, since a second random sample of 150 would give a slightly different mean; saying 150 pupils cannot be picked at random from 1,500 treats randomness as something only a whole population can have, when drawing names from the register is exactly how a random sample is taken; saying that only asking all 1,500 could show anything rejects sampling altogether, which would leave no way to study any population too large to count.
- (a) 1.5 — Method: with an even number of values the median is the mean of the two middle values, which for 12 values are the 6th and the 7th once the data are in order. Working: the results are already in order, and 12 ÷ 2 = 6, so the middle pair are the 6th value, 1, and the 7th value, 2; the median is (1 + 2) ÷ 2 = 1.5. Answer: 1.5 brothers and sisters. The distractors: 1 comes from reading the 6th value and stopping there instead of averaging the middle pair; 2 comes from working out the mean, 24 ÷ 12, instead of the median; 6 comes from working out the range, 6 − 0, which measures spread rather than centre.
- (d) 44 — Method: a cumulative frequency counts everything below a value, so the frequency of a class is the running total at the top of the class minus the running total at the bottom of it. Working: the running total below 20 kg is 96 and the running total below 10 kg is 52, so the number of boxes in the class 10 ≤ m < 20 is 96 − 52 = 44. Answer: 44 boxes. The distractors: 96 comes from quoting the running total at 20 kg itself, which counts every box below 20 kg rather than only those in this class; 34 comes from subtracting the wrong pair, 52 − 18, which gives the class 5 ≤ m < 10 instead; 54 comes from subtracting from the grand total, 150 − 96, which gives the boxes of 20 kg or more.
- (b) 13 — The total is 50, and the two known parts are 22 (tea) and 15 (coffee), so 50 − 22 − 15 = 13 hot chocolates. Choosing 28 comes from 50 − 22, subtracting only the tea and forgetting the coffee. Choosing 35 comes from 50 − 15, subtracting only the coffee and forgetting the tea. Choosing 37 comes from 22 + 15, which finds how many drinks were tea or coffee, not the number left over for hot chocolate.
- (c) The class with times from 10 up to 20 — Method: to find the median class from a histogram, first turn each bar's frequency density into a frequency using density × class width, build up the cumulative frequency, and find the first class whose cumulative frequency reaches or passes n ÷ 2. Working: the four classes have widths 10, 10, 20 and 20, so their frequencies are 5 × 10 = 50, 2 × 10 = 20, 1.5 × 20 = 30 and 1 × 20 = 20, which add to the 120 visitors stated. The median sits at position 120 ÷ 2 = 60. The cumulative frequency is 50 after the first class and 50 + 20 = 70 after the second, so the 60th visitor is reached during the second class. Answer: the median lies in the class 10 ≤ t < 20. Watch which class each shortcut lands on: the tallest bar belongs to the first class, with the highest frequency density, 5 — but the tallest bar shows where visitors are packed most densely, not where the middle visitor falls, and picking it lands one class too early, at 0 ≤ t < 10; taking half of the total TIME span instead of half of the total NUMBER of visitors, 60 minutes ÷ 2 = 30 minutes, lands in the class 20 ≤ t < 40, confusing a value on the horizontal axis with a position in the data; and using the full 120 visitors as the target position, rather than 120 ÷ 2 = 60, reaches all the way to the last class, 40 ≤ t < 60, treating the whole data set's size as though it were the position of a single middle value.
- (d) 648 kg — Method: to estimate a total from a sample, multiply the sample's mean by the number of items in the whole population, then check the units the question asks for. Working: 32.4 g × 20,000 = 648,000 g. Converting to kilograms, 648,000 ÷ 1,000 = 648 kg. This is only an estimate, not an exact total, because it assumes every one of the 20,000 packets has exactly the sample mean mass, when in reality individual packets vary above and below it. Giving 1.62 kg multiplies the mean by 50, the SAMPLE size, instead of by 20,000, the number of packets actually made that day — this finds the total mass of the 50 sampled packets, not the day's production. Giving 32.4 kg treats the sample mean itself, in grams, as if it already were the day's total mass in kilograms, skipping the scaling up altogether. Giving 648,000 kg correctly scales the mean up to the whole day's production but never converts the answer from grams to kilograms, leaving it 1,000 times too large. Always scale a sample's mean up by the SIZE OF THE WHOLE POPULATION, and always finish by checking the units the question asks for.
- (c) 28 — Method: multiply the number of whole symbols by the value of one symbol, then add the value of any half symbol shown. Working: 3 whole symbols represent 3 × 8 = 24 cars. The half symbol represents 4 cars. Total cars sold in March = 24 + 4 = 28. Leaving out the half symbol, 3 × 8 = 24, undercounts by exactly the value of that half symbol. Treating the half symbol as if it were a full symbol, 4 × 8 = 32, overcounts because it doubles the value the half symbol is worth. Giving 3.5 reports the number of symbols shown, not the number of cars they represent — the key still needs to be applied. Always apply the key to every symbol shown, including a half symbol, rather than reading off the symbol count itself.
- (c) The median, £160,000, as one very high price lifts the mean — Method: find both averages, then choose the one that sits closer to the bulk of the data. Working: in order the prices are 140,000, 150,000, 160,000, 170,000 and 580,000, so the median is the third of the five, £160,000. For the mean, 140,000 + 150,000 + 160,000 + 170,000 + 580,000 = 1,200,000 and 1,200,000 ÷ 5 = 240,000, so the mean is £240,000. Four of the five houses sold for £170,000 or less, so a reader told that a typical price is £240,000 would expect to pay at least £70,000 more than any of those four cost. Answer: the median, £160,000, as one very high price lifts the mean. The distractors: £580,000 is the middle value of the list as it is printed, which is the median only when the values have first been put in order; £240,000 is the mean, chosen on the ground that a median ignores three of the five prices, but a median uses all five to find which one is central and is then untroubled by how extreme the outer values are; £155,000 comes from deleting the £580,000 house and taking the mean of what is left, since 140,000 + 150,000 + 160,000 + 170,000 = 620,000 and 620,000 ÷ 4 = 155,000, but a real sale may not be thrown away merely for being large.
- (b) 5.5 kg — Method: with an even number of values the median is the mean of the two middle values, taken once the data are in order of size. Working: the eight masses are already in order and 8 ÷ 2 = 4, so the middle pair are the 4th and 5th values, 5 kg and 6 kg; the median is (5 + 6) ÷ 2 = 5.5 kg. Answer: 5.5 kg. The distractors: 5 kg comes from reading the 4th value and stopping there instead of averaging the middle pair; 8 kg comes from working out the range, 11 − 3, which measures spread rather than centre; 4 kg comes from writing down the modal mass, the only value that occurs twice, instead of the median.
- (a) Leeds has a higher median and a greater range than York. — In order, Leeds's temperatures are 14, 16, 18, 19 and 23, so the median is the middle value, 18, and the range is 23 − 14 = 9. York's temperatures in order are 15, 17, 17, 18 and 18, so the median is 17, and the range is 18 − 15 = 3. Since 18 is higher than 17, and 9 is greater than 3, Leeds has both the higher median and the greater range. Choosing 'Leeds has a higher median but a smaller range than York' gets the median comparison right but the range comparison backwards — Leeds's range of 9 is actually greater than York's range of 3. Choosing 'York has a higher median and a greater range than Leeds' reverses both comparisons. Choosing 'York has a higher median but a smaller range than Leeds' reverses the median comparison; York's median of 17 is lower than Leeds's 18, even though it is correct that York's range is the smaller one.
- (c) 76 — Method: the frequency of each class is the area of its bar, frequency density × class width, so work out all three frequencies and add them. Working: the widths are 5, 10 and 25 kg, so the frequencies are 4 × 5 = 20, 2.6 × 10 = 26 and 1.2 × 25 = 30; the total is 20 + 26 + 30 = 76. Answer: 76 dogs. The distractors: 7.8 comes from adding the three frequency densities, 4 + 2.6 + 1.2, treating each height as though it were a count; 107 comes from using each upper class boundary as the width, giving 4 × 5, 2.6 × 15 and 1.2 × 40; 39 comes from using the first class width, 5, for every bar, which ignores the unequal intervals and gives 20 + 13 + 6.
- (d) 18 minutes — Method: the lower quartile is the 80 ÷ 4 = 20th value and the upper quartile is the 3 × 80 ÷ 4 = 60th value; locate each inside its class by linear interpolation, then subtract. Working: the 20th value lies between the running totals 8 and 28, so it is in the class 10 ≤ t < 20, which holds 20 journeys across 10 minutes, and it is the 20 − 8 = 12th of them, giving 10 + (12 ÷ 20) × 10 = 16 minutes; the 60th value lies between the running totals 52 and 72, so it is in the class 30 ≤ t < 40, which also holds 20 journeys across 10 minutes, and it is the 60 − 52 = 8th of them, giving 30 + (8 ÷ 20) × 10 = 34 minutes; subtracting, 34 − 16 = 18. Answer: an estimated interquartile range of 18 minutes. The distractors: 20 minutes comes from taking the lower boundaries of the two quartile classes, 30 − 10, which locates the classes but never the values inside them; 40 minutes comes from subtracting the two positions, 60 − 20, instead of the two times; 22 minutes comes from interpolating downwards from each upper boundary rather than upwards from each lower boundary, giving 20 − 6 = 14 and 40 − 4 = 36.
- (b) 1 — Method: for data in a frequency table, find the position of the median using (n + 1) ÷ 2, then read off the value at that position from the cumulative frequencies. Working: there are 19 pupils, so the median is the 10th value. The cumulative frequencies are 7 (up to 0 pets), 10 (up to 1 pet), 14 (up to 2 pets) and 19 (up to 3 pets). The 10th value falls at the end of the '1 pet' group, so the median is 1 pet. Giving 0 pets is the mode — the category with the highest frequency, 7 — not the median. Giving 3, the highest number of pets minus the lowest, finds the range, a different statistic entirely. Giving 19 states the total number of pupils, not a number of pets at all. Find the middle POSITION first, then read off the value it belongs to — do not confuse it with the mode, the range or the total.
- (d) 38, and fairly confident, as 12 °C is inside the range — Method: substitute the forecast temperature into the equation of the line of best fit, then judge the prediction by where that temperature sits among the data the line was drawn from. Working: putting x = 12 into y = −3x + 74 gives −3 × 12 + 74 = 38, so the line predicts 38 hot chocolates. The recorded temperatures run from 4 °C to 18 °C, and 12 °C lies inside that interval, so this is interpolation, the safer kind of prediction. Answer: 38, and fairly confident, as 12 °C is inside the range; the owner should still expect the true figure to differ a little, since the points only lie near the line and not on it. The distractors: being completely certain treats a line of best fit as a rule that fixes each day's sales, when it describes a trend that individual days depart from; saying 12 °C is outside the range misreads the interval 4 °C to 18 °C, and the wrong warning would be attached to a sound prediction; 110 comes from −3 × 12 being taken as +36, giving 36 + 74 = 110, which loses the negative gradient and so predicts that a warm day sells more hot chocolate than a cold one.
- (c) 20 kg ≤ mass < 30 kg — The modal class is the class with the highest frequency. Reading the plotted points, the frequencies are 6, 10, 16, 6 and 2, so the highest frequency is 16, plotted at the midpoint 25. A class of width 10 centred on 25 runs from 25 − 5 = 20 to 25 + 5 = 30, so the modal class is 20 kg ≤ mass < 30 kg. Writing '25 kg' gives only the midpoint, not the class — the modal class is an interval, not a single value. '10 kg ≤ mass < 20 kg' is the class before the peak, centred on 15, which has frequency 10, not the highest. '30 kg ≤ mass < 40 kg' is the class after the peak, centred on 35, which has frequency 6, not the highest.
Build your own mix at the worksheet builder.