Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Non-calculator
Answer key: Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- (b) 55 — Method: count the classes that lie wholly above 8 minutes, then use linear interpolation for the class that 8 cuts through, assuming the delays in that class are spread evenly. Working: the class 10 ≤ d < 20 lies wholly above 8 and holds 25 buses; the value 8 lies in the class 5 ≤ d < 10, which is 5 minutes wide and holds 75 buses, and the part above 8 runs from 8 to 10, a width of 2, so the estimated share is (2 ÷ 5) × 75 = 30 buses; the estimate is 30 + 25 = 55. Answer: about 55 refunds. The distractors: 100 comes from adding the whole of the class 5 ≤ d < 10, 75 + 25, and so refunding buses only 5 minutes late; 25 comes from using only the class 10 ≤ d < 20 and ignoring the part class that 8 minutes cuts through; 70 comes from taking the part of the class from 5 up to 8 instead of from 8 up to 10, giving (3 ÷ 5) × 75 = 45 and then 45 + 25.
- (d) 44 — Method: a cumulative frequency counts everything below a value, so the frequency of a class is the running total at the top of the class minus the running total at the bottom of it. Working: the running total below 20 kg is 96 and the running total below 10 kg is 52, so the number of boxes in the class 10 ≤ m < 20 is 96 − 52 = 44. Answer: 44 boxes. The distractors: 96 comes from quoting the running total at 20 kg itself, which counts every box below 20 kg rather than only those in this class; 34 comes from subtracting the wrong pair, 52 − 18, which gives the class 5 ≤ m < 10 instead; 54 comes from subtracting from the grand total, 150 − 96, which gives the boxes of 20 kg or more.
- (d) 3 — Method: the height of a bar is its frequency density, frequency ÷ class width, so work out both heights and divide one by the other. Working: the class 0 ≤ t < 4 is 4 minutes wide and holds 30 visits, so its frequency density is 30 ÷ 4 = 7.5 per minute; the class 4 ≤ t < 20 is 16 minutes wide and holds 40 visits, so its frequency density is 40 ÷ 16 = 2.5 per minute; dividing the heights, 7.5 ÷ 2.5 = 3. Answer: the first bar is 3 times as tall. The distractors: 0.75 comes from comparing the frequencies, 30 ÷ 40, as though the frequencies were the heights, which is the mistake the unequal widths are there to expose; 4 comes from comparing the class widths, 16 ÷ 4, instead of the heights; 5 comes from subtracting the two frequency densities, 7.5 − 2.5, which answers how much taller rather than how many times taller.
- (b) 74 marks — Method: the two groups are different sizes, so their means cannot simply be averaged — rebuild each group's total mark, add the totals and divide by all 50 pupils. Working: Group A scored 20 × 80 = 1600 marks and Group B scored 30 × 70 = 2100 marks, giving 1600 + 2100 = 3700 marks altogether, so the overall mean is 3700 ÷ 50 = 74 marks. Answer: 74 marks, which sits nearer to 70 than to 80 because the larger group scored 70. The distractors: 75 marks comes from averaging the two group means, (80 + 70) ÷ 2, as though the groups were the same size; 76 marks comes from attaching each mean to the other group's size, (20 × 70 + 30 × 80) ÷ 50; 150 marks comes from adding the two means together and never dividing at all.
- (c) The data show a link only; a third factor may affect both — Method: a study of this kind measures two quantities and reports how they change together; deciding that one of them produces the other is a further claim, and it needs evidence that the measurements alone cannot give. Working: the study shows that more coffee goes with better concentration, which is a positive correlation; but a third factor that was never measured, such as how motivated someone is, could raise both the coffee drinking and the concentration, and the concentration could equally be what leads to the extra coffee. Answer: the data show a link only, because a third factor may be affecting both quantities, so no claim about cause can be made. The distractors: calling the conclusion safe because the correlation is positive treats the direction of a correlation as proof of cause, which no direction can give; calling it wrong because the correlation is negative misreads the direction of the relationship, since the study reports both quantities rising together; saying the two quantities are not linked denies the correlation the study actually found, when what fails is only the claim about cause.
- (c) 0.5 per gram — Method: turn the two known bars into frequencies using area, subtract from the total to find how many letters are left, then divide that frequency by the width of the last class to get its height. Working: the first bar covers 50 g at a frequency density of 1.2, giving 1.2 × 50 = 60 letters, and the second covers 50 g at 1.8, giving 1.8 × 50 = 90 letters; together that is 60 + 90 = 150 letters, so 200 − 150 = 50 letters remain; the class 100 ≤ m < 200 is 100 g wide, so its frequency density is 50 ÷ 100 = 0.5 per gram. Answer: 0.5 per gram. The distractors: 0.25 per gram comes from dividing the remaining 50 letters by the upper class boundary, 200, instead of by the class width of 100; 2 per gram comes from dividing the class width by the frequency, 100 ÷ 50, reversing the formula; 1.4 per gram comes from subtracting only the first bar's 60 letters, leaving 140, and then dividing by 100.
- (d) 16 — Method: on a box plot the interquartile range is the width of the box itself, upper quartile take away lower quartile. Working: the upper quartile is 38 and the lower quartile is 22, so 38 − 22 = 16. Answer: the interquartile range is 16 years. Watch which part of the box plot you are reading: the whole line from whisker to whisker gives the range, 61 − 15 = 46; the left half of the box alone gives median take away lower quartile, 29 − 22 = 7; and the right half of the box alone gives upper quartile take away median, 38 − 29 = 9 — neither half is the interquartile range on its own.
- (c) 156 cm — Method: to combine two groups' means, multiply each group's mean by its own number of pupils, add the two totals together, then divide by the total number of pupils in both groups. Working: 20 × 150 = 3,000 cm for the boys and 10 × 168 = 1,680 cm for the girls, giving a combined total of 3,000 + 1,680 = 4,680 cm. Dividing by all 30 pupils gives 4,680 ÷ 30 = 156 cm. Giving 159 cm averages the two means, (150 + 168) ÷ 2, treating the two groups as if they had the same number of pupils, when there are twice as many boys as girls. Giving 4,680 cm finds the correct combined total height but stops there, forgetting the final division by the 30 pupils. Giving 234 cm divides the combined total by 20, the number of boys only, forgetting that the total also includes the 10 girls. Always weight each mean by its own group size, and always divide by the TOTAL number of pupils in both groups combined.
- (a) The median, as the one very large wage does not move it — Method: an average describes a population well when it sits close to most of the values, so compare what each average does when one value lies far from the rest. Working: in order the wages are 420, 440, 460, 480 and 1,500, so the median is the third of the five, £460. The mean uses every wage: 420 + 440 + 460 + 480 + 1,500 = 3,300 and 3,300 ÷ 5 = 660, so the mean is £660. Four of the five people earn less than £660, and the nearest of those four wages is £180 below it, so £660 describes nobody at the garage; £460 sits inside the group of four similar wages. Answer: the median, as the one very large wage does not move it, while that same wage drags the mean £200 above the median. The distractors: saying the median is always larger than the mean is an invented rule, and here the median £460 is smaller than the mean £660; saying the mean is the only average that uses all five wages is true as far as it goes, but using a value and being dragged by it are the same thing when that value is £1,500; saying £660 lies between the smallest and largest wage is true of every mean ever calculated, so it proves nothing about whether this one is typical.
- (a) £16,000 — Method: first find the quarter with the highest sales figure, then subtract quarter 1's sales from it — remembering that every figure is given in THOUSANDS of pounds. Working: the highest sales figure is quarter 2, at £34,000 (34 thousand pounds). The increase from quarter 1 is £34,000 − £18,000 = £16,000. Giving £34,000 reads off the highest sales figure on its own, without subtracting quarter 1's sales — that is the highest quarter's total, not the increase. Giving £12,000 uses quarter 3's sales, 30, the SECOND-highest figure, instead of quarter 2's 34, the actual highest — 30 − 18 = 12, but quarter 3 is not the quarter with the highest sales. Giving £16 gets the subtraction right, 34 − 18 = 16, but forgets that every figure in the question is in thousands of pounds, so the increase is £16,000, not £16. Always identify the correct quarter FIRST, and always check the units the numbers are given in before writing your final answer.
- (c) Neither causes the other; sunshine links both. — Both ice cream sales and sunburn cases tend to rise on hot, sunny days, so the amount of sunshine is a third factor linked to both — neither variable causes the other. Saying ice cream sales cause the sunburn assumes a causal link in one direction that the correlation alone cannot establish. Saying sunburn cases cause the ice cream sales assumes the reverse causal link, which is no more justified. Saying a strong correlation always means causation is the general error this question is testing: correlation, however strong, does not by itself prove that one variable causes the other.
- (c) 33 — Method: multiply the mean by the number of tests to get the total marks, then subtract the marks that are already known. Working: four tests with a mean of 29 give a total of 29 × 4 = 116 marks; the first three marks total 31 + 26 + 26 = 83; so the fourth mark is 116 − 83 = 33. Answer: 33, and checking, (31 + 26 + 26 + 33) ÷ 4 = 116 ÷ 4 = 29. The distractors: 116 comes from stopping at the total for all four tests; 29 comes from assuming the missing mark must be the mean itself; 4 comes from multiplying the mean by 3, the number of marks given, leaving 87 − 83 = 4.
- (c) 145 g — Method: find the position of the median from the total frequency, locate the class that contains it, then use linear interpolation inside that class, assuming the apples in it are spread evenly. Working: the median is the 100 ÷ 2 = 50th apple; the running totals are 10, then 10 + 30 = 40, then 40 + 40 = 80, so the 50th apple lies in the class 140 ≤ m < 160; it is the 50 − 40 = 10th of the 40 apples in that class, and the class is 20 g wide, so the median is 140 + (10 ÷ 40) × 20 = 140 + 5 = 145. Answer: an estimated median of 145 g. The distractors: 150 g comes from giving the midpoint of the class that contains the median instead of interpolating inside it; 140 g comes from stopping at the lower boundary of that class, which locates the class but not the value; 155 g comes from measuring the 5 g step down from the upper boundary, 160 − 5, instead of up from the lower boundary.
- (a) The modal size, 9, bought by more customers than any other — Method: work out both averages from the frequencies, then choose the one the shop can act on. Working: for the mean, multiply each size by the number of pairs sold at it and add: 6 × 4 + 7 × 5 + 8 × 8 + 9 × 13 + 10 × 10 = 340, and 340 ÷ 40 = 8.5, so the mean size is 8.5. The largest frequency is 13, which belongs to size 9, so the modal size is 9. The mean 8.5 is a size no customer in the record asked for, so 40 pairs of it would sit unsold, while 13 of the 40 customers wanted size 9, more than wanted any other size. Answer: the modal size, 9, bought by more customers than any other. The distractors: the mean size 8.5 does take account of all 40 pairs, but a mean of sizes is a summary figure and not a size the month's customers were buying; the mean size 8 comes from averaging the five sizes on sale, 6 + 7 + 8 + 9 + 10 = 40 and 40 ÷ 5 = 8, which ignores how many pairs were sold at each size and so treats the 4 pairs of size 6 as equal in weight to the 13 pairs of size 9; the range 4 comes from 10 − 6 and measures spread, so it says how wide a set of sizes the shop must stock, not which size to stock most of.
- (a) No, 150 pupils are a tenth of the school, chosen at random — Method: judge a sample on two things, whether every member of the population had the same chance of being chosen, and whether the sample is large enough to carry a pattern. Working: the 150 pupils were drawn from the register of every pupil in the school, so no year group or set is shut out and no pupil chooses to take part; and 150 ÷ 1,500 = 0.1, so one pupil in ten has been asked. A random sample of that share is ample for an estimate of how long the school's pupils spend on homework. Answer: no, because 150 pupils are a tenth of the school and were chosen at random. The distractors: saying a random sample always gives the exact school figure reaches the same verdict for a reason that is false, since a second random sample of 150 would give a slightly different mean; saying 150 pupils cannot be picked at random from 1,500 treats randomness as something only a whole population can have, when drawing names from the register is exactly how a random sample is taken; saying that only asking all 1,500 could show anything rejects sampling altogether, which would leave no way to study any population too large to count.
Build your own mix at the worksheet builder.