Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Answer key: Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- (d) 25 — Method: the median is estimated at position n ÷ 2 in the cumulative frequency table, then interpolated across the class it falls in: lower boundary, plus the fraction of the way through the class, times the class width. Working: there are 80 sacks, so the median sits at position 80 ÷ 2 = 40. Before the class 20 ≤ m < 30 the cumulative frequency is 22, and by the end of it, it is 58, so this class holds the 40th sack; its frequency is 58 − 22 = 36 and its width is 30 − 20 = 10. The extra distance needed into the class is 40 − 22 = 18, and 18 ÷ 36 × 10 = 5, so the median is 20 + 5 = 25. Answer: the estimated median mass is 25 kg. Watch which numbers the interpolation uses: reading off just the lower boundary of the median class, 20, ignores how far into that class the 40th sack actually falls; treating n ÷ 2 = 40 itself as the median mass mistakes a position in the list for a mass in kilograms; and using the target position, 40, as the extra distance into the class instead of subtracting the sacks already counted changes the calculation to 20 + 40 ÷ 36 × 10. That comes to 20 + 11.1 = 31.1, overshooting the class because it never subtracts the 22 sacks already counted before it.
- (d) 19 kg — Method: estimate the mean of grouped data by multiplying each class's midpoint by its frequency, adding the four totals, then dividing by the total frequency. Working: the midpoints are 5, 15, 25 and 35 kg. The weighted totals are 11 × 5 = 55, 5 × 15 = 75, 5 × 25 = 125 and 9 × 35 = 315, which add to 570. Dividing by the 30 dogs gives an estimate of 570 ÷ 30 = 19 kg. Giving 5 kg reads off the midpoint of the modal class, 0 < m ≤ 10, the class with the most dogs — but the class with the most dogs is not where the mean falls, and neither is a substitute for actually calculating it. Giving 20 kg averages the four midpoints, (5 + 15 + 25 + 35) ÷ 4, treating every class as equally likely and ignoring that far more dogs are in the lightest and heaviest classes than in the middle two. Giving 570 kg stops after finding the correct weighted total and forgets the final division by the 30 dogs. Always weight each midpoint by its own frequency, and always finish by dividing by the total frequency, not the number of classes.
- (b) A person's shoe size and their favourite colour — A person's shoe size is not linked to which colour they prefer, so these two show no correlation. The other three pairs are all genuinely correlated: distance travelled and fuel used rise together, which is positive correlation; hours of revision and test score generally rise together, which is also positive correlation; and as outdoor temperature rises, fewer woolly hats are sold, which is negative correlation. Negative correlation is still a real relationship between two variables — it is not the same thing as no relationship at all, so the temperature and hats pair is not the answer to this question.
- (d) Correlation is a link; causation is one causing the other — Method: the two words describe different claims — one is about a pattern in the data, the other is about what produced that pattern. Working: correlation says only that two quantities tend to change together, which is something a scatter graph can display; causation says that a change in one quantity actually brings about the change in the other, which needs evidence a scatter graph cannot supply, because a third quantity may be driving both. Answer: correlation is a link between the quantities, while causation is one quantity causing the change in another. The distractors: the statement giving causation as the link and correlation as the cause simply swaps the two words over; the statement that the words mean the same thing is the classic error of reading a correlation as proof of cause; the statement that a scatter graph shows causation but not correlation reverses what a scatter graph can do, since the pattern it displays is exactly the correlation.
- (a) No, the size of the fire affects both of the quantities — Method: correlation says that two quantities change together; a claim that one of them produces the other is a further claim, and it needs evidence that a scatter graph on its own cannot give. Working: the graph does show strong positive correlation, so more engines did go with greater damage. But neither quantity was set by the researchers: both were decided by how large the fire was. A large blaze brings many appliances and also destroys a great deal, while a small one brings few and destroys little, so a third quantity is driving both of the recorded ones. Answer: no, because the size of the fire affects both of the quantities. The distractors: saying the correlation is negative contradicts the graph, which shows the two quantities rising together, and reaching the right verdict from a false reading of the data is not the reason the mark is for; saying that strong positive correlation shows one quantity causes the other is the assumption the question exists to test, and no strength of correlation can establish cause; saying the points lie close to the line of best fit describes how strong the correlation is, and strength and cause are different matters entirely.
- (d) Only internet users reach the website; others are excluded. — Method: a sample is biased when it systematically leaves out part of the population, or systematically over-represents another part. Working: anyone without internet access, or who does not visit the council's website, has NO chance of being included — the sample is drawn only from internet-using residents, which is not the whole town. Saying too many people might respond because the survey is free confuses bias with sample size — bias is about who CAN be reached, not how many respond. Saying people might lie describes a different problem, response honesty, not who was sampled in the first place. Saying online surveys cannot be anonymous is not a reason connected to bias at all. A sample is biased when part of the population has no chance of being included, whatever the reason for that.
- (d) £15.50 — 2 × 6 = 12, and 12 + 3.50 = 15.50, so the estimated fare is £15.50. Choosing £12.00 stops after 2 × 6 = 12 and forgets to add the £3.50. Choosing £19.00 adds the distance and the constant first and then multiplies: 6 + 3.50 = 9.50, and 9.50 × 2 = 19.00, applying the ×2 to the whole sum instead of only to the distance. Choosing £13.00 multiplies only the constant term by 2 instead of the distance: 2 × 3.50 = 7, and 7 + 6 = 13.00.
- (c) 0.5 per gram — Method: turn the two known bars into frequencies using area, subtract from the total to find how many letters are left, then divide that frequency by the width of the last class to get its height. Working: the first bar covers 50 g at a frequency density of 1.2, giving 1.2 × 50 = 60 letters, and the second covers 50 g at 1.8, giving 1.8 × 50 = 90 letters; together that is 60 + 90 = 150 letters, so 200 − 150 = 50 letters remain; the class 100 ≤ m < 200 is 100 g wide, so its frequency density is 50 ÷ 100 = 0.5 per gram. Answer: 0.5 per gram. The distractors: 0.25 per gram comes from dividing the remaining 50 letters by the upper class boundary, 200, instead of by the class width of 100; 2 per gram comes from dividing the class width by the frequency, 100 ÷ 50, reversing the formula; 1.4 per gram comes from subtracting only the first bar's 60 letters, leaving 140, and then dividing by 100.
- (c) Branch A waits longer, and Branch A is more consistent — Method: compare the two branches using a measure of location (the median) for who waits longer, and a measure of spread (the interquartile range) for who is more consistent — a smaller interquartile range means more consistent. Working: Branch A's median, 12 minutes, is higher than Branch B's, 9 minutes, so Branch A's customers wait longer on average. Branch A's interquartile range, 5 minutes, is smaller than Branch B's, 11 minutes, so Branch A's waiting times vary less. Answer: Branch A waits longer, and Branch A is also the more consistent of the two. Watch that each half of the comparison uses the right statistic and reads it correctly: swapping both readings gives Branch B the longer wait and the greater consistency, when neither is true; keeping the median comparison right but reading a larger interquartile range as 'more consistent' has the direction of spread backwards; and swapping only the median comparison keeps the correct branch for consistency but gives the wrong branch the longer wait.
- (d) 38, and fairly confident, as 12 °C is inside the range — Method: substitute the forecast temperature into the equation of the line of best fit, then judge the prediction by where that temperature sits among the data the line was drawn from. Working: putting x = 12 into y = −3x + 74 gives −3 × 12 + 74 = 38, so the line predicts 38 hot chocolates. The recorded temperatures run from 4 °C to 18 °C, and 12 °C lies inside that interval, so this is interpolation, the safer kind of prediction. Answer: 38, and fairly confident, as 12 °C is inside the range; the owner should still expect the true figure to differ a little, since the points only lie near the line and not on it. The distractors: being completely certain treats a line of best fit as a rule that fixes each day's sales, when it describes a trend that individual days depart from; saying 12 °C is outside the range misreads the interval 4 °C to 18 °C, and the wrong warning would be attached to a sound prediction; 110 comes from −3 × 12 being taken as +36, giving 36 + 74 = 110, which loses the negative gradient and so predicts that a warm day sells more hot chocolate than a cold one.
- (d) 23.57 °C — Method: add all seven temperatures, divide by the number of readings and round only at the end. Working: 22 + 24 + 23 + 25 + 26 + 21 + 24 = 165, and 165 ÷ 7 = 23.5714…, which rounds to 23.57 to 2 decimal places. Answer: 23.57 °C. The distractors: 24 °C comes from writing down the mode, the only temperature recorded twice, instead of the mean; 27.5 °C comes from dividing the total by 6 instead of by the 7 days recorded; 5 °C comes from working out the range, 26 − 21, which is a measure of spread and not an average.
- (c) 75 — Method: the number in a class is the area of its bar, frequency density × class width, so work out the frequency of each class that lies at or above 10 minutes and add them. Working: the class 10 ≤ t < 25 is 15 minutes wide with a frequency density of 3.2, giving 3.2 × 15 = 48 members; the class 25 ≤ t < 55 is 30 minutes wide with a frequency density of 0.9, giving 0.9 × 30 = 27 members; the total charged is 48 + 27 = 75. Answer: 75 members pay the extra charge. The distractors: 4.1 comes from adding the two frequency densities, 3.2 + 0.9, as though each height were a count; 93 comes from including the class 0 ≤ t < 10 as well, 1.8 × 10 = 18 added to 48 and 27, which charges every member; 27 comes from using only the class 25 ≤ t < 55 and forgetting that 10 ≤ t < 25 is also at or above 10 minutes.
- (c) 46 — Method: the cumulative frequency table gives the number of runners below each time; to find the number at or above a time, subtract that cumulative frequency from the total. Working: the cumulative frequency for t < 40 is 74, so 120 runners in total take away the 74 who finished in under 40 minutes: 120 − 74 = 46. Answer: 46 runners took 40 minutes or longer. Watch which boundary and which subtraction you use: reading off t < 50 instead of t < 40 and subtracting, 120 − 110 = 10, answers a different question, '50 minutes or longer'; giving 74 itself as the answer reports how many finished below 40 minutes, the opposite of what was asked; and subtracting the two nearby cumulative frequencies, 110 − 74 = 36, finds how many took between 40 and 50 minutes, not everyone from 40 minutes upward.
- (d) 25% — First find the number of fruit cakes: 80 − 34 − 26 = 20. Then write this as a percentage of the total: 20 ÷ 80 × 100 = 25%. Giving 20% comes from reporting the count of fruit cakes, 20, directly as a percentage, without dividing by the total of 80 first. Giving 32.5% computes the percentage of chocolate cakes instead of fruit cakes: 26 ÷ 80 × 100 = 32.5%. Giving 42.5% computes the percentage of sponge cakes instead of fruit cakes: 34 ÷ 80 × 100 = 42.5%.
- (c) Class X has the higher median and the wider spread — Method: compare the two box plots statistic by statistic — median for location, and the interquartile range for spread — checking the true value of each rather than assuming a pattern. Working: Class X's median is 60 and Class Y's is 58, so Class X's median is the higher one. Class X's interquartile range is 70 − 45 = 25 and Class Y's is 65 − 50 = 15 (and the ranges follow the same order: 95 − 20 = 75 against 80 − 35 = 45), so Class X also has the wider spread. Answer: Class X has both the higher median and the wider spread. Watch that each half of a compound statement is checked separately: claiming Class Y has the higher median and the wider spread gets both comparisons backwards; claiming Class X has the higher median but the narrower spread keeps the median right while reading the spread the wrong way round; and claiming Class Y has the higher median but the narrower spread swaps the median comparison while getting the spread right.
Build your own mix at the worksheet builder.