Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- 1.A scatter graph of taxi journeys in Bristol shows the distance, x miles, and the fare, y pounds. The line of best fit is y = 2x + 3.50. Work out the estimated fare for a journey of 6 miles, using the line of best fit.y = 2x + 3.5
- 2.A charity shop in Leicester made £2,400 in one month. A pie chart shows how this was spent: repairs took up an angle of 90°, wages took up an angle of 150°, and the rest was other costs. Work out how much was spent on wages.
- 3.A gym in Cardiff has a scatter graph of the number of training sessions, x, attended by each member and the weight lost, y kg. The line of best fit is y = 0.4x + 1. Members want to lose at least 9 kg. Using the line of best fit, work out the least whole number of sessions needed to reach this target.y = 0.4x + 1
- 4.A histogram shows the speeds, v mph, of 100 vehicles passing a checkpoint. The bar for 0 ≤ v < 20 has a frequency density of 1 vehicle per mph, the bar for 20 ≤ v < 30 has a frequency density of 3 vehicles per mph, the bar for 30 ≤ v < 50 has a frequency density of 2 vehicles per mph, and the bar for 50 ≤ v < 70 has a frequency density of 0.5 vehicles per mph. Estimate the mean speed of the vehicles.
- 5.A school has 1,000 pupils. A student wants to estimate how many of them walk to school, so she asks 8 pupils in her own class. She says her sample is large enough to give a reliable estimate for the whole school. Is she right? Give a reason for your answer.
- 6.Priya travels to work by one of two routes and records her journey times, in minutes, over several weeks. Route 1 has a median of 34 minutes and an interquartile range of 22 minutes. Route 2 has a median of 41 minutes and an interquartile range of 6 minutes. Priya wants the more reliable route for getting to an important meeting on time. Which route should she choose, and why?
- 7.A call centre records the length, t seconds, of 100 calls: 0 ≤ t < 20, 15 calls; 20 ≤ t < 30, 24 calls; 30 ≤ t < 50, 40 calls; 50 ≤ t < 80, 21 calls. The manager's target is for a call to be finished in under 35 seconds. Estimate the number of calls that met the target.
- 8.A school has 1200 pupils. A teacher wants to take a random sample of 60 of them. Write down which of these methods gives a random sample.
- 9.In a histogram of the lengths, x cm, of some rods, the bar for 10 ≤ x < 30 has a frequency density of 3 per cm. The bar for 30 ≤ x < 45 is twice as tall as the bar for 10 ≤ x < 30. Work out the number of rods with a length in the class 30 ≤ x < 45.
- 10.A scatter graph shows the number of years of experience, x, of 18 sales assistants and their monthly sales, y hundred pounds. The plotted points run from x = 1 to x = 12 years, and the line of best fit is y = 4x + 20. A new assistant has 25 years of experience. Use the line of best fit to estimate a value of y for this assistant, and decide whether the estimate would be reliable.y = 4x + 20
- 11.A box plot for the amount of pocket money, in pounds, saved by 30 students in a month is based on a lower quartile of £18 and an upper quartile of £42. A value above upper quartile + 1.5 × interquartile range is considered an outlier. Amara saved £80 in the month. Determine whether Amara's saving is an outlier.
- 12.A gym draws a histogram of the times, t minutes, that its members spend on one machine. The bar for 0 ≤ t < 10 has a frequency density of 1.8 per minute, the bar for 10 ≤ t < 25 has a frequency density of 3.2 per minute, and the bar for 25 ≤ t < 55 has a frequency density of 0.9 per minute. Members who spend 10 minutes or more on the machine pay an extra charge. Work out the number of members who pay the extra charge.
- 13.The delivery times, in minutes, of 15 parcels are given in order: 5, 8, 10, 12, 14, 15, 17, 19, 21, 23, 25, 27, 29, 31, 34. Work out the interquartile range of these times.
- 14.Forty pupils in class P and forty pupils in class Q each solved a puzzle. The times, in seconds, were summarised using cumulative frequency. For class P the lower quartile is 24, the median is 38 and the upper quartile is 46. For class Q the lower quartile is 30, the median is 35 and the upper quartile is 44. Write down the statement that correctly compares the two classes.
- 15.A vet records the masses, m kg, of the dogs seen in one week as a histogram. The bar for 0 ≤ m < 5 has a frequency density of 4 per kg, the bar for 5 ≤ m < 15 has a frequency density of 2.6 per kg, and the bar for 15 ≤ m < 40 has a frequency density of 1.2 per kg. Work out the total number of dogs seen that week.
Answer key
- (d) £15.50 — 2 × 6 = 12, and 12 + 3.50 = 15.50, so the estimated fare is £15.50. Choosing £12.00 stops after 2 × 6 = 12 and forgets to add the £3.50. Choosing £19.00 adds the distance and the constant first and then multiplies: 6 + 3.50 = 9.50, and 9.50 × 2 = 19.00, applying the ×2 to the whole sum instead of only to the distance. Choosing £13.00 multiplies only the constant term by 2 instead of the distance: 2 × 3.50 = 7, and 7 + 6 = 13.00.
- (b) £1,000 — Wages take up 150° out of 360°, so the amount spent on wages is 150 ÷ 360 × 2400 = £1,000. Choosing £600 uses the repairs angle, 90°, instead of the wages angle: 90 ÷ 360 × 2400 = 600. Choosing £3,600 treats the angle in degrees as if it were a percentage, 150 ÷ 100 × 2400 = 3600, instead of dividing by 360°. Choosing £800 uses the angle for the 'other costs' sector, 360 − 90 − 150 = 120°, instead of the wages sector: 120 ÷ 360 × 2400 = 800.
- (d) 20 — 0.4x + 1 = 9, so 0.4x = 9 − 1 = 8, and 8 ÷ 0.4 = 20, so 20 sessions are needed. Choosing 23 divides 9 by 0.4 without first subtracting the 1: 9 ÷ 0.4 = 22.5, rounded up to 23. Choosing 25 subtracts the wrong way, adding the 1 instead of taking it away: 9 + 1 = 10, and 10 ÷ 0.4 = 25. Choosing 2 misplaces the decimal point in the gradient, dividing by 4 instead of by 0.4: 8 ÷ 4 = 2.
- (a) 31.5 — Method: to estimate the mean from a histogram, first turn each bar into a frequency (frequency density × class width), then use mean = Σ(frequency × midpoint) ÷ Σfrequency, with the midpoint standing in for every value in that class. Working: the four classes have widths 20, 10, 20 and 20, so their frequencies are 1 × 20 = 20, 3 × 10 = 30, 2 × 20 = 40 and 0.5 × 20 = 10, which do add to the 100 vehicles stated. Their midpoints are 10, 25, 40 and 60, so Σfx = 20 × 10 + 30 × 25 + 40 × 40 + 10 × 60 = 200 + 750 + 1600 + 600 = 3150, and the mean is 3150 ÷ 100 = 31.5. Answer: the estimated mean speed is 31.5 mph. Watch which numbers you treat as the frequencies and which as the values: using the frequency densities themselves as the frequencies, without multiplying by the class widths first, gives 1 × 10 + 3 × 25 + 2 × 40 + 0.5 × 60 = 195 spread over 1 + 3 + 2 + 0.5 = 6.5, and 195 ÷ 6.5 = 30, a mean built from the wrong 'frequencies' altogether; averaging the four midpoints on their own, (10 + 25 + 40 + 60) ÷ 4 = 33.75, ignores how many vehicles are actually in each class; and using each class's lower boundary in place of its midpoint, 20 × 0 + 30 × 20 + 40 × 30 + 10 × 50 = 2300 and 2300 ÷ 100 = 23, systematically underestimates every class by roughly half its width.
- (c) No — 8 from one class is too small to represent the school. — Method: judge reliability by asking whether the sample is both large enough, and spread across the population, relative to what it is meant to represent. Working: 8 pupils is a tiny fraction of the school's 1,000 pupils, and all 8 come from a single class rather than a range of year groups, so the sample is both too small and too narrow to represent the whole school reliably. She is not right. Saying any sample size gives an equally reliable estimate ignores that reliability generally improves with a larger, more representative sample. Saying the method is unreliable because it was not done online is not a reason connected to sample size or representativeness at all. Saying 8 is reliable because it is more than half her class compares the sample to the wrong population — the school has 1,000 pupils, not one class. Always judge a sample's size against the population it is meant to represent, not against a smaller group within it.
- (c) Route 2, because its interquartile range is smaller — Method: for a journey where turning up on time matters, what matters is not the typical (median) time but how predictable it is — a smaller interquartile range means the middle half of journeys cluster closer together. Working: Route 1's median, 34 minutes, is in fact lower than Route 2's, 41 minutes, so Route 1 is faster on average; but Route 1's interquartile range, 22 minutes, is far larger than Route 2's, 6 minutes, so Route 1's times are much less predictable. Answer: Priya should choose Route 2, because its interquartile range is smaller, even though it is slower on average. Watch which statistic answers the question actually asked: Route 1 does not have the smaller interquartile range, Route 2 does, so picking Route 1 for that reason misreads the table; Route 1's median genuinely is the lower one, but a lower median answers 'which is faster', not 'which is more reliable'; and Route 2's median is not the lower one, so that claim about Route 2 is simply false.
- (a) 49 — Method: add the frequencies of the classes that lie wholly below 35 seconds, then use linear interpolation for the class that 35 cuts through, assuming the calls in that class are spread evenly across it. Working: below 30 seconds there are 15 + 24 = 39 calls; the value 35 lies in the class 30 ≤ t < 50, which is 20 seconds wide and holds 40 calls, and 35 is 35 − 30 = 5 seconds into it, so the estimated share is (5 ÷ 20) × 40 = 10 calls; the estimate is 39 + 10 = 49. Answer: about 49 calls met the target. The distractors: 79 comes from adding the whole of the class 30 ≤ t < 50, 39 + 40, and so counting calls of up to 50 seconds as being under 35; 39 comes from stopping at the class boundary 30 and ignoring the part class altogether; 69 comes from measuring the part of the class from 35 up to 50 instead of from 30 up to 35, giving (15 ÷ 20) × 40 = 30 and then 39 + 30.
- (a) Drawing 60 names at random from a list of all 1200 pupils — Method: a sample is random when every member of the population has the same chance of being chosen and nobody, including the pupils themselves, can influence who ends up in it; test each method against that. Working: drawing names from a list of all 1200 pupils gives each pupil the same chance, 60 out of 1200, whatever their year group, class or opinion, so the method is random. Answer: drawing 60 names at random from a list of all 1200 pupils. The distractors: asking the pupils who volunteer is self-selection, and the pupils with the strongest views volunteer first, so they decide the sample; asking the pupils nearest the door is convenience sampling, which reaches only those who happen to be in one place at one time; asking two Year 10 classes samples a cluster, so every pupil in the other year groups has no chance of being chosen at all.
- (c) 90 — Method: the height of a bar is its frequency density, so twice as tall means twice the frequency density — not twice the frequency, because the two classes have different widths. Then frequency = frequency density × class width. Working: the first bar has frequency density 3 per cm, so the second has frequency density 2 × 3 = 6 per cm; the class 30 ≤ x < 45 is 45 − 30 = 15 cm wide, so its frequency is 6 × 15 = 90. Answer: 90 rods. The distractors: 120 comes from doubling the first bar's frequency instead of its height — the first class holds 3 × 20 = 60 rods, and doubling that ignores the fact that the second class is narrower; 45 comes from using the first bar's frequency density, 3, for the second bar, 3 × 15, and so never using the information that it is twice as tall; 6 comes from stopping at the frequency density of the taller bar and quoting a height as though it were a count.
- (b) 120, unreliable — x = 25 is outside 1 to 12 — The line of best fit is y = 4x + 20. 4 × 25 = 100, and 100 + 20 = 120, so the estimate is y = 120. But x = 25 lies far outside the plotted range of 1 to 12 years, so this is an extrapolation, and the estimate is not reliable. Reaching 100 instead of 120 comes from 4 × 25 = 100 with the intercept of 20 left out — still correctly flagged as unreliable, but the wrong value. Calling the estimate reliable simply because it was calculated correctly, giving 120, wrongly assumes that a correct calculation is automatically trustworthy, ignoring that x = 25 lies far beyond the data actually collected. Reaching 68, from 4 × 12 = 48 and 48 + 20 = 68, substitutes x = 12, the top of the plotted range, instead of the assistant's actual x = 25, and wrongly calls that reliable because 12 lies inside the range.
- (c) Yes — £80 is above the boundary, £78 — Method: a value counts as an outlier when it lies more than 1.5 times the interquartile range beyond the nearer quartile; here that means checking it against upper quartile + 1.5 × interquartile range. Working: the interquartile range is 42 − 18 = 24. 1.5 × 24 = 36, and 42 + 36 = 78, so any saving above £78 is an outlier. Amara saved £80, and 80 is greater than 78. Answer: yes, Amara's saving is an outlier, because £80 is above the outlier boundary, £78. Watch how you build the boundary and what you compare it with: adding the two quartiles instead of subtracting them, 42 + 18 = 60, gives an interquartile range three times too big, and 42 + 1.5 × 60 = 42 + 90 = 132 puts the boundary so far out that £80 wrongly looks ordinary; comparing £80 with the upper quartile alone, £42, checks only that it lies in the top quarter of the data, which every value above £42 does, not that it lies unusually far beyond it; and adding the interquartile range on once instead of one and a half times, 42 + 24 = 66, uses the wrong multiplier, even though £80 still happens to clear that lower boundary too.
- (c) 75 — Method: the number in a class is the area of its bar, frequency density × class width, so work out the frequency of each class that lies at or above 10 minutes and add them. Working: the class 10 ≤ t < 25 is 15 minutes wide with a frequency density of 3.2, giving 3.2 × 15 = 48 members; the class 25 ≤ t < 55 is 30 minutes wide with a frequency density of 0.9, giving 0.9 × 30 = 27 members; the total charged is 48 + 27 = 75. Answer: 75 members pay the extra charge. The distractors: 4.1 comes from adding the two frequency densities, 3.2 + 0.9, as though each height were a count; 93 comes from including the class 0 ≤ t < 10 as well, 1.8 × 10 = 18 added to 48 and 27, which charges every member; 27 comes from using only the class 25 ≤ t < 55 and forgetting that 10 ≤ t < 25 is also at or above 10 minutes.
- (c) 15 — Method: the interquartile range is the upper quartile take away the lower quartile, IQR = Q3 − Q1. Working: n + 1 = 15 + 1 = 16, so the lower quartile sits at position 16 ÷ 4 = 4, the 4th value in the list, which is 12; the upper quartile sits at position 3 × 4 = 12, the 12th value, which is 27. So 27 − 12 = 15. Answer: the interquartile range is 15 minutes. Watch which values you use and which way round: taking the smallest time away from the largest, 34 − 5 = 29, finds the range, which uses every value between the extremes rather than just the middle half; taking the median away from the upper quartile instead of the lower quartile, 27 − 19 = 8, swaps the median in for the lower quartile; and reversing the subtraction, 12 − 27 = −15, finds the right two values but in the wrong order — an interquartile range is never negative.
- (b) Q was faster on average and more consistent — Method: compare the medians for the average and the interquartile ranges for the spread, remembering that a shorter time is faster and a smaller interquartile range means more consistent. Working: the median for class Q is 35 seconds against 38 seconds for class P, so class Q was faster on average; the interquartile range for class P is 46 − 24 = 22 seconds and for class Q it is 44 − 30 = 14 seconds, so class Q's times are more tightly grouped. Answer: class Q was faster on average and more consistent. The distractors: calling Q slower comes from comparing the lower quartiles, 30 against 24, as though a quartile were the average; calling Q less consistent comes from using the gap between the median and the upper quartile as the spread, 44 − 35 = 9 against 46 − 38 = 8, instead of the full interquartile range; the statement that Q was both slower and less consistent comes from making both of those mistakes together.
- (c) 76 — Method: the frequency of each class is the area of its bar, frequency density × class width, so work out all three frequencies and add them. Working: the widths are 5, 10 and 25 kg, so the frequencies are 4 × 5 = 20, 2.6 × 10 = 26 and 1.2 × 25 = 30; the total is 20 + 26 + 30 = 76. Answer: 76 dogs. The distractors: 7.8 comes from adding the three frequency densities, 4 + 2.6 + 1.2, treating each height as though it were a count; 107 comes from using each upper class boundary as the width, giving 4 × 5, 2.6 × 15 and 1.2 × 40; 39 comes from using the first class width, 5, for every bar, which ignores the unequal intervals and gives 20 + 13 + 6.
Build your own mix at the worksheet builder.