Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Non-calculator
Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- 1.Two classes sat the same maths test, both marked out of 100. Class A had a mean mark of 70 and a range of 30 marks. Class B had a mean mark of 70 and a range of 10 marks. Compare the marks of the two classes.
- 2.A dual bar chart shows the number of hours of rain recorded in Leeds and in Bristol on each of four days. Leeds: Monday 3 hours, Tuesday 5 hours, Wednesday 2 hours, Thursday 4 hours. Bristol: Monday 4 hours, Tuesday 4 hours, Wednesday 6 hours, Thursday 2 hours. Work out the greatest amount, in hours, by which Bristol's rainfall exceeded Leeds's rainfall on a single day.
- 3.The masses, m kg, of 60 parcels are grouped like this: 0 ≤ m < 5, 22 parcels; 5 ≤ m < 10, 20 parcels; 10 ≤ m < 20, 9 parcels; 20 ≤ m < 30, 5 parcels; 30 ≤ m < 50, 4 parcels. Write down the class interval that contains the median mass.
- 4.A school has 1,000 pupils. A student wants to estimate how many of them walk to school, so she asks 8 pupils in her own class. She says her sample is large enough to give a reliable estimate for the whole school. Is she right? Give a reason for your answer.
- 5.A stacked bar for one day at a café in Norwich shows total drink sales of 50 drinks, split into three parts: 22 were tea, 15 were coffee and the rest were hot chocolate. Work out the number of hot chocolates sold.
- 6.A council wants to find out what all residents of a town think about a new cycle lane. It posts a survey only on its website and asks people to fill it in online. Give a reason why this sample is likely to be biased.
- 7.A charity shop in Bath holds 2,000 books. Volunteer A checks a random sample of 50 books and finds 35 paperbacks. Volunteer B checks a different random sample of 50 books and finds 31 paperbacks. Work out the estimate each sample gives for the whole stock, and write down what the shop should do next.
- 8.Priya travels to work by one of two routes and records her journey times, in minutes, over several weeks. Route 1 has a median of 34 minutes and an interquartile range of 22 minutes. Route 2 has a median of 41 minutes and an interquartile range of 6 minutes. Priya wants the more reliable route for getting to an important meeting on time. Which route should she choose, and why?
- 9.A Year 10 class has 20 boys with a mean height of 150 cm and 10 girls with a mean height of 168 cm. Work out the mean height of all 30 pupils in the class.
- 10.A quality inspector weighs a random sample of 50 packets of crisps from one day's production and finds their mean mass is 32.4 g. The factory makes 20,000 packets that day. Work out an estimate for the total mass, in kg, of all the packets made that day.
- 11.In a histogram of the times, t minutes, taken by some people to complete a task, the class 15 ≤ t < 30 contains 24 people. Work out the frequency density for this class.
- 12.A histogram is drawn for the masses, m grams, of 200 letters. The bar for 0 ≤ m < 50 has a frequency density of 1.2 per gram and the bar for 50 ≤ m < 100 has a frequency density of 1.8 per gram. All the remaining letters lie in the class 100 ≤ m < 200. Work out the frequency density of the bar for 100 ≤ m < 200.
- 13.The weekly wages of the five people who work at a small garage in Norwich are £420, £440, £460, £480 and £1,500. Write down which average better describes a typical wage at this garage, and give a reason for your answer.
- 14.Work out the median of these six numbers: 13, 21, 22, 36, 37, 47
- 15.A box plot for the ages of 40 members of a gym is drawn from this five-number summary: minimum 15, lower quartile 22, median 29, upper quartile 38, maximum 61. Work out the interquartile range shown by this box plot.
Answer key
- (a) The means are equal, and Class B's marks are the more consistent because its range is smaller. — Method: comparing two distributions needs two things — a measure of average and a measure of spread — and each must be put into the context of the question. Working: both classes have a mean mark of 70, so on average the two classes scored the same; the range measures spread, and Class A's range of 30 marks is three times Class B's range of 10 marks, so Class B's marks sit closely around the mean while Class A's are far more spread out. Answer: the means are equal, and Class B's marks are the more consistent because its range is smaller. The distractors: the reply crediting Class A with more consistency reverses the meaning of the range, treating a larger range as tighter data when a larger range means more spread; the reply that Class A's mean mark is higher compares the wrong pair of figures, reading the range of 30 as an average; the reply that Class B's mean mark is higher reads the spread correctly but its claim about the means is false, since both means are 70.
- (b) 4 — The difference, Bristol minus Leeds, on each day is: Monday 4 − 3 = 1, Tuesday 4 − 5 = −1, Wednesday 6 − 2 = 4, Thursday 2 − 4 = −2. The greatest amount by which Bristol exceeded Leeds is 4 hours, on Wednesday. Choosing 1 takes Monday's smaller positive difference instead of the greatest one. Choosing 2 takes the size of Thursday's difference, but that is the amount by which Leeds exceeded Bristol, the opposite direction to the one asked for. Choosing 6 takes Bristol's raw figure on Wednesday without subtracting Leeds's 2 hours first.
- (b) 5 ≤ m < 10 — Method: with 60 values the median is the 60 ÷ 2 = 30th value in order, so build a running total until it first reaches 30. Working: the running totals are 22 after the first class, 22 + 20 = 42 after the second, 51 after the third, 56 after the fourth and 60 after the fifth; the 30th parcel is past 22 but not past 42, so it lies in the second class. Answer: the median lies in the class 5 ≤ m < 10. The distractors: 0 ≤ m < 5 comes from giving the class with the greatest frequency, 22, which is the modal class and not the median class; 10 ≤ m < 20 comes from choosing the middle class in the list of five instead of counting to the middle value; 20 ≤ m < 30 comes from halving the range of the data, 50 ÷ 2 = 25, and giving the class that contains 25 kg rather than the class that contains the 30th parcel.
- (c) No — 8 from one class is too small to represent the school. — Method: judge reliability by asking whether the sample is both large enough, and spread across the population, relative to what it is meant to represent. Working: 8 pupils is a tiny fraction of the school's 1,000 pupils, and all 8 come from a single class rather than a range of year groups, so the sample is both too small and too narrow to represent the whole school reliably. She is not right. Saying any sample size gives an equally reliable estimate ignores that reliability generally improves with a larger, more representative sample. Saying the method is unreliable because it was not done online is not a reason connected to sample size or representativeness at all. Saying 8 is reliable because it is more than half her class compares the sample to the wrong population — the school has 1,000 pupils, not one class. Always judge a sample's size against the population it is meant to represent, not against a smaller group within it.
- (b) 13 — The total is 50, and the two known parts are 22 (tea) and 15 (coffee), so 50 − 22 − 15 = 13 hot chocolates. Choosing 28 comes from 50 − 22, subtracting only the tea and forgetting the coffee. Choosing 35 comes from 50 − 15, subtracting only the coffee and forgetting the tea. Choosing 37 comes from 22 + 15, which finds how many drinks were tea or coffee, not the number left over for hot chocolate.
- (d) Only internet users reach the website; others are excluded. — Method: a sample is biased when it systematically leaves out part of the population, or systematically over-represents another part. Working: anyone without internet access, or who does not visit the council's website, has NO chance of being included — the sample is drawn only from internet-using residents, which is not the whole town. Saying too many people might respond because the survey is free confuses bias with sample size — bias is about who CAN be reached, not how many respond. Saying people might lie describes a different problem, response honesty, not who was sampled in the first place. Saying online surveys cannot be anonymous is not a reason connected to bias at all. A sample is biased when part of the population has no chance of being included, whatever the reason for that.
- (a) 1,400 and 1,240, so combine the samples for one estimate — Method: scale each sample up to the whole stock, then use the fact that a larger sample gives a more reliable estimate than a smaller one. Working: the first sample gives 35 ÷ 50 = 0.7 and 0.7 × 2,000 = 1,400 paperbacks; the second gives 31 ÷ 50 = 0.62 and 0.62 × 2,000 = 1,240 paperbacks. Two random samples of the same size are expected to differ a little, so neither estimate is wrong. Putting the two together gives 35 + 31 = 66 paperbacks in 100 books, and 66 ÷ 100 = 0.66 with 0.66 × 2,000 = 1,320, an estimate resting on twice as many books as either volunteer checked. Answer: 1,400 and 1,240, so combine the samples for one estimate. The distractors: keeping 1,400 because it is larger picks an estimate by its size, when both samples held 50 books and neither has a stronger claim; saying a volunteer must have miscounted assumes two random samples ought to agree exactly, which is precisely what random sampling does not promise; 1,750 and 1,550 come from 35 × 50 = 1,750 and 31 × 50 = 1,550, multiplying each count by the size of the sample instead of scaling by 2,000 ÷ 50.
- (c) Route 2, because its interquartile range is smaller — Method: for a journey where turning up on time matters, what matters is not the typical (median) time but how predictable it is — a smaller interquartile range means the middle half of journeys cluster closer together. Working: Route 1's median, 34 minutes, is in fact lower than Route 2's, 41 minutes, so Route 1 is faster on average; but Route 1's interquartile range, 22 minutes, is far larger than Route 2's, 6 minutes, so Route 1's times are much less predictable. Answer: Priya should choose Route 2, because its interquartile range is smaller, even though it is slower on average. Watch which statistic answers the question actually asked: Route 1 does not have the smaller interquartile range, Route 2 does, so picking Route 1 for that reason misreads the table; Route 1's median genuinely is the lower one, but a lower median answers 'which is faster', not 'which is more reliable'; and Route 2's median is not the lower one, so that claim about Route 2 is simply false.
- (c) 156 cm — Method: to combine two groups' means, multiply each group's mean by its own number of pupils, add the two totals together, then divide by the total number of pupils in both groups. Working: 20 × 150 = 3,000 cm for the boys and 10 × 168 = 1,680 cm for the girls, giving a combined total of 3,000 + 1,680 = 4,680 cm. Dividing by all 30 pupils gives 4,680 ÷ 30 = 156 cm. Giving 159 cm averages the two means, (150 + 168) ÷ 2, treating the two groups as if they had the same number of pupils, when there are twice as many boys as girls. Giving 4,680 cm finds the correct combined total height but stops there, forgetting the final division by the 30 pupils. Giving 234 cm divides the combined total by 20, the number of boys only, forgetting that the total also includes the 10 girls. Always weight each mean by its own group size, and always divide by the TOTAL number of pupils in both groups combined.
- (d) 648 kg — Method: to estimate a total from a sample, multiply the sample's mean by the number of items in the whole population, then check the units the question asks for. Working: 32.4 g × 20,000 = 648,000 g. Converting to kilograms, 648,000 ÷ 1,000 = 648 kg. This is only an estimate, not an exact total, because it assumes every one of the 20,000 packets has exactly the sample mean mass, when in reality individual packets vary above and below it. Giving 1.62 kg multiplies the mean by 50, the SAMPLE size, instead of by 20,000, the number of packets actually made that day — this finds the total mass of the 50 sampled packets, not the day's production. Giving 32.4 kg treats the sample mean itself, in grams, as if it already were the day's total mass in kilograms, skipping the scaling up altogether. Giving 648,000 kg correctly scales the mean up to the whole day's production but never converts the answer from grams to kilograms, leaving it 1,000 times too large. Always scale a sample's mean up by the SIZE OF THE WHOLE POPULATION, and always finish by checking the units the question asks for.
- (a) 1.6 — Method: on a histogram the height of a bar is the frequency density, and frequency density = frequency ÷ class width. Working: the class 15 ≤ t < 30 runs from 15 to 30, so its width is 30 − 15 = 15 minutes; the frequency is 24, so the frequency density is 24 ÷ 15 = 1.6. Answer: 1.6 people per minute. The distractors: 360 comes from multiplying the frequency by the class width, 24 × 15, which uses the area rule backwards — area gives the frequency, so the frequency must be divided by the width to give the height; 0.625 comes from dividing the class width by the frequency, 15 ÷ 24, reversing the formula; 0.8 comes from dividing by the upper class boundary, 24 ÷ 30, instead of by the width of the class.
- (c) 0.5 per gram — Method: turn the two known bars into frequencies using area, subtract from the total to find how many letters are left, then divide that frequency by the width of the last class to get its height. Working: the first bar covers 50 g at a frequency density of 1.2, giving 1.2 × 50 = 60 letters, and the second covers 50 g at 1.8, giving 1.8 × 50 = 90 letters; together that is 60 + 90 = 150 letters, so 200 − 150 = 50 letters remain; the class 100 ≤ m < 200 is 100 g wide, so its frequency density is 50 ÷ 100 = 0.5 per gram. Answer: 0.5 per gram. The distractors: 0.25 per gram comes from dividing the remaining 50 letters by the upper class boundary, 200, instead of by the class width of 100; 2 per gram comes from dividing the class width by the frequency, 100 ÷ 50, reversing the formula; 1.4 per gram comes from subtracting only the first bar's 60 letters, leaving 140, and then dividing by 100.
- (a) The median, as the one very large wage does not move it — Method: an average describes a population well when it sits close to most of the values, so compare what each average does when one value lies far from the rest. Working: in order the wages are 420, 440, 460, 480 and 1,500, so the median is the third of the five, £460. The mean uses every wage: 420 + 440 + 460 + 480 + 1,500 = 3,300 and 3,300 ÷ 5 = 660, so the mean is £660. Four of the five people earn less than £660, and the nearest of those four wages is £180 below it, so £660 describes nobody at the garage; £460 sits inside the group of four similar wages. Answer: the median, as the one very large wage does not move it, while that same wage drags the mean £200 above the median. The distractors: saying the median is always larger than the mean is an invented rule, and here the median £460 is smaller than the mean £660; saying the mean is the only average that uses all five wages is true as far as it goes, but using a value and being dragged by it are the same thing when that value is £1,500; saying £660 lies between the smallest and largest wage is true of every mean ever calculated, so it proves nothing about whether this one is typical.
- (d) 29 — Method: with an even number of values there is no single middle value, so the median is the mean of the two values either side of the middle. Working: the six numbers are already in order and 6 ÷ 2 = 3, so the middle pair are the third and fourth values, 22 and 36; their mean is (22 + 36) ÷ 2 = 58 ÷ 2 = 29. Answer: 29, which lies between the two middle values as a median of an even data set must. The distractors: 22 comes from taking the lower of the two middle values and stopping there instead of averaging the pair; 36 comes from taking the larger value of that pair because it sits just past the halfway point of the list; 34 comes from working out the range, 47 − 13, instead of a measure of centre.
- (d) 16 — Method: on a box plot the interquartile range is the width of the box itself, upper quartile take away lower quartile. Working: the upper quartile is 38 and the lower quartile is 22, so 38 − 22 = 16. Answer: the interquartile range is 16 years. Watch which part of the box plot you are reading: the whole line from whisker to whisker gives the range, 61 − 15 = 46; the left half of the box alone gives median take away lower quartile, 29 − 22 = 7; and the right half of the box alone gives upper quartile take away median, 38 − 29 = 9 — neither half is the interquartile range on its own.
Build your own mix at the worksheet builder.