Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Non-calculator
Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- 1.Grace asked 12 children how many brothers and sisters they have. Her results, in order, were 0, 0, 1, 1, 1, 1, 2, 2, 3, 3, 4, 6. Work out the median number of brothers and sisters.
- 2.The seven members of Team A took 20, 21, 22, 23, 24, 25 and 40 seconds to finish a task. The seven members of Team B took 20, 30, 32, 34, 36, 38 and 40 seconds. Tomás says that because the two teams have the same range, their times are spread out in the same way. Is Tomás right? Give a reason for your answer.
- 3.A school has 1200 pupils. A teacher wants to take a random sample of 60 of them. Write down which of these methods gives a random sample.
- 4.Work out the median of these six numbers: 13, 21, 22, 36, 37, 47
- 5.A pictogram shows the number of cars sold by a garage in Plymouth each month, where each whole car symbol represents 8 cars sold and a half symbol represents 4 cars. March shows 3 whole symbols and one half symbol. Work out how many cars were sold in March.
- 6.The heights, h cm, of 80 plants are grouped like this: 0 ≤ h < 20, 14 plants; 20 ≤ h < 40, 22 plants; 40 ≤ h < 50, 16 plants; 50 ≤ h < 80, 28 plants. Write down the class interval that contains the lower quartile.
- 7.A vet records the masses, m kg, of 30 dogs at a clinic in Preston: 0 < m ≤ 10 — 11 dogs, 10 < m ≤ 20 — 5 dogs, 20 < m ≤ 30 — 5 dogs, 30 < m ≤ 40 — 9 dogs. Work out an estimate for the mean mass, in kg, using the midpoint of each class interval.
- 8.The times, t minutes, of 80 journeys are summarised by these cumulative frequencies: t < 10, 8 journeys; t < 20, 28 journeys; t < 30, 52 journeys; t < 40, 72 journeys; t < 50, 80 journeys. Estimate the interquartile range.
- 9.In a spelling test the 20 pupils in Group A had a mean mark of 80, and the 30 pupils in Group B had a mean mark of 70. Work out the mean mark of all 50 pupils.
- 10.A vet records the masses, m kg, of the dogs seen in one week as a histogram. The bar for 0 ≤ m < 5 has a frequency density of 4 per kg, the bar for 5 ≤ m < 15 has a frequency density of 2.6 per kg, and the bar for 15 ≤ m < 40 has a frequency density of 1.2 per kg. Work out the total number of dogs seen that week.
- 11.A charity shop in Bath holds 2,000 books. Volunteer A checks a random sample of 50 books and finds 35 paperbacks. Volunteer B checks a different random sample of 50 books and finds 31 paperbacks. Work out the estimate each sample gives for the whole stock, and write down what the shop should do next.
- 12.A sports shop in Cardiff sold 40 pairs of football boots last month: 4 pairs of size 6, 5 pairs of size 7, 8 pairs of size 8, 13 pairs of size 9 and 10 pairs of size 10. The manager will order 40 pairs for next month and wants as many pairs as possible to be in a size customers will buy. Work out the mean size and the modal size, and write down which of the two he should use.
- 13.A café in York counts the number of customers in each of the nine hours it is open on one day: 4, 4, 4, 11, 13, 15, 18, 22 and 25. The owner says that a typical hour has about 4 customers, because 4 is the mode. Is the owner right? Give a reason for your answer.
- 14.A scatter graph shows the number of days, x, that each of 16 tomato plants was watered and its height, y cm. The line of best fit has equation y = 1.5x + 4. Write down what the 1.5 in this equation tells you about the plants.y = 1.5x + 4
- 15.A frequency polygon for the mass, in kg, of 40 parcels at a delivery depot is drawn by plotting one point at the midpoint of each class, joined by straight lines: (5, 6), (15, 10), (25, 16), (35, 6), (45, 2). Every class has a width of 10 kg. Write down the modal class.
Answer key
- (a) 1.5 — Method: with an even number of values the median is the mean of the two middle values, which for 12 values are the 6th and the 7th once the data are in order. Working: the results are already in order, and 12 ÷ 2 = 6, so the middle pair are the 6th value, 1, and the 7th value, 2; the median is (1 + 2) ÷ 2 = 1.5. Answer: 1.5 brothers and sisters. The distractors: 1 comes from reading the 6th value and stopping there instead of averaging the middle pair; 2 comes from working out the mean, 24 ÷ 12, instead of the median; 6 comes from working out the range, 6 − 0, which measures spread rather than centre.
- (d) No, the range uses only the fastest and slowest time — Method: check what the range is built from, then look at what it leaves out. Working: both teams have a fastest time of 20 seconds and a slowest of 40 seconds, so both ranges are 40 − 20 = 20 seconds and Tomás has that part right. But the range is calculated from those two values alone. Six of Team A's seven times lie between 20 and 25 seconds, with a single time far out at 40; Team B is the other way round, with six of its seven times at 30 seconds or more and a single time far out at 20. So Team A bunches at the fast end and Team B at the slow end. The two patterns are quite different, and the range cannot see the difference because the five middle times never enter the calculation. Answer: no, because the range uses only the fastest and slowest time. The distractors: comparing the means answers a different question, since a mean measures position rather than spread, and two sets with the same spread can have different means; saying that equal ranges mean equal spread is the very assumption that fails here; saying that seven times each forces the spreads to match confuses the size of a data set with how its values are arranged inside it.
- (a) Drawing 60 names at random from a list of all 1200 pupils — Method: a sample is random when every member of the population has the same chance of being chosen and nobody, including the pupils themselves, can influence who ends up in it; test each method against that. Working: drawing names from a list of all 1200 pupils gives each pupil the same chance, 60 out of 1200, whatever their year group, class or opinion, so the method is random. Answer: drawing 60 names at random from a list of all 1200 pupils. The distractors: asking the pupils who volunteer is self-selection, and the pupils with the strongest views volunteer first, so they decide the sample; asking the pupils nearest the door is convenience sampling, which reaches only those who happen to be in one place at one time; asking two Year 10 classes samples a cluster, so every pupil in the other year groups has no chance of being chosen at all.
- (d) 29 — Method: with an even number of values there is no single middle value, so the median is the mean of the two values either side of the middle. Working: the six numbers are already in order and 6 ÷ 2 = 3, so the middle pair are the third and fourth values, 22 and 36; their mean is (22 + 36) ÷ 2 = 58 ÷ 2 = 29. Answer: 29, which lies between the two middle values as a median of an even data set must. The distractors: 22 comes from taking the lower of the two middle values and stopping there instead of averaging the pair; 36 comes from taking the larger value of that pair because it sits just past the halfway point of the list; 34 comes from working out the range, 47 − 13, instead of a measure of centre.
- (c) 28 — Method: multiply the number of whole symbols by the value of one symbol, then add the value of any half symbol shown. Working: 3 whole symbols represent 3 × 8 = 24 cars. The half symbol represents 4 cars. Total cars sold in March = 24 + 4 = 28. Leaving out the half symbol, 3 × 8 = 24, undercounts by exactly the value of that half symbol. Treating the half symbol as if it were a full symbol, 4 × 8 = 32, overcounts because it doubles the value the half symbol is worth. Giving 3.5 reports the number of symbols shown, not the number of cars they represent — the key still needs to be applied. Always apply the key to every symbol shown, including a half symbol, rather than reading off the symbol count itself.
- (c) 20 ≤ h < 40 — Method: with 80 values the lower quartile is the 80 ÷ 4 = 20th value in order, so build a running total until it first reaches 20. Working: the running totals are 14, then 14 + 22 = 36, then 52, then 80; the 20th plant is past 14 but not past 36, so it lies in the second class. Answer: the lower quartile lies in the class 20 ≤ h < 40. The distractors: 0 ≤ h < 20 comes from believing that the bottom quarter of the data must all sit in the first class, when that class holds only 14 of the 80 plants; 40 ≤ h < 50 comes from using the position 80 ÷ 2 = 40 and so locating the median rather than the lower quartile; 50 ≤ h < 80 comes from counting 20 plants down from the tallest instead of up from the shortest, which locates the upper quartile at the 60th plant.
- (d) 19 kg — Method: estimate the mean of grouped data by multiplying each class's midpoint by its frequency, adding the four totals, then dividing by the total frequency. Working: the midpoints are 5, 15, 25 and 35 kg. The weighted totals are 11 × 5 = 55, 5 × 15 = 75, 5 × 25 = 125 and 9 × 35 = 315, which add to 570. Dividing by the 30 dogs gives an estimate of 570 ÷ 30 = 19 kg. Giving 5 kg reads off the midpoint of the modal class, 0 < m ≤ 10, the class with the most dogs — but the class with the most dogs is not where the mean falls, and neither is a substitute for actually calculating it. Giving 20 kg averages the four midpoints, (5 + 15 + 25 + 35) ÷ 4, treating every class as equally likely and ignoring that far more dogs are in the lightest and heaviest classes than in the middle two. Giving 570 kg stops after finding the correct weighted total and forgets the final division by the 30 dogs. Always weight each midpoint by its own frequency, and always finish by dividing by the total frequency, not the number of classes.
- (d) 18 minutes — Method: the lower quartile is the 80 ÷ 4 = 20th value and the upper quartile is the 3 × 80 ÷ 4 = 60th value; locate each inside its class by linear interpolation, then subtract. Working: the 20th value lies between the running totals 8 and 28, so it is in the class 10 ≤ t < 20, which holds 20 journeys across 10 minutes, and it is the 20 − 8 = 12th of them, giving 10 + (12 ÷ 20) × 10 = 16 minutes; the 60th value lies between the running totals 52 and 72, so it is in the class 30 ≤ t < 40, which also holds 20 journeys across 10 minutes, and it is the 60 − 52 = 8th of them, giving 30 + (8 ÷ 20) × 10 = 34 minutes; subtracting, 34 − 16 = 18. Answer: an estimated interquartile range of 18 minutes. The distractors: 20 minutes comes from taking the lower boundaries of the two quartile classes, 30 − 10, which locates the classes but never the values inside them; 40 minutes comes from subtracting the two positions, 60 − 20, instead of the two times; 22 minutes comes from interpolating downwards from each upper boundary rather than upwards from each lower boundary, giving 20 − 6 = 14 and 40 − 4 = 36.
- (b) 74 marks — Method: the two groups are different sizes, so their means cannot simply be averaged — rebuild each group's total mark, add the totals and divide by all 50 pupils. Working: Group A scored 20 × 80 = 1600 marks and Group B scored 30 × 70 = 2100 marks, giving 1600 + 2100 = 3700 marks altogether, so the overall mean is 3700 ÷ 50 = 74 marks. Answer: 74 marks, which sits nearer to 70 than to 80 because the larger group scored 70. The distractors: 75 marks comes from averaging the two group means, (80 + 70) ÷ 2, as though the groups were the same size; 76 marks comes from attaching each mean to the other group's size, (20 × 70 + 30 × 80) ÷ 50; 150 marks comes from adding the two means together and never dividing at all.
- (c) 76 — Method: the frequency of each class is the area of its bar, frequency density × class width, so work out all three frequencies and add them. Working: the widths are 5, 10 and 25 kg, so the frequencies are 4 × 5 = 20, 2.6 × 10 = 26 and 1.2 × 25 = 30; the total is 20 + 26 + 30 = 76. Answer: 76 dogs. The distractors: 7.8 comes from adding the three frequency densities, 4 + 2.6 + 1.2, treating each height as though it were a count; 107 comes from using each upper class boundary as the width, giving 4 × 5, 2.6 × 15 and 1.2 × 40; 39 comes from using the first class width, 5, for every bar, which ignores the unequal intervals and gives 20 + 13 + 6.
- (a) 1,400 and 1,240, so combine the samples for one estimate — Method: scale each sample up to the whole stock, then use the fact that a larger sample gives a more reliable estimate than a smaller one. Working: the first sample gives 35 ÷ 50 = 0.7 and 0.7 × 2,000 = 1,400 paperbacks; the second gives 31 ÷ 50 = 0.62 and 0.62 × 2,000 = 1,240 paperbacks. Two random samples of the same size are expected to differ a little, so neither estimate is wrong. Putting the two together gives 35 + 31 = 66 paperbacks in 100 books, and 66 ÷ 100 = 0.66 with 0.66 × 2,000 = 1,320, an estimate resting on twice as many books as either volunteer checked. Answer: 1,400 and 1,240, so combine the samples for one estimate. The distractors: keeping 1,400 because it is larger picks an estimate by its size, when both samples held 50 books and neither has a stronger claim; saying a volunteer must have miscounted assumes two random samples ought to agree exactly, which is precisely what random sampling does not promise; 1,750 and 1,550 come from 35 × 50 = 1,750 and 31 × 50 = 1,550, multiplying each count by the size of the sample instead of scaling by 2,000 ÷ 50.
- (a) The modal size, 9, bought by more customers than any other — Method: work out both averages from the frequencies, then choose the one the shop can act on. Working: for the mean, multiply each size by the number of pairs sold at it and add: 6 × 4 + 7 × 5 + 8 × 8 + 9 × 13 + 10 × 10 = 340, and 340 ÷ 40 = 8.5, so the mean size is 8.5. The largest frequency is 13, which belongs to size 9, so the modal size is 9. The mean 8.5 is a size no customer in the record asked for, so 40 pairs of it would sit unsold, while 13 of the 40 customers wanted size 9, more than wanted any other size. Answer: the modal size, 9, bought by more customers than any other. The distractors: the mean size 8.5 does take account of all 40 pairs, but a mean of sizes is a summary figure and not a size the month's customers were buying; the mean size 8 comes from averaging the five sizes on sale, 6 + 7 + 8 + 9 + 10 = 40 and 40 ÷ 5 = 8, which ignores how many pairs were sold at each size and so treats the 4 pairs of size 6 as equal in weight to the 13 pairs of size 9; the range 4 comes from 10 − 6 and measures spread, so it says how wide a set of sizes the shop must stock, not which size to stock most of.
- (c) No, the mode here is the lowest value of the nine — Method: an average is meant to stand for the data as a whole, so test any proposed average by asking how many values it sits near. Working: the value 4 appears three times and every other count appears once, so 4 is indeed the mode. But those three hours are the quiet ones at the start of the day, and the other six counts run from 11 up to 25; putting the nine counts in order, the middle one is the fifth, which is 13. So the mode sits at the very bottom of the data, with six of the nine hours far above it. Answer: no, because the mode here is the lowest value of the nine, so it describes the quiet opening hours rather than a typical hour. The distractors: saying the mode can only be used when no value repeats reverses the definition, since a mode exists only because a value does repeat; saying the mode is the value that occurs most often is a correct definition, but being the commonest value does not make a value typical when it lies at one end of the data; saying the mode is the best average for any list of numbers ignores the fact that mean, median and mode each describe a population well in different circumstances.
- (b) On average a plant grew 1.5 cm taller for each extra day — Method: in the equation of a line, the number multiplying x is the gradient, and a gradient states the change in y produced by an increase of 1 in x, read in the units of the two axes. Working: here x is measured in days and y in centimetres, so the gradient 1.5 carries the units centimetres per day. Testing it on the line, 5 days gives 1.5 × 5 + 4 = 11.5 cm and 6 days gives 1.5 × 6 + 4 = 13 cm, a rise of 1.5 cm for the one extra day. Answer: on average a plant grew 1.5 cm taller for each extra day of watering. The distractors: 1.5 cm as the height before any watering is the value of y when x is 0, which is the other number in the equation, 4 cm, so this swaps the gradient and the intercept; 1.5 cm as the gap between the tallest and the shortest plant reads the gradient as a range, when a range is a difference between two of the 16 plants and a gradient is a rate; 1.5 days for each extra centimetre inverts the rate, dividing days by centimetres instead of centimetres by days, and the line gives 1 cm of growth in two thirds of a day.
- (c) 20 kg ≤ mass < 30 kg — The modal class is the class with the highest frequency. Reading the plotted points, the frequencies are 6, 10, 16, 6 and 2, so the highest frequency is 16, plotted at the midpoint 25. A class of width 10 centred on 25 runs from 25 − 5 = 20 to 25 + 5 = 30, so the modal class is 20 kg ≤ mass < 30 kg. Writing '25 kg' gives only the midpoint, not the class — the modal class is an interval, not a single value. '10 kg ≤ mass < 20 kg' is the class before the peak, centred on 15, which has frequency 10, not the highest. '30 kg ≤ mass < 40 kg' is the class after the peak, centred on 35, which has frequency 6, not the highest.
Build your own mix at the worksheet builder.