Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- 1.A stacked bar for one day at a café in Norwich shows total drink sales of 50 drinks, split into three parts: 22 were tea, 15 were coffee and the rest were hot chocolate. Work out the number of hot chocolates sold.
- 2.Amelia's mean mark over four tests is 29. Her first three marks are 31, 26 and 26. Work out her fourth mark.
- 3.Five friends have heights, in cm, of 150, 152, 155, 158 and 160. A sixth friend, with a height of 170 cm, joins the group. Write down what happens to the mean and the range of the heights once this sixth friend is included.
- 4.The mean of the four numbers 10, 15, 20 and x is 18. Work out the value of x.
- 5.The times, t minutes, taken by 80 people to travel to work are grouped like this: 0 ≤ t < 10, 6 people; 10 ≤ t < 20, 14 people; 20 ≤ t < 30, 25 people; 30 ≤ t < 40, 20 people; 40 ≤ t < 50, 15 people. Work out the cumulative frequency for t < 30.
- 6.A scatter graph shows the height, x cm, and the mass, y kg, of 20 pupils in Year 10. The heights on the graph run from 150 cm to 180 cm, and the line of best fit is y = 0.9x − 85. Nadia puts x = 90 into this equation to estimate the mass of a two-year-old child who is 90 cm tall. Is her estimate reliable? Give a reason for your answer.y = 0.9x − 85
- 7.The mean mass of four parcels is 17 kg. Three of the parcels have masses 12 kg, 16 kg and 18 kg. Work out the mass of the fourth parcel.
- 8.A dual bar chart shows the number of hours of rain recorded in Leeds and in Bristol on each of four days. Leeds: Monday 3 hours, Tuesday 5 hours, Wednesday 2 hours, Thursday 4 hours. Bristol: Monday 4 hours, Tuesday 4 hours, Wednesday 6 hours, Thursday 2 hours. Work out the greatest amount, in hours, by which Bristol's rainfall exceeded Leeds's rainfall on a single day.
- 9.The midday temperature in Leeds was recorded on each day of one week: 22 °C, 24 °C, 23 °C, 25 °C, 26 °C, 21 °C, 24 °C. Work out the mean midday temperature, giving your answer to 2 decimal places.
- 10.The distances, d km, cycled by 180 riders in a charity sportive are summarised by these cumulative frequencies: d < 30, 20 riders; d < 60, 60 riders; d < 80, 120 riders; d < 100, 160 riders; d < 130, 180 riders. Use interpolation to estimate the median distance cycled.
- 11.A composite (stacked) bar for a charity bake sale in Durham shows the number of cakes sold, split into three types. The bar has a total height of 80 cakes: 34 were sponge cakes, 26 were chocolate cakes and the rest were fruit cakes. Work out the percentage of the cakes sold that were fruit cakes.
- 12.A café in York counts the number of customers in each of the nine hours it is open on one day: 4, 4, 4, 11, 13, 15, 18, 22 and 25. The owner says that a typical hour has about 4 customers, because 4 is the mode. Is the owner right? Give a reason for your answer.
- 13.A random sample of 50 pupils at a school were asked whether they prefer sport to music. 30 of the 50 pupils said they prefer sport. The school has 500 pupils altogether. Work out an estimate for the number of the 500 pupils who prefer sport.
- 14.A vet records the masses, m kg, of the dogs seen in one week as a histogram. The bar for 0 ≤ m < 5 has a frequency density of 4 per kg, the bar for 5 ≤ m < 15 has a frequency density of 2.6 per kg, and the bar for 15 ≤ m < 40 has a frequency density of 1.2 per kg. Work out the total number of dogs seen that week.
- 15.Four pairs of variables are listed below. Write down the pair that you would expect to show negative correlation.
Answer key
- (b) 13 — The total is 50, and the two known parts are 22 (tea) and 15 (coffee), so 50 − 22 − 15 = 13 hot chocolates. Choosing 28 comes from 50 − 22, subtracting only the tea and forgetting the coffee. Choosing 35 comes from 50 − 15, subtracting only the coffee and forgetting the tea. Choosing 37 comes from 22 + 15, which finds how many drinks were tea or coffee, not the number left over for hot chocolate.
- (c) 33 — Method: multiply the mean by the number of tests to get the total marks, then subtract the marks that are already known. Working: four tests with a mean of 29 give a total of 29 × 4 = 116 marks; the first three marks total 31 + 26 + 26 = 83; so the fourth mark is 116 − 83 = 33. Answer: 33, and checking, (31 + 26 + 26 + 33) ÷ 4 = 116 ÷ 4 = 29. The distractors: 116 comes from stopping at the total for all four tests; 29 comes from assuming the missing mark must be the mean itself; 4 comes from multiplying the mean by 3, the number of marks given, leaving 87 − 83 = 4.
- (d) Both the mean and the range increase. — The original mean is 150 + 152 + 155 + 158 + 160 = 775, and 775 ÷ 5 = 155 cm; the original range is 160 − 150 = 10 cm. Including the new height of 170 cm gives a new total of 775 + 170 = 945, and 945 ÷ 6 = 157.5 cm, which is higher than 155 cm, and a new range of 170 − 150 = 20 cm, which is higher than 10 cm, so both the mean and the range increase. Saying the range stays the same ignores that 170 cm is a new, higher maximum than the old 160 cm. Saying the mean stays the same ignores that 170 cm is above the original mean of 155 cm, which pulls the average up. Saying both decrease is the opposite of what happens here.
- (a) 27 — Method: turn the mean into a total using total = mean × number of values, then subtract the numbers that are already known. Working: four numbers with a mean of 18 have a total of 18 × 4 = 72; the three known numbers give 10 + 15 + 20 = 45; so x = 72 − 45 = 27. Answer: 27, and checking, (10 + 15 + 20 + 27) ÷ 4 = 72 ÷ 4 = 18. The distractors: 72 comes from stopping at the total the four numbers must reach and never subtracting the known three; 18 comes from assuming the missing number must equal the mean; 45 comes from stopping at the total of the three known numbers.
- (d) 45 — Method: a cumulative frequency is a running total — it counts everybody in every class up to and including the one that ends at the value given. Working: the classes that lie wholly below 30 minutes are 0 ≤ t < 10, 10 ≤ t < 20 and 20 ≤ t < 30, with frequencies 6, 14 and 25, so the running total is 6 + 14 = 20 and then 20 + 25 = 45. Answer: 45 people took less than 30 minutes. The distractors: 25 comes from quoting the frequency of the class 20 ≤ t < 30 on its own instead of the running total; 65 comes from accumulating one class too many and including 30 ≤ t < 40, which is 45 + 20; 35 comes from accumulating from the top downwards, 15 + 20, which counts the people who took 30 minutes or more rather than fewer.
- (c) No, 90 cm is far outside the heights on the graph — Method: a line of best fit describes the trend only across the stretch of data it was drawn through; predicting beyond that stretch is extrapolation, and nothing in the data supports it. Working: the heights used to draw this line run from 150 cm to 180 cm, all of them Year 10 pupils, while 90 cm is 60 cm below the shortest of them and belongs to a two-year-old child, whose build follows no trend the graph has measured. Substituting anyway gives 0.9 × 90 − 85 = −4, a mass of −4 kg, which cannot exist. Answer: no, because 90 cm is far outside the heights on the graph. The distractors: saying a line of best fit cannot be used to predict at all throws away its main purpose, since a prediction made between the plotted values is perfectly sound; saying the line passes through all 20 points misdescribes a line of best fit, which is drawn to follow the trend of the points and will normally pass through few of them; saying the equation works for any value put into it treats an equation fitted to Year 10 heights as a law of nature, and the mass of −4 kg shows what that assumption produces.
- (c) 22 kg — Method: multiply the mean by the number of parcels to rebuild the total mass, then subtract the masses that are known. Working: four parcels with a mean mass of 17 kg have a total mass of 17 × 4 = 68 kg; the three known parcels total 12 + 16 + 18 = 46 kg; so the fourth parcel has mass 68 − 46 = 22 kg. Answer: 22 kg. The distractors: 68 kg comes from stopping at the total mass of all four parcels; 17 kg comes from assuming the missing parcel must have the mean mass; 5 kg comes from multiplying the mean by 3, the number of parcels whose mass is given, leaving 51 − 46 = 5.
- (b) 4 — The difference, Bristol minus Leeds, on each day is: Monday 4 − 3 = 1, Tuesday 4 − 5 = −1, Wednesday 6 − 2 = 4, Thursday 2 − 4 = −2. The greatest amount by which Bristol exceeded Leeds is 4 hours, on Wednesday. Choosing 1 takes Monday's smaller positive difference instead of the greatest one. Choosing 2 takes the size of Thursday's difference, but that is the amount by which Leeds exceeded Bristol, the opposite direction to the one asked for. Choosing 6 takes Bristol's raw figure on Wednesday without subtracting Leeds's 2 hours first.
- (d) 23.57 °C — Method: add all seven temperatures, divide by the number of readings and round only at the end. Working: 22 + 24 + 23 + 25 + 26 + 21 + 24 = 165, and 165 ÷ 7 = 23.5714…, which rounds to 23.57 to 2 decimal places. Answer: 23.57 °C. The distractors: 24 °C comes from writing down the mode, the only temperature recorded twice, instead of the mean; 27.5 °C comes from dividing the total by 6 instead of by the 7 days recorded; 5 °C comes from working out the range, 26 − 21, which is a measure of spread and not an average.
- (c) 70 — Method: estimate the median from the cumulative frequency table by interpolation: find its position, n ÷ 2, locate the class it falls in, then add the fraction of the way through that class (adjusted for the cumulative frequency reached before it) to the class's lower boundary. Working: there are 180 riders, so the median is at position 180 ÷ 2 = 90. Before the class 60 ≤ d < 80 the cumulative frequency is 60, and by the end of it, 120, so the 90th rider falls in this class; its frequency is 120 − 60 = 60 and its width is 80 − 60 = 20. The extra distance needed into the class is 90 − 60 = 30, and 30 ÷ 60 × 20 = 10, so the median is 60 + 10 = 70. Answer: the estimated median distance is 70 km. Watch which numbers the interpolation actually uses: reading off just the class's lower boundary, 60, ignores how far into the class the 90th rider falls; using the target position, 90, as the extra distance instead of subtracting the 60 riders already counted before the class gives 90 ÷ 60 × 20 = 30, so 60 + 30 = 90, overshooting by treating the whole position as if none of it had already been counted; and using the total number of riders, 180, instead of half of it as the target position lands in the very last class, giving an estimate of 130 km — further than any rider is known to have ridden by that point in the table.
- (d) 25% — First find the number of fruit cakes: 80 − 34 − 26 = 20. Then write this as a percentage of the total: 20 ÷ 80 × 100 = 25%. Giving 20% comes from reporting the count of fruit cakes, 20, directly as a percentage, without dividing by the total of 80 first. Giving 32.5% computes the percentage of chocolate cakes instead of fruit cakes: 26 ÷ 80 × 100 = 32.5%. Giving 42.5% computes the percentage of sponge cakes instead of fruit cakes: 34 ÷ 80 × 100 = 42.5%.
- (c) No, the mode here is the lowest value of the nine — Method: an average is meant to stand for the data as a whole, so test any proposed average by asking how many values it sits near. Working: the value 4 appears three times and every other count appears once, so 4 is indeed the mode. But those three hours are the quiet ones at the start of the day, and the other six counts run from 11 up to 25; putting the nine counts in order, the middle one is the fifth, which is 13. So the mode sits at the very bottom of the data, with six of the nine hours far above it. Answer: no, because the mode here is the lowest value of the nine, so it describes the quiet opening hours rather than a typical hour. The distractors: saying the mode can only be used when no value repeats reverses the definition, since a mode exists only because a value does repeat; saying the mode is the value that occurs most often is a correct definition, but being the commonest value does not make a value typical when it lies at one end of the data; saying the mode is the best average for any list of numbers ignores the fact that mean, median and mode each describe a population well in different circumstances.
- (d) 300 pupils — Method: an estimate for a whole population is made by finding the proportion in the sample and applying that same proportion to the population. Working: in the sample 30 of the 50 pupils prefer sport, a proportion of 30 ÷ 50 = 0.6, and applying that proportion to the school gives 0.6 × 500 = 300 pupils. Answer: 300 pupils, and it is only an estimate, because a different random sample of 50 would give a slightly different figure. The distractors: 200 pupils comes from scaling up the 20 pupils in the sample who did not prefer sport, 20 × 10, which answers the opposite question; 150 pupils comes from reading 30 out of 50 as 30% and taking 30% of 500; 60 pupils comes from working out the proportion correctly as 60% and then writing the 60 down as a number of pupils instead of applying it to the 500.
- (c) 76 — Method: the frequency of each class is the area of its bar, frequency density × class width, so work out all three frequencies and add them. Working: the widths are 5, 10 and 25 kg, so the frequencies are 4 × 5 = 20, 2.6 × 10 = 26 and 1.2 × 25 = 30; the total is 20 + 26 + 30 = 76. Answer: 76 dogs. The distractors: 7.8 comes from adding the three frequency densities, 4 + 2.6 + 1.2, treating each height as though it were a count; 107 comes from using each upper class boundary as the width, giving 4 × 5, 2.6 × 15 and 1.2 × 40; 39 comes from using the first class width, 5, for every bar, which ignores the unequal intervals and gives 20 + 13 + 6.
- (a) Minutes a candle has burned and length remaining — As a candle burns for longer, less of it remains, so these two variables move in opposite directions as one increases — that is negative correlation. A pupil's shoe size generally increases as they get older, so age and shoe size show positive correlation, not negative, since both rise together. A football team's shirt colour is not a numerical quantity linked to how many matches it wins, so shirt colour and number of wins show no correlation at all. The number of letters in a pupil's name has no real connection to their ability in maths, so that pair also shows no correlation.
Build your own mix at the worksheet builder.