Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Non-calculator
Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- 1.The mean of 5 numbers is 8. Work out the total of the 5 numbers.
- 2.The delivery times, in minutes, of 15 parcels are given in order: 5, 8, 10, 12, 14, 15, 17, 19, 21, 23, 25, 27, 29, 31, 34. Work out the interquartile range of these times.
- 3.Five friends have heights, in cm, of 150, 152, 155, 158 and 160. A sixth friend, with a height of 170 cm, joins the group. Write down what happens to the mean and the range of the heights once this sixth friend is included.
- 4.A school has 1,500 pupils. The head teacher takes a random sample of 150 of them from the school register and asks how long they spend on homework. Rory says the sample is too small for the result to mean anything. Is Rory right? Give a reason for your answer.
- 5.The masses, m grams, of 100 apples are grouped like this: 100 ≤ m < 120, 10 apples; 120 ≤ m < 140, 30 apples; 140 ≤ m < 160, 40 apples; 160 ≤ m < 200, 20 apples. Estimate the median mass.
- 6.A scatter graph of the number of ice creams sold at a seaside kiosk in Bournemouth and the number of sunburn cases treated at a nearby pharmacy, recorded on the same 30 days, shows strong positive correlation. Which statement about this correlation is correct?
- 7.The marks scored by 11 pupils in a test are given in order: 12, 15, 18, 21, 24, 27, 30, 33, 36, 39, 42. Work out the lower quartile of these marks.
- 8.Forty pupils in class P and forty pupils in class Q each solved a puzzle. The times, in seconds, were summarised using cumulative frequency. For class P the lower quartile is 24, the median is 38 and the upper quartile is 46. For class Q the lower quartile is 30, the median is 35 and the upper quartile is 44. Write down the statement that correctly compares the two classes.
- 9.For each of 30 fires in one county, the number of fire engines sent and the cost of the damage were recorded. The scatter graph of the data shows strong positive correlation. A newspaper prints the headline "Sending more fire engines causes more damage". Is the newspaper right? Give a reason for your answer.
- 10.A scatter graph shows the midday temperature, x °C, and the number of ice creams sold at a seaside kiosk, y. The line of best fit is y = 3x − 20. Give a reason why the y-intercept of this line of best fit is not a sensible estimate of the number of ice creams sold.y = 3x − 20
- 11.A call centre records the length, t seconds, of 100 calls: 0 ≤ t < 20, 15 calls; 20 ≤ t < 30, 24 calls; 30 ≤ t < 50, 40 calls; 50 ≤ t < 80, 21 calls. The manager's target is for a call to be finished in under 35 seconds. Estimate the number of calls that met the target.
- 12.The mean mass of four parcels is 17 kg. Three of the parcels have masses 12 kg, 16 kg and 18 kg. Work out the mass of the fourth parcel.
- 13.In a scatter graph of the age, in years, and the wingspan, in cm, of 20 birds of the same species, all the points lie close to a rising line of best fit except one, which lies a long way below the line. That bird was later found to have a damaged wing. Give a reason why this point should not be used when drawing the line of best fit.
- 14.A bus company records the delay, d minutes, of 250 buses: 0 ≤ d < 2, 60 buses; 2 ≤ d < 5, 90 buses; 5 ≤ d < 10, 75 buses; 10 ≤ d < 20, 25 buses. The company refunds the fare whenever a bus is more than 8 minutes late. Estimate the number of refunds it must pay.
- 15.The masses of eight school bags, in kilograms, are 3, 4, 4, 5, 6, 7, 8 and 11. Work out the median mass.
Answer key
- (a) 40 — Method: the mean is the total divided by how many values there are, so rearranging gives total = mean × number of values. Working: the mean is 8 and there are 5 numbers, so the total is 8 × 5 = 40. Answer: 40, and checking, 40 ÷ 5 = 8, which is the mean given. The distractors: 13 comes from adding the mean and the count, 8 + 5, instead of multiplying them; 1.6 comes from dividing the mean by the count, 8 ÷ 5, which reverses the relationship; 8 comes from quoting the mean itself as the total, which is only true when there is a single number.
- (c) 15 — Method: the interquartile range is the upper quartile take away the lower quartile, IQR = Q3 − Q1. Working: n + 1 = 15 + 1 = 16, so the lower quartile sits at position 16 ÷ 4 = 4, the 4th value in the list, which is 12; the upper quartile sits at position 3 × 4 = 12, the 12th value, which is 27. So 27 − 12 = 15. Answer: the interquartile range is 15 minutes. Watch which values you use and which way round: taking the smallest time away from the largest, 34 − 5 = 29, finds the range, which uses every value between the extremes rather than just the middle half; taking the median away from the upper quartile instead of the lower quartile, 27 − 19 = 8, swaps the median in for the lower quartile; and reversing the subtraction, 12 − 27 = −15, finds the right two values but in the wrong order — an interquartile range is never negative.
- (d) Both the mean and the range increase. — The original mean is 150 + 152 + 155 + 158 + 160 = 775, and 775 ÷ 5 = 155 cm; the original range is 160 − 150 = 10 cm. Including the new height of 170 cm gives a new total of 775 + 170 = 945, and 945 ÷ 6 = 157.5 cm, which is higher than 155 cm, and a new range of 170 − 150 = 20 cm, which is higher than 10 cm, so both the mean and the range increase. Saying the range stays the same ignores that 170 cm is a new, higher maximum than the old 160 cm. Saying the mean stays the same ignores that 170 cm is above the original mean of 155 cm, which pulls the average up. Saying both decrease is the opposite of what happens here.
- (a) No, 150 pupils are a tenth of the school, chosen at random — Method: judge a sample on two things, whether every member of the population had the same chance of being chosen, and whether the sample is large enough to carry a pattern. Working: the 150 pupils were drawn from the register of every pupil in the school, so no year group or set is shut out and no pupil chooses to take part; and 150 ÷ 1,500 = 0.1, so one pupil in ten has been asked. A random sample of that share is ample for an estimate of how long the school's pupils spend on homework. Answer: no, because 150 pupils are a tenth of the school and were chosen at random. The distractors: saying a random sample always gives the exact school figure reaches the same verdict for a reason that is false, since a second random sample of 150 would give a slightly different mean; saying 150 pupils cannot be picked at random from 1,500 treats randomness as something only a whole population can have, when drawing names from the register is exactly how a random sample is taken; saying that only asking all 1,500 could show anything rejects sampling altogether, which would leave no way to study any population too large to count.
- (c) 145 g — Method: find the position of the median from the total frequency, locate the class that contains it, then use linear interpolation inside that class, assuming the apples in it are spread evenly. Working: the median is the 100 ÷ 2 = 50th apple; the running totals are 10, then 10 + 30 = 40, then 40 + 40 = 80, so the 50th apple lies in the class 140 ≤ m < 160; it is the 50 − 40 = 10th of the 40 apples in that class, and the class is 20 g wide, so the median is 140 + (10 ÷ 40) × 20 = 140 + 5 = 145. Answer: an estimated median of 145 g. The distractors: 150 g comes from giving the midpoint of the class that contains the median instead of interpolating inside it; 140 g comes from stopping at the lower boundary of that class, which locates the class but not the value; 155 g comes from measuring the 5 g step down from the upper boundary, 160 − 5, instead of up from the lower boundary.
- (c) Neither causes the other; sunshine links both. — Both ice cream sales and sunburn cases tend to rise on hot, sunny days, so the amount of sunshine is a third factor linked to both — neither variable causes the other. Saying ice cream sales cause the sunburn assumes a causal link in one direction that the correlation alone cannot establish. Saying sunburn cases cause the ice cream sales assumes the reverse causal link, which is no more justified. Saying a strong correlation always means causation is the general error this question is testing: correlation, however strong, does not by itself prove that one variable causes the other.
- (c) 18 — Method: for n ordered values, GCSE convention places the lower quartile at position (n + 1) ÷ 4, counting from the smallest value. Working: there are 11 marks, so n + 1 = 11 + 1 = 12 and 12 ÷ 4 = 3, so the lower quartile is the 3rd value in the ordered list 12, 15, 18, 21, 24, 27, 30, 33, 36, 39, 42, which is 18. Answer: the lower quartile is 18 marks. Watch the position you count to: dividing 11 ÷ 4 = 2.75 without adding 1 first, then rounding down, lands on the 2nd value, 15, not the 3rd; reaching for the middle of the whole list instead gives the median, 27, a different statistic; and averaging the 3rd and 4th values, 18 + 21 = 39 and 39 ÷ 2 = 19.5, borrows a method for an even split where it is not needed here.
- (b) Q was faster on average and more consistent — Method: compare the medians for the average and the interquartile ranges for the spread, remembering that a shorter time is faster and a smaller interquartile range means more consistent. Working: the median for class Q is 35 seconds against 38 seconds for class P, so class Q was faster on average; the interquartile range for class P is 46 − 24 = 22 seconds and for class Q it is 44 − 30 = 14 seconds, so class Q's times are more tightly grouped. Answer: class Q was faster on average and more consistent. The distractors: calling Q slower comes from comparing the lower quartiles, 30 against 24, as though a quartile were the average; calling Q less consistent comes from using the gap between the median and the upper quartile as the spread, 44 − 35 = 9 against 46 − 38 = 8, instead of the full interquartile range; the statement that Q was both slower and less consistent comes from making both of those mistakes together.
- (a) No, the size of the fire affects both of the quantities — Method: correlation says that two quantities change together; a claim that one of them produces the other is a further claim, and it needs evidence that a scatter graph on its own cannot give. Working: the graph does show strong positive correlation, so more engines did go with greater damage. But neither quantity was set by the researchers: both were decided by how large the fire was. A large blaze brings many appliances and also destroys a great deal, while a small one brings few and destroys little, so a third quantity is driving both of the recorded ones. Answer: no, because the size of the fire affects both of the quantities. The distractors: saying the correlation is negative contradicts the graph, which shows the two quantities rising together, and reaching the right verdict from a false reading of the data is not the reason the mark is for; saying that strong positive correlation shows one quantity causes the other is the assumption the question exists to test, and no strength of correlation can establish cause; saying the points lie close to the line of best fit describes how strong the correlation is, and strength and cause are different matters entirely.
- (a) x = 0 gives y = −20: a negative number sold — The y-intercept is the value the line predicts when x = 0: y = 3 × 0 − 20 = −20. A kiosk cannot sell a negative number of ice creams, so this is not a sensible estimate. The 3 in the equation is the gradient, not the intercept, so an option claiming x = 0 gives y = 3 has swapped the two numbers around — substituting x = 0 makes the 3x term equal 0, leaving −20, not 3. The danger of extrapolating to very high temperatures is a real issue with this line, but it is a different issue from the y-intercept, so it does not answer this question. And whether x = 0 could occur on a trading day is beside the point: the model still makes that prediction, and it is the prediction itself, −20, that is impossible.
- (a) 49 — Method: add the frequencies of the classes that lie wholly below 35 seconds, then use linear interpolation for the class that 35 cuts through, assuming the calls in that class are spread evenly across it. Working: below 30 seconds there are 15 + 24 = 39 calls; the value 35 lies in the class 30 ≤ t < 50, which is 20 seconds wide and holds 40 calls, and 35 is 35 − 30 = 5 seconds into it, so the estimated share is (5 ÷ 20) × 40 = 10 calls; the estimate is 39 + 10 = 49. Answer: about 49 calls met the target. The distractors: 79 comes from adding the whole of the class 30 ≤ t < 50, 39 + 40, and so counting calls of up to 50 seconds as being under 35; 39 comes from stopping at the class boundary 30 and ignoring the part class altogether; 69 comes from measuring the part of the class from 35 up to 50 instead of from 30 up to 35, giving (15 ÷ 20) × 40 = 30 and then 39 + 30.
- (c) 22 kg — Method: multiply the mean by the number of parcels to rebuild the total mass, then subtract the masses that are known. Working: four parcels with a mean mass of 17 kg have a total mass of 17 × 4 = 68 kg; the three known parcels total 12 + 16 + 18 = 46 kg; so the fourth parcel has mass 68 − 46 = 22 kg. Answer: 22 kg. The distractors: 68 kg comes from stopping at the total mass of all four parcels; 17 kg comes from assuming the missing parcel must have the mean mass; 5 kg comes from multiplying the mean by 3, the number of parcels whose mass is given, leaving 51 − 46 = 5.
- (d) An outlier from the damaged wing, not the trend. — That bird's point lies a long way from the rising trend followed by every other bird, and its low wingspan is explained by the damaged wing rather than by its age — it is an outlier caused by an unusual factor, not part of the general relationship between age and wingspan, so it should not be used when drawing the line of best fit. Saying every point must be used ignores that an outlier caused by a separate, identifiable factor can rightly be set aside. Saying it shows no correlation ignores that the other 19 points do show a clear rising trend; one outlier does not remove that. Saying it proves the line is inaccurate confuses one unusual bird with a fault in the line itself, when the line correctly describes the trend followed by the rest of the data.
- (b) 55 — Method: count the classes that lie wholly above 8 minutes, then use linear interpolation for the class that 8 cuts through, assuming the delays in that class are spread evenly. Working: the class 10 ≤ d < 20 lies wholly above 8 and holds 25 buses; the value 8 lies in the class 5 ≤ d < 10, which is 5 minutes wide and holds 75 buses, and the part above 8 runs from 8 to 10, a width of 2, so the estimated share is (2 ÷ 5) × 75 = 30 buses; the estimate is 30 + 25 = 55. Answer: about 55 refunds. The distractors: 100 comes from adding the whole of the class 5 ≤ d < 10, 75 + 25, and so refunding buses only 5 minutes late; 25 comes from using only the class 10 ≤ d < 20 and ignoring the part class that 8 minutes cuts through; 70 comes from taking the part of the class from 5 up to 8 instead of from 8 up to 10, giving (3 ÷ 5) × 75 = 45 and then 45 + 25.
- (b) 5.5 kg — Method: with an even number of values the median is the mean of the two middle values, taken once the data are in order of size. Working: the eight masses are already in order and 8 ÷ 2 = 4, so the middle pair are the 4th and 5th values, 5 kg and 6 kg; the median is (5 + 6) ÷ 2 = 5.5 kg. Answer: 5.5 kg. The distractors: 5 kg comes from reading the 4th value and stopping there instead of averaging the middle pair; 8 kg comes from working out the range, 11 − 3, which measures spread rather than centre; 4 kg comes from writing down the modal mass, the only value that occurs twice, instead of the median.
Build your own mix at the worksheet builder.