Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Answer key: Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- (d) No, the range uses only the fastest and slowest time — Method: check what the range is built from, then look at what it leaves out. Working: both teams have a fastest time of 20 seconds and a slowest of 40 seconds, so both ranges are 40 − 20 = 20 seconds and Tomás has that part right. But the range is calculated from those two values alone. Six of Team A's seven times lie between 20 and 25 seconds, with a single time far out at 40; Team B is the other way round, with six of its seven times at 30 seconds or more and a single time far out at 20. So Team A bunches at the fast end and Team B at the slow end. The two patterns are quite different, and the range cannot see the difference because the five middle times never enter the calculation. Answer: no, because the range uses only the fastest and slowest time. The distractors: comparing the means answers a different question, since a mean measures position rather than spread, and two sets with the same spread can have different means; saying that equal ranges mean equal spread is the very assumption that fails here; saying that seven times each forces the spreads to match confuses the size of a data set with how its values are arranged inside it.
- (b) 29 — There are 40 − 24 = 16 males, and 15 of them prefer cardio, so 16 − 15 = 1 male prefers weights. There are 24 females, and 10 prefer weights, so 24 − 10 = 14 females prefer cardio. Altogether, 15 + 14 = 29 people prefer cardio. Choosing 15 only counts the males who prefer cardio and forgets the females. Choosing 11 adds the two weights figures, 1 + 10 = 11, instead of the two cardio figures. Choosing 30 comes from 40 − 10, subtracting only the number of females who prefer weights from the grand total, rather than finding both cardio sub-totals separately.
- (a) Minutes a candle has burned and length remaining — As a candle burns for longer, less of it remains, so these two variables move in opposite directions as one increases — that is negative correlation. A pupil's shoe size generally increases as they get older, so age and shoe size show positive correlation, not negative, since both rise together. A football team's shirt colour is not a numerical quantity linked to how many matches it wins, so shirt colour and number of wins show no correlation at all. The number of letters in a pupil's name has no real connection to their ability in maths, so that pair also shows no correlation.
- (d) £15.50 — 2 × 6 = 12, and 12 + 3.50 = 15.50, so the estimated fare is £15.50. Choosing £12.00 stops after 2 × 6 = 12 and forgets to add the £3.50. Choosing £19.00 adds the distance and the constant first and then multiplies: 6 + 3.50 = 9.50, and 9.50 × 2 = 19.00, applying the ×2 to the whole sum instead of only to the distance. Choosing £13.00 multiplies only the constant term by 2 instead of the distance: 2 × 3.50 = 7, and 7 + 6 = 13.00.
- (b) 6 — Method: frequency density = frequency ÷ class width. Working: the class 12 ≤ h < 18 has width 18 − 12 = 6, so frequency density = 36 ÷ 6 = 6. Answer: the frequency density is 6 seedlings per cm. Watch which numbers you use: taking the lower bound, 12, as the width instead of 18 − 12 = 6 gives 36 ÷ 12 = 3; dividing the total number of seedlings, 90, rather than this class's frequency, 36, by the width gives 90 ÷ 6 = 15, a density that belongs to no single class; and multiplying instead of dividing gives 36 × 6 = 216, far too large a density for so narrow a class.
- (c) Neither causes the other; sunshine links both. — Both ice cream sales and sunburn cases tend to rise on hot, sunny days, so the amount of sunshine is a third factor linked to both — neither variable causes the other. Saying ice cream sales cause the sunburn assumes a causal link in one direction that the correlation alone cannot establish. Saying sunburn cases cause the ice cream sales assumes the reverse causal link, which is no more justified. Saying a strong correlation always means causation is the general error this question is testing: correlation, however strong, does not by itself prove that one variable causes the other.
- (b) Town B — higher median and smaller IQR — Method: 'higher and more consistent' needs two comparisons — the median for typical price, and the interquartile range for spread, with a smaller interquartile range meaning more consistent. Working: Town B's median, £235,000, is higher than Town A's, £220,000. Town A's interquartile range is 310 − 180 = 130 and Town B's is 260 − 175 = 85, so Town B's interquartile range is the smaller of the two. Answer: Town B has both the higher median and the smaller interquartile range, so it is the town with higher, more consistent prices. Watch which combination of median and interquartile range each statement claims, and for which town: claiming Town A has the higher median and the smaller interquartile range gets both comparisons wrong, since Town B leads on both; claiming Town A has the higher median (still wrong) but the larger interquartile range at least reads the spread correctly, without it rescuing the false median claim; and claiming Town B has the higher median (correct) but the larger interquartile range misreads the spread — Town B's interquartile range is the smaller of the two, not the larger.
- (d) An outlier from the damaged wing, not the trend. — That bird's point lies a long way from the rising trend followed by every other bird, and its low wingspan is explained by the damaged wing rather than by its age — it is an outlier caused by an unusual factor, not part of the general relationship between age and wingspan, so it should not be used when drawing the line of best fit. Saying every point must be used ignores that an outlier caused by a separate, identifiable factor can rightly be set aside. Saying it shows no correlation ignores that the other 19 points do show a clear rising trend; one outlier does not remove that. Saying it proves the line is inaccurate confuses one unusual bird with a fault in the line itself, when the line correctly describes the trend followed by the rest of the data.
- (c) 42 — Method: find the target cumulative frequency, 90% of the total, locate the class it falls in from the plotted points, then interpolate: lower boundary, plus the extra distance needed into the class divided by the class's frequency, times its width. Working: 90% of 320 is 0.9 × 320 = 288. The plotted points show a cumulative frequency of 280 at d = 40 and 320 at d = 50, so the class 40 ≤ d < 50 has frequency 320 − 280 = 40 and width 50 − 40 = 10, and 288 falls inside it. The extra distance needed into the class is 288 − 280 = 8, and 8 ÷ 40 × 10 = 2, so the diameter is 40 + 2 = 42. Answer: the estimated diameter is 42 mm. Watch which point and which class the interpolation actually uses: reading off d = 40, the plotted point just below the target, instead of interpolating the extra 8 ball bearings into the next 10 mm, stops one step short of the true answer; finding the diameter below which only 10% lie instead of 90% gives a target of 0.1 × 320 = 32, which falls in the class 10 ≤ d < 20 — the extra distance into that class is 32 − 30 = 2, and 2 ÷ 60 × 10 = 0.3, so this route gives 10 + 0.3 = 10.3, the bottom decile rather than the top 90%; and interpolating within the class 30 ≤ d < 40 instead of 40 ≤ d < 50, as though 288 had not yet reached a cumulative frequency of 280, treats the extra distance as 288 − 190 = 98, and 98 ÷ 90 × 10 = 10.9, giving 30 + 10.9 = 40.9, one class too early.
- (a) A vertical line chart (discrete numerical data) — The number of pets is discrete numerical data — whole-number values such as 0, 1, 2, 3 or 4 — recorded for one variable, so a vertical line chart is the chart specified for this kind of data. A bar chart is used for categorical data, such as favourite colour, not numerical values counted like this. A pie chart shows proportions of a whole and does not show the frequency of each separate value. A scatter graph compares two different variables against each other, and only one variable, the number of pets, is recorded here.
- (a) 26 — Method: a line of best fit lets one quantity be predicted from the other, so the score is substituted into the equation of the line and the resulting inequality is solved for the number of hours. Working: a score of at least 80 means 2.5x + 15 ≥ 80; taking 15 from both sides gives 2.5x ≥ 65, and dividing both sides by 2.5 gives x ≥ 26, so the least whole number of hours is 26. Checking, 2.5 × 26 + 15 = 80, which does reach the target. Answer: 26 hours — and this is only an estimate, because a line of best fit predicts a trend rather than an individual result, and a prediction made outside the range of hours the pupils actually revised for would be an extrapolation and less reliable still. The distractors: 27 comes from reaching 26 and then rounding up again, although 26 hours already gives a score of exactly 80; 32 comes from 80 ÷ 2.5, which ignores the 15 in the equation of the line; 38 comes from (80 + 15) ÷ 2.5, that is from adding the 15 instead of subtracting it when rearranging.
- (d) 648 kg — Method: to estimate a total from a sample, multiply the sample's mean by the number of items in the whole population, then check the units the question asks for. Working: 32.4 g × 20,000 = 648,000 g. Converting to kilograms, 648,000 ÷ 1,000 = 648 kg. This is only an estimate, not an exact total, because it assumes every one of the 20,000 packets has exactly the sample mean mass, when in reality individual packets vary above and below it. Giving 1.62 kg multiplies the mean by 50, the SAMPLE size, instead of by 20,000, the number of packets actually made that day — this finds the total mass of the 50 sampled packets, not the day's production. Giving 32.4 kg treats the sample mean itself, in grams, as if it already were the day's total mass in kilograms, skipping the scaling up altogether. Giving 648,000 kg correctly scales the mean up to the whole day's production but never converts the answer from grams to kilograms, leaving it 1,000 times too large. Always scale a sample's mean up by the SIZE OF THE WHOLE POPULATION, and always finish by checking the units the question asks for.
- (b) 55 — Method: count the classes that lie wholly above 8 minutes, then use linear interpolation for the class that 8 cuts through, assuming the delays in that class are spread evenly. Working: the class 10 ≤ d < 20 lies wholly above 8 and holds 25 buses; the value 8 lies in the class 5 ≤ d < 10, which is 5 minutes wide and holds 75 buses, and the part above 8 runs from 8 to 10, a width of 2, so the estimated share is (2 ÷ 5) × 75 = 30 buses; the estimate is 30 + 25 = 55. Answer: about 55 refunds. The distractors: 100 comes from adding the whole of the class 5 ≤ d < 10, 75 + 25, and so refunding buses only 5 minutes late; 25 comes from using only the class 10 ≤ d < 20 and ignoring the part class that 8 minutes cuts through; 70 comes from taking the part of the class from 5 up to 8 instead of from 8 up to 10, giving (3 ÷ 5) × 75 = 45 and then 45 + 25.
- (a) Drawing 60 names at random from a list of all 1200 pupils — Method: a sample is random when every member of the population has the same chance of being chosen and nobody, including the pupils themselves, can influence who ends up in it; test each method against that. Working: drawing names from a list of all 1200 pupils gives each pupil the same chance, 60 out of 1200, whatever their year group, class or opinion, so the method is random. Answer: drawing 60 names at random from a list of all 1200 pupils. The distractors: asking the pupils who volunteer is self-selection, and the pupils with the strongest views volunteer first, so they decide the sample; asking the pupils nearest the door is convenience sampling, which reaches only those who happen to be in one place at one time; asking two Year 10 classes samples a cluster, so every pupil in the other year groups has no chance of being chosen at all.
- (b) £1,000 — Wages take up 150° out of 360°, so the amount spent on wages is 150 ÷ 360 × 2400 = £1,000. Choosing £600 uses the repairs angle, 90°, instead of the wages angle: 90 ÷ 360 × 2400 = 600. Choosing £3,600 treats the angle in degrees as if it were a percentage, 150 ÷ 100 × 2400 = 3600, instead of dividing by 360°. Choosing £800 uses the angle for the 'other costs' sector, 360 − 90 − 150 = 120°, instead of the wages sector: 120 ÷ 360 × 2400 = 800.
Build your own mix at the worksheet builder.