Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Answer key: Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- (b) All 2,000 members of the sports centre. — Method: in a survey, the population is the whole group the survey is trying to find out about, and the sample is the smaller group actually asked. Working: this survey wants to know what the sports centre's members think, so the population is every one of the 2,000 members — whether or not they were personally asked. Saying the population is the 100 members who were asked names the sample, not the population; the sample is drawn FROM the population, so it is smaller than it, not the same as it. Saying the population is everybody who lives in Ipswich widens the group far beyond who the survey is actually about — plenty of Ipswich residents are not members of the sports centre at all, so they are outside this survey altogether. Saying the population is the members who say they are unhappy confuses the population with a result of the survey: whether a member turns out to be happy or unhappy is something the survey finds out, not part of the definition of who is being studied. The population is always the whole group the question is about, before any sampling or any results come in.
- (d) 19 kg — Method: estimate the mean of grouped data by multiplying each class's midpoint by its frequency, adding the four totals, then dividing by the total frequency. Working: the midpoints are 5, 15, 25 and 35 kg. The weighted totals are 11 × 5 = 55, 5 × 15 = 75, 5 × 25 = 125 and 9 × 35 = 315, which add to 570. Dividing by the 30 dogs gives an estimate of 570 ÷ 30 = 19 kg. Giving 5 kg reads off the midpoint of the modal class, 0 < m ≤ 10, the class with the most dogs — but the class with the most dogs is not where the mean falls, and neither is a substitute for actually calculating it. Giving 20 kg averages the four midpoints, (5 + 15 + 25 + 35) ÷ 4, treating every class as equally likely and ignoring that far more dogs are in the lightest and heaviest classes than in the middle two. Giving 570 kg stops after finding the correct weighted total and forgets the final division by the 30 dogs. Always weight each midpoint by its own frequency, and always finish by dividing by the total frequency, not the number of classes.
- (c) Because every pupil has an equal chance of being picked — Method: whether a sample represents its population is decided by the selection method, not by the size of the sample, so ask whether the method gives every member of the population the same chance of being chosen. Working: the names are drawn at random from a list of all 10,000 pupils, so each pupil has the same chance, 500 out of 10,000, of being drawn, and no group of pupils is more likely to appear than any other; that is what keeps bias out of the sample. Answer: because every pupil has an equal chance of being picked. The distractors: the reply about 5% treats the sampling fraction as the test of fairness, but a badly chosen 5% is still biased and a well chosen 1% is not; the reply about 500 being large enough makes size the test instead, which is the same mistake in another form, since a large sample drawn from one school would still misrepresent the city; the reply about the most willing pupils describes self-selection, which hands the choice of who is in the sample to the pupils who feel most strongly about the question.
- (c) No — median £505 at A vs £510 at B. — Branch A's seven wages in order are £480, £495, £500, £505, £510, £515 and £1,200, so the median, the 4th value, is £505. Branch B's in order are £480, £490, £500, £510, £520, £530 and £540, so the median is £510. Since £505 is lower than £510, the median wage is not higher at Branch A, so the claim is not fairly supported. Choosing 'Yes — mean £600.71 at A vs £510 at B' uses the mean: 480 + 495 + 500 + 505 + 510 + 515 + 1200 = 4205, and 4205 ÷ 7 = 600.71, a figure pulled upward by the £1,200 outlier that does not represent a typical wage. Choosing 'Yes — median £515 at A vs £510 at B' miscounts the middle position, taking the 6th wage, £515, instead of the correct 4th value, £505. Choosing 'Yes — highest wage £1,200 at A vs £540 at B' compares the highest wage at each branch rather than a measure of the typical, or average, wage.
- (c) Route 2, because its interquartile range is smaller — Method: for a journey where turning up on time matters, what matters is not the typical (median) time but how predictable it is — a smaller interquartile range means the middle half of journeys cluster closer together. Working: Route 1's median, 34 minutes, is in fact lower than Route 2's, 41 minutes, so Route 1 is faster on average; but Route 1's interquartile range, 22 minutes, is far larger than Route 2's, 6 minutes, so Route 1's times are much less predictable. Answer: Priya should choose Route 2, because its interquartile range is smaller, even though it is slower on average. Watch which statistic answers the question actually asked: Route 1 does not have the smaller interquartile range, Route 2 does, so picking Route 1 for that reason misreads the table; Route 1's median genuinely is the lower one, but a lower median answers 'which is faster', not 'which is more reliable'; and Route 2's median is not the lower one, so that claim about Route 2 is simply false.
- (b) 29 — There are 40 − 24 = 16 males, and 15 of them prefer cardio, so 16 − 15 = 1 male prefers weights. There are 24 females, and 10 prefer weights, so 24 − 10 = 14 females prefer cardio. Altogether, 15 + 14 = 29 people prefer cardio. Choosing 15 only counts the males who prefer cardio and forgets the females. Choosing 11 adds the two weights figures, 1 + 10 = 11, instead of the two cardio figures. Choosing 30 comes from 40 − 10, subtracting only the number of females who prefer weights from the grand total, rather than finding both cardio sub-totals separately.
- (c) 70 — Method: estimate the median from the cumulative frequency table by interpolation: find its position, n ÷ 2, locate the class it falls in, then add the fraction of the way through that class (adjusted for the cumulative frequency reached before it) to the class's lower boundary. Working: there are 180 riders, so the median is at position 180 ÷ 2 = 90. Before the class 60 ≤ d < 80 the cumulative frequency is 60, and by the end of it, 120, so the 90th rider falls in this class; its frequency is 120 − 60 = 60 and its width is 80 − 60 = 20. The extra distance needed into the class is 90 − 60 = 30, and 30 ÷ 60 × 20 = 10, so the median is 60 + 10 = 70. Answer: the estimated median distance is 70 km. Watch which numbers the interpolation actually uses: reading off just the class's lower boundary, 60, ignores how far into the class the 90th rider falls; using the target position, 90, as the extra distance instead of subtracting the 60 riders already counted before the class gives 90 ÷ 60 × 20 = 30, so 60 + 30 = 90, overshooting by treating the whole position as if none of it had already been counted; and using the total number of riders, 180, instead of half of it as the target position lands in the very last class, giving an estimate of 130 km — further than any rider is known to have ridden by that point in the table.
- (d) 4, 3, 2, 6 — Method: frequency density = frequency ÷ class width for each class in turn; do not assume the classes are all the same width. Working: the four classes have widths 10 − 0 = 10, 30 − 10 = 20, 45 − 30 = 15 and 50 − 45 = 5. Dividing each frequency by its own width gives 40 ÷ 10 = 4, 60 ÷ 20 = 3, 30 ÷ 15 = 2 and 30 ÷ 5 = 6. Answer: the frequency densities, in order, are 4, 3, 2 and 6. Watch the width of each class separately: treating the last class as if it were also 10 units wide, like the first, gives 30 ÷ 10 = 3 instead of 30 ÷ 5 = 6 — the classes here are deliberately unequal, so no width can be borrowed from another class; dividing the width by the frequency instead of the frequency by the width for the third class gives 15 ÷ 30 = 0.5 in place of 2, the formula the wrong way round; and reading the frequency column straight off the table, 40, 60, 30, 30, skips the division by width altogether and reports how many fish are in each class rather than how densely packed each bar is.
- (c) The data show a link only; a third factor may affect both — Method: a study of this kind measures two quantities and reports how they change together; deciding that one of them produces the other is a further claim, and it needs evidence that the measurements alone cannot give. Working: the study shows that more coffee goes with better concentration, which is a positive correlation; but a third factor that was never measured, such as how motivated someone is, could raise both the coffee drinking and the concentration, and the concentration could equally be what leads to the extra coffee. Answer: the data show a link only, because a third factor may be affecting both quantities, so no claim about cause can be made. The distractors: calling the conclusion safe because the correlation is positive treats the direction of a correlation as proof of cause, which no direction can give; calling it wrong because the correlation is negative misreads the direction of the relationship, since the study reports both quantities rising together; saying the two quantities are not linked denies the correlation the study actually found, when what fails is only the claim about cause.
- (d) 18 minutes — Method: the lower quartile is the 80 ÷ 4 = 20th value and the upper quartile is the 3 × 80 ÷ 4 = 60th value; locate each inside its class by linear interpolation, then subtract. Working: the 20th value lies between the running totals 8 and 28, so it is in the class 10 ≤ t < 20, which holds 20 journeys across 10 minutes, and it is the 20 − 8 = 12th of them, giving 10 + (12 ÷ 20) × 10 = 16 minutes; the 60th value lies between the running totals 52 and 72, so it is in the class 30 ≤ t < 40, which also holds 20 journeys across 10 minutes, and it is the 60 − 52 = 8th of them, giving 30 + (8 ÷ 20) × 10 = 34 minutes; subtracting, 34 − 16 = 18. Answer: an estimated interquartile range of 18 minutes. The distractors: 20 minutes comes from taking the lower boundaries of the two quartile classes, 30 − 10, which locates the classes but never the values inside them; 40 minutes comes from subtracting the two positions, 60 − 20, instead of the two times; 22 minutes comes from interpolating downwards from each upper boundary rather than upwards from each lower boundary, giving 20 − 6 = 14 and 40 − 4 = 36.
- (d) An outlier from the damaged wing, not the trend. — That bird's point lies a long way from the rising trend followed by every other bird, and its low wingspan is explained by the damaged wing rather than by its age — it is an outlier caused by an unusual factor, not part of the general relationship between age and wingspan, so it should not be used when drawing the line of best fit. Saying every point must be used ignores that an outlier caused by a separate, identifiable factor can rightly be set aside. Saying it shows no correlation ignores that the other 19 points do show a clear rising trend; one outlier does not remove that. Saying it proves the line is inaccurate confuses one unusual bird with a fault in the line itself, when the line correctly describes the trend followed by the rest of the data.
- (c) Class X has the higher median and the wider spread — Method: compare the two box plots statistic by statistic — median for location, and the interquartile range for spread — checking the true value of each rather than assuming a pattern. Working: Class X's median is 60 and Class Y's is 58, so Class X's median is the higher one. Class X's interquartile range is 70 − 45 = 25 and Class Y's is 65 − 50 = 15 (and the ranges follow the same order: 95 − 20 = 75 against 80 − 35 = 45), so Class X also has the wider spread. Answer: Class X has both the higher median and the wider spread. Watch that each half of a compound statement is checked separately: claiming Class Y has the higher median and the wider spread gets both comparisons backwards; claiming Class X has the higher median but the narrower spread keeps the median right while reading the spread the wrong way round; and claiming Class Y has the higher median but the narrower spread swaps the median comparison while getting the spread right.
- (c) 75 — Method: the number in a class is the area of its bar, frequency density × class width, so work out the frequency of each class that lies at or above 10 minutes and add them. Working: the class 10 ≤ t < 25 is 15 minutes wide with a frequency density of 3.2, giving 3.2 × 15 = 48 members; the class 25 ≤ t < 55 is 30 minutes wide with a frequency density of 0.9, giving 0.9 × 30 = 27 members; the total charged is 48 + 27 = 75. Answer: 75 members pay the extra charge. The distractors: 4.1 comes from adding the two frequency densities, 3.2 + 0.9, as though each height were a count; 93 comes from including the class 0 ≤ t < 10 as well, 1.8 × 10 = 18 added to 48 and 27, which charges every member; 27 comes from using only the class 25 ≤ t < 55 and forgetting that 10 ≤ t < 25 is also at or above 10 minutes.
- (c) Neither causes the other; sunshine links both. — Both ice cream sales and sunburn cases tend to rise on hot, sunny days, so the amount of sunshine is a third factor linked to both — neither variable causes the other. Saying ice cream sales cause the sunburn assumes a causal link in one direction that the correlation alone cannot establish. Saying sunburn cases cause the ice cream sales assumes the reverse causal link, which is no more justified. Saying a strong correlation always means causation is the general error this question is testing: correlation, however strong, does not by itself prove that one variable causes the other.
- (c) 69 — Method: multiply the mean by the number of values to find the total, then subtract the total of the known values. Working: the total of all five scores is 68 × 5 = 340. The total of the four known scores is 55 + 62 + 74 + 80 = 271. The fifth score is 340 − 271 = 69. Subtracting the other way round, 271 − 340 = −69, gives the right size answer with the wrong sign. Guessing that the missing score simply equals the mean, 68, ignores that the four known scores are not themselves centred on 68. Multiplying the mean by 4 instead of 5, 68 × 4 = 272, then 272 − 271 = 1, undercounts how many scores there are. Always multiply the mean by the TOTAL number of values before subtracting.
Build your own mix at the worksheet builder.