Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Answer key: Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- (b) 29 — There are 40 − 24 = 16 males, and 15 of them prefer cardio, so 16 − 15 = 1 male prefers weights. There are 24 females, and 10 prefer weights, so 24 − 10 = 14 females prefer cardio. Altogether, 15 + 14 = 29 people prefer cardio. Choosing 15 only counts the males who prefer cardio and forgets the females. Choosing 11 adds the two weights figures, 1 + 10 = 11, instead of the two cardio figures. Choosing 30 comes from 40 − 10, subtracting only the number of females who prefer weights from the grand total, rather than finding both cardio sub-totals separately.
- (d) 45 — Method: a cumulative frequency is a running total — it counts everybody in every class up to and including the one that ends at the value given. Working: the classes that lie wholly below 30 minutes are 0 ≤ t < 10, 10 ≤ t < 20 and 20 ≤ t < 30, with frequencies 6, 14 and 25, so the running total is 6 + 14 = 20 and then 20 + 25 = 45. Answer: 45 people took less than 30 minutes. The distractors: 25 comes from quoting the frequency of the class 20 ≤ t < 30 on its own instead of the running total; 65 comes from accumulating one class too many and including 30 ≤ t < 40, which is 45 + 20; 35 comes from accumulating from the top downwards, 15 + 20, which counts the people who took 30 minutes or more rather than fewer.
- (c) 156 cm — Method: to combine two groups' means, multiply each group's mean by its own number of pupils, add the two totals together, then divide by the total number of pupils in both groups. Working: 20 × 150 = 3,000 cm for the boys and 10 × 168 = 1,680 cm for the girls, giving a combined total of 3,000 + 1,680 = 4,680 cm. Dividing by all 30 pupils gives 4,680 ÷ 30 = 156 cm. Giving 159 cm averages the two means, (150 + 168) ÷ 2, treating the two groups as if they had the same number of pupils, when there are twice as many boys as girls. Giving 4,680 cm finds the correct combined total height but stops there, forgetting the final division by the 30 pupils. Giving 234 cm divides the combined total by 20, the number of boys only, forgetting that the total also includes the 10 girls. Always weight each mean by its own group size, and always divide by the TOTAL number of pupils in both groups combined.
- (a) No, 150 pupils are a tenth of the school, chosen at random — Method: judge a sample on two things, whether every member of the population had the same chance of being chosen, and whether the sample is large enough to carry a pattern. Working: the 150 pupils were drawn from the register of every pupil in the school, so no year group or set is shut out and no pupil chooses to take part; and 150 ÷ 1,500 = 0.1, so one pupil in ten has been asked. A random sample of that share is ample for an estimate of how long the school's pupils spend on homework. Answer: no, because 150 pupils are a tenth of the school and were chosen at random. The distractors: saying a random sample always gives the exact school figure reaches the same verdict for a reason that is false, since a second random sample of 150 would give a slightly different mean; saying 150 pupils cannot be picked at random from 1,500 treats randomness as something only a whole population can have, when drawing names from the register is exactly how a random sample is taken; saying that only asking all 1,500 could show anything rejects sampling altogether, which would leave no way to study any population too large to count.
- (d) 4, 3, 2, 6 — Method: frequency density = frequency ÷ class width for each class in turn; do not assume the classes are all the same width. Working: the four classes have widths 10 − 0 = 10, 30 − 10 = 20, 45 − 30 = 15 and 50 − 45 = 5. Dividing each frequency by its own width gives 40 ÷ 10 = 4, 60 ÷ 20 = 3, 30 ÷ 15 = 2 and 30 ÷ 5 = 6. Answer: the frequency densities, in order, are 4, 3, 2 and 6. Watch the width of each class separately: treating the last class as if it were also 10 units wide, like the first, gives 30 ÷ 10 = 3 instead of 30 ÷ 5 = 6 — the classes here are deliberately unequal, so no width can be borrowed from another class; dividing the width by the frequency instead of the frequency by the width for the third class gives 15 ÷ 30 = 0.5 in place of 2, the formula the wrong way round; and reading the frequency column straight off the table, 40, 60, 30, 30, skips the division by width altogether and reports how many fish are in each class rather than how densely packed each bar is.
- (b) 13 — The total is 50, and the two known parts are 22 (tea) and 15 (coffee), so 50 − 22 − 15 = 13 hot chocolates. Choosing 28 comes from 50 − 22, subtracting only the tea and forgetting the coffee. Choosing 35 comes from 50 − 15, subtracting only the coffee and forgetting the tea. Choosing 37 comes from 22 + 15, which finds how many drinks were tea or coffee, not the number left over for hot chocolate.
- (a) 26 — Method: a line of best fit lets one quantity be predicted from the other, so the score is substituted into the equation of the line and the resulting inequality is solved for the number of hours. Working: a score of at least 80 means 2.5x + 15 ≥ 80; taking 15 from both sides gives 2.5x ≥ 65, and dividing both sides by 2.5 gives x ≥ 26, so the least whole number of hours is 26. Checking, 2.5 × 26 + 15 = 80, which does reach the target. Answer: 26 hours — and this is only an estimate, because a line of best fit predicts a trend rather than an individual result, and a prediction made outside the range of hours the pupils actually revised for would be an extrapolation and less reliable still. The distractors: 27 comes from reaching 26 and then rounding up again, although 26 hours already gives a score of exactly 80; 32 comes from 80 ÷ 2.5, which ignores the 15 in the equation of the line; 38 comes from (80 + 15) ÷ 2.5, that is from adding the 15 instead of subtracting it when rearranging.
- (c) The mean, because every value counts towards it, so 100 pulls it from 7 up to about 20.3. — Method: work each measure out before the extra value is added and again afterwards, then compare the size of the two changes. Working: before, the six values total 42, so the mean is 42 ÷ 6 = 7, and the middle pair 6 and 8 give a median of (6 + 8) ÷ 2 = 7; after, the seven values total 142, so the mean is 142 ÷ 7 = 20.29 to 2 decimal places, while the median is now the 4th of the seven ordered values, which is 8; the mean has moved by about 13.3 and the median by 1. Answer: the mean, because every value counts towards it, so 100 pulls it from 7 up to about 20.3 — this is why the median is often preferred when a data set contains an outlier. The distractors: the reply that the mean rises by 100 adds the extra value to the mean instead of adding it to the total; the reply that the median moves to 12 takes the largest of the original values as the new middle instead of counting to the 4th of the seven values; the reply about even and odd counts quotes a rule that does not exist, since the median moved because a very large value was added, not because the count of values changed.
- (b) 60 — Method: for n ordered values, the upper quartile sits at position 3(n + 1) ÷ 4, counting from the smallest. Working: n + 1 = 11 + 1 = 12; 3 × 12 = 36 and 36 ÷ 4 = 9, so the upper quartile is the 9th value in the list 40, 42, 45, 47, 50, 52, 55, 58, 60, 63, 65, which is 60. Answer: the upper quartile is 60 seconds. Watch which quartile you find: counting to the 6th value gives the median, 52, not the upper quartile; using 3 × 11 = 33 and 33 ÷ 4 = 8.25 without adding 1 to n first, then rounding down, reaches the 8th value, 58, not the 9th; and counting to the 3rd value uses the lower quartile's position, 45, the wrong end of the list.
- (a) 22 — Method: the two subject totals overlap, because every pupil who passed both subjects has been counted once in the maths total and once again in the science total; adding the totals therefore counts those pupils twice, and the overlap has to be taken off once. Working: 18 + 12 = 30, and the 8 pupils who passed both have been counted twice in that 30, so the number who passed at least one subject is 30 − 8 = 22. Answer: 22 pupils, a count of pupils, and it is less than the 30 in the class, which leaves 8 pupils who passed neither. The distractors: 30 comes from adding the two subject totals and never removing the overlap, so it counts the 8 pupils twice; 14 comes from taking the 8 away twice, 18 + 12 − 8 − 8, removing an overlap that was only counted twice once too often; 18 comes from writing down the larger of the two subject totals on its own, which leaves out every pupil who passed science but not maths.
- (b) On average a plant grew 1.5 cm taller for each extra day — Method: in the equation of a line, the number multiplying x is the gradient, and a gradient states the change in y produced by an increase of 1 in x, read in the units of the two axes. Working: here x is measured in days and y in centimetres, so the gradient 1.5 carries the units centimetres per day. Testing it on the line, 5 days gives 1.5 × 5 + 4 = 11.5 cm and 6 days gives 1.5 × 6 + 4 = 13 cm, a rise of 1.5 cm for the one extra day. Answer: on average a plant grew 1.5 cm taller for each extra day of watering. The distractors: 1.5 cm as the height before any watering is the value of y when x is 0, which is the other number in the equation, 4 cm, so this swaps the gradient and the intercept; 1.5 cm as the gap between the tallest and the shortest plant reads the gradient as a range, when a range is a difference between two of the 16 plants and a gradient is a rate; 1.5 days for each extra centimetre inverts the rate, dividing days by centimetres instead of centimetres by days, and the line gives 1 cm of growth in two thirds of a day.
- (c) 72 — Method: on a histogram the frequency of a class is the area of its bar, so frequency = frequency density × class width. Working: the class 50 ≤ m < 80 has width 80 − 50 = 30 grams and a frequency density of 2.4 per gram, so the frequency is 2.4 × 30 = 72. Answer: 72 pebbles. The distractors: 192 comes from using the upper class boundary, 80, as the width, giving 2.4 × 80; 12.5 comes from dividing the width by the density, 30 ÷ 2.4, which reverses the area rule; 2.4 comes from reading the height of the bar as the frequency itself, the commonest mistake on histograms, where a height is a density and only an area is a count.
- (c) Yes — £80 is above the boundary, £78 — Method: a value counts as an outlier when it lies more than 1.5 times the interquartile range beyond the nearer quartile; here that means checking it against upper quartile + 1.5 × interquartile range. Working: the interquartile range is 42 − 18 = 24. 1.5 × 24 = 36, and 42 + 36 = 78, so any saving above £78 is an outlier. Amara saved £80, and 80 is greater than 78. Answer: yes, Amara's saving is an outlier, because £80 is above the outlier boundary, £78. Watch how you build the boundary and what you compare it with: adding the two quartiles instead of subtracting them, 42 + 18 = 60, gives an interquartile range three times too big, and 42 + 1.5 × 60 = 42 + 90 = 132 puts the boundary so far out that £80 wrongly looks ordinary; comparing £80 with the upper quartile alone, £42, checks only that it lies in the top quarter of the data, which every value above £42 does, not that it lies unusually far beyond it; and adding the interquartile range on once instead of one and a half times, 42 + 24 = 66, uses the wrong multiplier, even though £80 still happens to clear that lower boundary too.
- (a) No, the size of the fire affects both of the quantities — Method: correlation says that two quantities change together; a claim that one of them produces the other is a further claim, and it needs evidence that a scatter graph on its own cannot give. Working: the graph does show strong positive correlation, so more engines did go with greater damage. But neither quantity was set by the researchers: both were decided by how large the fire was. A large blaze brings many appliances and also destroys a great deal, while a small one brings few and destroys little, so a third quantity is driving both of the recorded ones. Answer: no, because the size of the fire affects both of the quantities. The distractors: saying the correlation is negative contradicts the graph, which shows the two quantities rising together, and reaching the right verdict from a false reading of the data is not the reason the mark is for; saying that strong positive correlation shows one quantity causes the other is the assumption the question exists to test, and no strength of correlation can establish cause; saying the points lie close to the line of best fit describes how strong the correlation is, and strength and cause are different matters entirely.
- (a) 1,400 and 1,240, so combine the samples for one estimate — Method: scale each sample up to the whole stock, then use the fact that a larger sample gives a more reliable estimate than a smaller one. Working: the first sample gives 35 ÷ 50 = 0.7 and 0.7 × 2,000 = 1,400 paperbacks; the second gives 31 ÷ 50 = 0.62 and 0.62 × 2,000 = 1,240 paperbacks. Two random samples of the same size are expected to differ a little, so neither estimate is wrong. Putting the two together gives 35 + 31 = 66 paperbacks in 100 books, and 66 ÷ 100 = 0.66 with 0.66 × 2,000 = 1,320, an estimate resting on twice as many books as either volunteer checked. Answer: 1,400 and 1,240, so combine the samples for one estimate. The distractors: keeping 1,400 because it is larger picks an estimate by its size, when both samples held 50 books and neither has a stronger claim; saying a volunteer must have miscounted assumes two random samples ought to agree exactly, which is precisely what random sampling does not promise; 1,750 and 1,550 come from 35 × 50 = 1,750 and 31 × 50 = 1,550, multiplying each count by the size of the sample instead of scaling by 2,000 ÷ 50.
Build your own mix at the worksheet builder.