Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Answer key: Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- (a) 1,400 and 1,240, so combine the samples for one estimate — Method: scale each sample up to the whole stock, then use the fact that a larger sample gives a more reliable estimate than a smaller one. Working: the first sample gives 35 ÷ 50 = 0.7 and 0.7 × 2,000 = 1,400 paperbacks; the second gives 31 ÷ 50 = 0.62 and 0.62 × 2,000 = 1,240 paperbacks. Two random samples of the same size are expected to differ a little, so neither estimate is wrong. Putting the two together gives 35 + 31 = 66 paperbacks in 100 books, and 66 ÷ 100 = 0.66 with 0.66 × 2,000 = 1,320, an estimate resting on twice as many books as either volunteer checked. Answer: 1,400 and 1,240, so combine the samples for one estimate. The distractors: keeping 1,400 because it is larger picks an estimate by its size, when both samples held 50 books and neither has a stronger claim; saying a volunteer must have miscounted assumes two random samples ought to agree exactly, which is precisely what random sampling does not promise; 1,750 and 1,550 come from 35 × 50 = 1,750 and 31 × 50 = 1,550, multiplying each count by the size of the sample instead of scaling by 2,000 ÷ 50.
- (c) Neither causes the other; sunshine links both. — Both ice cream sales and sunburn cases tend to rise on hot, sunny days, so the amount of sunshine is a third factor linked to both — neither variable causes the other. Saying ice cream sales cause the sunburn assumes a causal link in one direction that the correlation alone cannot establish. Saying sunburn cases cause the ice cream sales assumes the reverse causal link, which is no more justified. Saying a strong correlation always means causation is the general error this question is testing: correlation, however strong, does not by itself prove that one variable causes the other.
- (b) All 2,000 members of the sports centre. — Method: in a survey, the population is the whole group the survey is trying to find out about, and the sample is the smaller group actually asked. Working: this survey wants to know what the sports centre's members think, so the population is every one of the 2,000 members — whether or not they were personally asked. Saying the population is the 100 members who were asked names the sample, not the population; the sample is drawn FROM the population, so it is smaller than it, not the same as it. Saying the population is everybody who lives in Ipswich widens the group far beyond who the survey is actually about — plenty of Ipswich residents are not members of the sports centre at all, so they are outside this survey altogether. Saying the population is the members who say they are unhappy confuses the population with a result of the survey: whether a member turns out to be happy or unhappy is something the survey finds out, not part of the definition of who is being studied. The population is always the whole group the question is about, before any sampling or any results come in.
- (d) 16 — Method: on a box plot the interquartile range is the width of the box itself, upper quartile take away lower quartile. Working: the upper quartile is 38 and the lower quartile is 22, so 38 − 22 = 16. Answer: the interquartile range is 16 years. Watch which part of the box plot you are reading: the whole line from whisker to whisker gives the range, 61 − 15 = 46; the left half of the box alone gives median take away lower quartile, 29 − 22 = 7; and the right half of the box alone gives upper quartile take away median, 38 − 29 = 9 — neither half is the interquartile range on its own.
- (a) 15 — Year 11 has 50 − 28 = 22 pupils in total. Of the 22 pupils who walk in total, 15 are in Year 10, so 22 − 15 = 7 Year 11 pupils walk. Subtracting that from the Year 11 total gives 22 − 7 = 15 Year 11 pupils who are driven. Choosing 28 takes the whole school's driven total, 50 − 22 = 28, and treats it as if it were Year 11's alone, without separating the year groups. Choosing 7 correctly finds how many Year 11 pupils walk but stops there, giving that figure instead of the number who are driven. Choosing 35 comes from 50 − 15, subtracting the Year 10 walkers from the whole school total rather than working within Year 11.
- (a) Minutes a candle has burned and length remaining — As a candle burns for longer, less of it remains, so these two variables move in opposite directions as one increases — that is negative correlation. A pupil's shoe size generally increases as they get older, so age and shoe size show positive correlation, not negative, since both rise together. A football team's shirt colour is not a numerical quantity linked to how many matches it wins, so shirt colour and number of wins show no correlation at all. The number of letters in a pupil's name has no real connection to their ability in maths, so that pair also shows no correlation.
- (c) No — 8 from one class is too small to represent the school. — Method: judge reliability by asking whether the sample is both large enough, and spread across the population, relative to what it is meant to represent. Working: 8 pupils is a tiny fraction of the school's 1,000 pupils, and all 8 come from a single class rather than a range of year groups, so the sample is both too small and too narrow to represent the whole school reliably. She is not right. Saying any sample size gives an equally reliable estimate ignores that reliability generally improves with a larger, more representative sample. Saying the method is unreliable because it was not done online is not a reason connected to sample size or representativeness at all. Saying 8 is reliable because it is more than half her class compares the sample to the wrong population — the school has 1,000 pupils, not one class. Always judge a sample's size against the population it is meant to represent, not against a smaller group within it.
- (c) No, the mode here is the lowest value of the nine — Method: an average is meant to stand for the data as a whole, so test any proposed average by asking how many values it sits near. Working: the value 4 appears three times and every other count appears once, so 4 is indeed the mode. But those three hours are the quiet ones at the start of the day, and the other six counts run from 11 up to 25; putting the nine counts in order, the middle one is the fifth, which is 13. So the mode sits at the very bottom of the data, with six of the nine hours far above it. Answer: no, because the mode here is the lowest value of the nine, so it describes the quiet opening hours rather than a typical hour. The distractors: saying the mode can only be used when no value repeats reverses the definition, since a mode exists only because a value does repeat; saying the mode is the value that occurs most often is a correct definition, but being the commonest value does not make a value typical when it lies at one end of the data; saying the mode is the best average for any list of numbers ignores the fact that mean, median and mode each describe a population well in different circumstances.
- (c) The median, £160,000, as one very high price lifts the mean — Method: find both averages, then choose the one that sits closer to the bulk of the data. Working: in order the prices are 140,000, 150,000, 160,000, 170,000 and 580,000, so the median is the third of the five, £160,000. For the mean, 140,000 + 150,000 + 160,000 + 170,000 + 580,000 = 1,200,000 and 1,200,000 ÷ 5 = 240,000, so the mean is £240,000. Four of the five houses sold for £170,000 or less, so a reader told that a typical price is £240,000 would expect to pay at least £70,000 more than any of those four cost. Answer: the median, £160,000, as one very high price lifts the mean. The distractors: £580,000 is the middle value of the list as it is printed, which is the median only when the values have first been put in order; £240,000 is the mean, chosen on the ground that a median ignores three of the five prices, but a median uses all five to find which one is central and is then untroubled by how extreme the outer values are; £155,000 comes from deleting the £580,000 house and taking the mean of what is left, since 140,000 + 150,000 + 160,000 + 170,000 = 620,000 and 620,000 ÷ 4 = 155,000, but a real sale may not be thrown away merely for being large.
- (a) (30, 60) — Method: a cumulative frequency point is plotted at the upper boundary of its class, paired with the running total of all the frequencies up to and including that class. Working: the running totals are 7, then 7 + 19 = 26, then 26 + 34 = 60, then 60 + 40 = 100; the class 20 ≤ t < 30 has upper boundary 30, and the running total there is 60. Answer: the point for that class is plotted at 30 seconds against a cumulative frequency of 60. The distractors: (25, 60) comes from plotting at the class midpoint, which is what a frequency polygon uses and not what a cumulative frequency diagram uses; (30, 34) comes from plotting the class frequency, 34, rather than the running total; (20, 60) comes from plotting at the lower boundary of the class, which would claim that 60 calls took less than 20 seconds when only 26 did.
- (c) A histogram, with frequency density up the vertical axis — Method: decide which diagram makes area stand for frequency, which is the property the question asks for. Working: on a histogram the vertical axis is frequency density, so the area of a bar is frequency density × class width, and that product is the frequency; this is exactly what is wanted, and it is what allows classes of unequal width to be shown fairly. Answer: a histogram, with frequency density up the vertical axis. The distractors: a bar chart plots frequency as the height, so with unequal widths a wide class would cover far more area than a narrow class holding the same number of batteries, and area would measure nothing; a cumulative frequency diagram plots running totals against upper class boundaries, so a point on it gives how many lie below a value rather than how many lie in a class; a pie chart shows each class as a share of the whole 300 and loses the class widths entirely, so no area on it is tied to a scale of hours.
- (d) 25 — Method: the median is estimated at position n ÷ 2 in the cumulative frequency table, then interpolated across the class it falls in: lower boundary, plus the fraction of the way through the class, times the class width. Working: there are 80 sacks, so the median sits at position 80 ÷ 2 = 40. Before the class 20 ≤ m < 30 the cumulative frequency is 22, and by the end of it, it is 58, so this class holds the 40th sack; its frequency is 58 − 22 = 36 and its width is 30 − 20 = 10. The extra distance needed into the class is 40 − 22 = 18, and 18 ÷ 36 × 10 = 5, so the median is 20 + 5 = 25. Answer: the estimated median mass is 25 kg. Watch which numbers the interpolation uses: reading off just the lower boundary of the median class, 20, ignores how far into that class the 40th sack actually falls; treating n ÷ 2 = 40 itself as the median mass mistakes a position in the list for a mass in kilograms; and using the target position, 40, as the extra distance into the class instead of subtracting the sacks already counted changes the calculation to 20 + 40 ÷ 36 × 10. That comes to 20 + 11.1 = 31.1, overshooting the class because it never subtracts the 22 sacks already counted before it.
- (d) 25% — First find the number of fruit cakes: 80 − 34 − 26 = 20. Then write this as a percentage of the total: 20 ÷ 80 × 100 = 25%. Giving 20% comes from reporting the count of fruit cakes, 20, directly as a percentage, without dividing by the total of 80 first. Giving 32.5% computes the percentage of chocolate cakes instead of fruit cakes: 26 ÷ 80 × 100 = 32.5%. Giving 42.5% computes the percentage of sponge cakes instead of fruit cakes: 34 ÷ 80 × 100 = 42.5%.
- (d) 38, and fairly confident, as 12 °C is inside the range — Method: substitute the forecast temperature into the equation of the line of best fit, then judge the prediction by where that temperature sits among the data the line was drawn from. Working: putting x = 12 into y = −3x + 74 gives −3 × 12 + 74 = 38, so the line predicts 38 hot chocolates. The recorded temperatures run from 4 °C to 18 °C, and 12 °C lies inside that interval, so this is interpolation, the safer kind of prediction. Answer: 38, and fairly confident, as 12 °C is inside the range; the owner should still expect the true figure to differ a little, since the points only lie near the line and not on it. The distractors: being completely certain treats a line of best fit as a rule that fixes each day's sales, when it describes a trend that individual days depart from; saying 12 °C is outside the range misreads the interval 4 °C to 18 °C, and the wrong warning would be attached to a sound prediction; 110 comes from −3 × 12 being taken as +36, giving 36 + 74 = 110, which loses the negative gradient and so predicts that a warm day sells more hot chocolate than a cold one.
- (b) 6 — Method: frequency density = frequency ÷ class width. Working: the class 12 ≤ h < 18 has width 18 − 12 = 6, so frequency density = 36 ÷ 6 = 6. Answer: the frequency density is 6 seedlings per cm. Watch which numbers you use: taking the lower bound, 12, as the width instead of 18 − 12 = 6 gives 36 ÷ 12 = 3; dividing the total number of seedlings, 90, rather than this class's frequency, 36, by the width gives 90 ÷ 6 = 15, a density that belongs to no single class; and multiplying instead of dividing gives 36 × 6 = 216, far too large a density for so narrow a class.
Build your own mix at the worksheet builder.