Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- 1.The times, t minutes, taken by 120 runners to finish a fun run are summarised by these cumulative frequencies: t < 20, 8 runners; t < 30, 26 runners; t < 40, 74 runners; t < 50, 110 runners; t < 60, 120 runners. Work out the number of runners who took 40 minutes or longer to finish.
- 2.Two classes at a school in Coventry sit the same maths test, out of 20 marks. Class A has a mean mark of 14 and a range of 6. Class B has a mean mark of 14 and a range of 14. Write a sentence comparing the two classes, using the mean and the range.
- 3.Box plots summarise the test scores of two classes. Class X: minimum 20, lower quartile 45, median 60, upper quartile 70, maximum 95. Class Y: minimum 35, lower quartile 50, median 58, upper quartile 65, maximum 80. Which statement about the two classes is correct?
- 4.A box plot for the amount of pocket money, in pounds, saved by 30 students in a month is based on a lower quartile of £18 and an upper quartile of £42. A value above upper quartile + 1.5 × interquartile range is considered an outlier. Amara saved £80 in the month. Determine whether Amara's saving is an outlier.
- 5.A two-way table records whether each of 40 people at a gym in Sheffield prefers weights or cardio, and whether they are male or female. 24 of the 40 people are female. 15 of the males prefer cardio. 10 of the females prefer weights. Work out the total number of people who prefer cardio.
- 6.A gym draws a histogram of the times, t minutes, that its members spend on one machine. The bar for 0 ≤ t < 10 has a frequency density of 1.8 per minute, the bar for 10 ≤ t < 25 has a frequency density of 3.2 per minute, and the bar for 25 ≤ t < 55 has a frequency density of 0.9 per minute. Members who spend 10 minutes or more on the machine pay an extra charge. Work out the number of members who pay the extra charge.
- 7.A scatter graph of the number of hours, x, that pupils revised against their test score, y, has the line of best fit y = 2.5x + 15. Amelia wants a score of at least 80. Work out the least whole number of hours of revision the line of best fit suggests she needs.y = 2.5x + 15
- 8.A Year 10 class has 20 boys with a mean height of 150 cm and 10 girls with a mean height of 168 cm. Work out the mean height of all 30 pupils in the class.
- 9.A charity shop in Leicester made £2,400 in one month. A pie chart shows how this was spent: repairs took up an angle of 90°, wages took up an angle of 150°, and the rest was other costs. Work out how much was spent on wages.
- 10.A town council wants to find out about the eating habits of the people who live in the town. It asks only the members of a local sports club. Give a reason why this sample is biased.
- 11.A café owner in Brighton records the midday temperature, x °C, and the number of hot chocolates sold, y, on 12 days. The temperatures recorded run from 4 °C to 18 °C, and the line of best fit is y = −3x + 74. The forecast for tomorrow gives a midday temperature of 12 °C. Work out the number the line of best fit predicts, and write down how much confidence the owner can have in it.y = -3x + 74
- 12.Four pairs of variables are listed below. Write down the pair that you would expect to show no correlation.
- 13.A council wants to find out what all residents of a town think about a new cycle lane. It posts a survey only on its website and asks people to fill it in online. Give a reason why this sample is likely to be biased.
- 14.A factory makes a batch of 4,000 circuit boards. It checks a random sample of 50 boards and finds that 4 are faulty. The factory will scrap the whole batch if the estimated number of faulty boards in the batch is more than 250. Should the factory scrap the batch?
- 15.A scatter graph has 50 points. Most of them lie close to a rising line of best fit, but two of them lie a long way from that line. Write down how those two points should be treated.
Answer key
- (c) 46 — Method: the cumulative frequency table gives the number of runners below each time; to find the number at or above a time, subtract that cumulative frequency from the total. Working: the cumulative frequency for t < 40 is 74, so 120 runners in total take away the 74 who finished in under 40 minutes: 120 − 74 = 46. Answer: 46 runners took 40 minutes or longer. Watch which boundary and which subtraction you use: reading off t < 50 instead of t < 40 and subtracting, 120 − 110 = 10, answers a different question, '50 minutes or longer'; giving 74 itself as the answer reports how many finished below 40 minutes, the opposite of what was asked; and subtracting the two nearby cumulative frequencies, 110 − 74 = 36, finds how many took between 40 and 50 minutes, not everyone from 40 minutes upward.
- (a) Equal means; Class A is more consistent, smaller range. — Method: when two data sets share a measure of location, compare a measure of spread to say more about consistency. Working: both classes have the same mean mark, 14, so on average they performed equally well. Class A has the smaller range, 6, so its marks are more tightly grouped around 14 than Class B's marks, which vary by as much as 14. So Class A's marks were more consistent, even though neither class did better on average. Saying Class B did better because it has the bigger range confuses a wide spread with a high score — a big range describes variability, not performance. Saying Class A did better because it has the smaller range makes the same mistake in the other direction: the two classes are tied on the mean, so neither one 'did better'. Saying the classes cannot be compared because their means are equal misses the whole point of also comparing the range. Always compare both an average AND a spread before describing two data sets — either one alone tells only half the story.
- (c) Class X has the higher median and the wider spread — Method: compare the two box plots statistic by statistic — median for location, and the interquartile range for spread — checking the true value of each rather than assuming a pattern. Working: Class X's median is 60 and Class Y's is 58, so Class X's median is the higher one. Class X's interquartile range is 70 − 45 = 25 and Class Y's is 65 − 50 = 15 (and the ranges follow the same order: 95 − 20 = 75 against 80 − 35 = 45), so Class X also has the wider spread. Answer: Class X has both the higher median and the wider spread. Watch that each half of a compound statement is checked separately: claiming Class Y has the higher median and the wider spread gets both comparisons backwards; claiming Class X has the higher median but the narrower spread keeps the median right while reading the spread the wrong way round; and claiming Class Y has the higher median but the narrower spread swaps the median comparison while getting the spread right.
- (c) Yes — £80 is above the boundary, £78 — Method: a value counts as an outlier when it lies more than 1.5 times the interquartile range beyond the nearer quartile; here that means checking it against upper quartile + 1.5 × interquartile range. Working: the interquartile range is 42 − 18 = 24. 1.5 × 24 = 36, and 42 + 36 = 78, so any saving above £78 is an outlier. Amara saved £80, and 80 is greater than 78. Answer: yes, Amara's saving is an outlier, because £80 is above the outlier boundary, £78. Watch how you build the boundary and what you compare it with: adding the two quartiles instead of subtracting them, 42 + 18 = 60, gives an interquartile range three times too big, and 42 + 1.5 × 60 = 42 + 90 = 132 puts the boundary so far out that £80 wrongly looks ordinary; comparing £80 with the upper quartile alone, £42, checks only that it lies in the top quarter of the data, which every value above £42 does, not that it lies unusually far beyond it; and adding the interquartile range on once instead of one and a half times, 42 + 24 = 66, uses the wrong multiplier, even though £80 still happens to clear that lower boundary too.
- (b) 29 — There are 40 − 24 = 16 males, and 15 of them prefer cardio, so 16 − 15 = 1 male prefers weights. There are 24 females, and 10 prefer weights, so 24 − 10 = 14 females prefer cardio. Altogether, 15 + 14 = 29 people prefer cardio. Choosing 15 only counts the males who prefer cardio and forgets the females. Choosing 11 adds the two weights figures, 1 + 10 = 11, instead of the two cardio figures. Choosing 30 comes from 40 − 10, subtracting only the number of females who prefer weights from the grand total, rather than finding both cardio sub-totals separately.
- (c) 75 — Method: the number in a class is the area of its bar, frequency density × class width, so work out the frequency of each class that lies at or above 10 minutes and add them. Working: the class 10 ≤ t < 25 is 15 minutes wide with a frequency density of 3.2, giving 3.2 × 15 = 48 members; the class 25 ≤ t < 55 is 30 minutes wide with a frequency density of 0.9, giving 0.9 × 30 = 27 members; the total charged is 48 + 27 = 75. Answer: 75 members pay the extra charge. The distractors: 4.1 comes from adding the two frequency densities, 3.2 + 0.9, as though each height were a count; 93 comes from including the class 0 ≤ t < 10 as well, 1.8 × 10 = 18 added to 48 and 27, which charges every member; 27 comes from using only the class 25 ≤ t < 55 and forgetting that 10 ≤ t < 25 is also at or above 10 minutes.
- (a) 26 — Method: a line of best fit lets one quantity be predicted from the other, so the score is substituted into the equation of the line and the resulting inequality is solved for the number of hours. Working: a score of at least 80 means 2.5x + 15 ≥ 80; taking 15 from both sides gives 2.5x ≥ 65, and dividing both sides by 2.5 gives x ≥ 26, so the least whole number of hours is 26. Checking, 2.5 × 26 + 15 = 80, which does reach the target. Answer: 26 hours — and this is only an estimate, because a line of best fit predicts a trend rather than an individual result, and a prediction made outside the range of hours the pupils actually revised for would be an extrapolation and less reliable still. The distractors: 27 comes from reaching 26 and then rounding up again, although 26 hours already gives a score of exactly 80; 32 comes from 80 ÷ 2.5, which ignores the 15 in the equation of the line; 38 comes from (80 + 15) ÷ 2.5, that is from adding the 15 instead of subtracting it when rearranging.
- (c) 156 cm — Method: to combine two groups' means, multiply each group's mean by its own number of pupils, add the two totals together, then divide by the total number of pupils in both groups. Working: 20 × 150 = 3,000 cm for the boys and 10 × 168 = 1,680 cm for the girls, giving a combined total of 3,000 + 1,680 = 4,680 cm. Dividing by all 30 pupils gives 4,680 ÷ 30 = 156 cm. Giving 159 cm averages the two means, (150 + 168) ÷ 2, treating the two groups as if they had the same number of pupils, when there are twice as many boys as girls. Giving 4,680 cm finds the correct combined total height but stops there, forgetting the final division by the 30 pupils. Giving 234 cm divides the combined total by 20, the number of boys only, forgetting that the total also includes the 10 girls. Always weight each mean by its own group size, and always divide by the TOTAL number of pupils in both groups combined.
- (b) £1,000 — Wages take up 150° out of 360°, so the amount spent on wages is 150 ÷ 360 × 2400 = £1,000. Choosing £600 uses the repairs angle, 90°, instead of the wages angle: 90 ÷ 360 × 2400 = 600. Choosing £3,600 treats the angle in degrees as if it were a percentage, 150 ÷ 100 × 2400 = 3600, instead of dividing by 360°. Choosing £800 uses the angle for the 'other costs' sector, 360 − 90 − 150 = 120°, instead of the wages sector: 120 ÷ 360 × 2400 = 800.
- (c) Club members probably eat differently from most people — Method: a sample is biased when the group it is drawn from differs from the population in the very thing the survey is measuring, so compare the subgroup with the population on that quantity. Working: the survey measures eating habits, and people who join a sports club take more exercise than average and are known to eat differently from the town as a whole, so their replies pull the results away from the true picture for the town however many of them are asked. Answer: club members probably eat differently from most people. The distractors: the reply about the number of members treats bias as a question of size, but a large biased sample is still biased; the reply that the members were picked at random is false, since the council picked a club rather than picking residents, and it confuses bias with non-response; the reply that everyone asked lives in the town notes something true of the members but draws the false conclusion that the sample therefore covers the town, when a sample must reflect a population and not merely be taken from inside it.
- (d) 38, and fairly confident, as 12 °C is inside the range — Method: substitute the forecast temperature into the equation of the line of best fit, then judge the prediction by where that temperature sits among the data the line was drawn from. Working: putting x = 12 into y = −3x + 74 gives −3 × 12 + 74 = 38, so the line predicts 38 hot chocolates. The recorded temperatures run from 4 °C to 18 °C, and 12 °C lies inside that interval, so this is interpolation, the safer kind of prediction. Answer: 38, and fairly confident, as 12 °C is inside the range; the owner should still expect the true figure to differ a little, since the points only lie near the line and not on it. The distractors: being completely certain treats a line of best fit as a rule that fixes each day's sales, when it describes a trend that individual days depart from; saying 12 °C is outside the range misreads the interval 4 °C to 18 °C, and the wrong warning would be attached to a sound prediction; 110 comes from −3 × 12 being taken as +36, giving 36 + 74 = 110, which loses the negative gradient and so predicts that a warm day sells more hot chocolate than a cold one.
- (b) A person's shoe size and their favourite colour — A person's shoe size is not linked to which colour they prefer, so these two show no correlation. The other three pairs are all genuinely correlated: distance travelled and fuel used rise together, which is positive correlation; hours of revision and test score generally rise together, which is also positive correlation; and as outdoor temperature rises, fewer woolly hats are sold, which is negative correlation. Negative correlation is still a real relationship between two variables — it is not the same thing as no relationship at all, so the temperature and hats pair is not the answer to this question.
- (d) Only internet users reach the website; others are excluded. — Method: a sample is biased when it systematically leaves out part of the population, or systematically over-represents another part. Working: anyone without internet access, or who does not visit the council's website, has NO chance of being included — the sample is drawn only from internet-using residents, which is not the whole town. Saying too many people might respond because the survey is free confuses bias with sample size — bias is about who CAN be reached, not how many respond. Saying people might lie describes a different problem, response honesty, not who was sampled in the first place. Saying online surveys cannot be anonymous is not a reason connected to bias at all. A sample is biased when part of the population has no chance of being included, whatever the reason for that.
- (c) Yes — with an estimate of 320, above the 250 limit. — Method: scale the sample proportion up to the whole batch to get an estimate, then compare that estimate with the 250 limit to reach a decision. Working: in the sample, 4 out of 50 boards are faulty, a proportion of 4 ÷ 50 = 0.08. Applying that proportion to the batch of 4,000 gives an estimate of 0.08 × 4000 = 320 faulty boards. Since 320 is more than 250, the factory should scrap the batch. Inverting the proportion, 50 ÷ 4 = 12.5, and treating that as a percentage of the batch, 12.5% × 4000 = 500, still gives 'yes' but from the wrong fraction, so it overstates the estimate. Comparing the raw number of faulty boards found in the sample, 4, directly with the 250 limit skips the scaling up to the batch altogether, and 4 is nowhere near 250, so that route wrongly says 'no'. Dividing the batch by the sample size, 4000 ÷ 50 = 80, finds how many samples of 50 fit into the batch but stops before multiplying by the 4 faulty boards found, so it also wrongly says 'no'. Always find the proportion in the sample first, scale it up to the whole batch, and only then compare the estimate with the limit given.
- (b) Treat them as outliers and check them before deciding — Method: a point lying a long way from the pattern the rest of the data make is called an outlier, and an outlier is investigated before anything is done with it, because it may be an error in the data or it may be a genuine but unusual case. Working: 48 of the 50 points lie close to the rising line of best fit, so the trend is set by those 48; the two remaining points do not follow it, so they are identified as outliers and checked — a mistake in measuring or recording would be corrected, while a genuine reading would be kept and reported. Answer: treat them as outliers and check them before deciding what to do with them. The distractors: deleting them at once assumes that every point far from the line must be an error, which throws away real data; moving the line so that it passes through them assumes a line of best fit must touch particular points, when it is drawn to follow all 50; taking them as proof that there is no correlation lets two points overturn the pattern that the other 48 agree on.
Build your own mix at the worksheet builder.