Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- 1.A scatter graph shows the number of days, x, that each of 16 tomato plants was watered and its height, y cm. The line of best fit has equation y = 1.5x + 4. Write down what the 1.5 in this equation tells you about the plants.y = 1.5x + 4
- 2.In a scatter graph of the age, in years, and the wingspan, in cm, of 20 birds of the same species, all the points lie close to a rising line of best fit except one, which lies a long way below the line. That bird was later found to have a damaged wing. Give a reason why this point should not be used when drawing the line of best fit.
- 3.Four pairs of variables are listed below. Write down the pair that you would expect to show no correlation.
- 4.A two-way table records whether each of 40 people at a gym in Sheffield prefers weights or cardio, and whether they are male or female. 24 of the 40 people are female. 15 of the males prefer cardio. 10 of the females prefer weights. Work out the total number of people who prefer cardio.
- 5.Forty pupils in class P and forty pupils in class Q each solved a puzzle. The times, in seconds, were summarised using cumulative frequency. For class P the lower quartile is 24, the median is 38 and the upper quartile is 46. For class Q the lower quartile is 30, the median is 35 and the upper quartile is 44. Write down the statement that correctly compares the two classes.
- 6.In a histogram of the heights, h cm, of 90 seedlings, the class 12 ≤ h < 18 contains 36 seedlings. Work out the frequency density for this class.
- 7.A scatter graph of the number of ice creams sold at a seaside kiosk in Bournemouth and the number of sunburn cases treated at a nearby pharmacy, recorded on the same 30 days, shows strong positive correlation. Which statement about this correlation is correct?
- 8.A council wants to find out what all residents of a town think about a new cycle lane. It posts a survey only on its website and asks people to fill it in online. Give a reason why this sample is likely to be biased.
- 9.Priya travels to work by one of two routes and records her journey times, in minutes, over several weeks. Route 1 has a median of 34 minutes and an interquartile range of 22 minutes. Route 2 has a median of 41 minutes and an interquartile range of 6 minutes. Priya wants the more reliable route for getting to an important meeting on time. Which route should she choose, and why?
- 10.The numbers 2, 4, 6, 8, 10 and 12 have a mean of 7 and a median of 7. The value 100 is now added to the list. Which measure is changed more by adding 100, and why?
- 11.A scatter graph shows the number of years of experience, x, of 18 sales assistants and their monthly sales, y hundred pounds. The plotted points run from x = 1 to x = 12 years, and the line of best fit is y = 4x + 20. A new assistant has 25 years of experience. Use the line of best fit to estimate a value of y for this assistant, and decide whether the estimate would be reliable.y = 4x + 20
- 12.A scatter graph of the number of hours, x, that pupils revised against their test score, y, has the line of best fit y = 2.5x + 15. Amelia wants a score of at least 80. Work out the least whole number of hours of revision the line of best fit suggests she needs.y = 2.5x + 15
- 13.Write down the statement that correctly describes the difference between correlation and causation.
- 14.The times, t minutes, taken by 120 runners to finish a fun run are summarised by these cumulative frequencies: t < 20, 8 runners; t < 30, 26 runners; t < 40, 74 runners; t < 50, 110 runners; t < 60, 120 runners. Work out the number of runners who took 40 minutes or longer to finish.
- 15.A café in York counts the number of customers in each of the nine hours it is open on one day: 4, 4, 4, 11, 13, 15, 18, 22 and 25. The owner says that a typical hour has about 4 customers, because 4 is the mode. Is the owner right? Give a reason for your answer.
Answer key
- (b) On average a plant grew 1.5 cm taller for each extra day — Method: in the equation of a line, the number multiplying x is the gradient, and a gradient states the change in y produced by an increase of 1 in x, read in the units of the two axes. Working: here x is measured in days and y in centimetres, so the gradient 1.5 carries the units centimetres per day. Testing it on the line, 5 days gives 1.5 × 5 + 4 = 11.5 cm and 6 days gives 1.5 × 6 + 4 = 13 cm, a rise of 1.5 cm for the one extra day. Answer: on average a plant grew 1.5 cm taller for each extra day of watering. The distractors: 1.5 cm as the height before any watering is the value of y when x is 0, which is the other number in the equation, 4 cm, so this swaps the gradient and the intercept; 1.5 cm as the gap between the tallest and the shortest plant reads the gradient as a range, when a range is a difference between two of the 16 plants and a gradient is a rate; 1.5 days for each extra centimetre inverts the rate, dividing days by centimetres instead of centimetres by days, and the line gives 1 cm of growth in two thirds of a day.
- (d) An outlier from the damaged wing, not the trend. — That bird's point lies a long way from the rising trend followed by every other bird, and its low wingspan is explained by the damaged wing rather than by its age — it is an outlier caused by an unusual factor, not part of the general relationship between age and wingspan, so it should not be used when drawing the line of best fit. Saying every point must be used ignores that an outlier caused by a separate, identifiable factor can rightly be set aside. Saying it shows no correlation ignores that the other 19 points do show a clear rising trend; one outlier does not remove that. Saying it proves the line is inaccurate confuses one unusual bird with a fault in the line itself, when the line correctly describes the trend followed by the rest of the data.
- (b) A person's shoe size and their favourite colour — A person's shoe size is not linked to which colour they prefer, so these two show no correlation. The other three pairs are all genuinely correlated: distance travelled and fuel used rise together, which is positive correlation; hours of revision and test score generally rise together, which is also positive correlation; and as outdoor temperature rises, fewer woolly hats are sold, which is negative correlation. Negative correlation is still a real relationship between two variables — it is not the same thing as no relationship at all, so the temperature and hats pair is not the answer to this question.
- (b) 29 — There are 40 − 24 = 16 males, and 15 of them prefer cardio, so 16 − 15 = 1 male prefers weights. There are 24 females, and 10 prefer weights, so 24 − 10 = 14 females prefer cardio. Altogether, 15 + 14 = 29 people prefer cardio. Choosing 15 only counts the males who prefer cardio and forgets the females. Choosing 11 adds the two weights figures, 1 + 10 = 11, instead of the two cardio figures. Choosing 30 comes from 40 − 10, subtracting only the number of females who prefer weights from the grand total, rather than finding both cardio sub-totals separately.
- (b) Q was faster on average and more consistent — Method: compare the medians for the average and the interquartile ranges for the spread, remembering that a shorter time is faster and a smaller interquartile range means more consistent. Working: the median for class Q is 35 seconds against 38 seconds for class P, so class Q was faster on average; the interquartile range for class P is 46 − 24 = 22 seconds and for class Q it is 44 − 30 = 14 seconds, so class Q's times are more tightly grouped. Answer: class Q was faster on average and more consistent. The distractors: calling Q slower comes from comparing the lower quartiles, 30 against 24, as though a quartile were the average; calling Q less consistent comes from using the gap between the median and the upper quartile as the spread, 44 − 35 = 9 against 46 − 38 = 8, instead of the full interquartile range; the statement that Q was both slower and less consistent comes from making both of those mistakes together.
- (b) 6 — Method: frequency density = frequency ÷ class width. Working: the class 12 ≤ h < 18 has width 18 − 12 = 6, so frequency density = 36 ÷ 6 = 6. Answer: the frequency density is 6 seedlings per cm. Watch which numbers you use: taking the lower bound, 12, as the width instead of 18 − 12 = 6 gives 36 ÷ 12 = 3; dividing the total number of seedlings, 90, rather than this class's frequency, 36, by the width gives 90 ÷ 6 = 15, a density that belongs to no single class; and multiplying instead of dividing gives 36 × 6 = 216, far too large a density for so narrow a class.
- (c) Neither causes the other; sunshine links both. — Both ice cream sales and sunburn cases tend to rise on hot, sunny days, so the amount of sunshine is a third factor linked to both — neither variable causes the other. Saying ice cream sales cause the sunburn assumes a causal link in one direction that the correlation alone cannot establish. Saying sunburn cases cause the ice cream sales assumes the reverse causal link, which is no more justified. Saying a strong correlation always means causation is the general error this question is testing: correlation, however strong, does not by itself prove that one variable causes the other.
- (d) Only internet users reach the website; others are excluded. — Method: a sample is biased when it systematically leaves out part of the population, or systematically over-represents another part. Working: anyone without internet access, or who does not visit the council's website, has NO chance of being included — the sample is drawn only from internet-using residents, which is not the whole town. Saying too many people might respond because the survey is free confuses bias with sample size — bias is about who CAN be reached, not how many respond. Saying people might lie describes a different problem, response honesty, not who was sampled in the first place. Saying online surveys cannot be anonymous is not a reason connected to bias at all. A sample is biased when part of the population has no chance of being included, whatever the reason for that.
- (c) Route 2, because its interquartile range is smaller — Method: for a journey where turning up on time matters, what matters is not the typical (median) time but how predictable it is — a smaller interquartile range means the middle half of journeys cluster closer together. Working: Route 1's median, 34 minutes, is in fact lower than Route 2's, 41 minutes, so Route 1 is faster on average; but Route 1's interquartile range, 22 minutes, is far larger than Route 2's, 6 minutes, so Route 1's times are much less predictable. Answer: Priya should choose Route 2, because its interquartile range is smaller, even though it is slower on average. Watch which statistic answers the question actually asked: Route 1 does not have the smaller interquartile range, Route 2 does, so picking Route 1 for that reason misreads the table; Route 1's median genuinely is the lower one, but a lower median answers 'which is faster', not 'which is more reliable'; and Route 2's median is not the lower one, so that claim about Route 2 is simply false.
- (c) The mean, because every value counts towards it, so 100 pulls it from 7 up to about 20.3. — Method: work each measure out before the extra value is added and again afterwards, then compare the size of the two changes. Working: before, the six values total 42, so the mean is 42 ÷ 6 = 7, and the middle pair 6 and 8 give a median of (6 + 8) ÷ 2 = 7; after, the seven values total 142, so the mean is 142 ÷ 7 = 20.29 to 2 decimal places, while the median is now the 4th of the seven ordered values, which is 8; the mean has moved by about 13.3 and the median by 1. Answer: the mean, because every value counts towards it, so 100 pulls it from 7 up to about 20.3 — this is why the median is often preferred when a data set contains an outlier. The distractors: the reply that the mean rises by 100 adds the extra value to the mean instead of adding it to the total; the reply that the median moves to 12 takes the largest of the original values as the new middle instead of counting to the 4th of the seven values; the reply about even and odd counts quotes a rule that does not exist, since the median moved because a very large value was added, not because the count of values changed.
- (b) 120, unreliable — x = 25 is outside 1 to 12 — The line of best fit is y = 4x + 20. 4 × 25 = 100, and 100 + 20 = 120, so the estimate is y = 120. But x = 25 lies far outside the plotted range of 1 to 12 years, so this is an extrapolation, and the estimate is not reliable. Reaching 100 instead of 120 comes from 4 × 25 = 100 with the intercept of 20 left out — still correctly flagged as unreliable, but the wrong value. Calling the estimate reliable simply because it was calculated correctly, giving 120, wrongly assumes that a correct calculation is automatically trustworthy, ignoring that x = 25 lies far beyond the data actually collected. Reaching 68, from 4 × 12 = 48 and 48 + 20 = 68, substitutes x = 12, the top of the plotted range, instead of the assistant's actual x = 25, and wrongly calls that reliable because 12 lies inside the range.
- (a) 26 — Method: a line of best fit lets one quantity be predicted from the other, so the score is substituted into the equation of the line and the resulting inequality is solved for the number of hours. Working: a score of at least 80 means 2.5x + 15 ≥ 80; taking 15 from both sides gives 2.5x ≥ 65, and dividing both sides by 2.5 gives x ≥ 26, so the least whole number of hours is 26. Checking, 2.5 × 26 + 15 = 80, which does reach the target. Answer: 26 hours — and this is only an estimate, because a line of best fit predicts a trend rather than an individual result, and a prediction made outside the range of hours the pupils actually revised for would be an extrapolation and less reliable still. The distractors: 27 comes from reaching 26 and then rounding up again, although 26 hours already gives a score of exactly 80; 32 comes from 80 ÷ 2.5, which ignores the 15 in the equation of the line; 38 comes from (80 + 15) ÷ 2.5, that is from adding the 15 instead of subtracting it when rearranging.
- (d) Correlation is a link; causation is one causing the other — Method: the two words describe different claims — one is about a pattern in the data, the other is about what produced that pattern. Working: correlation says only that two quantities tend to change together, which is something a scatter graph can display; causation says that a change in one quantity actually brings about the change in the other, which needs evidence a scatter graph cannot supply, because a third quantity may be driving both. Answer: correlation is a link between the quantities, while causation is one quantity causing the change in another. The distractors: the statement giving causation as the link and correlation as the cause simply swaps the two words over; the statement that the words mean the same thing is the classic error of reading a correlation as proof of cause; the statement that a scatter graph shows causation but not correlation reverses what a scatter graph can do, since the pattern it displays is exactly the correlation.
- (c) 46 — Method: the cumulative frequency table gives the number of runners below each time; to find the number at or above a time, subtract that cumulative frequency from the total. Working: the cumulative frequency for t < 40 is 74, so 120 runners in total take away the 74 who finished in under 40 minutes: 120 − 74 = 46. Answer: 46 runners took 40 minutes or longer. Watch which boundary and which subtraction you use: reading off t < 50 instead of t < 40 and subtracting, 120 − 110 = 10, answers a different question, '50 minutes or longer'; giving 74 itself as the answer reports how many finished below 40 minutes, the opposite of what was asked; and subtracting the two nearby cumulative frequencies, 110 − 74 = 36, finds how many took between 40 and 50 minutes, not everyone from 40 minutes upward.
- (c) No, the mode here is the lowest value of the nine — Method: an average is meant to stand for the data as a whole, so test any proposed average by asking how many values it sits near. Working: the value 4 appears three times and every other count appears once, so 4 is indeed the mode. But those three hours are the quiet ones at the start of the day, and the other six counts run from 11 up to 25; putting the nine counts in order, the middle one is the fifth, which is 13. So the mode sits at the very bottom of the data, with six of the nine hours far above it. Answer: no, because the mode here is the lowest value of the nine, so it describes the quiet opening hours rather than a typical hour. The distractors: saying the mode can only be used when no value repeats reverses the definition, since a mode exists only because a value does repeat; saying the mode is the value that occurs most often is a correct definition, but being the commonest value does not make a value typical when it lies at one end of the data; saying the mode is the best average for any list of numbers ignores the fact that mean, median and mode each describe a population well in different circumstances.
Build your own mix at the worksheet builder.