Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Answer key: Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- (c) Route 2, because its interquartile range is smaller — Method: for a journey where turning up on time matters, what matters is not the typical (median) time but how predictable it is — a smaller interquartile range means the middle half of journeys cluster closer together. Working: Route 1's median, 34 minutes, is in fact lower than Route 2's, 41 minutes, so Route 1 is faster on average; but Route 1's interquartile range, 22 minutes, is far larger than Route 2's, 6 minutes, so Route 1's times are much less predictable. Answer: Priya should choose Route 2, because its interquartile range is smaller, even though it is slower on average. Watch which statistic answers the question actually asked: Route 1 does not have the smaller interquartile range, Route 2 does, so picking Route 1 for that reason misreads the table; Route 1's median genuinely is the lower one, but a lower median answers 'which is faster', not 'which is more reliable'; and Route 2's median is not the lower one, so that claim about Route 2 is simply false.
- (d) 20 — 0.4x + 1 = 9, so 0.4x = 9 − 1 = 8, and 8 ÷ 0.4 = 20, so 20 sessions are needed. Choosing 23 divides 9 by 0.4 without first subtracting the 1: 9 ÷ 0.4 = 22.5, rounded up to 23. Choosing 25 subtracts the wrong way, adding the 1 instead of taking it away: 9 + 1 = 10, and 10 ÷ 0.4 = 25. Choosing 2 misplaces the decimal point in the gradient, dividing by 4 instead of by 0.4: 8 ÷ 4 = 2.
- (d) 16 — Method: rearrange interquartile range = upper quartile − lower quartile to make the lower quartile the subject: lower quartile = upper quartile − interquartile range, then check the answer sits in the right place in the list. Working: there are 11 values, so 11 + 1 = 12; the upper quartile sits at position 3 × 12 ÷ 4 = 9, which is 35, and x sits at position 12 ÷ 4 = 3, which is the lower quartile. So x = 35 − 19 = 16, and 16 does sit between the 2nd value, 9, and the 4th value, 17, as it should. Answer: x = 16. Watch how you rearrange and where you count to: adding instead of subtracting, 35 + 19 = 54, treats the interquartile range as something added on rather than a gap taken away; subtracting in the wrong order, 19 − 35 = −16, finds the right two numbers but flips the sign; and counting to the 8th value instead of the 9th treats 31 as the upper quartile, giving 31 − 19 = 12, one position short of where the upper quartile actually sits.
- (a) 1.24 — Method: for data given as a frequency table, the mean is Σfx ÷ Σf — multiply each value by its frequency, add the results, then divide by the total frequency. Working: 0 × 6 = 0. 1 × 10 = 10. 2 × 6 = 12. 3 × 3 = 9. So Σfx = 0 + 10 + 12 + 9 = 31. The total frequency is Σf = 6 + 10 + 6 + 3 = 25. Mean = 31 ÷ 25 = 1.24 siblings. Averaging the frequency column itself, (6 + 10 + 6 + 3) ÷ 4 = 6.25, mixes up the frequencies with the values they belong to. Writing down 1, the number of siblings with the highest frequency, gives the mode, not the mean. Writing down 31 stops after finding Σfx and forgets to divide by the total frequency, 25. Always divide Σfx by Σf — never stop at the top of the fraction.
- (d) Only internet users reach the website; others are excluded. — Method: a sample is biased when it systematically leaves out part of the population, or systematically over-represents another part. Working: anyone without internet access, or who does not visit the council's website, has NO chance of being included — the sample is drawn only from internet-using residents, which is not the whole town. Saying too many people might respond because the survey is free confuses bias with sample size — bias is about who CAN be reached, not how many respond. Saying people might lie describes a different problem, response honesty, not who was sampled in the first place. Saying online surveys cannot be anonymous is not a reason connected to bias at all. A sample is biased when part of the population has no chance of being included, whatever the reason for that.
- (b) On average a plant grew 1.5 cm taller for each extra day — Method: in the equation of a line, the number multiplying x is the gradient, and a gradient states the change in y produced by an increase of 1 in x, read in the units of the two axes. Working: here x is measured in days and y in centimetres, so the gradient 1.5 carries the units centimetres per day. Testing it on the line, 5 days gives 1.5 × 5 + 4 = 11.5 cm and 6 days gives 1.5 × 6 + 4 = 13 cm, a rise of 1.5 cm for the one extra day. Answer: on average a plant grew 1.5 cm taller for each extra day of watering. The distractors: 1.5 cm as the height before any watering is the value of y when x is 0, which is the other number in the equation, 4 cm, so this swaps the gradient and the intercept; 1.5 cm as the gap between the tallest and the shortest plant reads the gradient as a range, when a range is a difference between two of the 16 plants and a gradient is a rate; 1.5 days for each extra centimetre inverts the rate, dividing days by centimetres instead of centimetres by days, and the line gives 1 cm of growth in two thirds of a day.
- (d) An outlier from the damaged wing, not the trend. — That bird's point lies a long way from the rising trend followed by every other bird, and its low wingspan is explained by the damaged wing rather than by its age — it is an outlier caused by an unusual factor, not part of the general relationship between age and wingspan, so it should not be used when drawing the line of best fit. Saying every point must be used ignores that an outlier caused by a separate, identifiable factor can rightly be set aside. Saying it shows no correlation ignores that the other 19 points do show a clear rising trend; one outlier does not remove that. Saying it proves the line is inaccurate confuses one unusual bird with a fault in the line itself, when the line correctly describes the trend followed by the rest of the data.
- (d) 45 — Method: convert the angle into a fraction of the full circle, 360°, then apply that fraction to the total number of shoppers. Working: the card sector is 90° out of 360°, a fraction of 90 ÷ 360 = 0.25. Applying that fraction to the 180 shoppers gives 0.25 × 180 = 45 shoppers. Giving 90 states the angle itself, not a number of shoppers — the angle first has to be converted into a fraction. Using the remaining angle, 360 − 90 = 270°, and scaling that, 270 ÷ 360 × 180 = 135, finds the number who did NOT pay by card, not the number who did. Dividing 360 by 90, 360 ÷ 90 = 4, finds how many equal 90° sectors fit in the circle, a fact about the pie chart's shape, not about the shoppers at all. Always convert the angle to a fraction of 360° first, and apply that same fraction to the total number of people.
- (b) 29 — There are 40 − 24 = 16 males, and 15 of them prefer cardio, so 16 − 15 = 1 male prefers weights. There are 24 females, and 10 prefer weights, so 24 − 10 = 14 females prefer cardio. Altogether, 15 + 14 = 29 people prefer cardio. Choosing 15 only counts the males who prefer cardio and forgets the females. Choosing 11 adds the two weights figures, 1 + 10 = 11, instead of the two cardio figures. Choosing 30 comes from 40 − 10, subtracting only the number of females who prefer weights from the grand total, rather than finding both cardio sub-totals separately.
- (b) £11.00 — 1.5 × 6 = 9, and 9 + 2 = 11, so the estimated cost is £11.00. Choosing £9.00 stops after 1.5 × 6 = 9 and forgets to add the £2. Choosing £12.00 adds the mass and the constant first and then multiplies: 6 + 2 = 8, and 8 × 1.5 = 12.00. Choosing £13.50 swaps the gradient and the intercept, using y = 2x + 1.5 instead: 2 × 6 = 12, and 12 + 1.5 = 13.50.
- (c) Class X has the higher median and the wider spread — Method: compare the two box plots statistic by statistic — median for location, and the interquartile range for spread — checking the true value of each rather than assuming a pattern. Working: Class X's median is 60 and Class Y's is 58, so Class X's median is the higher one. Class X's interquartile range is 70 − 45 = 25 and Class Y's is 65 − 50 = 15 (and the ranges follow the same order: 95 − 20 = 75 against 80 − 35 = 45), so Class X also has the wider spread. Answer: Class X has both the higher median and the wider spread. Watch that each half of a compound statement is checked separately: claiming Class Y has the higher median and the wider spread gets both comparisons backwards; claiming Class X has the higher median but the narrower spread keeps the median right while reading the spread the wrong way round; and claiming Class Y has the higher median but the narrower spread swaps the median comparison while getting the spread right.
- (d) 1100 — Method: to combine two samples of different sizes, add the faulty counts together and add the sample sizes together before scaling up, rather than treating the two samples separately. Working: the combined sample found 34 + 21 = 55 scratched cases out of 100 + 50 = 150 cases checked, a proportion of 55 ÷ 150. Applying that proportion to the week's production of 3,000 gives an estimate of 55 ÷ 150 × 3000 = 1100 scratched cases. Averaging the two shifts' proportions instead of combining their totals, (34 ÷ 100 + 21 ÷ 50) ÷ 2 = 0.38, gives 0.38 × 3000 = 1140 — this treats the two samples as equally weighted even though Shift A checked twice as many cases as Shift B. Using only Shift A's sample, 34 ÷ 100 × 3000 = 1020, ignores Shift B's cases completely. Using only Shift B's sample, 21 ÷ 50 × 3000 = 1260, ignores Shift A's cases completely. When two samples are different sizes, combine their totals before finding the proportion — do not average the two proportions, and do not use only one shift's sample.
- (a) A vertical line chart (discrete numerical data) — The number of pets is discrete numerical data — whole-number values such as 0, 1, 2, 3 or 4 — recorded for one variable, so a vertical line chart is the chart specified for this kind of data. A bar chart is used for categorical data, such as favourite colour, not numerical values counted like this. A pie chart shows proportions of a whole and does not show the frequency of each separate value. A scatter graph compares two different variables against each other, and only one variable, the number of pets, is recorded here.
- (a) 3/10 — There are 60 − 9 = 51 right-handed pupils. Of the 24 pupils who prefer Art, 6 are left-handed, so 24 − 6 = 18 are right-handed and prefer Art. The probability that a randomly chosen pupil is right-handed and prefers Art is 18/60, which simplifies to 3/10. Giving 2/5 is 24/60 simplified — the probability of preferring Art, ignoring the right-handed condition entirely. Giving 17/20 is 51/60 simplified — the probability of being right-handed, ignoring the Art condition entirely. Giving 1/10 is 6/60 simplified — the probability of being left-handed and preferring Art, the wrong hand condition.
- (a) 31.5 — Method: to estimate the mean from a histogram, first turn each bar into a frequency (frequency density × class width), then use mean = Σ(frequency × midpoint) ÷ Σfrequency, with the midpoint standing in for every value in that class. Working: the four classes have widths 20, 10, 20 and 20, so their frequencies are 1 × 20 = 20, 3 × 10 = 30, 2 × 20 = 40 and 0.5 × 20 = 10, which do add to the 100 vehicles stated. Their midpoints are 10, 25, 40 and 60, so Σfx = 20 × 10 + 30 × 25 + 40 × 40 + 10 × 60 = 200 + 750 + 1600 + 600 = 3150, and the mean is 3150 ÷ 100 = 31.5. Answer: the estimated mean speed is 31.5 mph. Watch which numbers you treat as the frequencies and which as the values: using the frequency densities themselves as the frequencies, without multiplying by the class widths first, gives 1 × 10 + 3 × 25 + 2 × 40 + 0.5 × 60 = 195 spread over 1 + 3 + 2 + 0.5 = 6.5, and 195 ÷ 6.5 = 30, a mean built from the wrong 'frequencies' altogether; averaging the four midpoints on their own, (10 + 25 + 40 + 60) ÷ 4 = 33.75, ignores how many vehicles are actually in each class; and using each class's lower boundary in place of its midpoint, 20 × 0 + 30 × 20 + 40 × 30 + 10 × 50 = 2300 and 2300 ÷ 100 = 23, systematically underestimates every class by roughly half its width.
Build your own mix at the worksheet builder.