Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- 1.Box plots summarise house prices, in thousands of pounds, in two towns. Town A: minimum 120, lower quartile 180, median 220, upper quartile 310, maximum 520. Town B: minimum 95, lower quartile 175, median 235, upper quartile 260, maximum 340. A buyer wants the town where house prices are typically higher but more consistent. Which statement is correct?
- 2.A charity shop in Leicester made £2,400 in one month. A pie chart shows how this was spent: repairs took up an angle of 90°, wages took up an angle of 150°, and the rest was other costs. Work out how much was spent on wages.
- 3.A quality inspector weighs a random sample of 50 packets of crisps from one day's production and finds their mean mass is 32.4 g. The factory makes 20,000 packets that day. Work out an estimate for the total mass, in kg, of all the packets made that day.
- 4.A vet records the masses, m kg, of the dogs seen in one week as a histogram. The bar for 0 ≤ m < 5 has a frequency density of 4 per kg, the bar for 5 ≤ m < 15 has a frequency density of 2.6 per kg, and the bar for 15 ≤ m < 40 has a frequency density of 1.2 per kg. Work out the total number of dogs seen that week.
- 5.A survey of 25 pupils in Derby records how many siblings each has: 0 siblings — 6 pupils, 1 sibling — 10 pupils, 2 siblings — 6 pupils, 3 siblings — 3 pupils. Calculate the mean number of siblings.
- 6.A histogram shows the speeds, v mph, of 100 vehicles passing a checkpoint. The bar for 0 ≤ v < 20 has a frequency density of 1 vehicle per mph, the bar for 20 ≤ v < 30 has a frequency density of 3 vehicles per mph, the bar for 30 ≤ v < 50 has a frequency density of 2 vehicles per mph, and the bar for 50 ≤ v < 70 has a frequency density of 0.5 vehicles per mph. Estimate the mean speed of the vehicles.
- 7.In a histogram of the lengths, x cm, of some rods, the bar for 10 ≤ x < 30 has a frequency density of 3 per cm. The bar for 30 ≤ x < 45 is twice as tall as the bar for 10 ≤ x < 30. Work out the number of rods with a length in the class 30 ≤ x < 45.
- 8.A scatter graph shows the number of years of experience, x, of 18 sales assistants and their monthly sales, y hundred pounds. The plotted points run from x = 1 to x = 12 years, and the line of best fit is y = 4x + 20. A new assistant has 25 years of experience. Use the line of best fit to estimate a value of y for this assistant, and decide whether the estimate would be reliable.y = 4x + 20
- 9.A school has 1,500 pupils. The head teacher takes a random sample of 150 of them from the school register and asks how long they spend on homework. Rory says the sample is too small for the result to mean anything. Is Rory right? Give a reason for your answer.
- 10.The masses, m kg, of 80 sacks of grain are summarised by these cumulative frequencies: m < 10, 6 sacks; m < 20, 22 sacks; m < 30, 58 sacks; m < 40, 74 sacks; m < 50, 80 sacks. Use interpolation to estimate the median mass.
- 11.A two-way table records whether each of 40 people at a gym in Sheffield prefers weights or cardio, and whether they are male or female. 24 of the 40 people are female. 15 of the males prefer cardio. 10 of the females prefer weights. Work out the total number of people who prefer cardio.
- 12.A café owner in Brighton records the midday temperature, x °C, and the number of hot chocolates sold, y, on 12 days. The temperatures recorded run from 4 °C to 18 °C, and the line of best fit is y = −3x + 74. The forecast for tomorrow gives a midday temperature of 12 °C. Work out the number the line of best fit predicts, and write down how much confidence the owner can have in it.y = -3x + 74
- 13.A two-way table records the favourite subject, Maths or Art, of 60 pupils in Year 10, and whether each pupil is left-handed or right-handed. 9 of the 60 pupils are left-handed, and 6 of those left-handed pupils prefer Art. In total, 24 of the 60 pupils prefer Art. A pupil is chosen at random from the 60. Work out the probability that the pupil is right-handed and prefers Art.
- 14.A pie chart shows how 180 shoppers at a supermarket paid for their shopping. The sector for card payments has an angle of 90° at the centre. Work out how many of the 180 shoppers paid by card.
- 15.A teacher records the number of pets owned by each of 25 pupils in a class; each pupil owns 0, 1, 2, 3 or 4 pets. The teacher wants to show how many pupils own each number of pets. Write down the most suitable type of chart for this data, and give a reason for your answer.
Answer key
- (b) Town B — higher median and smaller IQR — Method: 'higher and more consistent' needs two comparisons — the median for typical price, and the interquartile range for spread, with a smaller interquartile range meaning more consistent. Working: Town B's median, £235,000, is higher than Town A's, £220,000. Town A's interquartile range is 310 − 180 = 130 and Town B's is 260 − 175 = 85, so Town B's interquartile range is the smaller of the two. Answer: Town B has both the higher median and the smaller interquartile range, so it is the town with higher, more consistent prices. Watch which combination of median and interquartile range each statement claims, and for which town: claiming Town A has the higher median and the smaller interquartile range gets both comparisons wrong, since Town B leads on both; claiming Town A has the higher median (still wrong) but the larger interquartile range at least reads the spread correctly, without it rescuing the false median claim; and claiming Town B has the higher median (correct) but the larger interquartile range misreads the spread — Town B's interquartile range is the smaller of the two, not the larger.
- (b) £1,000 — Wages take up 150° out of 360°, so the amount spent on wages is 150 ÷ 360 × 2400 = £1,000. Choosing £600 uses the repairs angle, 90°, instead of the wages angle: 90 ÷ 360 × 2400 = 600. Choosing £3,600 treats the angle in degrees as if it were a percentage, 150 ÷ 100 × 2400 = 3600, instead of dividing by 360°. Choosing £800 uses the angle for the 'other costs' sector, 360 − 90 − 150 = 120°, instead of the wages sector: 120 ÷ 360 × 2400 = 800.
- (d) 648 kg — Method: to estimate a total from a sample, multiply the sample's mean by the number of items in the whole population, then check the units the question asks for. Working: 32.4 g × 20,000 = 648,000 g. Converting to kilograms, 648,000 ÷ 1,000 = 648 kg. This is only an estimate, not an exact total, because it assumes every one of the 20,000 packets has exactly the sample mean mass, when in reality individual packets vary above and below it. Giving 1.62 kg multiplies the mean by 50, the SAMPLE size, instead of by 20,000, the number of packets actually made that day — this finds the total mass of the 50 sampled packets, not the day's production. Giving 32.4 kg treats the sample mean itself, in grams, as if it already were the day's total mass in kilograms, skipping the scaling up altogether. Giving 648,000 kg correctly scales the mean up to the whole day's production but never converts the answer from grams to kilograms, leaving it 1,000 times too large. Always scale a sample's mean up by the SIZE OF THE WHOLE POPULATION, and always finish by checking the units the question asks for.
- (c) 76 — Method: the frequency of each class is the area of its bar, frequency density × class width, so work out all three frequencies and add them. Working: the widths are 5, 10 and 25 kg, so the frequencies are 4 × 5 = 20, 2.6 × 10 = 26 and 1.2 × 25 = 30; the total is 20 + 26 + 30 = 76. Answer: 76 dogs. The distractors: 7.8 comes from adding the three frequency densities, 4 + 2.6 + 1.2, treating each height as though it were a count; 107 comes from using each upper class boundary as the width, giving 4 × 5, 2.6 × 15 and 1.2 × 40; 39 comes from using the first class width, 5, for every bar, which ignores the unequal intervals and gives 20 + 13 + 6.
- (a) 1.24 — Method: for data given as a frequency table, the mean is Σfx ÷ Σf — multiply each value by its frequency, add the results, then divide by the total frequency. Working: 0 × 6 = 0. 1 × 10 = 10. 2 × 6 = 12. 3 × 3 = 9. So Σfx = 0 + 10 + 12 + 9 = 31. The total frequency is Σf = 6 + 10 + 6 + 3 = 25. Mean = 31 ÷ 25 = 1.24 siblings. Averaging the frequency column itself, (6 + 10 + 6 + 3) ÷ 4 = 6.25, mixes up the frequencies with the values they belong to. Writing down 1, the number of siblings with the highest frequency, gives the mode, not the mean. Writing down 31 stops after finding Σfx and forgets to divide by the total frequency, 25. Always divide Σfx by Σf — never stop at the top of the fraction.
- (a) 31.5 — Method: to estimate the mean from a histogram, first turn each bar into a frequency (frequency density × class width), then use mean = Σ(frequency × midpoint) ÷ Σfrequency, with the midpoint standing in for every value in that class. Working: the four classes have widths 20, 10, 20 and 20, so their frequencies are 1 × 20 = 20, 3 × 10 = 30, 2 × 20 = 40 and 0.5 × 20 = 10, which do add to the 100 vehicles stated. Their midpoints are 10, 25, 40 and 60, so Σfx = 20 × 10 + 30 × 25 + 40 × 40 + 10 × 60 = 200 + 750 + 1600 + 600 = 3150, and the mean is 3150 ÷ 100 = 31.5. Answer: the estimated mean speed is 31.5 mph. Watch which numbers you treat as the frequencies and which as the values: using the frequency densities themselves as the frequencies, without multiplying by the class widths first, gives 1 × 10 + 3 × 25 + 2 × 40 + 0.5 × 60 = 195 spread over 1 + 3 + 2 + 0.5 = 6.5, and 195 ÷ 6.5 = 30, a mean built from the wrong 'frequencies' altogether; averaging the four midpoints on their own, (10 + 25 + 40 + 60) ÷ 4 = 33.75, ignores how many vehicles are actually in each class; and using each class's lower boundary in place of its midpoint, 20 × 0 + 30 × 20 + 40 × 30 + 10 × 50 = 2300 and 2300 ÷ 100 = 23, systematically underestimates every class by roughly half its width.
- (c) 90 — Method: the height of a bar is its frequency density, so twice as tall means twice the frequency density — not twice the frequency, because the two classes have different widths. Then frequency = frequency density × class width. Working: the first bar has frequency density 3 per cm, so the second has frequency density 2 × 3 = 6 per cm; the class 30 ≤ x < 45 is 45 − 30 = 15 cm wide, so its frequency is 6 × 15 = 90. Answer: 90 rods. The distractors: 120 comes from doubling the first bar's frequency instead of its height — the first class holds 3 × 20 = 60 rods, and doubling that ignores the fact that the second class is narrower; 45 comes from using the first bar's frequency density, 3, for the second bar, 3 × 15, and so never using the information that it is twice as tall; 6 comes from stopping at the frequency density of the taller bar and quoting a height as though it were a count.
- (b) 120, unreliable — x = 25 is outside 1 to 12 — The line of best fit is y = 4x + 20. 4 × 25 = 100, and 100 + 20 = 120, so the estimate is y = 120. But x = 25 lies far outside the plotted range of 1 to 12 years, so this is an extrapolation, and the estimate is not reliable. Reaching 100 instead of 120 comes from 4 × 25 = 100 with the intercept of 20 left out — still correctly flagged as unreliable, but the wrong value. Calling the estimate reliable simply because it was calculated correctly, giving 120, wrongly assumes that a correct calculation is automatically trustworthy, ignoring that x = 25 lies far beyond the data actually collected. Reaching 68, from 4 × 12 = 48 and 48 + 20 = 68, substitutes x = 12, the top of the plotted range, instead of the assistant's actual x = 25, and wrongly calls that reliable because 12 lies inside the range.
- (a) No, 150 pupils are a tenth of the school, chosen at random — Method: judge a sample on two things, whether every member of the population had the same chance of being chosen, and whether the sample is large enough to carry a pattern. Working: the 150 pupils were drawn from the register of every pupil in the school, so no year group or set is shut out and no pupil chooses to take part; and 150 ÷ 1,500 = 0.1, so one pupil in ten has been asked. A random sample of that share is ample for an estimate of how long the school's pupils spend on homework. Answer: no, because 150 pupils are a tenth of the school and were chosen at random. The distractors: saying a random sample always gives the exact school figure reaches the same verdict for a reason that is false, since a second random sample of 150 would give a slightly different mean; saying 150 pupils cannot be picked at random from 1,500 treats randomness as something only a whole population can have, when drawing names from the register is exactly how a random sample is taken; saying that only asking all 1,500 could show anything rejects sampling altogether, which would leave no way to study any population too large to count.
- (d) 25 — Method: the median is estimated at position n ÷ 2 in the cumulative frequency table, then interpolated across the class it falls in: lower boundary, plus the fraction of the way through the class, times the class width. Working: there are 80 sacks, so the median sits at position 80 ÷ 2 = 40. Before the class 20 ≤ m < 30 the cumulative frequency is 22, and by the end of it, it is 58, so this class holds the 40th sack; its frequency is 58 − 22 = 36 and its width is 30 − 20 = 10. The extra distance needed into the class is 40 − 22 = 18, and 18 ÷ 36 × 10 = 5, so the median is 20 + 5 = 25. Answer: the estimated median mass is 25 kg. Watch which numbers the interpolation uses: reading off just the lower boundary of the median class, 20, ignores how far into that class the 40th sack actually falls; treating n ÷ 2 = 40 itself as the median mass mistakes a position in the list for a mass in kilograms; and using the target position, 40, as the extra distance into the class instead of subtracting the sacks already counted changes the calculation to 20 + 40 ÷ 36 × 10. That comes to 20 + 11.1 = 31.1, overshooting the class because it never subtracts the 22 sacks already counted before it.
- (b) 29 — There are 40 − 24 = 16 males, and 15 of them prefer cardio, so 16 − 15 = 1 male prefers weights. There are 24 females, and 10 prefer weights, so 24 − 10 = 14 females prefer cardio. Altogether, 15 + 14 = 29 people prefer cardio. Choosing 15 only counts the males who prefer cardio and forgets the females. Choosing 11 adds the two weights figures, 1 + 10 = 11, instead of the two cardio figures. Choosing 30 comes from 40 − 10, subtracting only the number of females who prefer weights from the grand total, rather than finding both cardio sub-totals separately.
- (d) 38, and fairly confident, as 12 °C is inside the range — Method: substitute the forecast temperature into the equation of the line of best fit, then judge the prediction by where that temperature sits among the data the line was drawn from. Working: putting x = 12 into y = −3x + 74 gives −3 × 12 + 74 = 38, so the line predicts 38 hot chocolates. The recorded temperatures run from 4 °C to 18 °C, and 12 °C lies inside that interval, so this is interpolation, the safer kind of prediction. Answer: 38, and fairly confident, as 12 °C is inside the range; the owner should still expect the true figure to differ a little, since the points only lie near the line and not on it. The distractors: being completely certain treats a line of best fit as a rule that fixes each day's sales, when it describes a trend that individual days depart from; saying 12 °C is outside the range misreads the interval 4 °C to 18 °C, and the wrong warning would be attached to a sound prediction; 110 comes from −3 × 12 being taken as +36, giving 36 + 74 = 110, which loses the negative gradient and so predicts that a warm day sells more hot chocolate than a cold one.
- (a) 3/10 — There are 60 − 9 = 51 right-handed pupils. Of the 24 pupils who prefer Art, 6 are left-handed, so 24 − 6 = 18 are right-handed and prefer Art. The probability that a randomly chosen pupil is right-handed and prefers Art is 18/60, which simplifies to 3/10. Giving 2/5 is 24/60 simplified — the probability of preferring Art, ignoring the right-handed condition entirely. Giving 17/20 is 51/60 simplified — the probability of being right-handed, ignoring the Art condition entirely. Giving 1/10 is 6/60 simplified — the probability of being left-handed and preferring Art, the wrong hand condition.
- (d) 45 — Method: convert the angle into a fraction of the full circle, 360°, then apply that fraction to the total number of shoppers. Working: the card sector is 90° out of 360°, a fraction of 90 ÷ 360 = 0.25. Applying that fraction to the 180 shoppers gives 0.25 × 180 = 45 shoppers. Giving 90 states the angle itself, not a number of shoppers — the angle first has to be converted into a fraction. Using the remaining angle, 360 − 90 = 270°, and scaling that, 270 ÷ 360 × 180 = 135, finds the number who did NOT pay by card, not the number who did. Dividing 360 by 90, 360 ÷ 90 = 4, finds how many equal 90° sectors fit in the circle, a fact about the pie chart's shape, not about the shoppers at all. Always convert the angle to a fraction of 360° first, and apply that same fraction to the total number of people.
- (a) A vertical line chart (discrete numerical data) — The number of pets is discrete numerical data — whole-number values such as 0, 1, 2, 3 or 4 — recorded for one variable, so a vertical line chart is the chart specified for this kind of data. A bar chart is used for categorical data, such as favourite colour, not numerical values counted like this. A pie chart shows proportions of a whole and does not show the frequency of each separate value. A scatter graph compares two different variables against each other, and only one variable, the number of pets, is recorded here.
Build your own mix at the worksheet builder.