Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- 1.A bus company records the delay, d minutes, of 250 buses: 0 ≤ d < 2, 60 buses; 2 ≤ d < 5, 90 buses; 5 ≤ d < 10, 75 buses; 10 ≤ d < 20, 25 buses. The company refunds the fare whenever a bus is more than 8 minutes late. Estimate the number of refunds it must pay.
- 2.A recruitment agency in Manchester compares the weekly pay, in pounds, of seven employees at two branches. Branch A: £480, £495, £500, £505, £510, £515, £1,200. Branch B: £480, £490, £500, £510, £520, £530, £540. An advert for the agency claims 'Branch A pays more on average.' Decide whether this claim is fairly supported, using an appropriate average, and choose the correct conclusion.
- 3.The midday temperature in Leeds was recorded on each day of one week: 22 °C, 24 °C, 23 °C, 25 °C, 26 °C, 21 °C, 24 °C. Work out the mean midday temperature, giving your answer to 2 decimal places.
- 4.In a random sample of 40 pupils at a school, 6 are left-handed. The school has 900 pupils. Work out an estimate for the number of left-handed pupils in the school.
- 5.A stacked bar for one day at a café in Norwich shows total drink sales of 50 drinks, split into three parts: 22 were tea, 15 were coffee and the rest were hot chocolate. Work out the number of hot chocolates sold.
- 6.A shop compares customer waiting times, in minutes, at two branches over one week. Branch A had a median of 12 minutes and an interquartile range of 5 minutes. Branch B had a median of 9 minutes and an interquartile range of 11 minutes. Compare the waiting times at the two branches.
- 7.A frequency polygon for the mass, in kg, of 40 parcels at a delivery depot is drawn by plotting one point at the midpoint of each class, joined by straight lines: (5, 6), (15, 10), (25, 16), (35, 6), (45, 2). Every class has a width of 10 kg. Write down the modal class.
- 8.The mean of the four numbers 10, 15, 20 and x is 18. Work out the value of x.
- 9.A scatter graph shows the number of days, x, that each of 16 tomato plants was watered and its height, y cm. The line of best fit has equation y = 1.5x + 4. Write down what the 1.5 in this equation tells you about the plants.y = 1.5x + 4
- 10.A scatter graph of the number of hours, x, that pupils revised against their test score, y, has the line of best fit y = 2.5x + 15. Amelia wants a score of at least 80. Work out the least whole number of hours of revision the line of best fit suggests she needs.y = 2.5x + 15
- 11.Work out the median of these six numbers: 13, 21, 22, 36, 37, 47
- 12.Five friends have heights, in cm, of 150, 152, 155, 158 and 160. A sixth friend, with a height of 170 cm, joins the group. Write down what happens to the mean and the range of the heights once this sixth friend is included.
- 13.A sports shop in Cardiff sold 40 pairs of football boots last month: 4 pairs of size 6, 5 pairs of size 7, 8 pairs of size 8, 13 pairs of size 9 and 10 pairs of size 10. The manager will order 40 pairs for next month and wants as many pairs as possible to be in a size customers will buy. Work out the mean size and the modal size, and write down which of the two he should use.
- 14.A two-way table records the favourite subject, Maths or Art, of 60 pupils in Year 10, and whether each pupil is left-handed or right-handed. 9 of the 60 pupils are left-handed, and 6 of those left-handed pupils prefer Art. In total, 24 of the 60 pupils prefer Art. A pupil is chosen at random from the 60. Work out the probability that the pupil is right-handed and prefers Art.
- 15.The masses, m kg, of 60 parcels are grouped like this: 0 ≤ m < 5, 22 parcels; 5 ≤ m < 10, 20 parcels; 10 ≤ m < 20, 9 parcels; 20 ≤ m < 30, 5 parcels; 30 ≤ m < 50, 4 parcels. Write down the class interval that contains the median mass.
Answer key
- (b) 55 — Method: count the classes that lie wholly above 8 minutes, then use linear interpolation for the class that 8 cuts through, assuming the delays in that class are spread evenly. Working: the class 10 ≤ d < 20 lies wholly above 8 and holds 25 buses; the value 8 lies in the class 5 ≤ d < 10, which is 5 minutes wide and holds 75 buses, and the part above 8 runs from 8 to 10, a width of 2, so the estimated share is (2 ÷ 5) × 75 = 30 buses; the estimate is 30 + 25 = 55. Answer: about 55 refunds. The distractors: 100 comes from adding the whole of the class 5 ≤ d < 10, 75 + 25, and so refunding buses only 5 minutes late; 25 comes from using only the class 10 ≤ d < 20 and ignoring the part class that 8 minutes cuts through; 70 comes from taking the part of the class from 5 up to 8 instead of from 8 up to 10, giving (3 ÷ 5) × 75 = 45 and then 45 + 25.
- (c) No — median £505 at A vs £510 at B. — Branch A's seven wages in order are £480, £495, £500, £505, £510, £515 and £1,200, so the median, the 4th value, is £505. Branch B's in order are £480, £490, £500, £510, £520, £530 and £540, so the median is £510. Since £505 is lower than £510, the median wage is not higher at Branch A, so the claim is not fairly supported. Choosing 'Yes — mean £600.71 at A vs £510 at B' uses the mean: 480 + 495 + 500 + 505 + 510 + 515 + 1200 = 4205, and 4205 ÷ 7 = 600.71, a figure pulled upward by the £1,200 outlier that does not represent a typical wage. Choosing 'Yes — median £515 at A vs £510 at B' miscounts the middle position, taking the 6th wage, £515, instead of the correct 4th value, £505. Choosing 'Yes — highest wage £1,200 at A vs £540 at B' compares the highest wage at each branch rather than a measure of the typical, or average, wage.
- (d) 23.57 °C — Method: add all seven temperatures, divide by the number of readings and round only at the end. Working: 22 + 24 + 23 + 25 + 26 + 21 + 24 = 165, and 165 ÷ 7 = 23.5714…, which rounds to 23.57 to 2 decimal places. Answer: 23.57 °C. The distractors: 24 °C comes from writing down the mode, the only temperature recorded twice, instead of the mean; 27.5 °C comes from dividing the total by 6 instead of by the 7 days recorded; 5 °C comes from working out the range, 26 − 21, which is a measure of spread and not an average.
- (b) 135 — Method: use the sample to find the PROPORTION of left-handed pupils, then apply that same proportion to the whole school population. Working: in the sample, 6 out of 40 pupils are left-handed, a proportion of 6 ÷ 40 = 0.15. Applying that proportion to the school's 900 pupils gives an estimate of 0.15 × 900 = 135 pupils. Giving 6 simply repeats the number of left-handed pupils IN THE SAMPLE, without scaling up to the whole school at all. Multiplying the population by the number of left-handed pupils in the sample without first dividing by the sample size, 900 × 6 = 5400, badly overestimates — that is more pupils than the whole school has. Dividing the population by the sample size but forgetting to multiply by the number of left-handed pupils found, 900 ÷ 40 = 22.5, finds the scale factor but stops one step short of using it. Always find the proportion in the sample first, then scale that same proportion up to the population.
- (b) 13 — The total is 50, and the two known parts are 22 (tea) and 15 (coffee), so 50 − 22 − 15 = 13 hot chocolates. Choosing 28 comes from 50 − 22, subtracting only the tea and forgetting the coffee. Choosing 35 comes from 50 − 15, subtracting only the coffee and forgetting the tea. Choosing 37 comes from 22 + 15, which finds how many drinks were tea or coffee, not the number left over for hot chocolate.
- (c) Branch A waits longer, and Branch A is more consistent — Method: compare the two branches using a measure of location (the median) for who waits longer, and a measure of spread (the interquartile range) for who is more consistent — a smaller interquartile range means more consistent. Working: Branch A's median, 12 minutes, is higher than Branch B's, 9 minutes, so Branch A's customers wait longer on average. Branch A's interquartile range, 5 minutes, is smaller than Branch B's, 11 minutes, so Branch A's waiting times vary less. Answer: Branch A waits longer, and Branch A is also the more consistent of the two. Watch that each half of the comparison uses the right statistic and reads it correctly: swapping both readings gives Branch B the longer wait and the greater consistency, when neither is true; keeping the median comparison right but reading a larger interquartile range as 'more consistent' has the direction of spread backwards; and swapping only the median comparison keeps the correct branch for consistency but gives the wrong branch the longer wait.
- (c) 20 kg ≤ mass < 30 kg — The modal class is the class with the highest frequency. Reading the plotted points, the frequencies are 6, 10, 16, 6 and 2, so the highest frequency is 16, plotted at the midpoint 25. A class of width 10 centred on 25 runs from 25 − 5 = 20 to 25 + 5 = 30, so the modal class is 20 kg ≤ mass < 30 kg. Writing '25 kg' gives only the midpoint, not the class — the modal class is an interval, not a single value. '10 kg ≤ mass < 20 kg' is the class before the peak, centred on 15, which has frequency 10, not the highest. '30 kg ≤ mass < 40 kg' is the class after the peak, centred on 35, which has frequency 6, not the highest.
- (a) 27 — Method: turn the mean into a total using total = mean × number of values, then subtract the numbers that are already known. Working: four numbers with a mean of 18 have a total of 18 × 4 = 72; the three known numbers give 10 + 15 + 20 = 45; so x = 72 − 45 = 27. Answer: 27, and checking, (10 + 15 + 20 + 27) ÷ 4 = 72 ÷ 4 = 18. The distractors: 72 comes from stopping at the total the four numbers must reach and never subtracting the known three; 18 comes from assuming the missing number must equal the mean; 45 comes from stopping at the total of the three known numbers.
- (b) On average a plant grew 1.5 cm taller for each extra day — Method: in the equation of a line, the number multiplying x is the gradient, and a gradient states the change in y produced by an increase of 1 in x, read in the units of the two axes. Working: here x is measured in days and y in centimetres, so the gradient 1.5 carries the units centimetres per day. Testing it on the line, 5 days gives 1.5 × 5 + 4 = 11.5 cm and 6 days gives 1.5 × 6 + 4 = 13 cm, a rise of 1.5 cm for the one extra day. Answer: on average a plant grew 1.5 cm taller for each extra day of watering. The distractors: 1.5 cm as the height before any watering is the value of y when x is 0, which is the other number in the equation, 4 cm, so this swaps the gradient and the intercept; 1.5 cm as the gap between the tallest and the shortest plant reads the gradient as a range, when a range is a difference between two of the 16 plants and a gradient is a rate; 1.5 days for each extra centimetre inverts the rate, dividing days by centimetres instead of centimetres by days, and the line gives 1 cm of growth in two thirds of a day.
- (a) 26 — Method: a line of best fit lets one quantity be predicted from the other, so the score is substituted into the equation of the line and the resulting inequality is solved for the number of hours. Working: a score of at least 80 means 2.5x + 15 ≥ 80; taking 15 from both sides gives 2.5x ≥ 65, and dividing both sides by 2.5 gives x ≥ 26, so the least whole number of hours is 26. Checking, 2.5 × 26 + 15 = 80, which does reach the target. Answer: 26 hours — and this is only an estimate, because a line of best fit predicts a trend rather than an individual result, and a prediction made outside the range of hours the pupils actually revised for would be an extrapolation and less reliable still. The distractors: 27 comes from reaching 26 and then rounding up again, although 26 hours already gives a score of exactly 80; 32 comes from 80 ÷ 2.5, which ignores the 15 in the equation of the line; 38 comes from (80 + 15) ÷ 2.5, that is from adding the 15 instead of subtracting it when rearranging.
- (d) 29 — Method: with an even number of values there is no single middle value, so the median is the mean of the two values either side of the middle. Working: the six numbers are already in order and 6 ÷ 2 = 3, so the middle pair are the third and fourth values, 22 and 36; their mean is (22 + 36) ÷ 2 = 58 ÷ 2 = 29. Answer: 29, which lies between the two middle values as a median of an even data set must. The distractors: 22 comes from taking the lower of the two middle values and stopping there instead of averaging the pair; 36 comes from taking the larger value of that pair because it sits just past the halfway point of the list; 34 comes from working out the range, 47 − 13, instead of a measure of centre.
- (d) Both the mean and the range increase. — The original mean is 150 + 152 + 155 + 158 + 160 = 775, and 775 ÷ 5 = 155 cm; the original range is 160 − 150 = 10 cm. Including the new height of 170 cm gives a new total of 775 + 170 = 945, and 945 ÷ 6 = 157.5 cm, which is higher than 155 cm, and a new range of 170 − 150 = 20 cm, which is higher than 10 cm, so both the mean and the range increase. Saying the range stays the same ignores that 170 cm is a new, higher maximum than the old 160 cm. Saying the mean stays the same ignores that 170 cm is above the original mean of 155 cm, which pulls the average up. Saying both decrease is the opposite of what happens here.
- (a) The modal size, 9, bought by more customers than any other — Method: work out both averages from the frequencies, then choose the one the shop can act on. Working: for the mean, multiply each size by the number of pairs sold at it and add: 6 × 4 + 7 × 5 + 8 × 8 + 9 × 13 + 10 × 10 = 340, and 340 ÷ 40 = 8.5, so the mean size is 8.5. The largest frequency is 13, which belongs to size 9, so the modal size is 9. The mean 8.5 is a size no customer in the record asked for, so 40 pairs of it would sit unsold, while 13 of the 40 customers wanted size 9, more than wanted any other size. Answer: the modal size, 9, bought by more customers than any other. The distractors: the mean size 8.5 does take account of all 40 pairs, but a mean of sizes is a summary figure and not a size the month's customers were buying; the mean size 8 comes from averaging the five sizes on sale, 6 + 7 + 8 + 9 + 10 = 40 and 40 ÷ 5 = 8, which ignores how many pairs were sold at each size and so treats the 4 pairs of size 6 as equal in weight to the 13 pairs of size 9; the range 4 comes from 10 − 6 and measures spread, so it says how wide a set of sizes the shop must stock, not which size to stock most of.
- (a) 3/10 — There are 60 − 9 = 51 right-handed pupils. Of the 24 pupils who prefer Art, 6 are left-handed, so 24 − 6 = 18 are right-handed and prefer Art. The probability that a randomly chosen pupil is right-handed and prefers Art is 18/60, which simplifies to 3/10. Giving 2/5 is 24/60 simplified — the probability of preferring Art, ignoring the right-handed condition entirely. Giving 17/20 is 51/60 simplified — the probability of being right-handed, ignoring the Art condition entirely. Giving 1/10 is 6/60 simplified — the probability of being left-handed and preferring Art, the wrong hand condition.
- (b) 5 ≤ m < 10 — Method: with 60 values the median is the 60 ÷ 2 = 30th value in order, so build a running total until it first reaches 30. Working: the running totals are 22 after the first class, 22 + 20 = 42 after the second, 51 after the third, 56 after the fourth and 60 after the fifth; the 30th parcel is past 22 but not past 42, so it lies in the second class. Answer: the median lies in the class 5 ≤ m < 10. The distractors: 0 ≤ m < 5 comes from giving the class with the greatest frequency, 22, which is the modal class and not the median class; 10 ≤ m < 20 comes from choosing the middle class in the list of five instead of counting to the middle value; 20 ≤ m < 30 comes from halving the range of the data, 50 ÷ 2 = 25, and giving the class that contains 25 kg rather than the class that contains the 30th parcel.
Build your own mix at the worksheet builder.