Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Non-calculator
Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- 1.Work out the median of these six numbers: 13, 21, 22, 36, 37, 47
- 2.A scatter graph shows the midday temperature, x °C, and the number of ice creams sold at a seaside kiosk, y. The line of best fit is y = 3x − 20. Give a reason why the y-intercept of this line of best fit is not a sensible estimate of the number of ice creams sold.y = 3x − 20
- 3.In a scatter graph of the age, in years, and the wingspan, in cm, of 20 birds of the same species, all the points lie close to a rising line of best fit except one, which lies a long way below the line. That bird was later found to have a damaged wing. Give a reason why this point should not be used when drawing the line of best fit.
- 4.The delivery times, in minutes, of 15 parcels are given in order: 5, 8, 10, 12, 14, 15, 17, 19, 21, 23, 25, 27, 29, 31, 34. Work out the interquartile range of these times.
- 5.The masses, m grams, of 100 apples are grouped like this: 100 ≤ m < 120, 10 apples; 120 ≤ m < 140, 30 apples; 140 ≤ m < 160, 40 apples; 160 ≤ m < 200, 20 apples. Estimate the median mass.
- 6.A sports shop in Cardiff sold 40 pairs of football boots last month: 4 pairs of size 6, 5 pairs of size 7, 8 pairs of size 8, 13 pairs of size 9 and 10 pairs of size 10. The manager will order 40 pairs for next month and wants as many pairs as possible to be in a size customers will buy. Work out the mean size and the modal size, and write down which of the two he should use.
- 7.The weekly wages of the five people who work at a small garage in Norwich are £420, £440, £460, £480 and £1,500. Write down which average better describes a typical wage at this garage, and give a reason for your answer.
- 8.A survey of 25 pupils in Derby records how many siblings each has: 0 siblings — 6 pupils, 1 sibling — 10 pupils, 2 siblings — 6 pupils, 3 siblings — 3 pupils. Calculate the mean number of siblings.
- 9.The times, in seconds, taken by 11 runners to finish a race are given in order: 40, 42, 45, 47, 50, 52, 55, 58, 60, 63, 65. Work out the upper quartile of these times.
- 10.A council wants to find out what all residents of a town think about a new cycle lane. It posts a survey only on its website and asks people to fill it in online. Give a reason why this sample is likely to be biased.
- 11.Five friends have heights, in cm, of 150, 152, 155, 158 and 160. A sixth friend, with a height of 170 cm, joins the group. Write down what happens to the mean and the range of the heights once this sixth friend is included.
- 12.A Year 10 class has 20 boys with a mean height of 150 cm and 10 girls with a mean height of 168 cm. Work out the mean height of all 30 pupils in the class.
- 13.A histogram is drawn for the masses, m grams, of 200 letters. The bar for 0 ≤ m < 50 has a frequency density of 1.2 per gram and the bar for 50 ≤ m < 100 has a frequency density of 1.8 per gram. All the remaining letters lie in the class 100 ≤ m < 200. Work out the frequency density of the bar for 100 ≤ m < 200.
- 14.The times, t minutes, of 80 journeys are summarised by these cumulative frequencies: t < 10, 8 journeys; t < 20, 28 journeys; t < 30, 52 journeys; t < 40, 72 journeys; t < 50, 80 journeys. Estimate the interquartile range.
- 15.In a histogram of the distances, d metres, thrown by some athletes, the bar covering 20 ≤ d < 60 has a constant frequency density of 1.8 per metre. Estimate the number of throws of at least 20 metres but less than 35 metres.
Answer key
- (d) 29 — Method: with an even number of values there is no single middle value, so the median is the mean of the two values either side of the middle. Working: the six numbers are already in order and 6 ÷ 2 = 3, so the middle pair are the third and fourth values, 22 and 36; their mean is (22 + 36) ÷ 2 = 58 ÷ 2 = 29. Answer: 29, which lies between the two middle values as a median of an even data set must. The distractors: 22 comes from taking the lower of the two middle values and stopping there instead of averaging the pair; 36 comes from taking the larger value of that pair because it sits just past the halfway point of the list; 34 comes from working out the range, 47 − 13, instead of a measure of centre.
- (a) x = 0 gives y = −20: a negative number sold — The y-intercept is the value the line predicts when x = 0: y = 3 × 0 − 20 = −20. A kiosk cannot sell a negative number of ice creams, so this is not a sensible estimate. The 3 in the equation is the gradient, not the intercept, so an option claiming x = 0 gives y = 3 has swapped the two numbers around — substituting x = 0 makes the 3x term equal 0, leaving −20, not 3. The danger of extrapolating to very high temperatures is a real issue with this line, but it is a different issue from the y-intercept, so it does not answer this question. And whether x = 0 could occur on a trading day is beside the point: the model still makes that prediction, and it is the prediction itself, −20, that is impossible.
- (d) An outlier from the damaged wing, not the trend. — That bird's point lies a long way from the rising trend followed by every other bird, and its low wingspan is explained by the damaged wing rather than by its age — it is an outlier caused by an unusual factor, not part of the general relationship between age and wingspan, so it should not be used when drawing the line of best fit. Saying every point must be used ignores that an outlier caused by a separate, identifiable factor can rightly be set aside. Saying it shows no correlation ignores that the other 19 points do show a clear rising trend; one outlier does not remove that. Saying it proves the line is inaccurate confuses one unusual bird with a fault in the line itself, when the line correctly describes the trend followed by the rest of the data.
- (c) 15 — Method: the interquartile range is the upper quartile take away the lower quartile, IQR = Q3 − Q1. Working: n + 1 = 15 + 1 = 16, so the lower quartile sits at position 16 ÷ 4 = 4, the 4th value in the list, which is 12; the upper quartile sits at position 3 × 4 = 12, the 12th value, which is 27. So 27 − 12 = 15. Answer: the interquartile range is 15 minutes. Watch which values you use and which way round: taking the smallest time away from the largest, 34 − 5 = 29, finds the range, which uses every value between the extremes rather than just the middle half; taking the median away from the upper quartile instead of the lower quartile, 27 − 19 = 8, swaps the median in for the lower quartile; and reversing the subtraction, 12 − 27 = −15, finds the right two values but in the wrong order — an interquartile range is never negative.
- (c) 145 g — Method: find the position of the median from the total frequency, locate the class that contains it, then use linear interpolation inside that class, assuming the apples in it are spread evenly. Working: the median is the 100 ÷ 2 = 50th apple; the running totals are 10, then 10 + 30 = 40, then 40 + 40 = 80, so the 50th apple lies in the class 140 ≤ m < 160; it is the 50 − 40 = 10th of the 40 apples in that class, and the class is 20 g wide, so the median is 140 + (10 ÷ 40) × 20 = 140 + 5 = 145. Answer: an estimated median of 145 g. The distractors: 150 g comes from giving the midpoint of the class that contains the median instead of interpolating inside it; 140 g comes from stopping at the lower boundary of that class, which locates the class but not the value; 155 g comes from measuring the 5 g step down from the upper boundary, 160 − 5, instead of up from the lower boundary.
- (a) The modal size, 9, bought by more customers than any other — Method: work out both averages from the frequencies, then choose the one the shop can act on. Working: for the mean, multiply each size by the number of pairs sold at it and add: 6 × 4 + 7 × 5 + 8 × 8 + 9 × 13 + 10 × 10 = 340, and 340 ÷ 40 = 8.5, so the mean size is 8.5. The largest frequency is 13, which belongs to size 9, so the modal size is 9. The mean 8.5 is a size no customer in the record asked for, so 40 pairs of it would sit unsold, while 13 of the 40 customers wanted size 9, more than wanted any other size. Answer: the modal size, 9, bought by more customers than any other. The distractors: the mean size 8.5 does take account of all 40 pairs, but a mean of sizes is a summary figure and not a size the month's customers were buying; the mean size 8 comes from averaging the five sizes on sale, 6 + 7 + 8 + 9 + 10 = 40 and 40 ÷ 5 = 8, which ignores how many pairs were sold at each size and so treats the 4 pairs of size 6 as equal in weight to the 13 pairs of size 9; the range 4 comes from 10 − 6 and measures spread, so it says how wide a set of sizes the shop must stock, not which size to stock most of.
- (a) The median, as the one very large wage does not move it — Method: an average describes a population well when it sits close to most of the values, so compare what each average does when one value lies far from the rest. Working: in order the wages are 420, 440, 460, 480 and 1,500, so the median is the third of the five, £460. The mean uses every wage: 420 + 440 + 460 + 480 + 1,500 = 3,300 and 3,300 ÷ 5 = 660, so the mean is £660. Four of the five people earn less than £660, and the nearest of those four wages is £180 below it, so £660 describes nobody at the garage; £460 sits inside the group of four similar wages. Answer: the median, as the one very large wage does not move it, while that same wage drags the mean £200 above the median. The distractors: saying the median is always larger than the mean is an invented rule, and here the median £460 is smaller than the mean £660; saying the mean is the only average that uses all five wages is true as far as it goes, but using a value and being dragged by it are the same thing when that value is £1,500; saying £660 lies between the smallest and largest wage is true of every mean ever calculated, so it proves nothing about whether this one is typical.
- (a) 1.24 — Method: for data given as a frequency table, the mean is Σfx ÷ Σf — multiply each value by its frequency, add the results, then divide by the total frequency. Working: 0 × 6 = 0. 1 × 10 = 10. 2 × 6 = 12. 3 × 3 = 9. So Σfx = 0 + 10 + 12 + 9 = 31. The total frequency is Σf = 6 + 10 + 6 + 3 = 25. Mean = 31 ÷ 25 = 1.24 siblings. Averaging the frequency column itself, (6 + 10 + 6 + 3) ÷ 4 = 6.25, mixes up the frequencies with the values they belong to. Writing down 1, the number of siblings with the highest frequency, gives the mode, not the mean. Writing down 31 stops after finding Σfx and forgets to divide by the total frequency, 25. Always divide Σfx by Σf — never stop at the top of the fraction.
- (b) 60 — Method: for n ordered values, the upper quartile sits at position 3(n + 1) ÷ 4, counting from the smallest. Working: n + 1 = 11 + 1 = 12; 3 × 12 = 36 and 36 ÷ 4 = 9, so the upper quartile is the 9th value in the list 40, 42, 45, 47, 50, 52, 55, 58, 60, 63, 65, which is 60. Answer: the upper quartile is 60 seconds. Watch which quartile you find: counting to the 6th value gives the median, 52, not the upper quartile; using 3 × 11 = 33 and 33 ÷ 4 = 8.25 without adding 1 to n first, then rounding down, reaches the 8th value, 58, not the 9th; and counting to the 3rd value uses the lower quartile's position, 45, the wrong end of the list.
- (d) Only internet users reach the website; others are excluded. — Method: a sample is biased when it systematically leaves out part of the population, or systematically over-represents another part. Working: anyone without internet access, or who does not visit the council's website, has NO chance of being included — the sample is drawn only from internet-using residents, which is not the whole town. Saying too many people might respond because the survey is free confuses bias with sample size — bias is about who CAN be reached, not how many respond. Saying people might lie describes a different problem, response honesty, not who was sampled in the first place. Saying online surveys cannot be anonymous is not a reason connected to bias at all. A sample is biased when part of the population has no chance of being included, whatever the reason for that.
- (d) Both the mean and the range increase. — The original mean is 150 + 152 + 155 + 158 + 160 = 775, and 775 ÷ 5 = 155 cm; the original range is 160 − 150 = 10 cm. Including the new height of 170 cm gives a new total of 775 + 170 = 945, and 945 ÷ 6 = 157.5 cm, which is higher than 155 cm, and a new range of 170 − 150 = 20 cm, which is higher than 10 cm, so both the mean and the range increase. Saying the range stays the same ignores that 170 cm is a new, higher maximum than the old 160 cm. Saying the mean stays the same ignores that 170 cm is above the original mean of 155 cm, which pulls the average up. Saying both decrease is the opposite of what happens here.
- (c) 156 cm — Method: to combine two groups' means, multiply each group's mean by its own number of pupils, add the two totals together, then divide by the total number of pupils in both groups. Working: 20 × 150 = 3,000 cm for the boys and 10 × 168 = 1,680 cm for the girls, giving a combined total of 3,000 + 1,680 = 4,680 cm. Dividing by all 30 pupils gives 4,680 ÷ 30 = 156 cm. Giving 159 cm averages the two means, (150 + 168) ÷ 2, treating the two groups as if they had the same number of pupils, when there are twice as many boys as girls. Giving 4,680 cm finds the correct combined total height but stops there, forgetting the final division by the 30 pupils. Giving 234 cm divides the combined total by 20, the number of boys only, forgetting that the total also includes the 10 girls. Always weight each mean by its own group size, and always divide by the TOTAL number of pupils in both groups combined.
- (c) 0.5 per gram — Method: turn the two known bars into frequencies using area, subtract from the total to find how many letters are left, then divide that frequency by the width of the last class to get its height. Working: the first bar covers 50 g at a frequency density of 1.2, giving 1.2 × 50 = 60 letters, and the second covers 50 g at 1.8, giving 1.8 × 50 = 90 letters; together that is 60 + 90 = 150 letters, so 200 − 150 = 50 letters remain; the class 100 ≤ m < 200 is 100 g wide, so its frequency density is 50 ÷ 100 = 0.5 per gram. Answer: 0.5 per gram. The distractors: 0.25 per gram comes from dividing the remaining 50 letters by the upper class boundary, 200, instead of by the class width of 100; 2 per gram comes from dividing the class width by the frequency, 100 ÷ 50, reversing the formula; 1.4 per gram comes from subtracting only the first bar's 60 letters, leaving 140, and then dividing by 100.
- (d) 18 minutes — Method: the lower quartile is the 80 ÷ 4 = 20th value and the upper quartile is the 3 × 80 ÷ 4 = 60th value; locate each inside its class by linear interpolation, then subtract. Working: the 20th value lies between the running totals 8 and 28, so it is in the class 10 ≤ t < 20, which holds 20 journeys across 10 minutes, and it is the 20 − 8 = 12th of them, giving 10 + (12 ÷ 20) × 10 = 16 minutes; the 60th value lies between the running totals 52 and 72, so it is in the class 30 ≤ t < 40, which also holds 20 journeys across 10 minutes, and it is the 60 − 52 = 8th of them, giving 30 + (8 ÷ 20) × 10 = 34 minutes; subtracting, 34 − 16 = 18. Answer: an estimated interquartile range of 18 minutes. The distractors: 20 minutes comes from taking the lower boundaries of the two quartile classes, 30 − 10, which locates the classes but never the values inside them; 40 minutes comes from subtracting the two positions, 60 − 20, instead of the two times; 22 minutes comes from interpolating downwards from each upper boundary rather than upwards from each lower boundary, giving 20 − 6 = 14 and 40 − 4 = 36.
- (c) 27 — Method: a frequency is the area of the part of the bar being asked about, so frequency = frequency density × the width of that part. Working: the part asked about runs from 20 to 35, so its width is 35 − 20 = 15 metres; the frequency density there is 1.8 per metre, so the estimate is 1.8 × 15 = 27. Answer: about 27 throws. The distractors: 72 comes from taking the whole bar, 1.8 × 40, and so counting every throw from 20 up to 60; 1.8 comes from reading the height of the bar as a frequency, when a height is a density and only an area is a count; 63 comes from using the upper value 35 as the width, 1.8 × 35, instead of the width 35 − 20.
Build your own mix at the worksheet builder.