Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- 1.A scatter graph has 50 points. Most of them lie close to a rising line of best fit, but two of them lie a long way from that line. Write down how those two points should be treated.
- 2.A histogram is drawn for the masses, m grams, of 200 letters. The bar for 0 ≤ m < 50 has a frequency density of 1.2 per gram and the bar for 50 ≤ m < 100 has a frequency density of 1.8 per gram. All the remaining letters lie in the class 100 ≤ m < 200. Work out the frequency density of the bar for 100 ≤ m < 200.
- 3.In a scatter graph of the age, in years, and the wingspan, in cm, of 20 birds of the same species, all the points lie close to a rising line of best fit except one, which lies a long way below the line. That bird was later found to have a damaged wing. Give a reason why this point should not be used when drawing the line of best fit.
- 4.A garden centre records the heights, in cm, of eleven seedlings. In order, the heights are 5, 9, x, 17, 20, 24, 28, 31, 35, 40, 44, where x is unknown. The interquartile range of the eleven heights is 19 cm. Work out the value of x.
- 5.In a histogram of the distances, d metres, thrown by some athletes, the bar covering 20 ≤ d < 60 has a constant frequency density of 1.8 per metre. Estimate the number of throws of at least 20 metres but less than 35 metres.
- 6.A café owner in Brighton records the midday temperature, x °C, and the number of hot chocolates sold, y, on 12 days. The temperatures recorded run from 4 °C to 18 °C, and the line of best fit is y = −3x + 74. The forecast for tomorrow gives a midday temperature of 12 °C. Work out the number the line of best fit predicts, and write down how much confidence the owner can have in it.y = -3x + 74
- 7.Write down the statement that correctly describes the difference between correlation and causation.
- 8.A café in York counts the number of customers in each of the nine hours it is open on one day: 4, 4, 4, 11, 13, 15, 18, 22 and 25. The owner says that a typical hour has about 4 customers, because 4 is the mode. Is the owner right? Give a reason for your answer.
- 9.A cumulative frequency graph for the diameters, d mm, of 320 ball bearings is plotted from these points (upper class boundary, cumulative frequency): (10, 30), (20, 90), (30, 190), (40, 280), (50, 320). Estimate the diameter below which 90% of the ball bearings measure.
- 10.A scatter graph shows the mass, x kg, of a parcel and the cost, y pounds, of posting it. The line of best fit is y = 1.5x + 2. Work out the estimated cost of posting a parcel with a mass of 6 kg, using the line of best fit.y = 1.5x + 2
- 11.The masses, m kg, of 80 sacks of grain are summarised by these cumulative frequencies: m < 10, 6 sacks; m < 20, 22 sacks; m < 30, 58 sacks; m < 40, 74 sacks; m < 50, 80 sacks. Use interpolation to estimate the median mass.
- 12.A study found that people who drink more coffee tend to concentrate better at work. A coffee company says that this shows that drinking coffee improves concentration. Give the reason why this conclusion cannot be drawn.
- 13.The mean of five test scores is 68. Four of the scores are 55, 62, 74 and 80. Work out the fifth score.
- 14.A Year 10 class has 20 boys with a mean height of 150 cm and 10 girls with a mean height of 168 cm. Work out the mean height of all 30 pupils in the class.
- 15.A council in Leeds wants to know what local people think about letting shops stay open later in the evening. It rings landline telephone numbers between 10 am and 2 pm on a Tuesday. Write down which group is most likely to be under-represented in the sample, and give a reason for your answer.
Answer key
- (b) Treat them as outliers and check them before deciding — Method: a point lying a long way from the pattern the rest of the data make is called an outlier, and an outlier is investigated before anything is done with it, because it may be an error in the data or it may be a genuine but unusual case. Working: 48 of the 50 points lie close to the rising line of best fit, so the trend is set by those 48; the two remaining points do not follow it, so they are identified as outliers and checked — a mistake in measuring or recording would be corrected, while a genuine reading would be kept and reported. Answer: treat them as outliers and check them before deciding what to do with them. The distractors: deleting them at once assumes that every point far from the line must be an error, which throws away real data; moving the line so that it passes through them assumes a line of best fit must touch particular points, when it is drawn to follow all 50; taking them as proof that there is no correlation lets two points overturn the pattern that the other 48 agree on.
- (c) 0.5 per gram — Method: turn the two known bars into frequencies using area, subtract from the total to find how many letters are left, then divide that frequency by the width of the last class to get its height. Working: the first bar covers 50 g at a frequency density of 1.2, giving 1.2 × 50 = 60 letters, and the second covers 50 g at 1.8, giving 1.8 × 50 = 90 letters; together that is 60 + 90 = 150 letters, so 200 − 150 = 50 letters remain; the class 100 ≤ m < 200 is 100 g wide, so its frequency density is 50 ÷ 100 = 0.5 per gram. Answer: 0.5 per gram. The distractors: 0.25 per gram comes from dividing the remaining 50 letters by the upper class boundary, 200, instead of by the class width of 100; 2 per gram comes from dividing the class width by the frequency, 100 ÷ 50, reversing the formula; 1.4 per gram comes from subtracting only the first bar's 60 letters, leaving 140, and then dividing by 100.
- (d) An outlier from the damaged wing, not the trend. — That bird's point lies a long way from the rising trend followed by every other bird, and its low wingspan is explained by the damaged wing rather than by its age — it is an outlier caused by an unusual factor, not part of the general relationship between age and wingspan, so it should not be used when drawing the line of best fit. Saying every point must be used ignores that an outlier caused by a separate, identifiable factor can rightly be set aside. Saying it shows no correlation ignores that the other 19 points do show a clear rising trend; one outlier does not remove that. Saying it proves the line is inaccurate confuses one unusual bird with a fault in the line itself, when the line correctly describes the trend followed by the rest of the data.
- (d) 16 — Method: rearrange interquartile range = upper quartile − lower quartile to make the lower quartile the subject: lower quartile = upper quartile − interquartile range, then check the answer sits in the right place in the list. Working: there are 11 values, so 11 + 1 = 12; the upper quartile sits at position 3 × 12 ÷ 4 = 9, which is 35, and x sits at position 12 ÷ 4 = 3, which is the lower quartile. So x = 35 − 19 = 16, and 16 does sit between the 2nd value, 9, and the 4th value, 17, as it should. Answer: x = 16. Watch how you rearrange and where you count to: adding instead of subtracting, 35 + 19 = 54, treats the interquartile range as something added on rather than a gap taken away; subtracting in the wrong order, 19 − 35 = −16, finds the right two numbers but flips the sign; and counting to the 8th value instead of the 9th treats 31 as the upper quartile, giving 31 − 19 = 12, one position short of where the upper quartile actually sits.
- (c) 27 — Method: a frequency is the area of the part of the bar being asked about, so frequency = frequency density × the width of that part. Working: the part asked about runs from 20 to 35, so its width is 35 − 20 = 15 metres; the frequency density there is 1.8 per metre, so the estimate is 1.8 × 15 = 27. Answer: about 27 throws. The distractors: 72 comes from taking the whole bar, 1.8 × 40, and so counting every throw from 20 up to 60; 1.8 comes from reading the height of the bar as a frequency, when a height is a density and only an area is a count; 63 comes from using the upper value 35 as the width, 1.8 × 35, instead of the width 35 − 20.
- (d) 38, and fairly confident, as 12 °C is inside the range — Method: substitute the forecast temperature into the equation of the line of best fit, then judge the prediction by where that temperature sits among the data the line was drawn from. Working: putting x = 12 into y = −3x + 74 gives −3 × 12 + 74 = 38, so the line predicts 38 hot chocolates. The recorded temperatures run from 4 °C to 18 °C, and 12 °C lies inside that interval, so this is interpolation, the safer kind of prediction. Answer: 38, and fairly confident, as 12 °C is inside the range; the owner should still expect the true figure to differ a little, since the points only lie near the line and not on it. The distractors: being completely certain treats a line of best fit as a rule that fixes each day's sales, when it describes a trend that individual days depart from; saying 12 °C is outside the range misreads the interval 4 °C to 18 °C, and the wrong warning would be attached to a sound prediction; 110 comes from −3 × 12 being taken as +36, giving 36 + 74 = 110, which loses the negative gradient and so predicts that a warm day sells more hot chocolate than a cold one.
- (d) Correlation is a link; causation is one causing the other — Method: the two words describe different claims — one is about a pattern in the data, the other is about what produced that pattern. Working: correlation says only that two quantities tend to change together, which is something a scatter graph can display; causation says that a change in one quantity actually brings about the change in the other, which needs evidence a scatter graph cannot supply, because a third quantity may be driving both. Answer: correlation is a link between the quantities, while causation is one quantity causing the change in another. The distractors: the statement giving causation as the link and correlation as the cause simply swaps the two words over; the statement that the words mean the same thing is the classic error of reading a correlation as proof of cause; the statement that a scatter graph shows causation but not correlation reverses what a scatter graph can do, since the pattern it displays is exactly the correlation.
- (c) No, the mode here is the lowest value of the nine — Method: an average is meant to stand for the data as a whole, so test any proposed average by asking how many values it sits near. Working: the value 4 appears three times and every other count appears once, so 4 is indeed the mode. But those three hours are the quiet ones at the start of the day, and the other six counts run from 11 up to 25; putting the nine counts in order, the middle one is the fifth, which is 13. So the mode sits at the very bottom of the data, with six of the nine hours far above it. Answer: no, because the mode here is the lowest value of the nine, so it describes the quiet opening hours rather than a typical hour. The distractors: saying the mode can only be used when no value repeats reverses the definition, since a mode exists only because a value does repeat; saying the mode is the value that occurs most often is a correct definition, but being the commonest value does not make a value typical when it lies at one end of the data; saying the mode is the best average for any list of numbers ignores the fact that mean, median and mode each describe a population well in different circumstances.
- (c) 42 — Method: find the target cumulative frequency, 90% of the total, locate the class it falls in from the plotted points, then interpolate: lower boundary, plus the extra distance needed into the class divided by the class's frequency, times its width. Working: 90% of 320 is 0.9 × 320 = 288. The plotted points show a cumulative frequency of 280 at d = 40 and 320 at d = 50, so the class 40 ≤ d < 50 has frequency 320 − 280 = 40 and width 50 − 40 = 10, and 288 falls inside it. The extra distance needed into the class is 288 − 280 = 8, and 8 ÷ 40 × 10 = 2, so the diameter is 40 + 2 = 42. Answer: the estimated diameter is 42 mm. Watch which point and which class the interpolation actually uses: reading off d = 40, the plotted point just below the target, instead of interpolating the extra 8 ball bearings into the next 10 mm, stops one step short of the true answer; finding the diameter below which only 10% lie instead of 90% gives a target of 0.1 × 320 = 32, which falls in the class 10 ≤ d < 20 — the extra distance into that class is 32 − 30 = 2, and 2 ÷ 60 × 10 = 0.3, so this route gives 10 + 0.3 = 10.3, the bottom decile rather than the top 90%; and interpolating within the class 30 ≤ d < 40 instead of 40 ≤ d < 50, as though 288 had not yet reached a cumulative frequency of 280, treats the extra distance as 288 − 190 = 98, and 98 ÷ 90 × 10 = 10.9, giving 30 + 10.9 = 40.9, one class too early.
- (b) £11.00 — 1.5 × 6 = 9, and 9 + 2 = 11, so the estimated cost is £11.00. Choosing £9.00 stops after 1.5 × 6 = 9 and forgets to add the £2. Choosing £12.00 adds the mass and the constant first and then multiplies: 6 + 2 = 8, and 8 × 1.5 = 12.00. Choosing £13.50 swaps the gradient and the intercept, using y = 2x + 1.5 instead: 2 × 6 = 12, and 12 + 1.5 = 13.50.
- (d) 25 — Method: the median is estimated at position n ÷ 2 in the cumulative frequency table, then interpolated across the class it falls in: lower boundary, plus the fraction of the way through the class, times the class width. Working: there are 80 sacks, so the median sits at position 80 ÷ 2 = 40. Before the class 20 ≤ m < 30 the cumulative frequency is 22, and by the end of it, it is 58, so this class holds the 40th sack; its frequency is 58 − 22 = 36 and its width is 30 − 20 = 10. The extra distance needed into the class is 40 − 22 = 18, and 18 ÷ 36 × 10 = 5, so the median is 20 + 5 = 25. Answer: the estimated median mass is 25 kg. Watch which numbers the interpolation uses: reading off just the lower boundary of the median class, 20, ignores how far into that class the 40th sack actually falls; treating n ÷ 2 = 40 itself as the median mass mistakes a position in the list for a mass in kilograms; and using the target position, 40, as the extra distance into the class instead of subtracting the sacks already counted changes the calculation to 20 + 40 ÷ 36 × 10. That comes to 20 + 11.1 = 31.1, overshooting the class because it never subtracts the 22 sacks already counted before it.
- (c) The data show a link only; a third factor may affect both — Method: a study of this kind measures two quantities and reports how they change together; deciding that one of them produces the other is a further claim, and it needs evidence that the measurements alone cannot give. Working: the study shows that more coffee goes with better concentration, which is a positive correlation; but a third factor that was never measured, such as how motivated someone is, could raise both the coffee drinking and the concentration, and the concentration could equally be what leads to the extra coffee. Answer: the data show a link only, because a third factor may be affecting both quantities, so no claim about cause can be made. The distractors: calling the conclusion safe because the correlation is positive treats the direction of a correlation as proof of cause, which no direction can give; calling it wrong because the correlation is negative misreads the direction of the relationship, since the study reports both quantities rising together; saying the two quantities are not linked denies the correlation the study actually found, when what fails is only the claim about cause.
- (c) 69 — Method: multiply the mean by the number of values to find the total, then subtract the total of the known values. Working: the total of all five scores is 68 × 5 = 340. The total of the four known scores is 55 + 62 + 74 + 80 = 271. The fifth score is 340 − 271 = 69. Subtracting the other way round, 271 − 340 = −69, gives the right size answer with the wrong sign. Guessing that the missing score simply equals the mean, 68, ignores that the four known scores are not themselves centred on 68. Multiplying the mean by 4 instead of 5, 68 × 4 = 272, then 272 − 271 = 1, undercounts how many scores there are. Always multiply the mean by the TOTAL number of values before subtracting.
- (c) 156 cm — Method: to combine two groups' means, multiply each group's mean by its own number of pupils, add the two totals together, then divide by the total number of pupils in both groups. Working: 20 × 150 = 3,000 cm for the boys and 10 × 168 = 1,680 cm for the girls, giving a combined total of 3,000 + 1,680 = 4,680 cm. Dividing by all 30 pupils gives 4,680 ÷ 30 = 156 cm. Giving 159 cm averages the two means, (150 + 168) ÷ 2, treating the two groups as if they had the same number of pupils, when there are twice as many boys as girls. Giving 4,680 cm finds the correct combined total height but stops there, forgetting the final division by the 30 pupils. Giving 234 cm divides the combined total by 20, the number of boys only, forgetting that the total also includes the 10 girls. Always weight each mean by its own group size, and always divide by the TOTAL number of pupils in both groups combined.
- (d) Full-time workers, as most are at work at that time — Method: a sample is biased when the method of contact makes part of the population much less likely to be reached, so test each group against where its members actually are between 10 am and 2 pm on a weekday, and test each stated reason against the facts. Working: those hours are the middle of the working day, so people in full-time employment are at work and not beside a landline telephone, while people who are retired and people who are unemployed are far more likely to be at home and are reached at the usual rate; the method therefore collects far fewer replies from full-time workers than their share of the adult population the council is consulting. Answer: full-time workers, as most are at work at that time. The distractors: the reply naming retired people rests on the false claim that most retired people are at work in the daytime, when in fact a daytime call reaches them more easily than anyone; the reply naming unemployed people rests on the false claim that they are out during the day, when they too are among the easiest people to reach by a daytime call; the reply naming children rests on the false claim that children are at home at 11 am on a Tuesday, when they are at school and so are not reached by the call at all, and school-age children are in any case not the adults whose views the council is collecting.
Build your own mix at the worksheet builder.