Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- 1.A call centre records the length, t seconds, of 100 calls: 0 ≤ t < 20, 15 calls; 20 ≤ t < 30, 24 calls; 30 ≤ t < 50, 40 calls; 50 ≤ t < 80, 21 calls. The manager's target is for a call to be finished in under 35 seconds. Estimate the number of calls that met the target.
- 2.The mean of the four numbers 10, 15, 20 and x is 18. Work out the value of x.
- 3.A café owner in Brighton records the midday temperature, x °C, and the number of hot chocolates sold, y, on 12 days. The temperatures recorded run from 4 °C to 18 °C, and the line of best fit is y = −3x + 74. The forecast for tomorrow gives a midday temperature of 12 °C. Work out the number the line of best fit predicts, and write down how much confidence the owner can have in it.y = -3x + 74
- 4.A scatter graph shows the number of years of experience, x, of 18 sales assistants and their monthly sales, y hundred pounds. The plotted points run from x = 1 to x = 12 years, and the line of best fit is y = 4x + 20. A new assistant has 25 years of experience. Use the line of best fit to estimate a value of y for this assistant, and decide whether the estimate would be reliable.y = 4x + 20
- 5.The mean mass of four parcels is 17 kg. Three of the parcels have masses 12 kg, 16 kg and 18 kg. Work out the mass of the fourth parcel.
- 6.A scientist has grouped the lifetimes, in hours, of 300 batteries into classes of unequal width. She wants a diagram in which the number of batteries in a class is given by the area of its bar. Write down the type of diagram she should draw.
- 7.A survey of 25 pupils in Derby records how many siblings each has: 0 siblings — 6 pupils, 1 sibling — 10 pupils, 2 siblings — 6 pupils, 3 siblings — 3 pupils. Calculate the mean number of siblings.
- 8.The delivery times, in minutes, of 15 parcels are given in order: 5, 8, 10, 12, 14, 15, 17, 19, 21, 23, 25, 27, 29, 31, 34. Work out the interquartile range of these times.
- 9.A garden centre's sales, in thousands of pounds, at the end of each quarter last year were: quarter 1 — 18, quarter 2 — 34, quarter 3 — 30, quarter 4 — 22. Work out the increase in sales from quarter 1 to the quarter with the highest sales.
- 10.A café in York counts the number of customers in each of the nine hours it is open on one day: 4, 4, 4, 11, 13, 15, 18, 22 and 25. The owner says that a typical hour has about 4 customers, because 4 is the mode. Is the owner right? Give a reason for your answer.
- 11.Two classes sat the same maths test, both marked out of 100. Class A had a mean mark of 70 and a range of 30 marks. Class B had a mean mark of 70 and a range of 10 marks. Compare the marks of the two classes.
- 12.Marta is drawing a cumulative frequency diagram for the times, t seconds, of 100 telephone calls. The grouped frequencies are: 0 ≤ t < 10, 7 calls; 10 ≤ t < 20, 19 calls; 20 ≤ t < 30, 34 calls; 30 ≤ t < 40, 40 calls. Write down the coordinates of the point Marta should plot for the class 20 ≤ t < 30.
- 13.In a random sample of 40 pupils at a school, 6 are left-handed. The school has 900 pupils. Work out an estimate for the number of left-handed pupils in the school.
- 14.The number of pets owned by each of 19 pupils in a class is recorded: 0 pets — 7 pupils, 1 pet — 3 pupils, 2 pets — 4 pupils, 3 pets — 5 pupils. Work out the median number of pets.
- 15.A scatter graph has 50 points. Most of them lie close to a rising line of best fit, but two of them lie a long way from that line. Write down how those two points should be treated.
Answer key
- (a) 49 — Method: add the frequencies of the classes that lie wholly below 35 seconds, then use linear interpolation for the class that 35 cuts through, assuming the calls in that class are spread evenly across it. Working: below 30 seconds there are 15 + 24 = 39 calls; the value 35 lies in the class 30 ≤ t < 50, which is 20 seconds wide and holds 40 calls, and 35 is 35 − 30 = 5 seconds into it, so the estimated share is (5 ÷ 20) × 40 = 10 calls; the estimate is 39 + 10 = 49. Answer: about 49 calls met the target. The distractors: 79 comes from adding the whole of the class 30 ≤ t < 50, 39 + 40, and so counting calls of up to 50 seconds as being under 35; 39 comes from stopping at the class boundary 30 and ignoring the part class altogether; 69 comes from measuring the part of the class from 35 up to 50 instead of from 30 up to 35, giving (15 ÷ 20) × 40 = 30 and then 39 + 30.
- (a) 27 — Method: turn the mean into a total using total = mean × number of values, then subtract the numbers that are already known. Working: four numbers with a mean of 18 have a total of 18 × 4 = 72; the three known numbers give 10 + 15 + 20 = 45; so x = 72 − 45 = 27. Answer: 27, and checking, (10 + 15 + 20 + 27) ÷ 4 = 72 ÷ 4 = 18. The distractors: 72 comes from stopping at the total the four numbers must reach and never subtracting the known three; 18 comes from assuming the missing number must equal the mean; 45 comes from stopping at the total of the three known numbers.
- (d) 38, and fairly confident, as 12 °C is inside the range — Method: substitute the forecast temperature into the equation of the line of best fit, then judge the prediction by where that temperature sits among the data the line was drawn from. Working: putting x = 12 into y = −3x + 74 gives −3 × 12 + 74 = 38, so the line predicts 38 hot chocolates. The recorded temperatures run from 4 °C to 18 °C, and 12 °C lies inside that interval, so this is interpolation, the safer kind of prediction. Answer: 38, and fairly confident, as 12 °C is inside the range; the owner should still expect the true figure to differ a little, since the points only lie near the line and not on it. The distractors: being completely certain treats a line of best fit as a rule that fixes each day's sales, when it describes a trend that individual days depart from; saying 12 °C is outside the range misreads the interval 4 °C to 18 °C, and the wrong warning would be attached to a sound prediction; 110 comes from −3 × 12 being taken as +36, giving 36 + 74 = 110, which loses the negative gradient and so predicts that a warm day sells more hot chocolate than a cold one.
- (b) 120, unreliable — x = 25 is outside 1 to 12 — The line of best fit is y = 4x + 20. 4 × 25 = 100, and 100 + 20 = 120, so the estimate is y = 120. But x = 25 lies far outside the plotted range of 1 to 12 years, so this is an extrapolation, and the estimate is not reliable. Reaching 100 instead of 120 comes from 4 × 25 = 100 with the intercept of 20 left out — still correctly flagged as unreliable, but the wrong value. Calling the estimate reliable simply because it was calculated correctly, giving 120, wrongly assumes that a correct calculation is automatically trustworthy, ignoring that x = 25 lies far beyond the data actually collected. Reaching 68, from 4 × 12 = 48 and 48 + 20 = 68, substitutes x = 12, the top of the plotted range, instead of the assistant's actual x = 25, and wrongly calls that reliable because 12 lies inside the range.
- (c) 22 kg — Method: multiply the mean by the number of parcels to rebuild the total mass, then subtract the masses that are known. Working: four parcels with a mean mass of 17 kg have a total mass of 17 × 4 = 68 kg; the three known parcels total 12 + 16 + 18 = 46 kg; so the fourth parcel has mass 68 − 46 = 22 kg. Answer: 22 kg. The distractors: 68 kg comes from stopping at the total mass of all four parcels; 17 kg comes from assuming the missing parcel must have the mean mass; 5 kg comes from multiplying the mean by 3, the number of parcels whose mass is given, leaving 51 − 46 = 5.
- (c) A histogram, with frequency density up the vertical axis — Method: decide which diagram makes area stand for frequency, which is the property the question asks for. Working: on a histogram the vertical axis is frequency density, so the area of a bar is frequency density × class width, and that product is the frequency; this is exactly what is wanted, and it is what allows classes of unequal width to be shown fairly. Answer: a histogram, with frequency density up the vertical axis. The distractors: a bar chart plots frequency as the height, so with unequal widths a wide class would cover far more area than a narrow class holding the same number of batteries, and area would measure nothing; a cumulative frequency diagram plots running totals against upper class boundaries, so a point on it gives how many lie below a value rather than how many lie in a class; a pie chart shows each class as a share of the whole 300 and loses the class widths entirely, so no area on it is tied to a scale of hours.
- (a) 1.24 — Method: for data given as a frequency table, the mean is Σfx ÷ Σf — multiply each value by its frequency, add the results, then divide by the total frequency. Working: 0 × 6 = 0. 1 × 10 = 10. 2 × 6 = 12. 3 × 3 = 9. So Σfx = 0 + 10 + 12 + 9 = 31. The total frequency is Σf = 6 + 10 + 6 + 3 = 25. Mean = 31 ÷ 25 = 1.24 siblings. Averaging the frequency column itself, (6 + 10 + 6 + 3) ÷ 4 = 6.25, mixes up the frequencies with the values they belong to. Writing down 1, the number of siblings with the highest frequency, gives the mode, not the mean. Writing down 31 stops after finding Σfx and forgets to divide by the total frequency, 25. Always divide Σfx by Σf — never stop at the top of the fraction.
- (c) 15 — Method: the interquartile range is the upper quartile take away the lower quartile, IQR = Q3 − Q1. Working: n + 1 = 15 + 1 = 16, so the lower quartile sits at position 16 ÷ 4 = 4, the 4th value in the list, which is 12; the upper quartile sits at position 3 × 4 = 12, the 12th value, which is 27. So 27 − 12 = 15. Answer: the interquartile range is 15 minutes. Watch which values you use and which way round: taking the smallest time away from the largest, 34 − 5 = 29, finds the range, which uses every value between the extremes rather than just the middle half; taking the median away from the upper quartile instead of the lower quartile, 27 − 19 = 8, swaps the median in for the lower quartile; and reversing the subtraction, 12 − 27 = −15, finds the right two values but in the wrong order — an interquartile range is never negative.
- (a) £16,000 — Method: first find the quarter with the highest sales figure, then subtract quarter 1's sales from it — remembering that every figure is given in THOUSANDS of pounds. Working: the highest sales figure is quarter 2, at £34,000 (34 thousand pounds). The increase from quarter 1 is £34,000 − £18,000 = £16,000. Giving £34,000 reads off the highest sales figure on its own, without subtracting quarter 1's sales — that is the highest quarter's total, not the increase. Giving £12,000 uses quarter 3's sales, 30, the SECOND-highest figure, instead of quarter 2's 34, the actual highest — 30 − 18 = 12, but quarter 3 is not the quarter with the highest sales. Giving £16 gets the subtraction right, 34 − 18 = 16, but forgets that every figure in the question is in thousands of pounds, so the increase is £16,000, not £16. Always identify the correct quarter FIRST, and always check the units the numbers are given in before writing your final answer.
- (c) No, the mode here is the lowest value of the nine — Method: an average is meant to stand for the data as a whole, so test any proposed average by asking how many values it sits near. Working: the value 4 appears three times and every other count appears once, so 4 is indeed the mode. But those three hours are the quiet ones at the start of the day, and the other six counts run from 11 up to 25; putting the nine counts in order, the middle one is the fifth, which is 13. So the mode sits at the very bottom of the data, with six of the nine hours far above it. Answer: no, because the mode here is the lowest value of the nine, so it describes the quiet opening hours rather than a typical hour. The distractors: saying the mode can only be used when no value repeats reverses the definition, since a mode exists only because a value does repeat; saying the mode is the value that occurs most often is a correct definition, but being the commonest value does not make a value typical when it lies at one end of the data; saying the mode is the best average for any list of numbers ignores the fact that mean, median and mode each describe a population well in different circumstances.
- (a) The means are equal, and Class B's marks are the more consistent because its range is smaller. — Method: comparing two distributions needs two things — a measure of average and a measure of spread — and each must be put into the context of the question. Working: both classes have a mean mark of 70, so on average the two classes scored the same; the range measures spread, and Class A's range of 30 marks is three times Class B's range of 10 marks, so Class B's marks sit closely around the mean while Class A's are far more spread out. Answer: the means are equal, and Class B's marks are the more consistent because its range is smaller. The distractors: the reply crediting Class A with more consistency reverses the meaning of the range, treating a larger range as tighter data when a larger range means more spread; the reply that Class A's mean mark is higher compares the wrong pair of figures, reading the range of 30 as an average; the reply that Class B's mean mark is higher reads the spread correctly but its claim about the means is false, since both means are 70.
- (a) (30, 60) — Method: a cumulative frequency point is plotted at the upper boundary of its class, paired with the running total of all the frequencies up to and including that class. Working: the running totals are 7, then 7 + 19 = 26, then 26 + 34 = 60, then 60 + 40 = 100; the class 20 ≤ t < 30 has upper boundary 30, and the running total there is 60. Answer: the point for that class is plotted at 30 seconds against a cumulative frequency of 60. The distractors: (25, 60) comes from plotting at the class midpoint, which is what a frequency polygon uses and not what a cumulative frequency diagram uses; (30, 34) comes from plotting the class frequency, 34, rather than the running total; (20, 60) comes from plotting at the lower boundary of the class, which would claim that 60 calls took less than 20 seconds when only 26 did.
- (b) 135 — Method: use the sample to find the PROPORTION of left-handed pupils, then apply that same proportion to the whole school population. Working: in the sample, 6 out of 40 pupils are left-handed, a proportion of 6 ÷ 40 = 0.15. Applying that proportion to the school's 900 pupils gives an estimate of 0.15 × 900 = 135 pupils. Giving 6 simply repeats the number of left-handed pupils IN THE SAMPLE, without scaling up to the whole school at all. Multiplying the population by the number of left-handed pupils in the sample without first dividing by the sample size, 900 × 6 = 5400, badly overestimates — that is more pupils than the whole school has. Dividing the population by the sample size but forgetting to multiply by the number of left-handed pupils found, 900 ÷ 40 = 22.5, finds the scale factor but stops one step short of using it. Always find the proportion in the sample first, then scale that same proportion up to the population.
- (b) 1 — Method: for data in a frequency table, find the position of the median using (n + 1) ÷ 2, then read off the value at that position from the cumulative frequencies. Working: there are 19 pupils, so the median is the 10th value. The cumulative frequencies are 7 (up to 0 pets), 10 (up to 1 pet), 14 (up to 2 pets) and 19 (up to 3 pets). The 10th value falls at the end of the '1 pet' group, so the median is 1 pet. Giving 0 pets is the mode — the category with the highest frequency, 7 — not the median. Giving 3, the highest number of pets minus the lowest, finds the range, a different statistic entirely. Giving 19 states the total number of pupils, not a number of pets at all. Find the middle POSITION first, then read off the value it belongs to — do not confuse it with the mode, the range or the total.
- (b) Treat them as outliers and check them before deciding — Method: a point lying a long way from the pattern the rest of the data make is called an outlier, and an outlier is investigated before anything is done with it, because it may be an error in the data or it may be a genuine but unusual case. Working: 48 of the 50 points lie close to the rising line of best fit, so the trend is set by those 48; the two remaining points do not follow it, so they are identified as outliers and checked — a mistake in measuring or recording would be corrected, while a genuine reading would be kept and reported. Answer: treat them as outliers and check them before deciding what to do with them. The distractors: deleting them at once assumes that every point far from the line must be an error, which throws away real data; moving the line so that it passes through them assumes a line of best fit must touch particular points, when it is drawn to follow all 50; taking them as proof that there is no correlation lets two points overturn the pattern that the other 48 agree on.
Build your own mix at the worksheet builder.