Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- 1.A scatter graph shows the midday temperature, x °C, and the number of ice creams sold at a seaside kiosk, y. The line of best fit is y = 3x − 20. Give a reason why the y-intercept of this line of best fit is not a sensible estimate of the number of ice creams sold.y = 3x − 20
- 2.Marta is drawing a cumulative frequency diagram for the times, t seconds, of 100 telephone calls. The grouped frequencies are: 0 ≤ t < 10, 7 calls; 10 ≤ t < 20, 19 calls; 20 ≤ t < 30, 34 calls; 30 ≤ t < 40, 40 calls. Write down the coordinates of the point Marta should plot for the class 20 ≤ t < 30.
- 3.A call centre records the length, t seconds, of 100 calls: 0 ≤ t < 20, 15 calls; 20 ≤ t < 30, 24 calls; 30 ≤ t < 50, 40 calls; 50 ≤ t < 80, 21 calls. The manager's target is for a call to be finished in under 35 seconds. Estimate the number of calls that met the target.
- 4.The heights, h cm, of 80 plants are grouped like this: 0 ≤ h < 20, 14 plants; 20 ≤ h < 40, 22 plants; 40 ≤ h < 50, 16 plants; 50 ≤ h < 80, 28 plants. Write down the class interval that contains the lower quartile.
- 5.A frequency polygon for the mass, in kg, of 40 parcels at a delivery depot is drawn by plotting one point at the midpoint of each class, joined by straight lines: (5, 6), (15, 10), (25, 16), (35, 6), (45, 2). Every class has a width of 10 kg. Write down the modal class.
- 6.The daily high temperatures, in °C, in Leeds over five days were 14, 16, 18, 19 and 23. In York over the same five days they were 15, 17, 17, 18 and 18. Write a sentence comparing Leeds and York using both the median and the range.
- 7.A scatter graph shows the number of years of experience, x, of 18 sales assistants and their monthly sales, y hundred pounds. The plotted points run from x = 1 to x = 12 years, and the line of best fit is y = 4x + 20. A new assistant has 25 years of experience. Use the line of best fit to estimate a value of y for this assistant, and decide whether the estimate would be reliable.y = 4x + 20
- 8.A council in Leeds wants to know what local people think about letting shops stay open later in the evening. It rings landline telephone numbers between 10 am and 2 pm on a Tuesday. Write down which group is most likely to be under-represented in the sample, and give a reason for your answer.
- 9.A school has 1200 pupils. A teacher wants to take a random sample of 60 of them. Write down which of these methods gives a random sample.
- 10.A recruitment agency in Manchester compares the weekly pay, in pounds, of seven employees at two branches. Branch A: £480, £495, £500, £505, £510, £515, £1,200. Branch B: £480, £490, £500, £510, £520, £530, £540. An advert for the agency claims 'Branch A pays more on average.' Decide whether this claim is fairly supported, using an appropriate average, and choose the correct conclusion.
- 11.A cumulative frequency graph for the diameters, d mm, of 320 ball bearings is plotted from these points (upper class boundary, cumulative frequency): (10, 30), (20, 90), (30, 190), (40, 280), (50, 320). Estimate the diameter below which 90% of the ball bearings measure.
- 12.A scatter graph has 50 points. Most of them lie close to a rising line of best fit, but two of them lie a long way from that line. Write down how those two points should be treated.
- 13.A dual bar chart shows the number of hours of rain recorded in Leeds and in Bristol on each of four days. Leeds: Monday 3 hours, Tuesday 5 hours, Wednesday 2 hours, Thursday 4 hours. Bristol: Monday 4 hours, Tuesday 4 hours, Wednesday 6 hours, Thursday 2 hours. Work out the greatest amount, in hours, by which Bristol's rainfall exceeded Leeds's rainfall on a single day.
- 14.The times taken, in minutes, by 30 runners in a Portsmouth fun run are grouped in this table: 0 < t ≤ 20 — 5 runners, 20 < t ≤ 40 — 10 runners, 40 < t ≤ 60 — 10 runners, 60 < t ≤ 80 — 5 runners. Work out an estimate for the mean time, in minutes.
- 15.A factory makes a batch of 4,000 circuit boards. It checks a random sample of 50 boards and finds that 4 are faulty. The factory will scrap the whole batch if the estimated number of faulty boards in the batch is more than 250. Should the factory scrap the batch?
Answer key
- (a) x = 0 gives y = −20: a negative number sold — The y-intercept is the value the line predicts when x = 0: y = 3 × 0 − 20 = −20. A kiosk cannot sell a negative number of ice creams, so this is not a sensible estimate. The 3 in the equation is the gradient, not the intercept, so an option claiming x = 0 gives y = 3 has swapped the two numbers around — substituting x = 0 makes the 3x term equal 0, leaving −20, not 3. The danger of extrapolating to very high temperatures is a real issue with this line, but it is a different issue from the y-intercept, so it does not answer this question. And whether x = 0 could occur on a trading day is beside the point: the model still makes that prediction, and it is the prediction itself, −20, that is impossible.
- (a) (30, 60) — Method: a cumulative frequency point is plotted at the upper boundary of its class, paired with the running total of all the frequencies up to and including that class. Working: the running totals are 7, then 7 + 19 = 26, then 26 + 34 = 60, then 60 + 40 = 100; the class 20 ≤ t < 30 has upper boundary 30, and the running total there is 60. Answer: the point for that class is plotted at 30 seconds against a cumulative frequency of 60. The distractors: (25, 60) comes from plotting at the class midpoint, which is what a frequency polygon uses and not what a cumulative frequency diagram uses; (30, 34) comes from plotting the class frequency, 34, rather than the running total; (20, 60) comes from plotting at the lower boundary of the class, which would claim that 60 calls took less than 20 seconds when only 26 did.
- (a) 49 — Method: add the frequencies of the classes that lie wholly below 35 seconds, then use linear interpolation for the class that 35 cuts through, assuming the calls in that class are spread evenly across it. Working: below 30 seconds there are 15 + 24 = 39 calls; the value 35 lies in the class 30 ≤ t < 50, which is 20 seconds wide and holds 40 calls, and 35 is 35 − 30 = 5 seconds into it, so the estimated share is (5 ÷ 20) × 40 = 10 calls; the estimate is 39 + 10 = 49. Answer: about 49 calls met the target. The distractors: 79 comes from adding the whole of the class 30 ≤ t < 50, 39 + 40, and so counting calls of up to 50 seconds as being under 35; 39 comes from stopping at the class boundary 30 and ignoring the part class altogether; 69 comes from measuring the part of the class from 35 up to 50 instead of from 30 up to 35, giving (15 ÷ 20) × 40 = 30 and then 39 + 30.
- (c) 20 ≤ h < 40 — Method: with 80 values the lower quartile is the 80 ÷ 4 = 20th value in order, so build a running total until it first reaches 20. Working: the running totals are 14, then 14 + 22 = 36, then 52, then 80; the 20th plant is past 14 but not past 36, so it lies in the second class. Answer: the lower quartile lies in the class 20 ≤ h < 40. The distractors: 0 ≤ h < 20 comes from believing that the bottom quarter of the data must all sit in the first class, when that class holds only 14 of the 80 plants; 40 ≤ h < 50 comes from using the position 80 ÷ 2 = 40 and so locating the median rather than the lower quartile; 50 ≤ h < 80 comes from counting 20 plants down from the tallest instead of up from the shortest, which locates the upper quartile at the 60th plant.
- (c) 20 kg ≤ mass < 30 kg — The modal class is the class with the highest frequency. Reading the plotted points, the frequencies are 6, 10, 16, 6 and 2, so the highest frequency is 16, plotted at the midpoint 25. A class of width 10 centred on 25 runs from 25 − 5 = 20 to 25 + 5 = 30, so the modal class is 20 kg ≤ mass < 30 kg. Writing '25 kg' gives only the midpoint, not the class — the modal class is an interval, not a single value. '10 kg ≤ mass < 20 kg' is the class before the peak, centred on 15, which has frequency 10, not the highest. '30 kg ≤ mass < 40 kg' is the class after the peak, centred on 35, which has frequency 6, not the highest.
- (a) Leeds has a higher median and a greater range than York. — In order, Leeds's temperatures are 14, 16, 18, 19 and 23, so the median is the middle value, 18, and the range is 23 − 14 = 9. York's temperatures in order are 15, 17, 17, 18 and 18, so the median is 17, and the range is 18 − 15 = 3. Since 18 is higher than 17, and 9 is greater than 3, Leeds has both the higher median and the greater range. Choosing 'Leeds has a higher median but a smaller range than York' gets the median comparison right but the range comparison backwards — Leeds's range of 9 is actually greater than York's range of 3. Choosing 'York has a higher median and a greater range than Leeds' reverses both comparisons. Choosing 'York has a higher median but a smaller range than Leeds' reverses the median comparison; York's median of 17 is lower than Leeds's 18, even though it is correct that York's range is the smaller one.
- (b) 120, unreliable — x = 25 is outside 1 to 12 — The line of best fit is y = 4x + 20. 4 × 25 = 100, and 100 + 20 = 120, so the estimate is y = 120. But x = 25 lies far outside the plotted range of 1 to 12 years, so this is an extrapolation, and the estimate is not reliable. Reaching 100 instead of 120 comes from 4 × 25 = 100 with the intercept of 20 left out — still correctly flagged as unreliable, but the wrong value. Calling the estimate reliable simply because it was calculated correctly, giving 120, wrongly assumes that a correct calculation is automatically trustworthy, ignoring that x = 25 lies far beyond the data actually collected. Reaching 68, from 4 × 12 = 48 and 48 + 20 = 68, substitutes x = 12, the top of the plotted range, instead of the assistant's actual x = 25, and wrongly calls that reliable because 12 lies inside the range.
- (d) Full-time workers, as most are at work at that time — Method: a sample is biased when the method of contact makes part of the population much less likely to be reached, so test each group against where its members actually are between 10 am and 2 pm on a weekday, and test each stated reason against the facts. Working: those hours are the middle of the working day, so people in full-time employment are at work and not beside a landline telephone, while people who are retired and people who are unemployed are far more likely to be at home and are reached at the usual rate; the method therefore collects far fewer replies from full-time workers than their share of the adult population the council is consulting. Answer: full-time workers, as most are at work at that time. The distractors: the reply naming retired people rests on the false claim that most retired people are at work in the daytime, when in fact a daytime call reaches them more easily than anyone; the reply naming unemployed people rests on the false claim that they are out during the day, when they too are among the easiest people to reach by a daytime call; the reply naming children rests on the false claim that children are at home at 11 am on a Tuesday, when they are at school and so are not reached by the call at all, and school-age children are in any case not the adults whose views the council is collecting.
- (a) Drawing 60 names at random from a list of all 1200 pupils — Method: a sample is random when every member of the population has the same chance of being chosen and nobody, including the pupils themselves, can influence who ends up in it; test each method against that. Working: drawing names from a list of all 1200 pupils gives each pupil the same chance, 60 out of 1200, whatever their year group, class or opinion, so the method is random. Answer: drawing 60 names at random from a list of all 1200 pupils. The distractors: asking the pupils who volunteer is self-selection, and the pupils with the strongest views volunteer first, so they decide the sample; asking the pupils nearest the door is convenience sampling, which reaches only those who happen to be in one place at one time; asking two Year 10 classes samples a cluster, so every pupil in the other year groups has no chance of being chosen at all.
- (c) No — median £505 at A vs £510 at B. — Branch A's seven wages in order are £480, £495, £500, £505, £510, £515 and £1,200, so the median, the 4th value, is £505. Branch B's in order are £480, £490, £500, £510, £520, £530 and £540, so the median is £510. Since £505 is lower than £510, the median wage is not higher at Branch A, so the claim is not fairly supported. Choosing 'Yes — mean £600.71 at A vs £510 at B' uses the mean: 480 + 495 + 500 + 505 + 510 + 515 + 1200 = 4205, and 4205 ÷ 7 = 600.71, a figure pulled upward by the £1,200 outlier that does not represent a typical wage. Choosing 'Yes — median £515 at A vs £510 at B' miscounts the middle position, taking the 6th wage, £515, instead of the correct 4th value, £505. Choosing 'Yes — highest wage £1,200 at A vs £540 at B' compares the highest wage at each branch rather than a measure of the typical, or average, wage.
- (c) 42 — Method: find the target cumulative frequency, 90% of the total, locate the class it falls in from the plotted points, then interpolate: lower boundary, plus the extra distance needed into the class divided by the class's frequency, times its width. Working: 90% of 320 is 0.9 × 320 = 288. The plotted points show a cumulative frequency of 280 at d = 40 and 320 at d = 50, so the class 40 ≤ d < 50 has frequency 320 − 280 = 40 and width 50 − 40 = 10, and 288 falls inside it. The extra distance needed into the class is 288 − 280 = 8, and 8 ÷ 40 × 10 = 2, so the diameter is 40 + 2 = 42. Answer: the estimated diameter is 42 mm. Watch which point and which class the interpolation actually uses: reading off d = 40, the plotted point just below the target, instead of interpolating the extra 8 ball bearings into the next 10 mm, stops one step short of the true answer; finding the diameter below which only 10% lie instead of 90% gives a target of 0.1 × 320 = 32, which falls in the class 10 ≤ d < 20 — the extra distance into that class is 32 − 30 = 2, and 2 ÷ 60 × 10 = 0.3, so this route gives 10 + 0.3 = 10.3, the bottom decile rather than the top 90%; and interpolating within the class 30 ≤ d < 40 instead of 40 ≤ d < 50, as though 288 had not yet reached a cumulative frequency of 280, treats the extra distance as 288 − 190 = 98, and 98 ÷ 90 × 10 = 10.9, giving 30 + 10.9 = 40.9, one class too early.
- (b) Treat them as outliers and check them before deciding — Method: a point lying a long way from the pattern the rest of the data make is called an outlier, and an outlier is investigated before anything is done with it, because it may be an error in the data or it may be a genuine but unusual case. Working: 48 of the 50 points lie close to the rising line of best fit, so the trend is set by those 48; the two remaining points do not follow it, so they are identified as outliers and checked — a mistake in measuring or recording would be corrected, while a genuine reading would be kept and reported. Answer: treat them as outliers and check them before deciding what to do with them. The distractors: deleting them at once assumes that every point far from the line must be an error, which throws away real data; moving the line so that it passes through them assumes a line of best fit must touch particular points, when it is drawn to follow all 50; taking them as proof that there is no correlation lets two points overturn the pattern that the other 48 agree on.
- (b) 4 — The difference, Bristol minus Leeds, on each day is: Monday 4 − 3 = 1, Tuesday 4 − 5 = −1, Wednesday 6 − 2 = 4, Thursday 2 − 4 = −2. The greatest amount by which Bristol exceeded Leeds is 4 hours, on Wednesday. Choosing 1 takes Monday's smaller positive difference instead of the greatest one. Choosing 2 takes the size of Thursday's difference, but that is the amount by which Leeds exceeded Bristol, the opposite direction to the one asked for. Choosing 6 takes Bristol's raw figure on Wednesday without subtracting Leeds's 2 hours first.
- (c) 40 minutes — Method: for grouped data, estimate the mean using the midpoint of each class — multiply each midpoint by its frequency, add the results, then divide by the total frequency. Working: the midpoints are 10, 30, 50 and 70 minutes. 10 × 5 = 50. 30 × 10 = 300. 50 × 10 = 500. 70 × 5 = 350. Σfx = 50 + 300 + 500 + 350 = 1200. Σf = 5 + 10 + 10 + 5 = 30. Estimated mean = 1200 ÷ 30 = 40 minutes. Using the upper boundary of each class instead of the midpoint — 20 × 5 = 100, 40 × 10 = 400, 60 × 10 = 600, 80 × 5 = 400 — gives a total of 1500 and an estimate of 1500 ÷ 30 = 50 minutes, too high because a boundary is not the middle of the class. Averaging the frequencies themselves, 5, 10, 10 and 5, ignores the times altogether and gives 7.5. Stopping after Σfx = 1200 without dividing by the total frequency gives a number far too large to be a time in minutes. Always find the midpoint of each class before multiplying by the frequency, and always divide by Σf at the end.
- (c) Yes — with an estimate of 320, above the 250 limit. — Method: scale the sample proportion up to the whole batch to get an estimate, then compare that estimate with the 250 limit to reach a decision. Working: in the sample, 4 out of 50 boards are faulty, a proportion of 4 ÷ 50 = 0.08. Applying that proportion to the batch of 4,000 gives an estimate of 0.08 × 4000 = 320 faulty boards. Since 320 is more than 250, the factory should scrap the batch. Inverting the proportion, 50 ÷ 4 = 12.5, and treating that as a percentage of the batch, 12.5% × 4000 = 500, still gives 'yes' but from the wrong fraction, so it overstates the estimate. Comparing the raw number of faulty boards found in the sample, 4, directly with the 250 limit skips the scaling up to the batch altogether, and 4 is nowhere near 250, so that route wrongly says 'no'. Dividing the batch by the sample size, 4000 ÷ 50 = 80, finds how many samples of 50 fit into the batch but stops before multiplying by the 4 faulty boards found, so it also wrongly says 'no'. Always find the proportion in the sample first, scale it up to the whole batch, and only then compare the estimate with the limit given.
Build your own mix at the worksheet builder.