Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- 1.A straight line is drawn on a scatter graph to show the trend of the points. Write down the name given to this line.
- 2.A scatter graph shows the mass, x kg, of a parcel and the cost, y pounds, of posting it. The line of best fit is y = 1.5x + 2. Work out the estimated cost of posting a parcel with a mass of 6 kg, using the line of best fit.y = 1.5x + 2
- 3.A factory makes a batch of 4,000 circuit boards. It checks a random sample of 50 boards and finds that 4 are faulty. The factory will scrap the whole batch if the estimated number of faulty boards in the batch is more than 250. Should the factory scrap the batch?
- 4.A two-way table records the favourite subject, Maths or Art, of 60 pupils in Year 10, and whether each pupil is left-handed or right-handed. 9 of the 60 pupils are left-handed, and 6 of those left-handed pupils prefer Art. In total, 24 of the 60 pupils prefer Art. A pupil is chosen at random from the 60. Work out the probability that the pupil is right-handed and prefers Art.
- 5.A two-way table records whether each of 50 pupils in Year 10 or Year 11 at a school walks to school or is driven. 28 of the 50 pupils are in Year 10. In total, 22 of the 50 pupils walk to school. Of the Year 10 pupils, 15 walk to school. Work out how many Year 11 pupils are driven to school.
- 6.A company makes 50,000 light bulbs a day and wants to check how long they last before they fail. Testing a bulb to find out how long it lasts destroys it. Give a reason why the company should test a sample of bulbs rather than every bulb it makes.
- 7.A scatter graph shows the height, x cm, and the mass, y kg, of 20 pupils in Year 10. The heights on the graph run from 150 cm to 180 cm, and the line of best fit is y = 0.9x − 85. Nadia puts x = 90 into this equation to estimate the mass of a two-year-old child who is 90 cm tall. Is her estimate reliable? Give a reason for your answer.y = 0.9x − 85
- 8.The mean of five test scores is 68. Four of the scores are 55, 62, 74 and 80. Work out the fifth score.
- 9.A scatter graph of taxi journeys in Bristol shows the distance, x miles, and the fare, y pounds. The line of best fit is y = 2x + 3.50. Work out the estimated fare for a journey of 6 miles, using the line of best fit.y = 2x + 3.5
- 10.A sports centre in Newcastle asks 60 members which activity they prefer: swimming 15 members, football 25 members, badminton 12 members and other activities 8 members. Work out the angle at the centre of the pie chart sector that represents badminton.
- 11.A café owner in Brighton records the midday temperature, x °C, and the number of hot chocolates sold, y, on 12 days. The temperatures recorded run from 4 °C to 18 °C, and the line of best fit is y = −3x + 74. The forecast for tomorrow gives a midday temperature of 12 °C. Work out the number the line of best fit predicts, and write down how much confidence the owner can have in it.y = -3x + 74
- 12.A school posts a questionnaire about school dinners to the families of all 200 pupils. Only 30 families send theirs back, and 27 of those 30 say they are unhappy with school dinners. Give a reason why this result may not represent all 200 families.
- 13.A school has 1200 pupils. A teacher wants to take a random sample of 60 of them. Write down which of these methods gives a random sample.
- 14.Four pairs of variables are listed below. Write down the pair that you would expect to show no correlation.
- 15.A survey of 25 pupils in Derby records how many siblings each has: 0 siblings — 6 pupils, 1 sibling — 10 pupils, 2 siblings — 6 pupils, 3 siblings — 3 pupils. Calculate the mean number of siblings.
Answer key
- (c) A line of best fit — Method: the straight line drawn on a scatter graph is named from the job it does — it is chosen so that it follows the whole set of points as closely as possible. Working: the line passes through the middle of the points, with roughly as many points above it as below it, and it need not pass through any of the plotted points at all; the name given to the straight line chosen in that way is a line of best fit. Answer: a line of best fit. The distractors: a line of symmetry comes from confusing a trend with symmetry, which is a property of a shape rather than of a set of data; a horizontal line through the mean comes from thinking the trend is shown by an average, when a horizontal line would say that the vertical quantity does not change and so show no correlation; a line joining the first and last points comes from thinking the line must join the two extreme points, which lets two points decide a trend that all of the points should share in.
- (b) £11.00 — 1.5 × 6 = 9, and 9 + 2 = 11, so the estimated cost is £11.00. Choosing £9.00 stops after 1.5 × 6 = 9 and forgets to add the £2. Choosing £12.00 adds the mass and the constant first and then multiplies: 6 + 2 = 8, and 8 × 1.5 = 12.00. Choosing £13.50 swaps the gradient and the intercept, using y = 2x + 1.5 instead: 2 × 6 = 12, and 12 + 1.5 = 13.50.
- (c) Yes — with an estimate of 320, above the 250 limit. — Method: scale the sample proportion up to the whole batch to get an estimate, then compare that estimate with the 250 limit to reach a decision. Working: in the sample, 4 out of 50 boards are faulty, a proportion of 4 ÷ 50 = 0.08. Applying that proportion to the batch of 4,000 gives an estimate of 0.08 × 4000 = 320 faulty boards. Since 320 is more than 250, the factory should scrap the batch. Inverting the proportion, 50 ÷ 4 = 12.5, and treating that as a percentage of the batch, 12.5% × 4000 = 500, still gives 'yes' but from the wrong fraction, so it overstates the estimate. Comparing the raw number of faulty boards found in the sample, 4, directly with the 250 limit skips the scaling up to the batch altogether, and 4 is nowhere near 250, so that route wrongly says 'no'. Dividing the batch by the sample size, 4000 ÷ 50 = 80, finds how many samples of 50 fit into the batch but stops before multiplying by the 4 faulty boards found, so it also wrongly says 'no'. Always find the proportion in the sample first, scale it up to the whole batch, and only then compare the estimate with the limit given.
- (a) 3/10 — There are 60 − 9 = 51 right-handed pupils. Of the 24 pupils who prefer Art, 6 are left-handed, so 24 − 6 = 18 are right-handed and prefer Art. The probability that a randomly chosen pupil is right-handed and prefers Art is 18/60, which simplifies to 3/10. Giving 2/5 is 24/60 simplified — the probability of preferring Art, ignoring the right-handed condition entirely. Giving 17/20 is 51/60 simplified — the probability of being right-handed, ignoring the Art condition entirely. Giving 1/10 is 6/60 simplified — the probability of being left-handed and preferring Art, the wrong hand condition.
- (a) 15 — Year 11 has 50 − 28 = 22 pupils in total. Of the 22 pupils who walk in total, 15 are in Year 10, so 22 − 15 = 7 Year 11 pupils walk. Subtracting that from the Year 11 total gives 22 − 7 = 15 Year 11 pupils who are driven. Choosing 28 takes the whole school's driven total, 50 − 22 = 28, and treats it as if it were Year 11's alone, without separating the year groups. Choosing 7 correctly finds how many Year 11 pupils walk but stops there, giving that figure instead of the number who are driven. Choosing 35 comes from 50 − 15, subtracting the Year 10 walkers from the whole school total rather than working within Year 11.
- (d) Testing destroys bulbs, so testing all leaves none to sell. — Method: testing every item in a population instead of a sample is a census — sensible only when testing does not use up or destroy what is being tested. Working: here, testing a bulb to find its lifespan destroys it, so testing all 50,000 bulbs would leave nothing left to sell — a sample lets the company estimate the typical lifespan without destroying its whole stock. Extra electricity used in testing is not the real reason a census is avoided here — it is the destruction of the product that matters. Saying a sample is always more accurate than a full census is the wrong way round: a census, if it could be carried out, gives the exact figure for the whole population — it is testing being destructive, not a lack of accuracy, that rules it out here. There is no law against testing every item a company makes — nothing in the question suggests that. When testing destroys the item being tested, sampling is necessary, not just convenient.
- (c) No, 90 cm is far outside the heights on the graph — Method: a line of best fit describes the trend only across the stretch of data it was drawn through; predicting beyond that stretch is extrapolation, and nothing in the data supports it. Working: the heights used to draw this line run from 150 cm to 180 cm, all of them Year 10 pupils, while 90 cm is 60 cm below the shortest of them and belongs to a two-year-old child, whose build follows no trend the graph has measured. Substituting anyway gives 0.9 × 90 − 85 = −4, a mass of −4 kg, which cannot exist. Answer: no, because 90 cm is far outside the heights on the graph. The distractors: saying a line of best fit cannot be used to predict at all throws away its main purpose, since a prediction made between the plotted values is perfectly sound; saying the line passes through all 20 points misdescribes a line of best fit, which is drawn to follow the trend of the points and will normally pass through few of them; saying the equation works for any value put into it treats an equation fitted to Year 10 heights as a law of nature, and the mass of −4 kg shows what that assumption produces.
- (c) 69 — Method: multiply the mean by the number of values to find the total, then subtract the total of the known values. Working: the total of all five scores is 68 × 5 = 340. The total of the four known scores is 55 + 62 + 74 + 80 = 271. The fifth score is 340 − 271 = 69. Subtracting the other way round, 271 − 340 = −69, gives the right size answer with the wrong sign. Guessing that the missing score simply equals the mean, 68, ignores that the four known scores are not themselves centred on 68. Multiplying the mean by 4 instead of 5, 68 × 4 = 272, then 272 − 271 = 1, undercounts how many scores there are. Always multiply the mean by the TOTAL number of values before subtracting.
- (d) £15.50 — 2 × 6 = 12, and 12 + 3.50 = 15.50, so the estimated fare is £15.50. Choosing £12.00 stops after 2 × 6 = 12 and forgets to add the £3.50. Choosing £19.00 adds the distance and the constant first and then multiplies: 6 + 3.50 = 9.50, and 9.50 × 2 = 19.00, applying the ×2 to the whole sum instead of only to the distance. Choosing £13.00 multiplies only the constant term by 2 instead of the distance: 2 × 3.50 = 7, and 7 + 6 = 13.00.
- (c) 72° — The angle is 12 ÷ 60 × 360 = 72°. Choosing 150° divides football's frequency of 25 instead of badminton's 12: 25 ÷ 60 × 360 = 150. Choosing 20° finds badminton as a percentage of the members, 12 ÷ 60 × 100 = 20, rather than an angle in degrees. Choosing 90° uses 48, the total of the other three activities, as the total instead of the full 60 members: 12 ÷ 48 × 360 = 90.
- (d) 38, and fairly confident, as 12 °C is inside the range — Method: substitute the forecast temperature into the equation of the line of best fit, then judge the prediction by where that temperature sits among the data the line was drawn from. Working: putting x = 12 into y = −3x + 74 gives −3 × 12 + 74 = 38, so the line predicts 38 hot chocolates. The recorded temperatures run from 4 °C to 18 °C, and 12 °C lies inside that interval, so this is interpolation, the safer kind of prediction. Answer: 38, and fairly confident, as 12 °C is inside the range; the owner should still expect the true figure to differ a little, since the points only lie near the line and not on it. The distractors: being completely certain treats a line of best fit as a rule that fixes each day's sales, when it describes a trend that individual days depart from; saying 12 °C is outside the range misreads the interval 4 °C to 18 °C, and the wrong warning would be attached to a sound prediction; 110 comes from −3 × 12 being taken as +36, giving 36 + 74 = 110, which loses the negative gradient and so predicts that a warm day sells more hot chocolate than a cold one.
- (d) Only families with strong feelings bothered to reply. — Method: a survey has non-response bias when only some of the people asked actually reply, and those who do are not a typical cross-section of everyone who was asked. Working: only 30 of the 200 families sent back their questionnaire, and 27 of those 30 — the great majority — said they were unhappy. Families who feel strongly about an issue, particularly those with a complaint, are far more likely to make the effort to reply than families who are simply satisfied and see no need to say anything, so the 30 replies over-represent unhappy families. Saying the families who replied were picked at random by the school gets the sampling the wrong way round: nobody picked them — they picked themselves by deciding to reply, and that is precisely why they are not a typical cross-section of all 200. Saying postal surveys always have low response rates restates that the response was low without explaining why a low response rate, on its own, makes a result unrepresentative — it is the reason FOR the low response, not the low response itself, that causes the bias here. Saying that the 27 unhappy replies show most families are unhappy is exactly the mistake the question is warning against: it treats the loudest 30 replies as if they stood for the other 170 who never sent theirs back. A low response rate is a warning sign only because the people who bother to reply are rarely typical of everyone who was asked.
- (a) Drawing 60 names at random from a list of all 1200 pupils — Method: a sample is random when every member of the population has the same chance of being chosen and nobody, including the pupils themselves, can influence who ends up in it; test each method against that. Working: drawing names from a list of all 1200 pupils gives each pupil the same chance, 60 out of 1200, whatever their year group, class or opinion, so the method is random. Answer: drawing 60 names at random from a list of all 1200 pupils. The distractors: asking the pupils who volunteer is self-selection, and the pupils with the strongest views volunteer first, so they decide the sample; asking the pupils nearest the door is convenience sampling, which reaches only those who happen to be in one place at one time; asking two Year 10 classes samples a cluster, so every pupil in the other year groups has no chance of being chosen at all.
- (b) A person's shoe size and their favourite colour — A person's shoe size is not linked to which colour they prefer, so these two show no correlation. The other three pairs are all genuinely correlated: distance travelled and fuel used rise together, which is positive correlation; hours of revision and test score generally rise together, which is also positive correlation; and as outdoor temperature rises, fewer woolly hats are sold, which is negative correlation. Negative correlation is still a real relationship between two variables — it is not the same thing as no relationship at all, so the temperature and hats pair is not the answer to this question.
- (a) 1.24 — Method: for data given as a frequency table, the mean is Σfx ÷ Σf — multiply each value by its frequency, add the results, then divide by the total frequency. Working: 0 × 6 = 0. 1 × 10 = 10. 2 × 6 = 12. 3 × 3 = 9. So Σfx = 0 + 10 + 12 + 9 = 31. The total frequency is Σf = 6 + 10 + 6 + 3 = 25. Mean = 31 ÷ 25 = 1.24 siblings. Averaging the frequency column itself, (6 + 10 + 6 + 3) ÷ 4 = 6.25, mixes up the frequencies with the values they belong to. Writing down 1, the number of siblings with the highest frequency, gives the mode, not the mean. Writing down 31 stops after finding Σfx and forgets to divide by the total frequency, 25. Always divide Σfx by Σf — never stop at the top of the fraction.
Build your own mix at the worksheet builder.