Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Answer key: Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- (b) Interpolation — The salary is being estimated for a value of x between 1 and 15, which is inside the range of x-values that were actually plotted, so this is interpolation. Extrapolation would apply if the estimate used a value of x below 1 or above 15, outside the plotted range. Correlation describes the relationship between the two variables, not the reliability of an estimate, and causation describes one variable actually causing a change in the other, which is a different idea altogether — neither is the word being asked for here.
- (d) Correlation is a link; causation is one causing the other — Method: the two words describe different claims — one is about a pattern in the data, the other is about what produced that pattern. Working: correlation says only that two quantities tend to change together, which is something a scatter graph can display; causation says that a change in one quantity actually brings about the change in the other, which needs evidence a scatter graph cannot supply, because a third quantity may be driving both. Answer: correlation is a link between the quantities, while causation is one quantity causing the change in another. The distractors: the statement giving causation as the link and correlation as the cause simply swaps the two words over; the statement that the words mean the same thing is the classic error of reading a correlation as proof of cause; the statement that a scatter graph shows causation but not correlation reverses what a scatter graph can do, since the pattern it displays is exactly the correlation.
- (c) No, the mode here is the lowest value of the nine — Method: an average is meant to stand for the data as a whole, so test any proposed average by asking how many values it sits near. Working: the value 4 appears three times and every other count appears once, so 4 is indeed the mode. But those three hours are the quiet ones at the start of the day, and the other six counts run from 11 up to 25; putting the nine counts in order, the middle one is the fifth, which is 13. So the mode sits at the very bottom of the data, with six of the nine hours far above it. Answer: no, because the mode here is the lowest value of the nine, so it describes the quiet opening hours rather than a typical hour. The distractors: saying the mode can only be used when no value repeats reverses the definition, since a mode exists only because a value does repeat; saying the mode is the value that occurs most often is a correct definition, but being the commonest value does not make a value typical when it lies at one end of the data; saying the mode is the best average for any list of numbers ignores the fact that mean, median and mode each describe a population well in different circumstances.
- (d) 38, and fairly confident, as 12 °C is inside the range — Method: substitute the forecast temperature into the equation of the line of best fit, then judge the prediction by where that temperature sits among the data the line was drawn from. Working: putting x = 12 into y = −3x + 74 gives −3 × 12 + 74 = 38, so the line predicts 38 hot chocolates. The recorded temperatures run from 4 °C to 18 °C, and 12 °C lies inside that interval, so this is interpolation, the safer kind of prediction. Answer: 38, and fairly confident, as 12 °C is inside the range; the owner should still expect the true figure to differ a little, since the points only lie near the line and not on it. The distractors: being completely certain treats a line of best fit as a rule that fixes each day's sales, when it describes a trend that individual days depart from; saying 12 °C is outside the range misreads the interval 4 °C to 18 °C, and the wrong warning would be attached to a sound prediction; 110 comes from −3 × 12 being taken as +36, giving 36 + 74 = 110, which loses the negative gradient and so predicts that a warm day sells more hot chocolate than a cold one.
- (b) All 2,000 members of the sports centre. — Method: in a survey, the population is the whole group the survey is trying to find out about, and the sample is the smaller group actually asked. Working: this survey wants to know what the sports centre's members think, so the population is every one of the 2,000 members — whether or not they were personally asked. Saying the population is the 100 members who were asked names the sample, not the population; the sample is drawn FROM the population, so it is smaller than it, not the same as it. Saying the population is everybody who lives in Ipswich widens the group far beyond who the survey is actually about — plenty of Ipswich residents are not members of the sports centre at all, so they are outside this survey altogether. Saying the population is the members who say they are unhappy confuses the population with a result of the survey: whether a member turns out to be happy or unhappy is something the survey finds out, not part of the definition of who is being studied. The population is always the whole group the question is about, before any sampling or any results come in.
- (b) 120, unreliable — x = 25 is outside 1 to 12 — The line of best fit is y = 4x + 20. 4 × 25 = 100, and 100 + 20 = 120, so the estimate is y = 120. But x = 25 lies far outside the plotted range of 1 to 12 years, so this is an extrapolation, and the estimate is not reliable. Reaching 100 instead of 120 comes from 4 × 25 = 100 with the intercept of 20 left out — still correctly flagged as unreliable, but the wrong value. Calling the estimate reliable simply because it was calculated correctly, giving 120, wrongly assumes that a correct calculation is automatically trustworthy, ignoring that x = 25 lies far beyond the data actually collected. Reaching 68, from 4 × 12 = 48 and 48 + 20 = 68, substitutes x = 12, the top of the plotted range, instead of the assistant's actual x = 25, and wrongly calls that reliable because 12 lies inside the range.
- (d) £15.50 — 2 × 6 = 12, and 12 + 3.50 = 15.50, so the estimated fare is £15.50. Choosing £12.00 stops after 2 × 6 = 12 and forgets to add the £3.50. Choosing £19.00 adds the distance and the constant first and then multiplies: 6 + 3.50 = 9.50, and 9.50 × 2 = 19.00, applying the ×2 to the whole sum instead of only to the distance. Choosing £13.00 multiplies only the constant term by 2 instead of the distance: 2 × 3.50 = 7, and 7 + 6 = 13.00.
- (d) Testing destroys bulbs, so testing all leaves none to sell. — Method: testing every item in a population instead of a sample is a census — sensible only when testing does not use up or destroy what is being tested. Working: here, testing a bulb to find its lifespan destroys it, so testing all 50,000 bulbs would leave nothing left to sell — a sample lets the company estimate the typical lifespan without destroying its whole stock. Extra electricity used in testing is not the real reason a census is avoided here — it is the destruction of the product that matters. Saying a sample is always more accurate than a full census is the wrong way round: a census, if it could be carried out, gives the exact figure for the whole population — it is testing being destructive, not a lack of accuracy, that rules it out here. There is no law against testing every item a company makes — nothing in the question suggests that. When testing destroys the item being tested, sampling is necessary, not just convenient.
- (c) A line of best fit — Method: the straight line drawn on a scatter graph is named from the job it does — it is chosen so that it follows the whole set of points as closely as possible. Working: the line passes through the middle of the points, with roughly as many points above it as below it, and it need not pass through any of the plotted points at all; the name given to the straight line chosen in that way is a line of best fit. Answer: a line of best fit. The distractors: a line of symmetry comes from confusing a trend with symmetry, which is a property of a shape rather than of a set of data; a horizontal line through the mean comes from thinking the trend is shown by an average, when a horizontal line would say that the vertical quantity does not change and so show no correlation; a line joining the first and last points comes from thinking the line must join the two extreme points, which lets two points decide a trend that all of the points should share in.
- (a) A vertical line chart (discrete numerical data) — The number of pets is discrete numerical data — whole-number values such as 0, 1, 2, 3 or 4 — recorded for one variable, so a vertical line chart is the chart specified for this kind of data. A bar chart is used for categorical data, such as favourite colour, not numerical values counted like this. A pie chart shows proportions of a whole and does not show the frequency of each separate value. A scatter graph compares two different variables against each other, and only one variable, the number of pets, is recorded here.
- (a) 3/10 — There are 60 − 9 = 51 right-handed pupils. Of the 24 pupils who prefer Art, 6 are left-handed, so 24 − 6 = 18 are right-handed and prefer Art. The probability that a randomly chosen pupil is right-handed and prefers Art is 18/60, which simplifies to 3/10. Giving 2/5 is 24/60 simplified — the probability of preferring Art, ignoring the right-handed condition entirely. Giving 17/20 is 51/60 simplified — the probability of being right-handed, ignoring the Art condition entirely. Giving 1/10 is 6/60 simplified — the probability of being left-handed and preferring Art, the wrong hand condition.
- (d) 24 — Method: for a sample in proportion to the population, apply the same fraction that each group makes up of the whole population to the size of the sample. Working: women make up 300 out of the 500 members, a fraction of 300 ÷ 500 = 0.6. Applying that fraction to the sample of 40 gives 0.6 × 40 = 24 women. Splitting the sample evenly, 40 ÷ 2 = 20, ignores that the club has more women than men and treats the two groups as equal in size, which they are not. Misreading the sample size as 50 instead of 40, then applying the 3:2 ratio of women to men, 3 ÷ 5 × 50 = 30, uses the right ratio but the wrong sample total. Working out the number of MEN instead of women, 200 ÷ 500 × 40 = 16, answers a different question — how many men, not how many women, belong in the sample. Always apply each group's own share of the population to the sample size, and check which group the question is actually asking about.
- (d) Only internet users reach the website; others are excluded. — Method: a sample is biased when it systematically leaves out part of the population, or systematically over-represents another part. Working: anyone without internet access, or who does not visit the council's website, has NO chance of being included — the sample is drawn only from internet-using residents, which is not the whole town. Saying too many people might respond because the survey is free confuses bias with sample size — bias is about who CAN be reached, not how many respond. Saying people might lie describes a different problem, response honesty, not who was sampled in the first place. Saying online surveys cannot be anonymous is not a reason connected to bias at all. A sample is biased when part of the population has no chance of being included, whatever the reason for that.
- (d) 1100 — Method: to combine two samples of different sizes, add the faulty counts together and add the sample sizes together before scaling up, rather than treating the two samples separately. Working: the combined sample found 34 + 21 = 55 scratched cases out of 100 + 50 = 150 cases checked, a proportion of 55 ÷ 150. Applying that proportion to the week's production of 3,000 gives an estimate of 55 ÷ 150 × 3000 = 1100 scratched cases. Averaging the two shifts' proportions instead of combining their totals, (34 ÷ 100 + 21 ÷ 50) ÷ 2 = 0.38, gives 0.38 × 3000 = 1140 — this treats the two samples as equally weighted even though Shift A checked twice as many cases as Shift B. Using only Shift A's sample, 34 ÷ 100 × 3000 = 1020, ignores Shift B's cases completely. Using only Shift B's sample, 21 ÷ 50 × 3000 = 1260, ignores Shift A's cases completely. When two samples are different sizes, combine their totals before finding the proportion — do not average the two proportions, and do not use only one shift's sample.
- (d) Only families with strong feelings bothered to reply. — Method: a survey has non-response bias when only some of the people asked actually reply, and those who do are not a typical cross-section of everyone who was asked. Working: only 30 of the 200 families sent back their questionnaire, and 27 of those 30 — the great majority — said they were unhappy. Families who feel strongly about an issue, particularly those with a complaint, are far more likely to make the effort to reply than families who are simply satisfied and see no need to say anything, so the 30 replies over-represent unhappy families. Saying the families who replied were picked at random by the school gets the sampling the wrong way round: nobody picked them — they picked themselves by deciding to reply, and that is precisely why they are not a typical cross-section of all 200. Saying postal surveys always have low response rates restates that the response was low without explaining why a low response rate, on its own, makes a result unrepresentative — it is the reason FOR the low response, not the low response itself, that causes the bias here. Saying that the 27 unhappy replies show most families are unhappy is exactly the mistake the question is warning against: it treats the loudest 30 replies as if they stood for the other 170 who never sent theirs back. A low response rate is a warning sign only because the people who bother to reply are rarely typical of everyone who was asked.
Build your own mix at the worksheet builder.