Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Answer key: Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- (d) 6 — Method: the mode, or modal value, is the value that occurs most often in the data set, and it is a value from the data rather than a count. Working: size 4 occurs twice, size 5 occurs once, size 6 occurs three times and size 9 occurs once, so the highest frequency is three and the size it belongs to is 6. Answer: 6. The distractors: 3 comes from writing down the frequency of the most common size instead of the size itself; 9 comes from picking the largest size in the list, which confuses the mode with the maximum; 4 comes from stopping at the first size that repeats rather than checking which size repeats most often.
- (c) The data show a link only; a third factor may affect both — Method: a study of this kind measures two quantities and reports how they change together; deciding that one of them produces the other is a further claim, and it needs evidence that the measurements alone cannot give. Working: the study shows that more coffee goes with better concentration, which is a positive correlation; but a third factor that was never measured, such as how motivated someone is, could raise both the coffee drinking and the concentration, and the concentration could equally be what leads to the extra coffee. Answer: the data show a link only, because a third factor may be affecting both quantities, so no claim about cause can be made. The distractors: calling the conclusion safe because the correlation is positive treats the direction of a correlation as proof of cause, which no direction can give; calling it wrong because the correlation is negative misreads the direction of the relationship, since the study reports both quantities rising together; saying the two quantities are not linked denies the correlation the study actually found, when what fails is only the claim about cause.
- (d) About 510 of the 600 bulbs are likely to have flowered — Method: the proportion found in a random sample is used as an estimate of the proportion in the whole population, and the conclusion is stated as an estimate, never as a fact about every member. Working: 17 of the 20 bulbs dug up had flowered, so the sample proportion is 17 ÷ 20 = 0.85, and applying that proportion to the whole planting gives 0.85 × 600 = 510 bulbs. A different random sample of 20 would very probably give a slightly different figure, so 510 is an estimate. Answer: about 510 of the 600 bulbs are likely to have flowered. The distractors: saying exactly 510 have flowered takes an estimate from a sample of 20 as a count of all 600, which no sample can deliver; saying exactly 17 of the 600 have flowered reports the sample count as though it were the population count, leaving the other 580 bulbs out of the answer altogether; saying about 20 have flowered uses the size of the sample as the estimate, when 20 is the number of bulbs she dug up rather than a number that flowered.
- (a) The median, as the one very large wage does not move it — Method: an average describes a population well when it sits close to most of the values, so compare what each average does when one value lies far from the rest. Working: in order the wages are 420, 440, 460, 480 and 1,500, so the median is the third of the five, £460. The mean uses every wage: 420 + 440 + 460 + 480 + 1,500 = 3,300 and 3,300 ÷ 5 = 660, so the mean is £660. Four of the five people earn less than £660, and the nearest of those four wages is £180 below it, so £660 describes nobody at the garage; £460 sits inside the group of four similar wages. Answer: the median, as the one very large wage does not move it, while that same wage drags the mean £200 above the median. The distractors: saying the median is always larger than the mean is an invented rule, and here the median £460 is smaller than the mean £660; saying the mean is the only average that uses all five wages is true as far as it goes, but using a value and being dragged by it are the same thing when that value is £1,500; saying £660 lies between the smallest and largest wage is true of every mean ever calculated, so it proves nothing about whether this one is typical.
- (c) Because every pupil has an equal chance of being picked — Method: whether a sample represents its population is decided by the selection method, not by the size of the sample, so ask whether the method gives every member of the population the same chance of being chosen. Working: the names are drawn at random from a list of all 10,000 pupils, so each pupil has the same chance, 500 out of 10,000, of being drawn, and no group of pupils is more likely to appear than any other; that is what keeps bias out of the sample. Answer: because every pupil has an equal chance of being picked. The distractors: the reply about 5% treats the sampling fraction as the test of fairness, but a badly chosen 5% is still biased and a well chosen 1% is not; the reply about 500 being large enough makes size the test instead, which is the same mistake in another form, since a large sample drawn from one school would still misrepresent the city; the reply about the most willing pupils describes self-selection, which hands the choice of who is in the sample to the pupils who feel most strongly about the question.
- (a) No correlation — Shoe size has no real relationship with spelling ability, and the points here are scattered with no rising or falling trend, so this is no correlation. A positive correlation would show the points rising together, and a negative correlation would show them falling as one increases; neither pattern is present here. Strong correlation is not correct either, since strength only applies once a positive or negative trend exists, and there isn't one.
- (a) 62.5 — Method: a mean cannot be averaged with a new value — rebuild the total, add the new value to it, then divide by the new count. Working: three numbers with a mean of 50 have a total of 50 × 3 = 150; adding 100 makes the total 150 + 100 = 250; there are now 4 numbers, so the new mean is 250 ÷ 4 = 62.5. Answer: 62.5. The distractors: 75 comes from averaging the old mean with the new value, (50 + 100) ÷ 2, which ignores that three numbers pull against one; 50 comes from assuming an extra value leaves the mean unchanged; 37.5 comes from dividing the old total of 150 by the new count of 4, adding the new value to the count but not to the total.
- (c) 33 — Method: multiply the mean by the number of tests to get the total marks, then subtract the marks that are already known. Working: four tests with a mean of 29 give a total of 29 × 4 = 116 marks; the first three marks total 31 + 26 + 26 = 83; so the fourth mark is 116 − 83 = 33. Answer: 33, and checking, (31 + 26 + 26 + 33) ÷ 4 = 116 ÷ 4 = 29. The distractors: 116 comes from stopping at the total for all four tests; 29 comes from assuming the missing mark must be the mean itself; 4 comes from multiplying the mean by 3, the number of marks given, leaving 87 − 83 = 4.
- (d) 24 — Method: for a sample in proportion to the population, apply the same fraction that each group makes up of the whole population to the size of the sample. Working: women make up 300 out of the 500 members, a fraction of 300 ÷ 500 = 0.6. Applying that fraction to the sample of 40 gives 0.6 × 40 = 24 women. Splitting the sample evenly, 40 ÷ 2 = 20, ignores that the club has more women than men and treats the two groups as equal in size, which they are not. Misreading the sample size as 50 instead of 40, then applying the 3:2 ratio of women to men, 3 ÷ 5 × 50 = 30, uses the right ratio but the wrong sample total. Working out the number of MEN instead of women, 200 ÷ 500 × 40 = 16, answers a different question — how many men, not how many women, belong in the sample. Always apply each group's own share of the population to the sample size, and check which group the question is actually asking about.
- (b) 74 marks — Method: the two groups are different sizes, so their means cannot simply be averaged — rebuild each group's total mark, add the totals and divide by all 50 pupils. Working: Group A scored 20 × 80 = 1600 marks and Group B scored 30 × 70 = 2100 marks, giving 1600 + 2100 = 3700 marks altogether, so the overall mean is 3700 ÷ 50 = 74 marks. Answer: 74 marks, which sits nearer to 70 than to 80 because the larger group scored 70. The distractors: 75 marks comes from averaging the two group means, (80 + 70) ÷ 2, as though the groups were the same size; 76 marks comes from attaching each mean to the other group's size, (20 × 70 + 30 × 80) ÷ 50; 150 marks comes from adding the two means together and never dividing at all.
- (d) Full-time workers, as most are at work at that time — Method: a sample is biased when the method of contact makes part of the population much less likely to be reached, so test each group against where its members actually are between 10 am and 2 pm on a weekday, and test each stated reason against the facts. Working: those hours are the middle of the working day, so people in full-time employment are at work and not beside a landline telephone, while people who are retired and people who are unemployed are far more likely to be at home and are reached at the usual rate; the method therefore collects far fewer replies from full-time workers than their share of the adult population the council is consulting. Answer: full-time workers, as most are at work at that time. The distractors: the reply naming retired people rests on the false claim that most retired people are at work in the daytime, when in fact a daytime call reaches them more easily than anyone; the reply naming unemployed people rests on the false claim that they are out during the day, when they too are among the easiest people to reach by a daytime call; the reply naming children rests on the false claim that children are at home at 11 am on a Tuesday, when they are at school and so are not reached by the call at all, and school-age children are in any case not the adults whose views the council is collecting.
- (c) Neither causes the other; sunshine links both. — Both ice cream sales and sunburn cases tend to rise on hot, sunny days, so the amount of sunshine is a third factor linked to both — neither variable causes the other. Saying ice cream sales cause the sunburn assumes a causal link in one direction that the correlation alone cannot establish. Saying sunburn cases cause the ice cream sales assumes the reverse causal link, which is no more justified. Saying a strong correlation always means causation is the general error this question is testing: correlation, however strong, does not by itself prove that one variable causes the other.
- (b) Treat them as outliers and check them before deciding — Method: a point lying a long way from the pattern the rest of the data make is called an outlier, and an outlier is investigated before anything is done with it, because it may be an error in the data or it may be a genuine but unusual case. Working: 48 of the 50 points lie close to the rising line of best fit, so the trend is set by those 48; the two remaining points do not follow it, so they are identified as outliers and checked — a mistake in measuring or recording would be corrected, while a genuine reading would be kept and reported. Answer: treat them as outliers and check them before deciding what to do with them. The distractors: deleting them at once assumes that every point far from the line must be an error, which throws away real data; moving the line so that it passes through them assumes a line of best fit must touch particular points, when it is drawn to follow all 50; taking them as proof that there is no correlation lets two points overturn the pattern that the other 48 agree on.
- (a) Equal means; Class A is more consistent, smaller range. — Method: when two data sets share a measure of location, compare a measure of spread to say more about consistency. Working: both classes have the same mean mark, 14, so on average they performed equally well. Class A has the smaller range, 6, so its marks are more tightly grouped around 14 than Class B's marks, which vary by as much as 14. So Class A's marks were more consistent, even though neither class did better on average. Saying Class B did better because it has the bigger range confuses a wide spread with a high score — a big range describes variability, not performance. Saying Class A did better because it has the smaller range makes the same mistake in the other direction: the two classes are tied on the mean, so neither one 'did better'. Saying the classes cannot be compared because their means are equal misses the whole point of also comparing the range. Always compare both an average AND a spread before describing two data sets — either one alone tells only half the story.
- (c) Strong negative correlation — Method: correlation is described by two things — the direction the points take as the graph is read from left to right, and how closely the points lie to a single straight line. Working: reading the pairs in order of age, the ages rise 14, 18, 23, 27, 31, 36, 42, 49 while the scores fall 92, 88, 85, 80, 78, 74, 70, 65; the score falls at every single step, with no reversal anywhere, so the points fall from left to right and lie close to a straight line. Answer: strong negative correlation — negative for the falling direction, strong because every point follows the pattern. The distractors: strong positive correlation comes from noticing a clear pattern and calling any clear pattern positive, without checking the direction; weak negative correlation comes from reading the direction correctly but judging points that do not lie exactly on a straight line to be only loosely related, when these eight fall without a single exception; no correlation comes from reading a falling trend as though it showed no relationship at all, when a falling trend is itself a relationship.
Build your own mix at the worksheet builder.