Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Answer key: Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- (d) Only families with strong feelings bothered to reply. — Method: a survey has non-response bias when only some of the people asked actually reply, and those who do are not a typical cross-section of everyone who was asked. Working: only 30 of the 200 families sent back their questionnaire, and 27 of those 30 — the great majority — said they were unhappy. Families who feel strongly about an issue, particularly those with a complaint, are far more likely to make the effort to reply than families who are simply satisfied and see no need to say anything, so the 30 replies over-represent unhappy families. Saying the families who replied were picked at random by the school gets the sampling the wrong way round: nobody picked them — they picked themselves by deciding to reply, and that is precisely why they are not a typical cross-section of all 200. Saying postal surveys always have low response rates restates that the response was low without explaining why a low response rate, on its own, makes a result unrepresentative — it is the reason FOR the low response, not the low response itself, that causes the bias here. Saying that the 27 unhappy replies show most families are unhappy is exactly the mistake the question is warning against: it treats the loudest 30 replies as if they stood for the other 170 who never sent theirs back. A low response rate is a warning sign only because the people who bother to reply are rarely typical of everyone who was asked.
- (a) Minutes a candle has burned and length remaining — As a candle burns for longer, less of it remains, so these two variables move in opposite directions as one increases — that is negative correlation. A pupil's shoe size generally increases as they get older, so age and shoe size show positive correlation, not negative, since both rise together. A football team's shirt colour is not a numerical quantity linked to how many matches it wins, so shirt colour and number of wins show no correlation at all. The number of letters in a pupil's name has no real connection to their ability in maths, so that pair also shows no correlation.
- (c) No, the mode here is the lowest value of the nine — Method: an average is meant to stand for the data as a whole, so test any proposed average by asking how many values it sits near. Working: the value 4 appears three times and every other count appears once, so 4 is indeed the mode. But those three hours are the quiet ones at the start of the day, and the other six counts run from 11 up to 25; putting the nine counts in order, the middle one is the fifth, which is 13. So the mode sits at the very bottom of the data, with six of the nine hours far above it. Answer: no, because the mode here is the lowest value of the nine, so it describes the quiet opening hours rather than a typical hour. The distractors: saying the mode can only be used when no value repeats reverses the definition, since a mode exists only because a value does repeat; saying the mode is the value that occurs most often is a correct definition, but being the commonest value does not make a value typical when it lies at one end of the data; saying the mode is the best average for any list of numbers ignores the fact that mean, median and mode each describe a population well in different circumstances.
- (c) The data show a link only; a third factor may affect both — Method: a study of this kind measures two quantities and reports how they change together; deciding that one of them produces the other is a further claim, and it needs evidence that the measurements alone cannot give. Working: the study shows that more coffee goes with better concentration, which is a positive correlation; but a third factor that was never measured, such as how motivated someone is, could raise both the coffee drinking and the concentration, and the concentration could equally be what leads to the extra coffee. Answer: the data show a link only, because a third factor may be affecting both quantities, so no claim about cause can be made. The distractors: calling the conclusion safe because the correlation is positive treats the direction of a correlation as proof of cause, which no direction can give; calling it wrong because the correlation is negative misreads the direction of the relationship, since the study reports both quantities rising together; saying the two quantities are not linked denies the correlation the study actually found, when what fails is only the claim about cause.
- (a) 150 bulbs — Method: assume the proportion faulty in a random sample is the proportion faulty in the whole day's output, and scale the sample up to the population. Working: the sample of 80 has to be scaled up to 4,000 bulbs, and 4,000 ÷ 80 = 50, so the day's output is 50 sample-sized batches. Each batch is expected to contain the same 3 faulty bulbs, so the estimate is 3 × 50 = 150. Answer: 150 bulbs, and it is an estimate, because another sample of 80 would probably contain a different number of faulty bulbs. The distractors: 50 bulbs is the scale factor 4,000 ÷ 80 written down as though it were the answer, so it reports how many batches there are rather than how many faulty bulbs; 120 bulbs comes from reading 3 out of 80 as 3%, then taking 0.03 × 4,000 = 120, but 3 out of 80 is 3.75%; 240 bulbs comes from 3 × 80 = 240, multiplying the faulty bulbs by the size of the sample instead of by the scale factor, which uses the 80 twice and the 4,000 not at all.
- (a) 144° — Method: a pie chart shares the 360° at its centre between the categories in proportion to their frequencies, so the angle of a sector is that category's fraction of the total multiplied by 360°. Working: the sector stands for 40 pupils out of 100, which is the fraction 40/100; one pupil is worth 360 ÷ 100 = 3.6°, so 40 pupils are worth 40 × 3.6° = 144°. Answer: 144°, and the answer is an angle in degrees, not a number of pupils. The distractors: 40° comes from sharing out 100 instead of 360, so the percentage is written straight down as a number of degrees; 216° comes from working out the angle for the other 60 pupils, 60 × 3.6°, which is the rest of the pie chart; 180° comes from assuming that the tallest line must stand for half of the pupils and so take half of the chart.
- (c) 67 — Method: the mean is the total of the values divided by how many values there are, so add first and divide second. Working: the total is 50 + 83 + 68 = 201 points and three matches were played, so the mean is 201 ÷ 3 = 67 points. Answer: 67. The distractors: 68 comes from writing down the median, the middle value of 50, 68, 83, instead of the mean; 33 comes from working out the range, 83 − 50, which measures spread and not centre; 100.5 comes from dividing the total by 2 instead of by the 3 matches played.
- (c) 33 — Method: multiply the mean by the number of tests to get the total marks, then subtract the marks that are already known. Working: four tests with a mean of 29 give a total of 29 × 4 = 116 marks; the first three marks total 31 + 26 + 26 = 83; so the fourth mark is 116 − 83 = 33. Answer: 33, and checking, (31 + 26 + 26 + 33) ÷ 4 = 116 ÷ 4 = 29. The distractors: 116 comes from stopping at the total for all four tests; 29 comes from assuming the missing mark must be the mean itself; 4 comes from multiplying the mean by 3, the number of marks given, leaving 87 − 83 = 4.
- (a) 23 minutes — Method: the equation of a line of best fit converts a value of x into a predicted value of y, so substitute the known number of pages for x and evaluate. Working: x is the number of pages, so put x = 10 into y = 2x + 3. Multiplication is carried out before addition, so 2 × 10 + 3 = 23. Answer: 23 minutes, and it is a prediction of the trend rather than a promise about any one chapter. The distractors: 20 minutes comes from working out 2 × 10 and stopping there, leaving out the 3 that the line adds; 26 minutes comes from reading the equation as y = 2(x + 3), adding first and then doubling, so 2 × 13 = 26; 13 minutes comes from adding 10 and 3 and never using the gradient at all, which treats the 2 as though it were not there.
- (d) Testing destroys bulbs, so testing all leaves none to sell. — Method: testing every item in a population instead of a sample is a census — sensible only when testing does not use up or destroy what is being tested. Working: here, testing a bulb to find its lifespan destroys it, so testing all 50,000 bulbs would leave nothing left to sell — a sample lets the company estimate the typical lifespan without destroying its whole stock. Extra electricity used in testing is not the real reason a census is avoided here — it is the destruction of the product that matters. Saying a sample is always more accurate than a full census is the wrong way round: a census, if it could be carried out, gives the exact figure for the whole population — it is testing being destructive, not a lack of accuracy, that rules it out here. There is no law against testing every item a company makes — nothing in the question suggests that. When testing destroys the item being tested, sampling is necessary, not just convenient.
- (c) The mean, because every value counts towards it, so 100 pulls it from 7 up to about 20.3. — Method: work each measure out before the extra value is added and again afterwards, then compare the size of the two changes. Working: before, the six values total 42, so the mean is 42 ÷ 6 = 7, and the middle pair 6 and 8 give a median of (6 + 8) ÷ 2 = 7; after, the seven values total 142, so the mean is 142 ÷ 7 = 20.29 to 2 decimal places, while the median is now the 4th of the seven ordered values, which is 8; the mean has moved by about 13.3 and the median by 1. Answer: the mean, because every value counts towards it, so 100 pulls it from 7 up to about 20.3 — this is why the median is often preferred when a data set contains an outlier. The distractors: the reply that the mean rises by 100 adds the extra value to the mean instead of adding it to the total; the reply that the median moves to 12 takes the largest of the original values as the new middle instead of counting to the 4th of the seven values; the reply about even and odd counts quotes a rule that does not exist, since the median moved because a very large value was added, not because the count of values changed.
- (d) Route 1, as its times vary by 6 minutes rather than 20 — Method: work out an average and a measure of spread for each route, then decide which matters to a commuter who must arrive on time every day. Working: for Route 1, 22 + 23 + 24 + 24 + 25 + 25 + 26 + 26 + 27 + 28 = 250 and 250 ÷ 10 = 25, so the mean is 25 minutes, and the range is 28 − 22 = 6 minutes. For Route 2, 18 + 19 + 20 + 20 + 21 + 22 + 26 + 30 + 36 + 38 = 250 and 250 ÷ 10 = 25, so the mean is also 25 minutes, but the range is 38 − 18 = 20 minutes. The means give no reason to prefer either route; the spreads do, because a commuter who must never be late has to allow for the worst day, which is 28 minutes on Route 1 and 38 minutes on Route 2. Answer: Route 1, as its times vary by 6 minutes rather than 20. The distractors: saying Route 2 has the lower mean assumes that its quicker-looking early times must pull the average down, when both routes total 250 minutes over the ten days; choosing Route 2 for its fastest journey of 18 minutes judges a route by its best day, and the commuter has to survive its worst; saying either route will do uses the equal means and ignores the spread altogether, which is the one thing that separates the two routes.
- (d) 23.57 °C — Method: add all seven temperatures, divide by the number of readings and round only at the end. Working: 22 + 24 + 23 + 25 + 26 + 21 + 24 = 165, and 165 ÷ 7 = 23.5714…, which rounds to 23.57 to 2 decimal places. Answer: 23.57 °C. The distractors: 24 °C comes from writing down the mode, the only temperature recorded twice, instead of the mean; 27.5 °C comes from dividing the total by 6 instead of by the 7 days recorded; 5 °C comes from working out the range, 26 − 21, which is a measure of spread and not an average.
- (d) 0 — Method: the range is the largest value minus the smallest value, whatever those two values turn out to be. Working: every value is 10, so the largest value is 10 and the smallest value is 10 as well, and the range is 10 − 10 = 0. Answer: 0 — a range of nothing says the data do not vary at all. The distractors: 10 comes from writing down the repeated value itself instead of the difference between the extremes; 20 comes from adding the largest and the smallest, 10 + 10, instead of subtracting; 40 comes from adding all four values, which gives the total sold and not a measure of spread.
- (a) 30 — Method: an outlier is a value that lies far away from the pattern set by the rest of the data, so compare each value with the group the others form. Working: five of the counts, 5, 7, 8, 9 and 11, lie within 6 of one another and the steps between them are 2, 1, 1 and 2; the remaining count of 30 is 19 above the nearest of them, so it is the value that does not belong to the pattern. Answer: 30. The distractors: 5 comes from picking the smallest value, on the idea that the odd one out must be at the bottom of the list; 11 comes from ordering the data and stopping one value short, taking the largest of the counts that sit close together; 8.5 comes from working out the median, (8 + 9) ÷ 2, and giving a measure of centre where a value standing apart was asked for.
Build your own mix at the worksheet builder.