Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Non-calculator
Answer key: Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- (b) 20 — Method: each sector of a pie chart stands for a percentage of the whole set of data, so the number of pupils in a sector is that percentage of the total number of pupils. Working: the football sector is 50% and the whole pie chart stands for all 40 pupils, so the number of pupils is 50% of 40 = 0.5 × 40 = 20. Answer: 20 pupils, which is a count of pupils rather than a percentage. The distractors: 10 comes from taking the netball sector, 25% of 40, instead of the football sector; 80 comes from doubling the 40 pupils instead of halving them, which is what multiplying by 50% would do the wrong way round; 40 comes from writing down the total number of pupils in the class and never applying the percentage at all.
- (b) On average a plant grew 1.5 cm taller for each extra day — Method: in the equation of a line, the number multiplying x is the gradient, and a gradient states the change in y produced by an increase of 1 in x, read in the units of the two axes. Working: here x is measured in days and y in centimetres, so the gradient 1.5 carries the units centimetres per day. Testing it on the line, 5 days gives 1.5 × 5 + 4 = 11.5 cm and 6 days gives 1.5 × 6 + 4 = 13 cm, a rise of 1.5 cm for the one extra day. Answer: on average a plant grew 1.5 cm taller for each extra day of watering. The distractors: 1.5 cm as the height before any watering is the value of y when x is 0, which is the other number in the equation, 4 cm, so this swaps the gradient and the intercept; 1.5 cm as the gap between the tallest and the shortest plant reads the gradient as a range, when a range is a difference between two of the 16 plants and a gradient is a rate; 1.5 days for each extra centimetre inverts the rate, dividing days by centimetres instead of centimetres by days, and the line gives 1 cm of growth in two thirds of a day.
- (a) The modal size, 9, bought by more customers than any other — Method: work out both averages from the frequencies, then choose the one the shop can act on. Working: for the mean, multiply each size by the number of pairs sold at it and add: 6 × 4 + 7 × 5 + 8 × 8 + 9 × 13 + 10 × 10 = 340, and 340 ÷ 40 = 8.5, so the mean size is 8.5. The largest frequency is 13, which belongs to size 9, so the modal size is 9. The mean 8.5 is a size no customer in the record asked for, so 40 pairs of it would sit unsold, while 13 of the 40 customers wanted size 9, more than wanted any other size. Answer: the modal size, 9, bought by more customers than any other. The distractors: the mean size 8.5 does take account of all 40 pairs, but a mean of sizes is a summary figure and not a size the month's customers were buying; the mean size 8 comes from averaging the five sizes on sale, 6 + 7 + 8 + 9 + 10 = 40 and 40 ÷ 5 = 8, which ignores how many pairs were sold at each size and so treats the 4 pairs of size 6 as equal in weight to the 13 pairs of size 9; the range 4 comes from 10 − 6 and measures spread, so it says how wide a set of sizes the shop must stock, not which size to stock most of.
- (d) 25 — In order, the nine ages are 14, 18, 20, 23, 25, 29, 31, 36 and 42, and with 9 values the median is the 5th one, which is 25. Choosing 23 takes the 4th value instead of the 5th. Choosing 29 takes the 6th value instead of the 5th. Choosing 26 comes from averaging the 4th and 6th values, 23 + 29 = 52, and 52 ÷ 2 = 26, a method that is only needed when there is an even number of values.
- (a) 26 — Method: a line of best fit lets one quantity be predicted from the other, so the score is substituted into the equation of the line and the resulting inequality is solved for the number of hours. Working: a score of at least 80 means 2.5x + 15 ≥ 80; taking 15 from both sides gives 2.5x ≥ 65, and dividing both sides by 2.5 gives x ≥ 26, so the least whole number of hours is 26. Checking, 2.5 × 26 + 15 = 80, which does reach the target. Answer: 26 hours — and this is only an estimate, because a line of best fit predicts a trend rather than an individual result, and a prediction made outside the range of hours the pupils actually revised for would be an extrapolation and less reliable still. The distractors: 27 comes from reaching 26 and then rounding up again, although 26 hours already gives a score of exactly 80; 32 comes from 80 ÷ 2.5, which ignores the 15 in the equation of the line; 38 comes from (80 + 15) ÷ 2.5, that is from adding the 15 instead of subtracting it when rearranging.
- (b) A person's shoe size and their favourite colour — A person's shoe size is not linked to which colour they prefer, so these two show no correlation. The other three pairs are all genuinely correlated: distance travelled and fuel used rise together, which is positive correlation; hours of revision and test score generally rise together, which is also positive correlation; and as outdoor temperature rises, fewer woolly hats are sold, which is negative correlation. Negative correlation is still a real relationship between two variables — it is not the same thing as no relationship at all, so the temperature and hats pair is not the answer to this question.
- (d) 24 — Method: for a sample in proportion to the population, apply the same fraction that each group makes up of the whole population to the size of the sample. Working: women make up 300 out of the 500 members, a fraction of 300 ÷ 500 = 0.6. Applying that fraction to the sample of 40 gives 0.6 × 40 = 24 women. Splitting the sample evenly, 40 ÷ 2 = 20, ignores that the club has more women than men and treats the two groups as equal in size, which they are not. Misreading the sample size as 50 instead of 40, then applying the 3:2 ratio of women to men, 3 ÷ 5 × 50 = 30, uses the right ratio but the wrong sample total. Working out the number of MEN instead of women, 200 ÷ 500 × 40 = 16, answers a different question — how many men, not how many women, belong in the sample. Always apply each group's own share of the population to the sample size, and check which group the question is actually asking about.
- (c) 5 — Method: the mode is the value that occurs most often, so count how many times each different value appears and compare the counts. Working: 3 appears twice, 5 appears three times, 7 appears once and 8 appears once, so the highest frequency is three and the value carrying it is 5. Answer: 5. The distractors: 3 comes from writing down the frequency of the most common answer instead of the answer itself; 7 comes from taking the middle number of the list as it was written, which applies the median without ordering the data and without answering the question asked; 8 comes from picking the largest value, which confuses the mode with the maximum.
- (a) 62.5 — Method: a mean cannot be averaged with a new value — rebuild the total, add the new value to it, then divide by the new count. Working: three numbers with a mean of 50 have a total of 50 × 3 = 150; adding 100 makes the total 150 + 100 = 250; there are now 4 numbers, so the new mean is 250 ÷ 4 = 62.5. Answer: 62.5. The distractors: 75 comes from averaging the old mean with the new value, (50 + 100) ÷ 2, which ignores that three numbers pull against one; 50 comes from assuming an extra value leaves the mean unchanged; 37.5 comes from dividing the old total of 150 by the new count of 4, adding the new value to the count but not to the total.
- (d) Testing destroys bulbs, so testing all leaves none to sell. — Method: testing every item in a population instead of a sample is a census — sensible only when testing does not use up or destroy what is being tested. Working: here, testing a bulb to find its lifespan destroys it, so testing all 50,000 bulbs would leave nothing left to sell — a sample lets the company estimate the typical lifespan without destroying its whole stock. Extra electricity used in testing is not the real reason a census is avoided here — it is the destruction of the product that matters. Saying a sample is always more accurate than a full census is the wrong way round: a census, if it could be carried out, gives the exact figure for the whole population — it is testing being destructive, not a lack of accuracy, that rules it out here. There is no law against testing every item a company makes — nothing in the question suggests that. When testing destroys the item being tested, sampling is necessary, not just convenient.
- (d) About 510 of the 600 bulbs are likely to have flowered — Method: the proportion found in a random sample is used as an estimate of the proportion in the whole population, and the conclusion is stated as an estimate, never as a fact about every member. Working: 17 of the 20 bulbs dug up had flowered, so the sample proportion is 17 ÷ 20 = 0.85, and applying that proportion to the whole planting gives 0.85 × 600 = 510 bulbs. A different random sample of 20 would very probably give a slightly different figure, so 510 is an estimate. Answer: about 510 of the 600 bulbs are likely to have flowered. The distractors: saying exactly 510 have flowered takes an estimate from a sample of 20 as a count of all 600, which no sample can deliver; saying exactly 17 of the 600 have flowered reports the sample count as though it were the population count, leaving the other 580 bulbs out of the answer altogether; saying about 20 have flowered uses the size of the sample as the estimate, when 20 is the number of bulbs she dug up rather than a number that flowered.
- (a) Negative correlation — As the age of the car increases, the points fall towards a lower value, so the value decreases as the age increases. This falling pattern is a negative correlation. A positive correlation would show the points rising together instead. No correlation would apply only if the points showed no pattern at all, and correlation is not the same as causation — strong causation is not a type of correlation.
- (c) The data show a link only; a third factor may affect both — Method: a study of this kind measures two quantities and reports how they change together; deciding that one of them produces the other is a further claim, and it needs evidence that the measurements alone cannot give. Working: the study shows that more coffee goes with better concentration, which is a positive correlation; but a third factor that was never measured, such as how motivated someone is, could raise both the coffee drinking and the concentration, and the concentration could equally be what leads to the extra coffee. Answer: the data show a link only, because a third factor may be affecting both quantities, so no claim about cause can be made. The distractors: calling the conclusion safe because the correlation is positive treats the direction of a correlation as proof of cause, which no direction can give; calling it wrong because the correlation is negative misreads the direction of the relationship, since the study reports both quantities rising together; saying the two quantities are not linked denies the correlation the study actually found, when what fails is only the claim about cause.
- (a) A vertical line chart (discrete numerical data) — The number of pets is discrete numerical data — whole-number values such as 0, 1, 2, 3 or 4 — recorded for one variable, so a vertical line chart is the chart specified for this kind of data. A bar chart is used for categorical data, such as favourite colour, not numerical values counted like this. A pie chart shows proportions of a whole and does not show the frequency of each separate value. A scatter graph compares two different variables against each other, and only one variable, the number of pets, is recorded here.
- (c) The median, £160,000, as one very high price lifts the mean — Method: find both averages, then choose the one that sits closer to the bulk of the data. Working: in order the prices are 140,000, 150,000, 160,000, 170,000 and 580,000, so the median is the third of the five, £160,000. For the mean, 140,000 + 150,000 + 160,000 + 170,000 + 580,000 = 1,200,000 and 1,200,000 ÷ 5 = 240,000, so the mean is £240,000. Four of the five houses sold for £170,000 or less, so a reader told that a typical price is £240,000 would expect to pay at least £70,000 more than any of those four cost. Answer: the median, £160,000, as one very high price lifts the mean. The distractors: £580,000 is the middle value of the list as it is printed, which is the median only when the values have first been put in order; £240,000 is the mean, chosen on the ground that a median ignores three of the five prices, but a median uses all five to find which one is central and is then untroubled by how extreme the outer values are; £155,000 comes from deleting the £580,000 house and taking the mean of what is left, since 140,000 + 150,000 + 160,000 + 170,000 = 620,000 and 620,000 ÷ 4 = 155,000, but a real sale may not be thrown away merely for being large.
Build your own mix at the worksheet builder.