Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Answer key: Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- (b) Treat them as outliers and check them before deciding — Method: a point lying a long way from the pattern the rest of the data make is called an outlier, and an outlier is investigated before anything is done with it, because it may be an error in the data or it may be a genuine but unusual case. Working: 48 of the 50 points lie close to the rising line of best fit, so the trend is set by those 48; the two remaining points do not follow it, so they are identified as outliers and checked — a mistake in measuring or recording would be corrected, while a genuine reading would be kept and reported. Answer: treat them as outliers and check them before deciding what to do with them. The distractors: deleting them at once assumes that every point far from the line must be an error, which throws away real data; moving the line so that it passes through them assumes a line of best fit must touch particular points, when it is drawn to follow all 50; taking them as proof that there is no correlation lets two points overturn the pattern that the other 48 agree on.
- (c) It is not representative, as she picked her own friends — Method: judge a sample by asking whether the pupils in it were chosen in a way that gives the whole school a fair chance of being heard. Working: Isla's friends are a group she formed herself, and friends tend to share tastes, so their favourite programme is likely to match hers rather than the school's, and pupils in other year groups and other friendship groups had no chance at all of being asked; the fault lies in how the pupils were selected, not in how many of them there were. Answer: it is not representative, as she picked her own friends. The distractors: the reply blaming the size claims the pupils were picked at random, which is false here, and it is the common mistake of thinking a biased sample can be cured by making it bigger; the reply calling the sample too large is false in the other direction, as a survey is never spoilt by collecting more replies; the reply that the sample is fine treats attending the school as enough, which would make any group of pupils in the building a fair sample.
- (b) 29 — There are 40 − 24 = 16 males, and 15 of them prefer cardio, so 16 − 15 = 1 male prefers weights. There are 24 females, and 10 prefer weights, so 24 − 10 = 14 females prefer cardio. Altogether, 15 + 14 = 29 people prefer cardio. Choosing 15 only counts the males who prefer cardio and forgets the females. Choosing 11 adds the two weights figures, 1 + 10 = 11, instead of the two cardio figures. Choosing 30 comes from 40 − 10, subtracting only the number of females who prefer weights from the grand total, rather than finding both cardio sub-totals separately.
- (d) The modal class, as the class with most pupils is shown — Method: a grouped frequency table records how many values fall into each class, but not the values themselves, so any average that needs the individual times can only be estimated from it. Working: the four frequencies are 8, 12, 6 and 4, and 8 + 12 + 6 + 4 = 30, so every pupil is counted. The largest frequency is 12, which belongs to the class 10 < t ≤ 20, and that class can be written down exactly, because finding it needs nothing but the counts the table already gives. Answer: the modal class, as the class with most pupils is shown. The distractors: the mean is said to use all 30 times, but the table does not hold them; the usual method replaces each class by its midpoint, 5, 15, 25 and 35, which gives an estimate of the mean and not its true value; the median is said to be shown, but the table locates only the class holding the 15th and 16th times, which is 10 < t ≤ 20, without saying what either time was; the range is said to be shown, but 0 and 40 are the boundaries of the first and last classes, not the fastest and slowest times actually recorded.
- (c) No — 8 from one class is too small to represent the school. — Method: judge reliability by asking whether the sample is both large enough, and spread across the population, relative to what it is meant to represent. Working: 8 pupils is a tiny fraction of the school's 1,000 pupils, and all 8 come from a single class rather than a range of year groups, so the sample is both too small and too narrow to represent the whole school reliably. She is not right. Saying any sample size gives an equally reliable estimate ignores that reliability generally improves with a larger, more representative sample. Saying the method is unreliable because it was not done online is not a reason connected to sample size or representativeness at all. Saying 8 is reliable because it is more than half her class compares the sample to the wrong population — the school has 1,000 pupils, not one class. Always judge a sample's size against the population it is meant to represent, not against a smaller group within it.
- (a) No, 150 pupils are a tenth of the school, chosen at random — Method: judge a sample on two things, whether every member of the population had the same chance of being chosen, and whether the sample is large enough to carry a pattern. Working: the 150 pupils were drawn from the register of every pupil in the school, so no year group or set is shut out and no pupil chooses to take part; and 150 ÷ 1,500 = 0.1, so one pupil in ten has been asked. A random sample of that share is ample for an estimate of how long the school's pupils spend on homework. Answer: no, because 150 pupils are a tenth of the school and were chosen at random. The distractors: saying a random sample always gives the exact school figure reaches the same verdict for a reason that is false, since a second random sample of 150 would give a slightly different mean; saying 150 pupils cannot be picked at random from 1,500 treats randomness as something only a whole population can have, when drawing names from the register is exactly how a random sample is taken; saying that only asking all 1,500 could show anything rejects sampling altogether, which would leave no way to study any population too large to count.
- (c) 69 — Method: multiply the mean by the number of values to find the total, then subtract the total of the known values. Working: the total of all five scores is 68 × 5 = 340. The total of the four known scores is 55 + 62 + 74 + 80 = 271. The fifth score is 340 − 271 = 69. Subtracting the other way round, 271 − 340 = −69, gives the right size answer with the wrong sign. Guessing that the missing score simply equals the mean, 68, ignores that the four known scores are not themselves centred on 68. Multiplying the mean by 4 instead of 5, 68 × 4 = 272, then 272 − 271 = 1, undercounts how many scores there are. Always multiply the mean by the TOTAL number of values before subtracting.
- (b) On average a plant grew 1.5 cm taller for each extra day — Method: in the equation of a line, the number multiplying x is the gradient, and a gradient states the change in y produced by an increase of 1 in x, read in the units of the two axes. Working: here x is measured in days and y in centimetres, so the gradient 1.5 carries the units centimetres per day. Testing it on the line, 5 days gives 1.5 × 5 + 4 = 11.5 cm and 6 days gives 1.5 × 6 + 4 = 13 cm, a rise of 1.5 cm for the one extra day. Answer: on average a plant grew 1.5 cm taller for each extra day of watering. The distractors: 1.5 cm as the height before any watering is the value of y when x is 0, which is the other number in the equation, 4 cm, so this swaps the gradient and the intercept; 1.5 cm as the gap between the tallest and the shortest plant reads the gradient as a range, when a range is a difference between two of the 16 plants and a gradient is a rate; 1.5 days for each extra centimetre inverts the rate, dividing days by centimetres instead of centimetres by days, and the line gives 1 cm of growth in two thirds of a day.
- (b) £11.00 — 1.5 × 6 = 9, and 9 + 2 = 11, so the estimated cost is £11.00. Choosing £9.00 stops after 1.5 × 6 = 9 and forgets to add the £2. Choosing £12.00 adds the mass and the constant first and then multiplies: 6 + 2 = 8, and 8 × 1.5 = 12.00. Choosing £13.50 swaps the gradient and the intercept, using y = 2x + 1.5 instead: 2 × 6 = 12, and 12 + 1.5 = 13.50.
- (d) £15.50 — 2 × 6 = 12, and 12 + 3.50 = 15.50, so the estimated fare is £15.50. Choosing £12.00 stops after 2 × 6 = 12 and forgets to add the £3.50. Choosing £19.00 adds the distance and the constant first and then multiplies: 6 + 3.50 = 9.50, and 9.50 × 2 = 19.00, applying the ×2 to the whole sum instead of only to the distance. Choosing £13.00 multiplies only the constant term by 2 instead of the distance: 2 × 3.50 = 7, and 7 + 6 = 13.00.
- (d) 1100 — Method: to combine two samples of different sizes, add the faulty counts together and add the sample sizes together before scaling up, rather than treating the two samples separately. Working: the combined sample found 34 + 21 = 55 scratched cases out of 100 + 50 = 150 cases checked, a proportion of 55 ÷ 150. Applying that proportion to the week's production of 3,000 gives an estimate of 55 ÷ 150 × 3000 = 1100 scratched cases. Averaging the two shifts' proportions instead of combining their totals, (34 ÷ 100 + 21 ÷ 50) ÷ 2 = 0.38, gives 0.38 × 3000 = 1140 — this treats the two samples as equally weighted even though Shift A checked twice as many cases as Shift B. Using only Shift A's sample, 34 ÷ 100 × 3000 = 1020, ignores Shift B's cases completely. Using only Shift B's sample, 21 ÷ 50 × 3000 = 1260, ignores Shift A's cases completely. When two samples are different sizes, combine their totals before finding the proportion — do not average the two proportions, and do not use only one shift's sample.
- (d) How much the temperatures varied over the seven days — Method: the range of a set of values is the largest value take away the smallest, so it is built from two values only and it measures the gap they leave between them. Working: the largest of the seven readings is 7 °C and the smallest is 2 °C, so the range is 7 − 2 = 5 °C. That figure says the week's readings covered a band 5 °C wide; it names no particular day and no particular reading. Answer: the range describes how much the temperatures varied over the seven days. The distractors: the temperature that occurred most often is the mode, which here is 3 °C, and a mode counts repeats instead of measuring a gap; the temperature typical of the week is an average, and the range is not an average, since it throws away every value lying between the two extremes; the number of different temperatures recorded is 6, a count of how many distinct values appear, while the range is a difference between two of them.
- (b) £1,000 — Wages take up 150° out of 360°, so the amount spent on wages is 150 ÷ 360 × 2400 = £1,000. Choosing £600 uses the repairs angle, 90°, instead of the wages angle: 90 ÷ 360 × 2400 = 600. Choosing £3,600 treats the angle in degrees as if it were a percentage, 150 ÷ 100 × 2400 = 3600, instead of dividing by 360°. Choosing £800 uses the angle for the 'other costs' sector, 360 − 90 − 150 = 120°, instead of the wages sector: 120 ÷ 360 × 2400 = 800.
- (a) Chloe's marks are far more spread out than Ben's — Method: a mean reports where a set of values sits, and two sets can sit in the same place while behaving quite differently, so a measure of spread has to be worked out as well. Working: Ben's marks add to 62 + 64 + 65 + 66 + 68 = 325 and 325 ÷ 5 = 65; Chloe's add to 40 + 52 + 65 + 78 + 90 = 325 and 325 ÷ 5 = 65, so the two means agree, as the question says. The ranges do not: Ben's is 68 − 62 = 6 marks, while Chloe's is 90 − 40 = 50 marks. Ben's five marks all sit within 3 marks of 65; Chloe's lowest is 25 marks below it and her highest 25 marks above it. Answer: Chloe's marks are far more spread out than Ben's, which is exactly what the mean cannot show. The distractors: saying Ben's marks are more spread out comes from subtracting in the order the values are written, 62 − 68 = −6 against 40 − 90 = −50, and then reading −6 as the larger spread; saying Chloe scored far more marks in total assumes a wider set of marks must add to more, when both totals are 325; saying the two sets vary by the same amount assumes that equal means force equal spread, when the two ranges are 6 and 50.
- (d) An outlier from the damaged wing, not the trend. — That bird's point lies a long way from the rising trend followed by every other bird, and its low wingspan is explained by the damaged wing rather than by its age — it is an outlier caused by an unusual factor, not part of the general relationship between age and wingspan, so it should not be used when drawing the line of best fit. Saying every point must be used ignores that an outlier caused by a separate, identifiable factor can rightly be set aside. Saying it shows no correlation ignores that the other 19 points do show a clear rising trend; one outlier does not remove that. Saying it proves the line is inaccurate confuses one unusual bird with a fault in the line itself, when the line correctly describes the trend followed by the rest of the data.
Build your own mix at the worksheet builder.