Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Non-calculator
Answer key: Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- (d) An outlier from the damaged wing, not the trend. — That bird's point lies a long way from the rising trend followed by every other bird, and its low wingspan is explained by the damaged wing rather than by its age — it is an outlier caused by an unusual factor, not part of the general relationship between age and wingspan, so it should not be used when drawing the line of best fit. Saying every point must be used ignores that an outlier caused by a separate, identifiable factor can rightly be set aside. Saying it shows no correlation ignores that the other 19 points do show a clear rising trend; one outlier does not remove that. Saying it proves the line is inaccurate confuses one unusual bird with a fault in the line itself, when the line correctly describes the trend followed by the rest of the data.
- (b) 5.5 kg — Method: with an even number of values the median is the mean of the two middle values, taken once the data are in order of size. Working: the eight masses are already in order and 8 ÷ 2 = 4, so the middle pair are the 4th and 5th values, 5 kg and 6 kg; the median is (5 + 6) ÷ 2 = 5.5 kg. Answer: 5.5 kg. The distractors: 5 kg comes from reading the 4th value and stopping there instead of averaging the middle pair; 8 kg comes from working out the range, 11 − 3, which measures spread rather than centre; 4 kg comes from writing down the modal mass, the only value that occurs twice, instead of the median.
- (d) Only internet users reach the website; others are excluded. — Method: a sample is biased when it systematically leaves out part of the population, or systematically over-represents another part. Working: anyone without internet access, or who does not visit the council's website, has NO chance of being included — the sample is drawn only from internet-using residents, which is not the whole town. Saying too many people might respond because the survey is free confuses bias with sample size — bias is about who CAN be reached, not how many respond. Saying people might lie describes a different problem, response honesty, not who was sampled in the first place. Saying online surveys cannot be anonymous is not a reason connected to bias at all. A sample is biased when part of the population has no chance of being included, whatever the reason for that.
- (c) 7 — Method: the median is the middle value when the data are written in order of size, and with an odd number of values there is exactly one middle value. Working: the numbers are already in order, 2, 4, 7, 12, 26, and there are 5 of them, so the middle position is the third and the value sitting there is 7. Answer: 7, with two values below it and two above it. The distractors: 10.2 comes from working out the mean, 51 ÷ 5, instead of the median; 14 comes from taking the value halfway between the smallest and the largest, (2 + 26) ÷ 2; 24 comes from working out the range, 26 − 2, which measures spread rather than centre.
- (d) 0 — Method: the range is the largest value minus the smallest value, whatever those two values turn out to be. Working: every value is 10, so the largest value is 10 and the smallest value is 10 as well, and the range is 10 − 10 = 0. Answer: 0 — a range of nothing says the data do not vary at all. The distractors: 10 comes from writing down the repeated value itself instead of the difference between the extremes; 20 comes from adding the largest and the smallest, 10 + 10, instead of subtracting; 40 comes from adding all four values, which gives the total sold and not a measure of spread.
- (c) mean = 8, median = 8, mode = 8 — Method: work out each measure separately — the mean is the total divided by how many values there are, the median is the middle value once the data are in order, and the mode is the value that occurs most often. Working: the total is 8 + 8 + 8 + 8 = 32 and there are 4 marks, so the mean is 32 ÷ 4 = 8; in order the marks read 8, 8, 8, 8, and the mean of the middle pair is (8 + 8) ÷ 2 = 8; the value 8 occurs 4 times and no other value occurs at all, so the mode is 8. Answer: mean = 8, median = 8, mode = 8 — when every value in a data set is the same, all three measures of central tendency take that value. The distractors: a mean of 32 comes from stopping at the total and never dividing by 4; a mode of 4 comes from writing down how many times 8 occurs instead of the value that occurs; a mean of 2 comes from dividing a single value, 8, by the 4 marks instead of dividing the total by 4.
- (c) 72 — Method: when a pie chart is divided into equal sectors, each sector stands for the same share of the people asked, so the fraction of the sectors that are shaded is also the fraction of the people. Working: 6 sectors out of 10 are purple, which is the fraction 6/10 of the whole pie chart; one tenth of the 120 people is 120 ÷ 10 = 12 people, so six tenths is 6 × 12 = 72 people. Answer: 72 people, a count of people rather than a number of sectors. The distractors: 48 comes from working with the 4 sectors that are not purple, 4 × 12, and so answering for the wrong part of the chart; 60 comes from turning the fraction 6/10 into 60% and then writing the 60 down as though it were a number of people; 6 comes from writing down the number of purple sectors instead of the number of people those sectors stand for.
- (d) 44 — Method: the frequency of a category in a frequency table is the number of times that category was counted, and it is read from the row for that category. Working: the rows of the table pair each colour with its count, and the row for black is paired with the count 44, so the frequency of black cars is 44. Answer: 44 cars — a frequency is a count of cars, not a colour and not a percentage. The distractors: 37 comes from reading the count paired with silver, that is from reading the wrong row of the table; 137 comes from adding every count in the table, 44 + 37 + 26 + 18 + 12, which gives the total number of cars rather than the frequency of one colour; 5 comes from counting how many different colours the table lists instead of how many cars were black.
- (a) x = 0 gives y = −20: a negative number sold — The y-intercept is the value the line predicts when x = 0: y = 3 × 0 − 20 = −20. A kiosk cannot sell a negative number of ice creams, so this is not a sensible estimate. The 3 in the equation is the gradient, not the intercept, so an option claiming x = 0 gives y = 3 has swapped the two numbers around — substituting x = 0 makes the 3x term equal 0, leaving −20, not 3. The danger of extrapolating to very high temperatures is a real issue with this line, but it is a different issue from the y-intercept, so it does not answer this question. And whether x = 0 could occur on a trading day is beside the point: the model still makes that prediction, and it is the prediction itself, −20, that is impossible.
- (b) No — in order the numbers are 1, 3, 5, 7, 9, so the median is 5. — Method: the median is the middle value of the data in order of size, so the data must be sorted before any position is read off. Working: Noah's list 9, 3, 7, 1, 5 is not in order; sorted it becomes 1, 3, 5, 7, 9, and with 5 values the middle position is the third, which now holds 5 rather than 7. Noah has read the third value of the unsorted list. Answer: no — in order the numbers are 1, 3, 5, 7, 9, so the median is 5. The distractors: the reply giving 3 as the median sorts the data correctly but then reads the value in the second place instead of the third; the reply that 7 is the third number he wrote accepts a position in the unsorted list, which is exactly the mistake the question is about; the reply using the mean claims a value of 7 for it, but the mean is 25 ÷ 5 = 5, so that reasoning is false as well.
- (c) The median, £160,000, as one very high price lifts the mean — Method: find both averages, then choose the one that sits closer to the bulk of the data. Working: in order the prices are 140,000, 150,000, 160,000, 170,000 and 580,000, so the median is the third of the five, £160,000. For the mean, 140,000 + 150,000 + 160,000 + 170,000 + 580,000 = 1,200,000 and 1,200,000 ÷ 5 = 240,000, so the mean is £240,000. Four of the five houses sold for £170,000 or less, so a reader told that a typical price is £240,000 would expect to pay at least £70,000 more than any of those four cost. Answer: the median, £160,000, as one very high price lifts the mean. The distractors: £580,000 is the middle value of the list as it is printed, which is the median only when the values have first been put in order; £240,000 is the mean, chosen on the ground that a median ignores three of the five prices, but a median uses all five to find which one is central and is then untroubled by how extreme the outer values are; £155,000 comes from deleting the £580,000 house and taking the mean of what is left, since 140,000 + 150,000 + 160,000 + 170,000 = 620,000 and 620,000 ÷ 4 = 155,000, but a real sale may not be thrown away merely for being large.
- (c) 33 — Method: multiply the mean by the number of tests to get the total marks, then subtract the marks that are already known. Working: four tests with a mean of 29 give a total of 29 × 4 = 116 marks; the first three marks total 31 + 26 + 26 = 83; so the fourth mark is 116 − 83 = 33. Answer: 33, and checking, (31 + 26 + 26 + 33) ÷ 4 = 116 ÷ 4 = 29. The distractors: 116 comes from stopping at the total for all four tests; 29 comes from assuming the missing mark must be the mean itself; 4 comes from multiplying the mean by 3, the number of marks given, leaving 87 − 83 = 4.
- (b) Ethan, 10 seconds — Method: over the same distance the fastest runner is the one who takes the least time, so the smallest time in the table is found first and the name is then read from the same row. Working: the four times are 12 seconds, 15 seconds, 10 seconds and 14 seconds; in order of size these are 10, 12, 14 and 15, so the least time is 10 seconds, and the row holding 10 seconds is the row for Ethan. Answer: Ethan, 10 seconds — the time is in seconds, and a smaller time means a faster runner. The distractors: Grace with 15 seconds comes from taking the largest number in the table to mean the fastest runner, which reverses the relationship between time and speed over a fixed distance; Oliver with 12 seconds comes from writing down the first row of the table without comparing the four times; Ethan with 15 seconds comes from identifying the right runner but then reading the time from a different row of the table.
- (d) How much the temperatures varied over the seven days — Method: the range of a set of values is the largest value take away the smallest, so it is built from two values only and it measures the gap they leave between them. Working: the largest of the seven readings is 7 °C and the smallest is 2 °C, so the range is 7 − 2 = 5 °C. That figure says the week's readings covered a band 5 °C wide; it names no particular day and no particular reading. Answer: the range describes how much the temperatures varied over the seven days. The distractors: the temperature that occurred most often is the mode, which here is 3 °C, and a mode counts repeats instead of measuring a gap; the temperature typical of the week is an average, and the range is not an average, since it throws away every value lying between the two extremes; the number of different temperatures recorded is 6, a count of how many distinct values appear, while the range is a difference between two of them.
- (c) 20 kg ≤ mass < 30 kg — The modal class is the class with the highest frequency. Reading the plotted points, the frequencies are 6, 10, 16, 6 and 2, so the highest frequency is 16, plotted at the midpoint 25. A class of width 10 centred on 25 runs from 25 − 5 = 20 to 25 + 5 = 30, so the modal class is 20 kg ≤ mass < 30 kg. Writing '25 kg' gives only the midpoint, not the class — the modal class is an interval, not a single value. '10 kg ≤ mass < 20 kg' is the class before the peak, centred on 15, which has frequency 10, not the highest. '30 kg ≤ mass < 40 kg' is the class after the peak, centred on 35, which has frequency 6, not the highest.
Build your own mix at the worksheet builder.