Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Answer key: Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- (b) A line graph — Sales recorded at the end of each of the twelve months are time series data, and a line graph is the chart built to show how a value changes over time, with the points usually joined in order. A pie chart is for showing categorical data as shares of a whole, not a trend over time. A pictogram shows a frequency for separate categories using symbols, not a continuous trend. A scatter graph is for comparing two different variables against each other, not one variable over time.
- (b) 1 — Method: for data in a frequency table, find the position of the median using (n + 1) ÷ 2, then read off the value at that position from the cumulative frequencies. Working: there are 19 pupils, so the median is the 10th value. The cumulative frequencies are 7 (up to 0 pets), 10 (up to 1 pet), 14 (up to 2 pets) and 19 (up to 3 pets). The 10th value falls at the end of the '1 pet' group, so the median is 1 pet. Giving 0 pets is the mode — the category with the highest frequency, 7 — not the median. Giving 3, the highest number of pets minus the lowest, finds the range, a different statistic entirely. Giving 19 states the total number of pupils, not a number of pets at all. Find the middle POSITION first, then read off the value it belongs to — do not confuse it with the mode, the range or the total.
- (a) 26 — Method: a line of best fit lets one quantity be predicted from the other, so the score is substituted into the equation of the line and the resulting inequality is solved for the number of hours. Working: a score of at least 80 means 2.5x + 15 ≥ 80; taking 15 from both sides gives 2.5x ≥ 65, and dividing both sides by 2.5 gives x ≥ 26, so the least whole number of hours is 26. Checking, 2.5 × 26 + 15 = 80, which does reach the target. Answer: 26 hours — and this is only an estimate, because a line of best fit predicts a trend rather than an individual result, and a prediction made outside the range of hours the pupils actually revised for would be an extrapolation and less reliable still. The distractors: 27 comes from reaching 26 and then rounding up again, although 26 hours already gives a score of exactly 80; 32 comes from 80 ÷ 2.5, which ignores the 15 in the equation of the line; 38 comes from (80 + 15) ÷ 2.5, that is from adding the 15 instead of subtracting it when rearranging.
- (b) 1 — The four frequencies are 4, 7, 6 and 3 matches, and the largest of these is 7, which corresponds to 1 goal, so the modal number of goals is 1. Choosing 7 confuses the frequency, how many matches, with the number of goals itself. Choosing 2 uses the second-largest frequency, 6 matches, instead of the largest. Choosing 3 uses the smallest frequency, which belongs to the fewest matches, not the most.
- (c) 67 — Method: the mean is the total of the values divided by how many values there are, so add first and divide second. Working: the total is 50 + 83 + 68 = 201 points and three matches were played, so the mean is 201 ÷ 3 = 67 points. Answer: 67. The distractors: 68 comes from writing down the median, the middle value of 50, 68, 83, instead of the mean; 33 comes from working out the range, 83 − 50, which measures spread and not centre; 100.5 comes from dividing the total by 2 instead of by the 3 matches played.
- (a) Priya — A line of best fit should be drawn so that the plotted points are roughly balanced above and below it. Work out the difference between the two counts for each pupil: Amir 10 − 2 = 8; Kofi 10 − 2 = 8 (10 below and 2 above); Leah 12 − 0 = 12; Priya 6 − 6 = 0. Priya's line has the smallest difference, an exact balance of 6 above and 6 below, so her line is drawn correctly. Amir's line has 10 of the 12 points above it, so it is drawn too low. Kofi's line has 10 of the 12 points below it, so it is drawn too high. Leah's line has every single point below it, so it is not a line of best fit at all.
- (a) The median, as the one very large wage does not move it — Method: an average describes a population well when it sits close to most of the values, so compare what each average does when one value lies far from the rest. Working: in order the wages are 420, 440, 460, 480 and 1,500, so the median is the third of the five, £460. The mean uses every wage: 420 + 440 + 460 + 480 + 1,500 = 3,300 and 3,300 ÷ 5 = 660, so the mean is £660. Four of the five people earn less than £660, and the nearest of those four wages is £180 below it, so £660 describes nobody at the garage; £460 sits inside the group of four similar wages. Answer: the median, as the one very large wage does not move it, while that same wage drags the mean £200 above the median. The distractors: saying the median is always larger than the mean is an invented rule, and here the median £460 is smaller than the mean £660; saying the mean is the only average that uses all five wages is true as far as it goes, but using a value and being dragged by it are the same thing when that value is £1,500; saying £660 lies between the smallest and largest wage is true of every mean ever calculated, so it proves nothing about whether this one is typical.
- (b) Interpolation — The salary is being estimated for a value of x between 1 and 15, which is inside the range of x-values that were actually plotted, so this is interpolation. Extrapolation would apply if the estimate used a value of x below 1 or above 15, outside the plotted range. Correlation describes the relationship between the two variables, not the reliability of an estimate, and causation describes one variable actually causing a change in the other, which is a different idea altogether — neither is the word being asked for here.
- (a) 15 — Year 11 has 50 − 28 = 22 pupils in total. Of the 22 pupils who walk in total, 15 are in Year 10, so 22 − 15 = 7 Year 11 pupils walk. Subtracting that from the Year 11 total gives 22 − 7 = 15 Year 11 pupils who are driven. Choosing 28 takes the whole school's driven total, 50 − 22 = 28, and treats it as if it were Year 11's alone, without separating the year groups. Choosing 7 correctly finds how many Year 11 pupils walk but stops there, giving that figure instead of the number who are driven. Choosing 35 comes from 50 − 15, subtracting the Year 10 walkers from the whole school total rather than working within Year 11.
- (d) 29 — Method: with an even number of values there is no single middle value, so the median is the mean of the two values either side of the middle. Working: the six numbers are already in order and 6 ÷ 2 = 3, so the middle pair are the third and fourth values, 22 and 36; their mean is (22 + 36) ÷ 2 = 58 ÷ 2 = 29. Answer: 29, which lies between the two middle values as a median of an even data set must. The distractors: 22 comes from taking the lower of the two middle values and stopping there instead of averaging the pair; 36 comes from taking the larger value of that pair because it sits just past the halfway point of the list; 34 comes from working out the range, 47 − 13, instead of a measure of centre.
- (b) 120, unreliable — x = 25 is outside 1 to 12 — The line of best fit is y = 4x + 20. 4 × 25 = 100, and 100 + 20 = 120, so the estimate is y = 120. But x = 25 lies far outside the plotted range of 1 to 12 years, so this is an extrapolation, and the estimate is not reliable. Reaching 100 instead of 120 comes from 4 × 25 = 100 with the intercept of 20 left out — still correctly flagged as unreliable, but the wrong value. Calling the estimate reliable simply because it was calculated correctly, giving 120, wrongly assumes that a correct calculation is automatically trustworthy, ignoring that x = 25 lies far beyond the data actually collected. Reaching 68, from 4 × 12 = 48 and 48 + 20 = 68, substitutes x = 12, the top of the plotted range, instead of the assistant's actual x = 25, and wrongly calls that reliable because 12 lies inside the range.
- (a) A vertical line chart (discrete numerical data) — The number of pets is discrete numerical data — whole-number values such as 0, 1, 2, 3 or 4 — recorded for one variable, so a vertical line chart is the chart specified for this kind of data. A bar chart is used for categorical data, such as favourite colour, not numerical values counted like this. A pie chart shows proportions of a whole and does not show the frequency of each separate value. A scatter graph compares two different variables against each other, and only one variable, the number of pets, is recorded here.
- (b) 74 marks — Method: the two groups are different sizes, so their means cannot simply be averaged — rebuild each group's total mark, add the totals and divide by all 50 pupils. Working: Group A scored 20 × 80 = 1600 marks and Group B scored 30 × 70 = 2100 marks, giving 1600 + 2100 = 3700 marks altogether, so the overall mean is 3700 ÷ 50 = 74 marks. Answer: 74 marks, which sits nearer to 70 than to 80 because the larger group scored 70. The distractors: 75 marks comes from averaging the two group means, (80 + 70) ÷ 2, as though the groups were the same size; 76 marks comes from attaching each mean to the other group's size, (20 × 70 + 30 × 80) ÷ 50; 150 marks comes from adding the two means together and never dividing at all.
- (a) Equal means; Class A is more consistent, smaller range. — Method: when two data sets share a measure of location, compare a measure of spread to say more about consistency. Working: both classes have the same mean mark, 14, so on average they performed equally well. Class A has the smaller range, 6, so its marks are more tightly grouped around 14 than Class B's marks, which vary by as much as 14. So Class A's marks were more consistent, even though neither class did better on average. Saying Class B did better because it has the bigger range confuses a wide spread with a high score — a big range describes variability, not performance. Saying Class A did better because it has the smaller range makes the same mistake in the other direction: the two classes are tied on the mean, so neither one 'did better'. Saying the classes cannot be compared because their means are equal misses the whole point of also comparing the range. Always compare both an average AND a spread before describing two data sets — either one alone tells only half the story.
- (d) Correlation is a link; causation is one causing the other — Method: the two words describe different claims — one is about a pattern in the data, the other is about what produced that pattern. Working: correlation says only that two quantities tend to change together, which is something a scatter graph can display; causation says that a change in one quantity actually brings about the change in the other, which needs evidence a scatter graph cannot supply, because a third quantity may be driving both. Answer: correlation is a link between the quantities, while causation is one quantity causing the change in another. The distractors: the statement giving causation as the link and correlation as the cause simply swaps the two words over; the statement that the words mean the same thing is the classic error of reading a correlation as proof of cause; the statement that a scatter graph shows causation but not correlation reverses what a scatter graph can do, since the pattern it displays is exactly the correlation.
Build your own mix at the worksheet builder.