Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Non-calculator
Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- 1.Four pairs of variables are listed below. Write down the pair that you would expect to show no correlation.
- 2.Harry counted the coins he found on each of six days: 5, 7, 8, 9, 11, 30. Write down the outlier.
- 3.Two classes sat the same test. The 30 pupils in Class A had a mean mark of 72. The 20 pupils in Class B had a mean mark of 82. Work out the mean mark of all 50 pupils.
- 4.A garden centre records its total sales, in pounds, at the end of each of the twelve months of one year, to see how sales change over time. Write down the most suitable type of chart to show this time series data.
- 5.A scatter graph plots the number of years, x, that 20 employees have worked at a company against their salary, y. All the plotted points lie between x = 1 and x = 15. Write down the word used to describe an estimate for y made using a value of x that lies between 1 and 15.
- 6.A vertical line chart shows the number of goals scored by a football team in each of its 20 matches: 0 goals in 4 matches, 1 goal in 7 matches, 2 goals in 6 matches and 3 goals in 3 matches. Write down the modal number of goals.
- 7.A table shows the time each of four pupils took to run 100 m: Oliver 12 seconds, Grace 15 seconds, Ethan 10 seconds, Freya 14 seconds. Write down the name of the fastest runner and the time taken.
- 8.The masses of eight school bags, in kilograms, are 3, 4, 4, 5, 6, 7, 8 and 11. Work out the median mass.
- 9.A bar chart shows the highest temperature recorded on each of five days. The temperatures were Monday 20 °C, Tuesday 22 °C, Wednesday 18 °C, Thursday 25 °C and Friday 23 °C. Write down the day on which the temperature was highest.
- 10.In a scatter graph of the age, in years, and the wingspan, in cm, of 20 birds of the same species, all the points lie close to a rising line of best fit except one, which lies a long way below the line. That bird was later found to have a damaged wing. Give a reason why this point should not be used when drawing the line of best fit.
- 11.A shop sold seven pairs of shoes in these sizes: 4, 4, 5, 6, 6, 6, 9. Write down the modal size.
- 12.A pictogram shows the number of cars sold by a garage in Plymouth each month, where each whole car symbol represents 8 cars sold and a half symbol represents 4 cars. March shows 3 whole symbols and one half symbol. Work out how many cars were sold in March.
- 13.The mean of three numbers is 50. A fourth number, 100, is added to the set. Work out the mean of the four numbers.
- 14.A bus company runs two routes into the centre of Exeter. On ten weekdays the journey time on Route 1 was, in minutes: 22, 23, 24, 24, 25, 25, 26, 26, 27 and 28. On Route 2 it was: 18, 19, 20, 20, 21, 22, 26, 30, 36 and 38. A commuter must reach the centre on time every day. Work out the mean and the range for each route, and write down which route she should take.
- 15.Write down what a scatter graph is used to show.
Answer key
- (b) A person's shoe size and their favourite colour — A person's shoe size is not linked to which colour they prefer, so these two show no correlation. The other three pairs are all genuinely correlated: distance travelled and fuel used rise together, which is positive correlation; hours of revision and test score generally rise together, which is also positive correlation; and as outdoor temperature rises, fewer woolly hats are sold, which is negative correlation. Negative correlation is still a real relationship between two variables — it is not the same thing as no relationship at all, so the temperature and hats pair is not the answer to this question.
- (a) 30 — Method: an outlier is a value that lies far away from the pattern set by the rest of the data, so compare each value with the group the others form. Working: five of the counts, 5, 7, 8, 9 and 11, lie within 6 of one another and the steps between them are 2, 1, 1 and 2; the remaining count of 30 is 19 above the nearest of them, so it is the value that does not belong to the pattern. Answer: 30. The distractors: 5 comes from picking the smallest value, on the idea that the odd one out must be at the bottom of the list; 11 comes from ordering the data and stopping one value short, taking the largest of the counts that sit close together; 8.5 comes from working out the median, (8 + 9) ÷ 2, and giving a measure of centre where a value standing apart was asked for.
- (c) 76 marks — Method: a mean of means only works when the groups are the same size, so rebuild each class's total mark, add the totals and divide by the number of pupils altogether. Working: Class A scored 30 × 72 = 2160 marks and Class B scored 20 × 82 = 1640 marks, giving 2160 + 1640 = 3800 marks between 50 pupils, so the overall mean is 3800 ÷ 50 = 76 marks. Answer: 76 marks. The distractors: 77 marks comes from averaging the two class means, (72 + 82) ÷ 2, which ignores the different class sizes; 78 marks comes from attaching each mean to the other class's size, (30 × 82 + 20 × 72) ÷ 50; 3800 marks comes from stopping at the combined total and never dividing by 50.
- (b) A line graph — Sales recorded at the end of each of the twelve months are time series data, and a line graph is the chart built to show how a value changes over time, with the points usually joined in order. A pie chart is for showing categorical data as shares of a whole, not a trend over time. A pictogram shows a frequency for separate categories using symbols, not a continuous trend. A scatter graph is for comparing two different variables against each other, not one variable over time.
- (b) Interpolation — The salary is being estimated for a value of x between 1 and 15, which is inside the range of x-values that were actually plotted, so this is interpolation. Extrapolation would apply if the estimate used a value of x below 1 or above 15, outside the plotted range. Correlation describes the relationship between the two variables, not the reliability of an estimate, and causation describes one variable actually causing a change in the other, which is a different idea altogether — neither is the word being asked for here.
- (b) 1 — The four frequencies are 4, 7, 6 and 3 matches, and the largest of these is 7, which corresponds to 1 goal, so the modal number of goals is 1. Choosing 7 confuses the frequency, how many matches, with the number of goals itself. Choosing 2 uses the second-largest frequency, 6 matches, instead of the largest. Choosing 3 uses the smallest frequency, which belongs to the fewest matches, not the most.
- (b) Ethan, 10 seconds — Method: over the same distance the fastest runner is the one who takes the least time, so the smallest time in the table is found first and the name is then read from the same row. Working: the four times are 12 seconds, 15 seconds, 10 seconds and 14 seconds; in order of size these are 10, 12, 14 and 15, so the least time is 10 seconds, and the row holding 10 seconds is the row for Ethan. Answer: Ethan, 10 seconds — the time is in seconds, and a smaller time means a faster runner. The distractors: Grace with 15 seconds comes from taking the largest number in the table to mean the fastest runner, which reverses the relationship between time and speed over a fixed distance; Oliver with 12 seconds comes from writing down the first row of the table without comparing the four times; Ethan with 15 seconds comes from identifying the right runner but then reading the time from a different row of the table.
- (b) 5.5 kg — Method: with an even number of values the median is the mean of the two middle values, taken once the data are in order of size. Working: the eight masses are already in order and 8 ÷ 2 = 4, so the middle pair are the 4th and 5th values, 5 kg and 6 kg; the median is (5 + 6) ÷ 2 = 5.5 kg. Answer: 5.5 kg. The distractors: 5 kg comes from reading the 4th value and stopping there instead of averaging the middle pair; 8 kg comes from working out the range, 11 − 3, which measures spread rather than centre; 4 kg comes from writing down the modal mass, the only value that occurs twice, instead of the median.
- (c) Thursday — Method: on a bar chart the tallest bar belongs to the greatest value, so compare the five temperatures and then read off the day that the greatest one belongs to. Working: the temperatures are 20 °C, 22 °C, 18 °C, 25 °C and 23 °C; in order of size these are 18, 20, 22, 23 and 25, so the greatest temperature is 25 °C, and the day recorded with 25 °C is Thursday. Answer: Thursday — the answer is a day, not a temperature. The distractors: Friday comes from stopping at the final temperature listed instead of comparing all five; Wednesday comes from picking out the shortest bar, 18 °C, and so answering for the lowest temperature rather than the highest; Monday comes from writing down the first day in the chart without comparing any of the temperatures at all.
- (d) An outlier from the damaged wing, not the trend. — That bird's point lies a long way from the rising trend followed by every other bird, and its low wingspan is explained by the damaged wing rather than by its age — it is an outlier caused by an unusual factor, not part of the general relationship between age and wingspan, so it should not be used when drawing the line of best fit. Saying every point must be used ignores that an outlier caused by a separate, identifiable factor can rightly be set aside. Saying it shows no correlation ignores that the other 19 points do show a clear rising trend; one outlier does not remove that. Saying it proves the line is inaccurate confuses one unusual bird with a fault in the line itself, when the line correctly describes the trend followed by the rest of the data.
- (d) 6 — Method: the mode, or modal value, is the value that occurs most often in the data set, and it is a value from the data rather than a count. Working: size 4 occurs twice, size 5 occurs once, size 6 occurs three times and size 9 occurs once, so the highest frequency is three and the size it belongs to is 6. Answer: 6. The distractors: 3 comes from writing down the frequency of the most common size instead of the size itself; 9 comes from picking the largest size in the list, which confuses the mode with the maximum; 4 comes from stopping at the first size that repeats rather than checking which size repeats most often.
- (c) 28 — Method: multiply the number of whole symbols by the value of one symbol, then add the value of any half symbol shown. Working: 3 whole symbols represent 3 × 8 = 24 cars. The half symbol represents 4 cars. Total cars sold in March = 24 + 4 = 28. Leaving out the half symbol, 3 × 8 = 24, undercounts by exactly the value of that half symbol. Treating the half symbol as if it were a full symbol, 4 × 8 = 32, overcounts because it doubles the value the half symbol is worth. Giving 3.5 reports the number of symbols shown, not the number of cars they represent — the key still needs to be applied. Always apply the key to every symbol shown, including a half symbol, rather than reading off the symbol count itself.
- (a) 62.5 — Method: a mean cannot be averaged with a new value — rebuild the total, add the new value to it, then divide by the new count. Working: three numbers with a mean of 50 have a total of 50 × 3 = 150; adding 100 makes the total 150 + 100 = 250; there are now 4 numbers, so the new mean is 250 ÷ 4 = 62.5. Answer: 62.5. The distractors: 75 comes from averaging the old mean with the new value, (50 + 100) ÷ 2, which ignores that three numbers pull against one; 50 comes from assuming an extra value leaves the mean unchanged; 37.5 comes from dividing the old total of 150 by the new count of 4, adding the new value to the count but not to the total.
- (d) Route 1, as its times vary by 6 minutes rather than 20 — Method: work out an average and a measure of spread for each route, then decide which matters to a commuter who must arrive on time every day. Working: for Route 1, 22 + 23 + 24 + 24 + 25 + 25 + 26 + 26 + 27 + 28 = 250 and 250 ÷ 10 = 25, so the mean is 25 minutes, and the range is 28 − 22 = 6 minutes. For Route 2, 18 + 19 + 20 + 20 + 21 + 22 + 26 + 30 + 36 + 38 = 250 and 250 ÷ 10 = 25, so the mean is also 25 minutes, but the range is 38 − 18 = 20 minutes. The means give no reason to prefer either route; the spreads do, because a commuter who must never be late has to allow for the worst day, which is 28 minutes on Route 1 and 38 minutes on Route 2. Answer: Route 1, as its times vary by 6 minutes rather than 20. The distractors: saying Route 2 has the lower mean assumes that its quicker-looking early times must pull the average down, when both routes total 250 minutes over the ten days; choosing Route 2 for its fastest journey of 18 minutes judges a route by its best day, and the commuter has to survive its worst; saying either route will do uses the equal means and ignores the spread altogether, which is the one thing that separates the two routes.
- (a) The relationship between two variables — Method: what a diagram shows is decided by what has to be known before a single mark can be plotted on it. Working: every point on a scatter graph is plotted from a pair of measurements taken from the same person or object, one read on the horizontal axis and one on the vertical axis; having two measurements for each point is what makes it possible to look for a pattern between them, and the pattern between two variables is what the graph displays. Answer: a scatter graph shows the relationship between two variables. The distractors: the frequency of each single value is what a bar chart or a vertical line chart shows, and it needs only one list of values; how a total is shared between categories is what a pie chart shows; how one quantity changes over time is what a time series line graph shows, in which one of the two axes is always time.
Build your own mix at the worksheet builder.