Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Non-calculator
Answer key: Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- (a) £16,000 — Method: first find the quarter with the highest sales figure, then subtract quarter 1's sales from it — remembering that every figure is given in THOUSANDS of pounds. Working: the highest sales figure is quarter 2, at £34,000 (34 thousand pounds). The increase from quarter 1 is £34,000 − £18,000 = £16,000. Giving £34,000 reads off the highest sales figure on its own, without subtracting quarter 1's sales — that is the highest quarter's total, not the increase. Giving £12,000 uses quarter 3's sales, 30, the SECOND-highest figure, instead of quarter 2's 34, the actual highest — 30 − 18 = 12, but quarter 3 is not the quarter with the highest sales. Giving £16 gets the subtraction right, 34 − 18 = 16, but forgets that every figure in the question is in thousands of pounds, so the increase is £16,000, not £16. Always identify the correct quarter FIRST, and always check the units the numbers are given in before writing your final answer.
- (a) Chloe's marks are far more spread out than Ben's — Method: a mean reports where a set of values sits, and two sets can sit in the same place while behaving quite differently, so a measure of spread has to be worked out as well. Working: Ben's marks add to 62 + 64 + 65 + 66 + 68 = 325 and 325 ÷ 5 = 65; Chloe's add to 40 + 52 + 65 + 78 + 90 = 325 and 325 ÷ 5 = 65, so the two means agree, as the question says. The ranges do not: Ben's is 68 − 62 = 6 marks, while Chloe's is 90 − 40 = 50 marks. Ben's five marks all sit within 3 marks of 65; Chloe's lowest is 25 marks below it and her highest 25 marks above it. Answer: Chloe's marks are far more spread out than Ben's, which is exactly what the mean cannot show. The distractors: saying Ben's marks are more spread out comes from subtracting in the order the values are written, 62 − 68 = −6 against 40 − 90 = −50, and then reading −6 as the larger spread; saying Chloe scored far more marks in total assumes a wider set of marks must add to more, when both totals are 325; saying the two sets vary by the same amount assumes that equal means force equal spread, when the two ranges are 6 and 50.
- (c) The frequency of that category — Method: a bar chart for categorical data has one bar for each category, and the vertical scale on which the bars are measured is a count. Working: a bar drawn twice as tall as another tells you that twice as many items of data fell into its category, so the height measures how many items of data belong to that one category, which is exactly what a frequency is. Answer: the height of each bar is the frequency of that category — a count of items of data. The distractors: the number of different categories comes from reading the vertical scale as though it counted the bars, which is shown along the horizontal axis instead; the total of all the data comes from treating one bar as though it stood for the whole data set rather than for one category; the mean of all the data comes from confusing a bar chart with a measure of average, which no single bar can show.
- (d) 0 — Method: the range is the largest value minus the smallest value, whatever those two values turn out to be. Working: every value is 10, so the largest value is 10 and the smallest value is 10 as well, and the range is 10 − 10 = 0. Answer: 0 — a range of nothing says the data do not vary at all. The distractors: 10 comes from writing down the repeated value itself instead of the difference between the extremes; 20 comes from adding the largest and the smallest, 10 + 10, instead of subtracting; 40 comes from adding all four values, which gives the total sold and not a measure of spread.
- (c) The data show a link only; a third factor may affect both — Method: a study of this kind measures two quantities and reports how they change together; deciding that one of them produces the other is a further claim, and it needs evidence that the measurements alone cannot give. Working: the study shows that more coffee goes with better concentration, which is a positive correlation; but a third factor that was never measured, such as how motivated someone is, could raise both the coffee drinking and the concentration, and the concentration could equally be what leads to the extra coffee. Answer: the data show a link only, because a third factor may be affecting both quantities, so no claim about cause can be made. The distractors: calling the conclusion safe because the correlation is positive treats the direction of a correlation as proof of cause, which no direction can give; calling it wrong because the correlation is negative misreads the direction of the relationship, since the study reports both quantities rising together; saying the two quantities are not linked denies the correlation the study actually found, when what fails is only the claim about cause.
- (b) 1 — The four frequencies are 4, 7, 6 and 3 matches, and the largest of these is 7, which corresponds to 1 goal, so the modal number of goals is 1. Choosing 7 confuses the frequency, how many matches, with the number of goals itself. Choosing 2 uses the second-largest frequency, 6 matches, instead of the largest. Choosing 3 uses the smallest frequency, which belongs to the fewest matches, not the most.
- (a) The temperature is rising steadily — Method: the trend of a line graph is the overall direction of the readings as time goes on, found by comparing each reading with the one before it. Working: from 10 °C to 14 °C is a rise of 4 °C, and the same comparison from 14 °C to 18 °C, from 18 °C to 22 °C, from 22 °C to 26 °C and from 26 °C to 30 °C gives a rise of 4 °C every time; every reading is greater than the one before it and none of them falls. Answer: the temperature is rising steadily — steadily because the rise is the same size each hour. The distractors: falling steadily comes from reading the six values from right to left, which reverses the direction of time; stays the same comes from noticing that the step of 4 °C is the same each hour and describing the step as constant instead of the temperature; rises and then falls comes from assuming that a line graph has to turn at some point rather than reading the values that are actually given.
- (b) A line graph — Sales recorded at the end of each of the twelve months are time series data, and a line graph is the chart built to show how a value changes over time, with the points usually joined in order. A pie chart is for showing categorical data as shares of a whole, not a trend over time. A pictogram shows a frequency for separate categories using symbols, not a continuous trend. A scatter graph is for comparing two different variables against each other, not one variable over time.
- (b) 13 — The total is 50, and the two known parts are 22 (tea) and 15 (coffee), so 50 − 22 − 15 = 13 hot chocolates. Choosing 28 comes from 50 − 22, subtracting only the tea and forgetting the coffee. Choosing 35 comes from 50 − 15, subtracting only the coffee and forgetting the tea. Choosing 37 comes from 22 + 15, which finds how many drinks were tea or coffee, not the number left over for hot chocolate.
- (a) Equal means; Class A is more consistent, smaller range. — Method: when two data sets share a measure of location, compare a measure of spread to say more about consistency. Working: both classes have the same mean mark, 14, so on average they performed equally well. Class A has the smaller range, 6, so its marks are more tightly grouped around 14 than Class B's marks, which vary by as much as 14. So Class A's marks were more consistent, even though neither class did better on average. Saying Class B did better because it has the bigger range confuses a wide spread with a high score — a big range describes variability, not performance. Saying Class A did better because it has the smaller range makes the same mistake in the other direction: the two classes are tied on the mean, so neither one 'did better'. Saying the classes cannot be compared because their means are equal misses the whole point of also comparing the range. Always compare both an average AND a spread before describing two data sets — either one alone tells only half the story.
- (c) 5 — Method: the mode is the value that occurs most often, so count how many times each different value appears and compare the counts. Working: 3 appears twice, 5 appears three times, 7 appears once and 8 appears once, so the highest frequency is three and the value carrying it is 5. Answer: 5. The distractors: 3 comes from writing down the frequency of the most common answer instead of the answer itself; 7 comes from taking the middle number of the list as it was written, which applies the median without ordering the data and without answering the question asked; 8 comes from picking the largest value, which confuses the mode with the maximum.
- (c) 90° — Method: the sectors of a pie chart share the 360° at the centre of the chart in the same proportion as the data, so a sector's angle is its share of the data multiplied by 360°. Working: a share of 25% is the fraction 25/100, which is one quarter of the whole chart, and one quarter of 360° is 360 ÷ 4 = 90°. Answer: 90°, an angle in degrees rather than a percentage. The distractors: 25° comes from sharing out 100 instead of 360, so the percentage is written straight down as a number of degrees; 14.4° comes from dividing 360 by 25 instead of multiplying 360 by the fraction 25/100, which is the division done the wrong way round; 45° comes from taking a quarter of 180° instead of a quarter of 360°, treating the pie chart as a semicircle.
- (c) 16 kg — Method: sort the seven weights before finding the middle value. Working: in order, the weights are 10, 12, 14, 16, 18, 20 and 50 kg. There are 7 values, so the median is the 4th one: 16 kg. Reading off the 4th weight in the order the vet recorded them, 10 kg, skips the sorting step and is not the median. Working out the mean, 140 ÷ 7 = 20 kg, finds a different average altogether. Working out the range, 50 − 10 = 40 kg, finds the spread, not the middle value. Always sort your data first — the median lives in the ordered list, not the collection order.
- (a) 1,400 and 1,240, so combine the samples for one estimate — Method: scale each sample up to the whole stock, then use the fact that a larger sample gives a more reliable estimate than a smaller one. Working: the first sample gives 35 ÷ 50 = 0.7 and 0.7 × 2,000 = 1,400 paperbacks; the second gives 31 ÷ 50 = 0.62 and 0.62 × 2,000 = 1,240 paperbacks. Two random samples of the same size are expected to differ a little, so neither estimate is wrong. Putting the two together gives 35 + 31 = 66 paperbacks in 100 books, and 66 ÷ 100 = 0.66 with 0.66 × 2,000 = 1,320, an estimate resting on twice as many books as either volunteer checked. Answer: 1,400 and 1,240, so combine the samples for one estimate. The distractors: keeping 1,400 because it is larger picks an estimate by its size, when both samples held 50 books and neither has a stronger claim; saying a volunteer must have miscounted assumes two random samples ought to agree exactly, which is precisely what random sampling does not promise; 1,750 and 1,550 come from 35 × 50 = 1,750 and 31 × 50 = 1,550, multiplying each count by the size of the sample instead of scaling by 2,000 ÷ 50.
- (b) Ethan, 10 seconds — Method: over the same distance the fastest runner is the one who takes the least time, so the smallest time in the table is found first and the name is then read from the same row. Working: the four times are 12 seconds, 15 seconds, 10 seconds and 14 seconds; in order of size these are 10, 12, 14 and 15, so the least time is 10 seconds, and the row holding 10 seconds is the row for Ethan. Answer: Ethan, 10 seconds — the time is in seconds, and a smaller time means a faster runner. The distractors: Grace with 15 seconds comes from taking the largest number in the table to mean the fastest runner, which reverses the relationship between time and speed over a fixed distance; Oliver with 12 seconds comes from writing down the first row of the table without comparing the four times; Ethan with 15 seconds comes from identifying the right runner but then reading the time from a different row of the table.
Build your own mix at the worksheet builder.