Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Answer key: Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- (a) 15 — Year 11 has 50 − 28 = 22 pupils in total. Of the 22 pupils who walk in total, 15 are in Year 10, so 22 − 15 = 7 Year 11 pupils walk. Subtracting that from the Year 11 total gives 22 − 7 = 15 Year 11 pupils who are driven. Choosing 28 takes the whole school's driven total, 50 − 22 = 28, and treats it as if it were Year 11's alone, without separating the year groups. Choosing 7 correctly finds how many Year 11 pupils walk but stops there, giving that figure instead of the number who are driven. Choosing 35 comes from 50 − 15, subtracting the Year 10 walkers from the whole school total rather than working within Year 11.
- (a) The relationship between two variables — Method: what a diagram shows is decided by what has to be known before a single mark can be plotted on it. Working: every point on a scatter graph is plotted from a pair of measurements taken from the same person or object, one read on the horizontal axis and one on the vertical axis; having two measurements for each point is what makes it possible to look for a pattern between them, and the pattern between two variables is what the graph displays. Answer: a scatter graph shows the relationship between two variables. The distractors: the frequency of each single value is what a bar chart or a vertical line chart shows, and it needs only one list of values; how a total is shared between categories is what a pie chart shows; how one quantity changes over time is what a time series line graph shows, in which one of the two axes is always time.
- (d) 6 — Method: the mode, or modal value, is the value that occurs most often in the data set, and it is a value from the data rather than a count. Working: size 4 occurs twice, size 5 occurs once, size 6 occurs three times and size 9 occurs once, so the highest frequency is three and the size it belongs to is 6. Answer: 6. The distractors: 3 comes from writing down the frequency of the most common size instead of the size itself; 9 comes from picking the largest size in the list, which confuses the mode with the maximum; 4 comes from stopping at the first size that repeats rather than checking which size repeats most often.
- (a) Leeds has a higher median and a greater range than York. — In order, Leeds's temperatures are 14, 16, 18, 19 and 23, so the median is the middle value, 18, and the range is 23 − 14 = 9. York's temperatures in order are 15, 17, 17, 18 and 18, so the median is 17, and the range is 18 − 15 = 3. Since 18 is higher than 17, and 9 is greater than 3, Leeds has both the higher median and the greater range. Choosing 'Leeds has a higher median but a smaller range than York' gets the median comparison right but the range comparison backwards — Leeds's range of 9 is actually greater than York's range of 3. Choosing 'York has a higher median and a greater range than Leeds' reverses both comparisons. Choosing 'York has a higher median but a smaller range than Leeds' reverses the median comparison; York's median of 17 is lower than Leeds's 18, even though it is correct that York's range is the smaller one.
- (c) 80 — Method: the range is a measure of spread and is found by subtracting the smallest value from the largest. Working: the largest number sold is 100 and the smallest is 20, so the range is 100 − 20 = 80. Answer: 80. The distractors: 100 comes from writing down the largest value and never subtracting the smallest; 60 comes from working out the median, the middle value of 20, 40, 60, 70, 100, instead of the range; 58 comes from working out the mean, 290 ÷ 5, which measures centre rather than spread.
- (c) 22 kg — Method: multiply the mean by the number of parcels to rebuild the total mass, then subtract the masses that are known. Working: four parcels with a mean mass of 17 kg have a total mass of 17 × 4 = 68 kg; the three known parcels total 12 + 16 + 18 = 46 kg; so the fourth parcel has mass 68 − 46 = 22 kg. Answer: 22 kg. The distractors: 68 kg comes from stopping at the total mass of all four parcels; 17 kg comes from assuming the missing parcel must have the mean mass; 5 kg comes from multiplying the mean by 3, the number of parcels whose mass is given, leaving 51 − 46 = 5.
- (c) mean = 8, median = 8, mode = 8 — Method: work out each measure separately — the mean is the total divided by how many values there are, the median is the middle value once the data are in order, and the mode is the value that occurs most often. Working: the total is 8 + 8 + 8 + 8 = 32 and there are 4 marks, so the mean is 32 ÷ 4 = 8; in order the marks read 8, 8, 8, 8, and the mean of the middle pair is (8 + 8) ÷ 2 = 8; the value 8 occurs 4 times and no other value occurs at all, so the mode is 8. Answer: mean = 8, median = 8, mode = 8 — when every value in a data set is the same, all three measures of central tendency take that value. The distractors: a mean of 32 comes from stopping at the total and never dividing by 4; a mode of 4 comes from writing down how many times 8 occurs instead of the value that occurs; a mean of 2 comes from dividing a single value, 8, by the 4 marks instead of dividing the total by 4.
- (d) The mode, because colours cannot be added or ordered — Method: an average can only be used on data that supports the operation it needs. A mean needs the values to be added and divided, a median needs them to be placed in order, and a range needs one value to be taken away from another; a mode needs only counting, so it is the average available when the data are categories rather than numbers. Working: the data collected here are colours, silver, black, blue and red. The numbers 74, 52, 40 and 34 count the cars of each colour, they do not measure them, and 74 + 52 + 40 + 34 = 200 simply returns the size of the survey. No colour can be added to another, and there is no order that puts blue before red, so of the four averages only the one found by counting survives. Answer: the mode, because colours cannot be added or ordered, and the mode is silver. The distractors: the mean is said to use all 200 colours, and a mean of the four frequencies, 200 ÷ 4 = 50, is a number of cars rather than a colour, so it describes nothing about a typical car; the median is said to put the colours in order, but ordering the frequencies 34, 40, 52, 74 orders the counts, not the colours, and gives 46, again a number of cars; the range is not an average at all, and 74 − 34 = 40 measures the gap between the commonest and rarest counts, which is a measure of spread.
- (d) 0 — Method: the range is the largest value minus the smallest value, whatever those two values turn out to be. Working: every value is 10, so the largest value is 10 and the smallest value is 10 as well, and the range is 10 − 10 = 0. Answer: 0 — a range of nothing says the data do not vary at all. The distractors: 10 comes from writing down the repeated value itself instead of the difference between the extremes; 20 comes from adding the largest and the smallest, 10 + 10, instead of subtracting; 40 comes from adding all four values, which gives the total sold and not a measure of spread.
- (a) £16,000 — Method: first find the quarter with the highest sales figure, then subtract quarter 1's sales from it — remembering that every figure is given in THOUSANDS of pounds. Working: the highest sales figure is quarter 2, at £34,000 (34 thousand pounds). The increase from quarter 1 is £34,000 − £18,000 = £16,000. Giving £34,000 reads off the highest sales figure on its own, without subtracting quarter 1's sales — that is the highest quarter's total, not the increase. Giving £12,000 uses quarter 3's sales, 30, the SECOND-highest figure, instead of quarter 2's 34, the actual highest — 30 − 18 = 12, but quarter 3 is not the quarter with the highest sales. Giving £16 gets the subtraction right, 34 − 18 = 16, but forgets that every figure in the question is in thousands of pounds, so the increase is £16,000, not £16. Always identify the correct quarter FIRST, and always check the units the numbers are given in before writing your final answer.
- (b) 13 — The total is 50, and the two known parts are 22 (tea) and 15 (coffee), so 50 − 22 − 15 = 13 hot chocolates. Choosing 28 comes from 50 − 22, subtracting only the tea and forgetting the coffee. Choosing 35 comes from 50 − 15, subtracting only the coffee and forgetting the tea. Choosing 37 comes from 22 + 15, which finds how many drinks were tea or coffee, not the number left over for hot chocolate.
- (c) Thursday — Method: on a bar chart the tallest bar belongs to the greatest value, so compare the five temperatures and then read off the day that the greatest one belongs to. Working: the temperatures are 20 °C, 22 °C, 18 °C, 25 °C and 23 °C; in order of size these are 18, 20, 22, 23 and 25, so the greatest temperature is 25 °C, and the day recorded with 25 °C is Thursday. Answer: Thursday — the answer is a day, not a temperature. The distractors: Friday comes from stopping at the final temperature listed instead of comparing all five; Wednesday comes from picking out the shortest bar, 18 °C, and so answering for the lowest temperature rather than the highest; Monday comes from writing down the first day in the chart without comparing any of the temperatures at all.
- (c) Yes — with an estimate of 320, above the 250 limit. — Method: scale the sample proportion up to the whole batch to get an estimate, then compare that estimate with the 250 limit to reach a decision. Working: in the sample, 4 out of 50 boards are faulty, a proportion of 4 ÷ 50 = 0.08. Applying that proportion to the batch of 4,000 gives an estimate of 0.08 × 4000 = 320 faulty boards. Since 320 is more than 250, the factory should scrap the batch. Inverting the proportion, 50 ÷ 4 = 12.5, and treating that as a percentage of the batch, 12.5% × 4000 = 500, still gives 'yes' but from the wrong fraction, so it overstates the estimate. Comparing the raw number of faulty boards found in the sample, 4, directly with the 250 limit skips the scaling up to the batch altogether, and 4 is nowhere near 250, so that route wrongly says 'no'. Dividing the batch by the sample size, 4000 ÷ 50 = 80, finds how many samples of 50 fit into the batch but stops before multiplying by the 4 faulty boards found, so it also wrongly says 'no'. Always find the proportion in the sample first, scale it up to the whole batch, and only then compare the estimate with the limit given.
- (a) x = 0 gives y = −20: a negative number sold — The y-intercept is the value the line predicts when x = 0: y = 3 × 0 − 20 = −20. A kiosk cannot sell a negative number of ice creams, so this is not a sensible estimate. The 3 in the equation is the gradient, not the intercept, so an option claiming x = 0 gives y = 3 has swapped the two numbers around — substituting x = 0 makes the 3x term equal 0, leaving −20, not 3. The danger of extrapolating to very high temperatures is a real issue with this line, but it is a different issue from the y-intercept, so it does not answer this question. And whether x = 0 could occur on a trading day is beside the point: the model still makes that prediction, and it is the prediction itself, −20, that is impossible.
- (a) No, the size of the fire affects both of the quantities — Method: correlation says that two quantities change together; a claim that one of them produces the other is a further claim, and it needs evidence that a scatter graph on its own cannot give. Working: the graph does show strong positive correlation, so more engines did go with greater damage. But neither quantity was set by the researchers: both were decided by how large the fire was. A large blaze brings many appliances and also destroys a great deal, while a small one brings few and destroys little, so a third quantity is driving both of the recorded ones. Answer: no, because the size of the fire affects both of the quantities. The distractors: saying the correlation is negative contradicts the graph, which shows the two quantities rising together, and reaching the right verdict from a false reading of the data is not the reason the mark is for; saying that strong positive correlation shows one quantity causes the other is the assumption the question exists to test, and no strength of correlation can establish cause; saying the points lie close to the line of best fit describes how strong the correlation is, and strength and cause are different matters entirely.
Build your own mix at the worksheet builder.