Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Answer key: Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- (d) The modal class, as the class with most pupils is shown — Method: a grouped frequency table records how many values fall into each class, but not the values themselves, so any average that needs the individual times can only be estimated from it. Working: the four frequencies are 8, 12, 6 and 4, and 8 + 12 + 6 + 4 = 30, so every pupil is counted. The largest frequency is 12, which belongs to the class 10 < t ≤ 20, and that class can be written down exactly, because finding it needs nothing but the counts the table already gives. Answer: the modal class, as the class with most pupils is shown. The distractors: the mean is said to use all 30 times, but the table does not hold them; the usual method replaces each class by its midpoint, 5, 15, 25 and 35, which gives an estimate of the mean and not its true value; the median is said to be shown, but the table locates only the class holding the 15th and 16th times, which is 10 < t ≤ 20, without saying what either time was; the range is said to be shown, but 0 and 40 are the boundaries of the first and last classes, not the fastest and slowest times actually recorded.
- (a) x = 0 gives y = −20: a negative number sold — The y-intercept is the value the line predicts when x = 0: y = 3 × 0 − 20 = −20. A kiosk cannot sell a negative number of ice creams, so this is not a sensible estimate. The 3 in the equation is the gradient, not the intercept, so an option claiming x = 0 gives y = 3 has swapped the two numbers around — substituting x = 0 makes the 3x term equal 0, leaving −20, not 3. The danger of extrapolating to very high temperatures is a real issue with this line, but it is a different issue from the y-intercept, so it does not answer this question. And whether x = 0 could occur on a trading day is beside the point: the model still makes that prediction, and it is the prediction itself, −20, that is impossible.
- (d) 1100 — Method: to combine two samples of different sizes, add the faulty counts together and add the sample sizes together before scaling up, rather than treating the two samples separately. Working: the combined sample found 34 + 21 = 55 scratched cases out of 100 + 50 = 150 cases checked, a proportion of 55 ÷ 150. Applying that proportion to the week's production of 3,000 gives an estimate of 55 ÷ 150 × 3000 = 1100 scratched cases. Averaging the two shifts' proportions instead of combining their totals, (34 ÷ 100 + 21 ÷ 50) ÷ 2 = 0.38, gives 0.38 × 3000 = 1140 — this treats the two samples as equally weighted even though Shift A checked twice as many cases as Shift B. Using only Shift A's sample, 34 ÷ 100 × 3000 = 1020, ignores Shift B's cases completely. Using only Shift B's sample, 21 ÷ 50 × 3000 = 1260, ignores Shift A's cases completely. When two samples are different sizes, combine their totals before finding the proportion — do not average the two proportions, and do not use only one shift's sample.
- (b) 120, unreliable — x = 25 is outside 1 to 12 — The line of best fit is y = 4x + 20. 4 × 25 = 100, and 100 + 20 = 120, so the estimate is y = 120. But x = 25 lies far outside the plotted range of 1 to 12 years, so this is an extrapolation, and the estimate is not reliable. Reaching 100 instead of 120 comes from 4 × 25 = 100 with the intercept of 20 left out — still correctly flagged as unreliable, but the wrong value. Calling the estimate reliable simply because it was calculated correctly, giving 120, wrongly assumes that a correct calculation is automatically trustworthy, ignoring that x = 25 lies far beyond the data actually collected. Reaching 68, from 4 × 12 = 48 and 48 + 20 = 68, substitutes x = 12, the top of the plotted range, instead of the assistant's actual x = 25, and wrongly calls that reliable because 12 lies inside the range.
- (d) 648 kg — Method: to estimate a total from a sample, multiply the sample's mean by the number of items in the whole population, then check the units the question asks for. Working: 32.4 g × 20,000 = 648,000 g. Converting to kilograms, 648,000 ÷ 1,000 = 648 kg. This is only an estimate, not an exact total, because it assumes every one of the 20,000 packets has exactly the sample mean mass, when in reality individual packets vary above and below it. Giving 1.62 kg multiplies the mean by 50, the SAMPLE size, instead of by 20,000, the number of packets actually made that day — this finds the total mass of the 50 sampled packets, not the day's production. Giving 32.4 kg treats the sample mean itself, in grams, as if it already were the day's total mass in kilograms, skipping the scaling up altogether. Giving 648,000 kg correctly scales the mean up to the whole day's production but never converts the answer from grams to kilograms, leaving it 1,000 times too large. Always scale a sample's mean up by the SIZE OF THE WHOLE POPULATION, and always finish by checking the units the question asks for.
- (d) 23.57 °C — Method: add all seven temperatures, divide by the number of readings and round only at the end. Working: 22 + 24 + 23 + 25 + 26 + 21 + 24 = 165, and 165 ÷ 7 = 23.5714…, which rounds to 23.57 to 2 decimal places. Answer: 23.57 °C. The distractors: 24 °C comes from writing down the mode, the only temperature recorded twice, instead of the mean; 27.5 °C comes from dividing the total by 6 instead of by the 7 days recorded; 5 °C comes from working out the range, 26 − 21, which is a measure of spread and not an average.
- (c) 40 minutes — Method: for grouped data, estimate the mean using the midpoint of each class — multiply each midpoint by its frequency, add the results, then divide by the total frequency. Working: the midpoints are 10, 30, 50 and 70 minutes. 10 × 5 = 50. 30 × 10 = 300. 50 × 10 = 500. 70 × 5 = 350. Σfx = 50 + 300 + 500 + 350 = 1200. Σf = 5 + 10 + 10 + 5 = 30. Estimated mean = 1200 ÷ 30 = 40 minutes. Using the upper boundary of each class instead of the midpoint — 20 × 5 = 100, 40 × 10 = 400, 60 × 10 = 600, 80 × 5 = 400 — gives a total of 1500 and an estimate of 1500 ÷ 30 = 50 minutes, too high because a boundary is not the middle of the class. Averaging the frequencies themselves, 5, 10, 10 and 5, ignores the times altogether and gives 7.5. Stopping after Σfx = 1200 without dividing by the total frequency gives a number far too large to be a time in minutes. Always find the midpoint of each class before multiplying by the frequency, and always divide by Σf at the end.
- (a) The relationship between two variables — Method: what a diagram shows is decided by what has to be known before a single mark can be plotted on it. Working: every point on a scatter graph is plotted from a pair of measurements taken from the same person or object, one read on the horizontal axis and one on the vertical axis; having two measurements for each point is what makes it possible to look for a pattern between them, and the pattern between two variables is what the graph displays. Answer: a scatter graph shows the relationship between two variables. The distractors: the frequency of each single value is what a bar chart or a vertical line chart shows, and it needs only one list of values; how a total is shared between categories is what a pie chart shows; how one quantity changes over time is what a time series line graph shows, in which one of the two axes is always time.
- (d) 24 — Method: for a sample in proportion to the population, apply the same fraction that each group makes up of the whole population to the size of the sample. Working: women make up 300 out of the 500 members, a fraction of 300 ÷ 500 = 0.6. Applying that fraction to the sample of 40 gives 0.6 × 40 = 24 women. Splitting the sample evenly, 40 ÷ 2 = 20, ignores that the club has more women than men and treats the two groups as equal in size, which they are not. Misreading the sample size as 50 instead of 40, then applying the 3:2 ratio of women to men, 3 ÷ 5 × 50 = 30, uses the right ratio but the wrong sample total. Working out the number of MEN instead of women, 200 ÷ 500 × 40 = 16, answers a different question — how many men, not how many women, belong in the sample. Always apply each group's own share of the population to the sample size, and check which group the question is actually asking about.
- (c) No — 8 from one class is too small to represent the school. — Method: judge reliability by asking whether the sample is both large enough, and spread across the population, relative to what it is meant to represent. Working: 8 pupils is a tiny fraction of the school's 1,000 pupils, and all 8 come from a single class rather than a range of year groups, so the sample is both too small and too narrow to represent the whole school reliably. She is not right. Saying any sample size gives an equally reliable estimate ignores that reliability generally improves with a larger, more representative sample. Saying the method is unreliable because it was not done online is not a reason connected to sample size or representativeness at all. Saying 8 is reliable because it is more than half her class compares the sample to the wrong population — the school has 1,000 pupils, not one class. Always judge a sample's size against the population it is meant to represent, not against a smaller group within it.
- (c) 69 — Method: multiply the mean by the number of values to find the total, then subtract the total of the known values. Working: the total of all five scores is 68 × 5 = 340. The total of the four known scores is 55 + 62 + 74 + 80 = 271. The fifth score is 340 − 271 = 69. Subtracting the other way round, 271 − 340 = −69, gives the right size answer with the wrong sign. Guessing that the missing score simply equals the mean, 68, ignores that the four known scores are not themselves centred on 68. Multiplying the mean by 4 instead of 5, 68 × 4 = 272, then 272 − 271 = 1, undercounts how many scores there are. Always multiply the mean by the TOTAL number of values before subtracting.
- (c) The data show a link only; a third factor may affect both — Method: a study of this kind measures two quantities and reports how they change together; deciding that one of them produces the other is a further claim, and it needs evidence that the measurements alone cannot give. Working: the study shows that more coffee goes with better concentration, which is a positive correlation; but a third factor that was never measured, such as how motivated someone is, could raise both the coffee drinking and the concentration, and the concentration could equally be what leads to the extra coffee. Answer: the data show a link only, because a third factor may be affecting both quantities, so no claim about cause can be made. The distractors: calling the conclusion safe because the correlation is positive treats the direction of a correlation as proof of cause, which no direction can give; calling it wrong because the correlation is negative misreads the direction of the relationship, since the study reports both quantities rising together; saying the two quantities are not linked denies the correlation the study actually found, when what fails is only the claim about cause.
- (d) 45 — Method: convert the angle into a fraction of the full circle, 360°, then apply that fraction to the total number of shoppers. Working: the card sector is 90° out of 360°, a fraction of 90 ÷ 360 = 0.25. Applying that fraction to the 180 shoppers gives 0.25 × 180 = 45 shoppers. Giving 90 states the angle itself, not a number of shoppers — the angle first has to be converted into a fraction. Using the remaining angle, 360 − 90 = 270°, and scaling that, 270 ÷ 360 × 180 = 135, finds the number who did NOT pay by card, not the number who did. Dividing 360 by 90, 360 ÷ 90 = 4, finds how many equal 90° sectors fit in the circle, a fact about the pie chart's shape, not about the shoppers at all. Always convert the angle to a fraction of 360° first, and apply that same fraction to the total number of people.
- (b) At 0 guests, the model predicts 2 m of table — The y-intercept of a line of best fit y = mx + c is the value of y when x = 0. Here y = 0.5 × 0 + 2 = 2, so the line predicts a table length of 2 m when there are 0 guests. The 2 m does not grow as more guests arrive — that role belongs to the gradient, 0.5 — so an option saying each extra guest adds 2 m has swapped the two numbers around. The 2 is a length in metres, not a number of guests, so an option requiring 2 guests before set-up has misread its units. And the table length does change with x, since it is 0.5x + 2 and not a fixed value, so an option claiming the table is always 2 m ignores the 0.5x term completely.
- (b) Interpolation — The salary is being estimated for a value of x between 1 and 15, which is inside the range of x-values that were actually plotted, so this is interpolation. Extrapolation would apply if the estimate used a value of x below 1 or above 15, outside the plotted range. Correlation describes the relationship between the two variables, not the reliability of an estimate, and causation describes one variable actually causing a change in the other, which is a different idea altogether — neither is the word being asked for here.
Build your own mix at the worksheet builder.