Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Answer key: Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- (a) A straight line sloping down from left to right — Method: a line of best fit is a straight line drawn to follow the trend of the points, so its slope is decided by the direction of the relationship between the two quantities. Working: as the price rises, the satisfaction score falls, so the points start high on the left of the graph and finish low on the right; the straight line that follows them therefore slopes downwards as the graph is read from left to right, which is the line of a negative correlation. Answer: a straight line sloping down from left to right. The distractors: a line sloping up comes from reading a falling relationship as a rising one; a horizontal line comes from expecting no correlation, since a horizontal line says the satisfaction score does not change as the price changes; a curve passing through every point comes from thinking a line of best fit has to touch all of the plotted points, when it is a single straight line drawn through the middle of them.
- (d) 45 — Method: convert the angle into a fraction of the full circle, 360°, then apply that fraction to the total number of shoppers. Working: the card sector is 90° out of 360°, a fraction of 90 ÷ 360 = 0.25. Applying that fraction to the 180 shoppers gives 0.25 × 180 = 45 shoppers. Giving 90 states the angle itself, not a number of shoppers — the angle first has to be converted into a fraction. Using the remaining angle, 360 − 90 = 270°, and scaling that, 270 ÷ 360 × 180 = 135, finds the number who did NOT pay by card, not the number who did. Dividing 360 by 90, 360 ÷ 90 = 4, finds how many equal 90° sectors fit in the circle, a fact about the pie chart's shape, not about the shoppers at all. Always convert the angle to a fraction of 360° first, and apply that same fraction to the total number of people.
- (b) A line graph — Sales recorded at the end of each of the twelve months are time series data, and a line graph is the chart built to show how a value changes over time, with the points usually joined in order. A pie chart is for showing categorical data as shares of a whole, not a trend over time. A pictogram shows a frequency for separate categories using symbols, not a continuous trend. A scatter graph is for comparing two different variables against each other, not one variable over time.
- (a) 26 — Method: a line of best fit lets one quantity be predicted from the other, so the score is substituted into the equation of the line and the resulting inequality is solved for the number of hours. Working: a score of at least 80 means 2.5x + 15 ≥ 80; taking 15 from both sides gives 2.5x ≥ 65, and dividing both sides by 2.5 gives x ≥ 26, so the least whole number of hours is 26. Checking, 2.5 × 26 + 15 = 80, which does reach the target. Answer: 26 hours — and this is only an estimate, because a line of best fit predicts a trend rather than an individual result, and a prediction made outside the range of hours the pupils actually revised for would be an extrapolation and less reliable still. The distractors: 27 comes from reaching 26 and then rounding up again, although 26 hours already gives a score of exactly 80; 32 comes from 80 ÷ 2.5, which ignores the 15 in the equation of the line; 38 comes from (80 + 15) ÷ 2.5, that is from adding the 15 instead of subtracting it when rearranging.
- (d) The modal class, as the class with most pupils is shown — Method: a grouped frequency table records how many values fall into each class, but not the values themselves, so any average that needs the individual times can only be estimated from it. Working: the four frequencies are 8, 12, 6 and 4, and 8 + 12 + 6 + 4 = 30, so every pupil is counted. The largest frequency is 12, which belongs to the class 10 < t ≤ 20, and that class can be written down exactly, because finding it needs nothing but the counts the table already gives. Answer: the modal class, as the class with most pupils is shown. The distractors: the mean is said to use all 30 times, but the table does not hold them; the usual method replaces each class by its midpoint, 5, 15, 25 and 35, which gives an estimate of the mean and not its true value; the median is said to be shown, but the table locates only the class holding the 15th and 16th times, which is 10 < t ≤ 20, without saying what either time was; the range is said to be shown, but 0 and 40 are the boundaries of the first and last classes, not the fastest and slowest times actually recorded.
- (c) 40 minutes — Method: for grouped data, estimate the mean using the midpoint of each class — multiply each midpoint by its frequency, add the results, then divide by the total frequency. Working: the midpoints are 10, 30, 50 and 70 minutes. 10 × 5 = 50. 30 × 10 = 300. 50 × 10 = 500. 70 × 5 = 350. Σfx = 50 + 300 + 500 + 350 = 1200. Σf = 5 + 10 + 10 + 5 = 30. Estimated mean = 1200 ÷ 30 = 40 minutes. Using the upper boundary of each class instead of the midpoint — 20 × 5 = 100, 40 × 10 = 400, 60 × 10 = 600, 80 × 5 = 400 — gives a total of 1500 and an estimate of 1500 ÷ 30 = 50 minutes, too high because a boundary is not the middle of the class. Averaging the frequencies themselves, 5, 10, 10 and 5, ignores the times altogether and gives 7.5. Stopping after Σfx = 1200 without dividing by the total frequency gives a number far too large to be a time in minutes. Always find the midpoint of each class before multiplying by the frequency, and always divide by Σf at the end.
- (d) Full-time workers, as most are at work at that time — Method: a sample is biased when the method of contact makes part of the population much less likely to be reached, so test each group against where its members actually are between 10 am and 2 pm on a weekday, and test each stated reason against the facts. Working: those hours are the middle of the working day, so people in full-time employment are at work and not beside a landline telephone, while people who are retired and people who are unemployed are far more likely to be at home and are reached at the usual rate; the method therefore collects far fewer replies from full-time workers than their share of the adult population the council is consulting. Answer: full-time workers, as most are at work at that time. The distractors: the reply naming retired people rests on the false claim that most retired people are at work in the daytime, when in fact a daytime call reaches them more easily than anyone; the reply naming unemployed people rests on the false claim that they are out during the day, when they too are among the easiest people to reach by a daytime call; the reply naming children rests on the false claim that children are at home at 11 am on a Tuesday, when they are at school and so are not reached by the call at all, and school-age children are in any case not the adults whose views the council is collecting.
- (c) No — 8 from one class is too small to represent the school. — Method: judge reliability by asking whether the sample is both large enough, and spread across the population, relative to what it is meant to represent. Working: 8 pupils is a tiny fraction of the school's 1,000 pupils, and all 8 come from a single class rather than a range of year groups, so the sample is both too small and too narrow to represent the whole school reliably. She is not right. Saying any sample size gives an equally reliable estimate ignores that reliability generally improves with a larger, more representative sample. Saying the method is unreliable because it was not done online is not a reason connected to sample size or representativeness at all. Saying 8 is reliable because it is more than half her class compares the sample to the wrong population — the school has 1,000 pupils, not one class. Always judge a sample's size against the population it is meant to represent, not against a smaller group within it.
- (d) 19 kg — Method: estimate the mean of grouped data by multiplying each class's midpoint by its frequency, adding the four totals, then dividing by the total frequency. Working: the midpoints are 5, 15, 25 and 35 kg. The weighted totals are 11 × 5 = 55, 5 × 15 = 75, 5 × 25 = 125 and 9 × 35 = 315, which add to 570. Dividing by the 30 dogs gives an estimate of 570 ÷ 30 = 19 kg. Giving 5 kg reads off the midpoint of the modal class, 0 < m ≤ 10, the class with the most dogs — but the class with the most dogs is not where the mean falls, and neither is a substitute for actually calculating it. Giving 20 kg averages the four midpoints, (5 + 15 + 25 + 35) ÷ 4, treating every class as equally likely and ignoring that far more dogs are in the lightest and heaviest classes than in the middle two. Giving 570 kg stops after finding the correct weighted total and forgets the final division by the 30 dogs. Always weight each midpoint by its own frequency, and always finish by dividing by the total frequency, not the number of classes.
- (b) A person's shoe size and their favourite colour — A person's shoe size is not linked to which colour they prefer, so these two show no correlation. The other three pairs are all genuinely correlated: distance travelled and fuel used rise together, which is positive correlation; hours of revision and test score generally rise together, which is also positive correlation; and as outdoor temperature rises, fewer woolly hats are sold, which is negative correlation. Negative correlation is still a real relationship between two variables — it is not the same thing as no relationship at all, so the temperature and hats pair is not the answer to this question.
- (c) A line of best fit — Method: the straight line drawn on a scatter graph is named from the job it does — it is chosen so that it follows the whole set of points as closely as possible. Working: the line passes through the middle of the points, with roughly as many points above it as below it, and it need not pass through any of the plotted points at all; the name given to the straight line chosen in that way is a line of best fit. Answer: a line of best fit. The distractors: a line of symmetry comes from confusing a trend with symmetry, which is a property of a shape rather than of a set of data; a horizontal line through the mean comes from thinking the trend is shown by an average, when a horizontal line would say that the vertical quantity does not change and so show no correlation; a line joining the first and last points comes from thinking the line must join the two extreme points, which lets two points decide a trend that all of the points should share in.
- (a) 5 pupils are far too few to represent 900 pupils — Method: a sample can only support a claim about a population if it is chosen fairly and if it is large enough for the pattern in it to be more than chance. Working: Priya's method of choosing is fair, because the 5 pupils were picked at random, so every pupil had the same chance of being asked. The difficulty is the size: 900 ÷ 5 = 180, so each pupil she asks stands for 180 pupils. If two of the five happen to play in the same netball team, netball takes 40% of her sample on the strength of two answers, and a second sample of 5 could easily give a different favourite sport. Answer: 5 pupils are far too few to represent 900 pupils. The distractors: saying the pupils were not chosen at random contradicts the question, which states that they were; saying the 5 may each name a different sport describes what often happens in a small sample, but disagreement is not the fault, since 5 pupils who all named the same sport would be just as weak a basis for a claim about 900; saying a sample must hold at least half of the population is an invented rule, and a properly chosen sample of a few hundred can describe a population of many thousands.
- (d) Only internet users reach the website; others are excluded. — Method: a sample is biased when it systematically leaves out part of the population, or systematically over-represents another part. Working: anyone without internet access, or who does not visit the council's website, has NO chance of being included — the sample is drawn only from internet-using residents, which is not the whole town. Saying too many people might respond because the survey is free confuses bias with sample size — bias is about who CAN be reached, not how many respond. Saying people might lie describes a different problem, response honesty, not who was sampled in the first place. Saying online surveys cannot be anonymous is not a reason connected to bias at all. A sample is biased when part of the population has no chance of being included, whatever the reason for that.
- (d) Testing destroys bulbs, so testing all leaves none to sell. — Method: testing every item in a population instead of a sample is a census — sensible only when testing does not use up or destroy what is being tested. Working: here, testing a bulb to find its lifespan destroys it, so testing all 50,000 bulbs would leave nothing left to sell — a sample lets the company estimate the typical lifespan without destroying its whole stock. Extra electricity used in testing is not the real reason a census is avoided here — it is the destruction of the product that matters. Saying a sample is always more accurate than a full census is the wrong way round: a census, if it could be carried out, gives the exact figure for the whole population — it is testing being destructive, not a lack of accuracy, that rules it out here. There is no law against testing every item a company makes — nothing in the question suggests that. When testing destroys the item being tested, sampling is necessary, not just convenient.
- (d) 20 — 0.4x + 1 = 9, so 0.4x = 9 − 1 = 8, and 8 ÷ 0.4 = 20, so 20 sessions are needed. Choosing 23 divides 9 by 0.4 without first subtracting the 1: 9 ÷ 0.4 = 22.5, rounded up to 23. Choosing 25 subtracts the wrong way, adding the 1 instead of taking it away: 9 + 1 = 10, and 10 ÷ 0.4 = 25. Choosing 2 misplaces the decimal point in the gradient, dividing by 4 instead of by 0.4: 8 ÷ 4 = 2.
Build your own mix at the worksheet builder.