Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Answer key: Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- (d) 45 — Method: convert the angle into a fraction of the full circle, 360°, then apply that fraction to the total number of shoppers. Working: the card sector is 90° out of 360°, a fraction of 90 ÷ 360 = 0.25. Applying that fraction to the 180 shoppers gives 0.25 × 180 = 45 shoppers. Giving 90 states the angle itself, not a number of shoppers — the angle first has to be converted into a fraction. Using the remaining angle, 360 − 90 = 270°, and scaling that, 270 ÷ 360 × 180 = 135, finds the number who did NOT pay by card, not the number who did. Dividing 360 by 90, 360 ÷ 90 = 4, finds how many equal 90° sectors fit in the circle, a fact about the pie chart's shape, not about the shoppers at all. Always convert the angle to a fraction of 360° first, and apply that same fraction to the total number of people.
- (d) 23.57 °C — Method: add all seven temperatures, divide by the number of readings and round only at the end. Working: 22 + 24 + 23 + 25 + 26 + 21 + 24 = 165, and 165 ÷ 7 = 23.5714…, which rounds to 23.57 to 2 decimal places. Answer: 23.57 °C. The distractors: 24 °C comes from writing down the mode, the only temperature recorded twice, instead of the mean; 27.5 °C comes from dividing the total by 6 instead of by the 7 days recorded; 5 °C comes from working out the range, 26 − 21, which is a measure of spread and not an average.
- (c) 40 minutes — Method: for grouped data, estimate the mean using the midpoint of each class — multiply each midpoint by its frequency, add the results, then divide by the total frequency. Working: the midpoints are 10, 30, 50 and 70 minutes. 10 × 5 = 50. 30 × 10 = 300. 50 × 10 = 500. 70 × 5 = 350. Σfx = 50 + 300 + 500 + 350 = 1200. Σf = 5 + 10 + 10 + 5 = 30. Estimated mean = 1200 ÷ 30 = 40 minutes. Using the upper boundary of each class instead of the midpoint — 20 × 5 = 100, 40 × 10 = 400, 60 × 10 = 600, 80 × 5 = 400 — gives a total of 1500 and an estimate of 1500 ÷ 30 = 50 minutes, too high because a boundary is not the middle of the class. Averaging the frequencies themselves, 5, 10, 10 and 5, ignores the times altogether and gives 7.5. Stopping after Σfx = 1200 without dividing by the total frequency gives a number far too large to be a time in minutes. Always find the midpoint of each class before multiplying by the frequency, and always divide by Σf at the end.
- (b) Treat them as outliers and check them before deciding — Method: a point lying a long way from the pattern the rest of the data make is called an outlier, and an outlier is investigated before anything is done with it, because it may be an error in the data or it may be a genuine but unusual case. Working: 48 of the 50 points lie close to the rising line of best fit, so the trend is set by those 48; the two remaining points do not follow it, so they are identified as outliers and checked — a mistake in measuring or recording would be corrected, while a genuine reading would be kept and reported. Answer: treat them as outliers and check them before deciding what to do with them. The distractors: deleting them at once assumes that every point far from the line must be an error, which throws away real data; moving the line so that it passes through them assumes a line of best fit must touch particular points, when it is drawn to follow all 50; taking them as proof that there is no correlation lets two points overturn the pattern that the other 48 agree on.
- (a) 5 pupils are far too few to represent 900 pupils — Method: a sample can only support a claim about a population if it is chosen fairly and if it is large enough for the pattern in it to be more than chance. Working: Priya's method of choosing is fair, because the 5 pupils were picked at random, so every pupil had the same chance of being asked. The difficulty is the size: 900 ÷ 5 = 180, so each pupil she asks stands for 180 pupils. If two of the five happen to play in the same netball team, netball takes 40% of her sample on the strength of two answers, and a second sample of 5 could easily give a different favourite sport. Answer: 5 pupils are far too few to represent 900 pupils. The distractors: saying the pupils were not chosen at random contradicts the question, which states that they were; saying the 5 may each name a different sport describes what often happens in a small sample, but disagreement is not the fault, since 5 pupils who all named the same sport would be just as weak a basis for a claim about 900; saying a sample must hold at least half of the population is an invented rule, and a properly chosen sample of a few hundred can describe a population of many thousands.
- (d) The mode, because colours cannot be added or ordered — Method: an average can only be used on data that supports the operation it needs. A mean needs the values to be added and divided, a median needs them to be placed in order, and a range needs one value to be taken away from another; a mode needs only counting, so it is the average available when the data are categories rather than numbers. Working: the data collected here are colours, silver, black, blue and red. The numbers 74, 52, 40 and 34 count the cars of each colour, they do not measure them, and 74 + 52 + 40 + 34 = 200 simply returns the size of the survey. No colour can be added to another, and there is no order that puts blue before red, so of the four averages only the one found by counting survives. Answer: the mode, because colours cannot be added or ordered, and the mode is silver. The distractors: the mean is said to use all 200 colours, and a mean of the four frequencies, 200 ÷ 4 = 50, is a number of cars rather than a colour, so it describes nothing about a typical car; the median is said to put the colours in order, but ordering the frequencies 34, 40, 52, 74 orders the counts, not the colours, and gives 46, again a number of cars; the range is not an average at all, and 74 − 34 = 40 measures the gap between the commonest and rarest counts, which is a measure of spread.
- (d) Testing destroys bulbs, so testing all leaves none to sell. — Method: testing every item in a population instead of a sample is a census — sensible only when testing does not use up or destroy what is being tested. Working: here, testing a bulb to find its lifespan destroys it, so testing all 50,000 bulbs would leave nothing left to sell — a sample lets the company estimate the typical lifespan without destroying its whole stock. Extra electricity used in testing is not the real reason a census is avoided here — it is the destruction of the product that matters. Saying a sample is always more accurate than a full census is the wrong way round: a census, if it could be carried out, gives the exact figure for the whole population — it is testing being destructive, not a lack of accuracy, that rules it out here. There is no law against testing every item a company makes — nothing in the question suggests that. When testing destroys the item being tested, sampling is necessary, not just convenient.
- (a) No, 150 pupils are a tenth of the school, chosen at random — Method: judge a sample on two things, whether every member of the population had the same chance of being chosen, and whether the sample is large enough to carry a pattern. Working: the 150 pupils were drawn from the register of every pupil in the school, so no year group or set is shut out and no pupil chooses to take part; and 150 ÷ 1,500 = 0.1, so one pupil in ten has been asked. A random sample of that share is ample for an estimate of how long the school's pupils spend on homework. Answer: no, because 150 pupils are a tenth of the school and were chosen at random. The distractors: saying a random sample always gives the exact school figure reaches the same verdict for a reason that is false, since a second random sample of 150 would give a slightly different mean; saying 150 pupils cannot be picked at random from 1,500 treats randomness as something only a whole population can have, when drawing names from the register is exactly how a random sample is taken; saying that only asking all 1,500 could show anything rejects sampling altogether, which would leave no way to study any population too large to count.
- (d) Full-time workers, as most are at work at that time — Method: a sample is biased when the method of contact makes part of the population much less likely to be reached, so test each group against where its members actually are between 10 am and 2 pm on a weekday, and test each stated reason against the facts. Working: those hours are the middle of the working day, so people in full-time employment are at work and not beside a landline telephone, while people who are retired and people who are unemployed are far more likely to be at home and are reached at the usual rate; the method therefore collects far fewer replies from full-time workers than their share of the adult population the council is consulting. Answer: full-time workers, as most are at work at that time. The distractors: the reply naming retired people rests on the false claim that most retired people are at work in the daytime, when in fact a daytime call reaches them more easily than anyone; the reply naming unemployed people rests on the false claim that they are out during the day, when they too are among the easiest people to reach by a daytime call; the reply naming children rests on the false claim that children are at home at 11 am on a Tuesday, when they are at school and so are not reached by the call at all, and school-age children are in any case not the adults whose views the council is collecting.
- (d) 210 — 350 − 150 = 200. 200 ÷ 10 = 20, so the gradient is 20. Using the point (5, 150): 20 × 5 = 100, so 150 − 100 = 50 is the intercept, giving the line y = 20x + 50. At x = 8: 20 × 8 = 160, and 160 + 50 = 210, so the estimated number of visitors is 210. Choosing 160 stops after 20 × 8 = 160 and forgets to add the intercept of 50. Choosing 250 comes from averaging the two given y-values: 150 + 350 = 500, and 500 ÷ 2 = 250, instead of using the line's equation. Choosing 240 assumes the visitors are directly proportional to the hours of sunshine using the first point, 150 × 8 ÷ 5 = 240, which ignores that the line does not pass through the origin.
- (d) 648 kg — Method: to estimate a total from a sample, multiply the sample's mean by the number of items in the whole population, then check the units the question asks for. Working: 32.4 g × 20,000 = 648,000 g. Converting to kilograms, 648,000 ÷ 1,000 = 648 kg. This is only an estimate, not an exact total, because it assumes every one of the 20,000 packets has exactly the sample mean mass, when in reality individual packets vary above and below it. Giving 1.62 kg multiplies the mean by 50, the SAMPLE size, instead of by 20,000, the number of packets actually made that day — this finds the total mass of the 50 sampled packets, not the day's production. Giving 32.4 kg treats the sample mean itself, in grams, as if it already were the day's total mass in kilograms, skipping the scaling up altogether. Giving 648,000 kg correctly scales the mean up to the whole day's production but never converts the answer from grams to kilograms, leaving it 1,000 times too large. Always scale a sample's mean up by the SIZE OF THE WHOLE POPULATION, and always finish by checking the units the question asks for.
- (c) 156 cm — Method: to combine two groups' means, multiply each group's mean by its own number of pupils, add the two totals together, then divide by the total number of pupils in both groups. Working: 20 × 150 = 3,000 cm for the boys and 10 × 168 = 1,680 cm for the girls, giving a combined total of 3,000 + 1,680 = 4,680 cm. Dividing by all 30 pupils gives 4,680 ÷ 30 = 156 cm. Giving 159 cm averages the two means, (150 + 168) ÷ 2, treating the two groups as if they had the same number of pupils, when there are twice as many boys as girls. Giving 4,680 cm finds the correct combined total height but stops there, forgetting the final division by the 30 pupils. Giving 234 cm divides the combined total by 20, the number of boys only, forgetting that the total also includes the 10 girls. Always weight each mean by its own group size, and always divide by the TOTAL number of pupils in both groups combined.
- (a) A vertical line chart (discrete numerical data) — The number of pets is discrete numerical data — whole-number values such as 0, 1, 2, 3 or 4 — recorded for one variable, so a vertical line chart is the chart specified for this kind of data. A bar chart is used for categorical data, such as favourite colour, not numerical values counted like this. A pie chart shows proportions of a whole and does not show the frequency of each separate value. A scatter graph compares two different variables against each other, and only one variable, the number of pets, is recorded here.
- (b) Interpolation — The salary is being estimated for a value of x between 1 and 15, which is inside the range of x-values that were actually plotted, so this is interpolation. Extrapolation would apply if the estimate used a value of x below 1 or above 15, outside the plotted range. Correlation describes the relationship between the two variables, not the reliability of an estimate, and causation describes one variable actually causing a change in the other, which is a different idea altogether — neither is the word being asked for here.
- (b) All 2,000 members of the sports centre. — Method: in a survey, the population is the whole group the survey is trying to find out about, and the sample is the smaller group actually asked. Working: this survey wants to know what the sports centre's members think, so the population is every one of the 2,000 members — whether or not they were personally asked. Saying the population is the 100 members who were asked names the sample, not the population; the sample is drawn FROM the population, so it is smaller than it, not the same as it. Saying the population is everybody who lives in Ipswich widens the group far beyond who the survey is actually about — plenty of Ipswich residents are not members of the sports centre at all, so they are outside this survey altogether. Saying the population is the members who say they are unhappy confuses the population with a result of the survey: whether a member turns out to be happy or unhappy is something the survey finds out, not part of the definition of who is being studied. The population is always the whole group the question is about, before any sampling or any results come in.
Build your own mix at the worksheet builder.