Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Answer key: Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- (c) 72 — Method: when a pie chart is divided into equal sectors, each sector stands for the same share of the people asked, so the fraction of the sectors that are shaded is also the fraction of the people. Working: 6 sectors out of 10 are purple, which is the fraction 6/10 of the whole pie chart; one tenth of the 120 people is 120 ÷ 10 = 12 people, so six tenths is 6 × 12 = 72 people. Answer: 72 people, a count of people rather than a number of sectors. The distractors: 48 comes from working with the 4 sectors that are not purple, 4 × 12, and so answering for the wrong part of the chart; 60 comes from turning the fraction 6/10 into 60% and then writing the 60 down as though it were a number of people; 6 comes from writing down the number of purple sectors instead of the number of people those sectors stand for.
- (c) Club members probably eat differently from most people — Method: a sample is biased when the group it is drawn from differs from the population in the very thing the survey is measuring, so compare the subgroup with the population on that quantity. Working: the survey measures eating habits, and people who join a sports club take more exercise than average and are known to eat differently from the town as a whole, so their replies pull the results away from the true picture for the town however many of them are asked. Answer: club members probably eat differently from most people. The distractors: the reply about the number of members treats bias as a question of size, but a large biased sample is still biased; the reply that the members were picked at random is false, since the council picked a club rather than picking residents, and it confuses bias with non-response; the reply that everyone asked lives in the town notes something true of the members but draws the false conclusion that the sample therefore covers the town, when a sample must reflect a population and not merely be taken from inside it.
- (c) No — median £505 at A vs £510 at B. — Branch A's seven wages in order are £480, £495, £500, £505, £510, £515 and £1,200, so the median, the 4th value, is £505. Branch B's in order are £480, £490, £500, £510, £520, £530 and £540, so the median is £510. Since £505 is lower than £510, the median wage is not higher at Branch A, so the claim is not fairly supported. Choosing 'Yes — mean £600.71 at A vs £510 at B' uses the mean: 480 + 495 + 500 + 505 + 510 + 515 + 1200 = 4205, and 4205 ÷ 7 = 600.71, a figure pulled upward by the £1,200 outlier that does not represent a typical wage. Choosing 'Yes — median £515 at A vs £510 at B' miscounts the middle position, taking the 6th wage, £515, instead of the correct 4th value, £505. Choosing 'Yes — highest wage £1,200 at A vs £540 at B' compares the highest wage at each branch rather than a measure of the typical, or average, wage.
- (b) Positive correlation — Method: the type of correlation is named from the direction the points take as the scatter graph is read from left to right. Working: the points rise from left to right, so as the arm span read on the horizontal axis increases, the height read on the vertical axis increases as well; two quantities that increase together show positive correlation. Answer: positive correlation. The distractors: negative correlation comes from naming the direction the wrong way round, since a negative correlation needs the points to fall as the graph is read from left to right; no correlation comes from treating points that are spread out rather than sitting exactly on a line as though they showed no relationship; direct proportion comes from confusing a rising trend with proportion, which would additionally need the line through the points to pass through the origin and would mean doubling one quantity doubles the other.
- (d) £15.50 — 2 × 6 = 12, and 12 + 3.50 = 15.50, so the estimated fare is £15.50. Choosing £12.00 stops after 2 × 6 = 12 and forgets to add the £3.50. Choosing £19.00 adds the distance and the constant first and then multiplies: 6 + 3.50 = 9.50, and 9.50 × 2 = 19.00, applying the ×2 to the whole sum instead of only to the distance. Choosing £13.00 multiplies only the constant term by 2 instead of the distance: 2 × 3.50 = 7, and 7 + 6 = 13.00.
- (b) 4 — The difference, Bristol minus Leeds, on each day is: Monday 4 − 3 = 1, Tuesday 4 − 5 = −1, Wednesday 6 − 2 = 4, Thursday 2 − 4 = −2. The greatest amount by which Bristol exceeded Leeds is 4 hours, on Wednesday. Choosing 1 takes Monday's smaller positive difference instead of the greatest one. Choosing 2 takes the size of Thursday's difference, but that is the amount by which Leeds exceeded Bristol, the opposite direction to the one asked for. Choosing 6 takes Bristol's raw figure on Wednesday without subtracting Leeds's 2 hours first.
- (d) 25% — First find the number of fruit cakes: 80 − 34 − 26 = 20. Then write this as a percentage of the total: 20 ÷ 80 × 100 = 25%. Giving 20% comes from reporting the count of fruit cakes, 20, directly as a percentage, without dividing by the total of 80 first. Giving 32.5% computes the percentage of chocolate cakes instead of fruit cakes: 26 ÷ 80 × 100 = 32.5%. Giving 42.5% computes the percentage of sponge cakes instead of fruit cakes: 34 ÷ 80 × 100 = 42.5%.
- (c) Yes — with an estimate of 320, above the 250 limit. — Method: scale the sample proportion up to the whole batch to get an estimate, then compare that estimate with the 250 limit to reach a decision. Working: in the sample, 4 out of 50 boards are faulty, a proportion of 4 ÷ 50 = 0.08. Applying that proportion to the batch of 4,000 gives an estimate of 0.08 × 4000 = 320 faulty boards. Since 320 is more than 250, the factory should scrap the batch. Inverting the proportion, 50 ÷ 4 = 12.5, and treating that as a percentage of the batch, 12.5% × 4000 = 500, still gives 'yes' but from the wrong fraction, so it overstates the estimate. Comparing the raw number of faulty boards found in the sample, 4, directly with the 250 limit skips the scaling up to the batch altogether, and 4 is nowhere near 250, so that route wrongly says 'no'. Dividing the batch by the sample size, 4000 ÷ 50 = 80, finds how many samples of 50 fit into the batch but stops before multiplying by the 4 faulty boards found, so it also wrongly says 'no'. Always find the proportion in the sample first, scale it up to the whole batch, and only then compare the estimate with the limit given.
- (a) 22 — Method: the two subject totals overlap, because every pupil who passed both subjects has been counted once in the maths total and once again in the science total; adding the totals therefore counts those pupils twice, and the overlap has to be taken off once. Working: 18 + 12 = 30, and the 8 pupils who passed both have been counted twice in that 30, so the number who passed at least one subject is 30 − 8 = 22. Answer: 22 pupils, a count of pupils, and it is less than the 30 in the class, which leaves 8 pupils who passed neither. The distractors: 30 comes from adding the two subject totals and never removing the overlap, so it counts the 8 pupils twice; 14 comes from taking the 8 away twice, 18 + 12 − 8 − 8, removing an overlap that was only counted twice once too often; 18 comes from writing down the larger of the two subject totals on its own, which leaves out every pupil who passed science but not maths.
- (d) Both the mean and the range increase. — The original mean is 150 + 152 + 155 + 158 + 160 = 775, and 775 ÷ 5 = 155 cm; the original range is 160 − 150 = 10 cm. Including the new height of 170 cm gives a new total of 775 + 170 = 945, and 945 ÷ 6 = 157.5 cm, which is higher than 155 cm, and a new range of 170 − 150 = 20 cm, which is higher than 10 cm, so both the mean and the range increase. Saying the range stays the same ignores that 170 cm is a new, higher maximum than the old 160 cm. Saying the mean stays the same ignores that 170 cm is above the original mean of 155 cm, which pulls the average up. Saying both decrease is the opposite of what happens here.
- (c) Height 170 cm, mass 65 kg — Method: a point on a scatter graph is written as a pair of coordinates in which the horizontal value is written first and the vertical value second, so each value is matched to the quantity named on its own axis. Working: in (170, 65) the value 170 is the horizontal coordinate and the horizontal axis shows height in centimetres, so the height is 170 cm; the value 65 is the vertical coordinate and the vertical axis shows mass in kilograms, so the mass is 65 kg. Answer: height 170 cm, mass 65 kg, each with the unit named on its own axis. The distractors: height 65 cm and mass 170 kg come from reading the pair the wrong way round, which would describe an impossible person; height 170 cm and mass 170 kg come from reading the horizontal coordinate for both quantities and never using the second number; height 235 cm and mass 105 kg come from combining the two coordinates, 170 + 65 and 170 − 65, instead of reading them separately.
- (a) Chloe's marks are far more spread out than Ben's — Method: a mean reports where a set of values sits, and two sets can sit in the same place while behaving quite differently, so a measure of spread has to be worked out as well. Working: Ben's marks add to 62 + 64 + 65 + 66 + 68 = 325 and 325 ÷ 5 = 65; Chloe's add to 40 + 52 + 65 + 78 + 90 = 325 and 325 ÷ 5 = 65, so the two means agree, as the question says. The ranges do not: Ben's is 68 − 62 = 6 marks, while Chloe's is 90 − 40 = 50 marks. Ben's five marks all sit within 3 marks of 65; Chloe's lowest is 25 marks below it and her highest 25 marks above it. Answer: Chloe's marks are far more spread out than Ben's, which is exactly what the mean cannot show. The distractors: saying Ben's marks are more spread out comes from subtracting in the order the values are written, 62 − 68 = −6 against 40 − 90 = −50, and then reading −6 as the larger spread; saying Chloe scored far more marks in total assumes a wider set of marks must add to more, when both totals are 325; saying the two sets vary by the same amount assumes that equal means force equal spread, when the two ranges are 6 and 50.
- (d) Full-time workers, as most are at work at that time — Method: a sample is biased when the method of contact makes part of the population much less likely to be reached, so test each group against where its members actually are between 10 am and 2 pm on a weekday, and test each stated reason against the facts. Working: those hours are the middle of the working day, so people in full-time employment are at work and not beside a landline telephone, while people who are retired and people who are unemployed are far more likely to be at home and are reached at the usual rate; the method therefore collects far fewer replies from full-time workers than their share of the adult population the council is consulting. Answer: full-time workers, as most are at work at that time. The distractors: the reply naming retired people rests on the false claim that most retired people are at work in the daytime, when in fact a daytime call reaches them more easily than anyone; the reply naming unemployed people rests on the false claim that they are out during the day, when they too are among the easiest people to reach by a daytime call; the reply naming children rests on the false claim that children are at home at 11 am on a Tuesday, when they are at school and so are not reached by the call at all, and school-age children are in any case not the adults whose views the council is collecting.
- (d) 6 — Method: the mode, or modal value, is the value that occurs most often in the data set, and it is a value from the data rather than a count. Working: size 4 occurs twice, size 5 occurs once, size 6 occurs three times and size 9 occurs once, so the highest frequency is three and the size it belongs to is 6. Answer: 6. The distractors: 3 comes from writing down the frequency of the most common size instead of the size itself; 9 comes from picking the largest size in the list, which confuses the mode with the maximum; 4 comes from stopping at the first size that repeats rather than checking which size repeats most often.
- (d) 648 kg — Method: to estimate a total from a sample, multiply the sample's mean by the number of items in the whole population, then check the units the question asks for. Working: 32.4 g × 20,000 = 648,000 g. Converting to kilograms, 648,000 ÷ 1,000 = 648 kg. This is only an estimate, not an exact total, because it assumes every one of the 20,000 packets has exactly the sample mean mass, when in reality individual packets vary above and below it. Giving 1.62 kg multiplies the mean by 50, the SAMPLE size, instead of by 20,000, the number of packets actually made that day — this finds the total mass of the 50 sampled packets, not the day's production. Giving 32.4 kg treats the sample mean itself, in grams, as if it already were the day's total mass in kilograms, skipping the scaling up altogether. Giving 648,000 kg correctly scales the mean up to the whole day's production but never converts the answer from grams to kilograms, leaving it 1,000 times too large. Always scale a sample's mean up by the SIZE OF THE WHOLE POPULATION, and always finish by checking the units the question asks for.
Build your own mix at the worksheet builder.