Printable · GCSE Higher · ages 14-16
Sampling and inference about populations worksheet — GCSE Higher
Fifteen questions on "sampling and inference about populations" — DfE statement S1. Print it, or print three versions so neighbours cannot copy by letter; the key gives the letter for each version.
Sampling and inference about populations worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- 1.A school has 1,000 pupils. A student wants to estimate how many of them walk to school, so she asks 8 pupils in her own class. She says her sample is large enough to give a reliable estimate for the whole school. Is she right? Give a reason for your answer.
- 2.A factory makes a batch of 4,000 circuit boards. It checks a random sample of 50 boards and finds that 4 are faulty. The factory will scrap the whole batch if the estimated number of faulty boards in the batch is more than 250. Should the factory scrap the batch?
- 3.A random sample of 50 pupils at a school were asked whether they prefer sport to music. 30 of the 50 pupils said they prefer sport. The school has 500 pupils altogether. Work out an estimate for the number of the 500 pupils who prefer sport.
- 4.A charity shop in Bath holds 2,000 books. Volunteer A checks a random sample of 50 books and finds 35 paperbacks. Volunteer B checks a different random sample of 50 books and finds 31 paperbacks. Work out the estimate each sample gives for the whole stock, and write down what the shop should do next.
- 5.A town council wants to find out about the eating habits of the people who live in the town. It asks only the members of a local sports club. Give a reason why this sample is biased.
- 6.In a random sample of 40 pupils at a school, 6 are left-handed. The school has 900 pupils. Work out an estimate for the number of left-handed pupils in the school.
- 7.A school has 1200 pupils. A teacher wants to take a random sample of 60 of them. Write down which of these methods gives a random sample.
- 8.A factory made 3,000 phone cases last week. Shift A checked a random sample of 100 cases and found that 34 were scratched. Shift B checked a different random sample of 50 cases and found that 21 were scratched. Using the COMBINED results from both shifts, work out an estimate for the number of scratched cases made last week.
- 9.A sports centre in Ipswich has 2,000 members. It wants to know what its members think of its opening hours, so it asks a random sample of 100 of them. Write down what the population is in this survey.
- 10.A school has 1,500 pupils. The head teacher takes a random sample of 150 of them from the school register and asks how long they spend on homework. Rory says the sample is too small for the result to mean anything. Is Rory right? Give a reason for your answer.
- 11.A factory makes 4,000 light bulbs a day. In a random sample of 80 of one day's bulbs, 3 were faulty. Work out an estimate for the number of faulty bulbs the factory makes in a day.
- 12.A council wants to find out what all residents of a town think about a new cycle lane. It posts a survey only on its website and asks people to fill it in online. Give a reason why this sample is likely to be biased.
- 13.There are 10,000 pupils in a city. A researcher picks a sample of 500 of them by drawing names at random from a list of every pupil in the city. Give the reason why this method gives a representative sample.
- 14.A quality inspector weighs a random sample of 50 packets of crisps from one day's production and finds their mean mass is 32.4 g. The factory makes 20,000 packets that day. Work out an estimate for the total mass, in kg, of all the packets made that day.
- 15.A council in Leeds wants to know what local people think about letting shops stay open later in the evening. It rings landline telephone numbers between 10 am and 2 pm on a Tuesday. Write down which group is most likely to be under-represented in the sample, and give a reason for your answer.
Answer key
- (c) No — 8 from one class is too small to represent the school. — Method: judge reliability by asking whether the sample is both large enough, and spread across the population, relative to what it is meant to represent. Working: 8 pupils is a tiny fraction of the school's 1,000 pupils, and all 8 come from a single class rather than a range of year groups, so the sample is both too small and too narrow to represent the whole school reliably. She is not right. Saying any sample size gives an equally reliable estimate ignores that reliability generally improves with a larger, more representative sample. Saying the method is unreliable because it was not done online is not a reason connected to sample size or representativeness at all. Saying 8 is reliable because it is more than half her class compares the sample to the wrong population — the school has 1,000 pupils, not one class. Always judge a sample's size against the population it is meant to represent, not against a smaller group within it.
- (c) Yes — with an estimate of 320, above the 250 limit. — Method: scale the sample proportion up to the whole batch to get an estimate, then compare that estimate with the 250 limit to reach a decision. Working: in the sample, 4 out of 50 boards are faulty, a proportion of 4 ÷ 50 = 0.08. Applying that proportion to the batch of 4,000 gives an estimate of 0.08 × 4000 = 320 faulty boards. Since 320 is more than 250, the factory should scrap the batch. Inverting the proportion, 50 ÷ 4 = 12.5, and treating that as a percentage of the batch, 12.5% × 4000 = 500, still gives 'yes' but from the wrong fraction, so it overstates the estimate. Comparing the raw number of faulty boards found in the sample, 4, directly with the 250 limit skips the scaling up to the batch altogether, and 4 is nowhere near 250, so that route wrongly says 'no'. Dividing the batch by the sample size, 4000 ÷ 50 = 80, finds how many samples of 50 fit into the batch but stops before multiplying by the 4 faulty boards found, so it also wrongly says 'no'. Always find the proportion in the sample first, scale it up to the whole batch, and only then compare the estimate with the limit given.
- (d) 300 pupils — Method: an estimate for a whole population is made by finding the proportion in the sample and applying that same proportion to the population. Working: in the sample 30 of the 50 pupils prefer sport, a proportion of 30 ÷ 50 = 0.6, and applying that proportion to the school gives 0.6 × 500 = 300 pupils. Answer: 300 pupils, and it is only an estimate, because a different random sample of 50 would give a slightly different figure. The distractors: 200 pupils comes from scaling up the 20 pupils in the sample who did not prefer sport, 20 × 10, which answers the opposite question; 150 pupils comes from reading 30 out of 50 as 30% and taking 30% of 500; 60 pupils comes from working out the proportion correctly as 60% and then writing the 60 down as a number of pupils instead of applying it to the 500.
- (a) 1,400 and 1,240, so combine the samples for one estimate — Method: scale each sample up to the whole stock, then use the fact that a larger sample gives a more reliable estimate than a smaller one. Working: the first sample gives 35 ÷ 50 = 0.7 and 0.7 × 2,000 = 1,400 paperbacks; the second gives 31 ÷ 50 = 0.62 and 0.62 × 2,000 = 1,240 paperbacks. Two random samples of the same size are expected to differ a little, so neither estimate is wrong. Putting the two together gives 35 + 31 = 66 paperbacks in 100 books, and 66 ÷ 100 = 0.66 with 0.66 × 2,000 = 1,320, an estimate resting on twice as many books as either volunteer checked. Answer: 1,400 and 1,240, so combine the samples for one estimate. The distractors: keeping 1,400 because it is larger picks an estimate by its size, when both samples held 50 books and neither has a stronger claim; saying a volunteer must have miscounted assumes two random samples ought to agree exactly, which is precisely what random sampling does not promise; 1,750 and 1,550 come from 35 × 50 = 1,750 and 31 × 50 = 1,550, multiplying each count by the size of the sample instead of scaling by 2,000 ÷ 50.
- (c) Club members probably eat differently from most people — Method: a sample is biased when the group it is drawn from differs from the population in the very thing the survey is measuring, so compare the subgroup with the population on that quantity. Working: the survey measures eating habits, and people who join a sports club take more exercise than average and are known to eat differently from the town as a whole, so their replies pull the results away from the true picture for the town however many of them are asked. Answer: club members probably eat differently from most people. The distractors: the reply about the number of members treats bias as a question of size, but a large biased sample is still biased; the reply that the members were picked at random is false, since the council picked a club rather than picking residents, and it confuses bias with non-response; the reply that everyone asked lives in the town notes something true of the members but draws the false conclusion that the sample therefore covers the town, when a sample must reflect a population and not merely be taken from inside it.
- (b) 135 — Method: use the sample to find the PROPORTION of left-handed pupils, then apply that same proportion to the whole school population. Working: in the sample, 6 out of 40 pupils are left-handed, a proportion of 6 ÷ 40 = 0.15. Applying that proportion to the school's 900 pupils gives an estimate of 0.15 × 900 = 135 pupils. Giving 6 simply repeats the number of left-handed pupils IN THE SAMPLE, without scaling up to the whole school at all. Multiplying the population by the number of left-handed pupils in the sample without first dividing by the sample size, 900 × 6 = 5400, badly overestimates — that is more pupils than the whole school has. Dividing the population by the sample size but forgetting to multiply by the number of left-handed pupils found, 900 ÷ 40 = 22.5, finds the scale factor but stops one step short of using it. Always find the proportion in the sample first, then scale that same proportion up to the population.
- (a) Drawing 60 names at random from a list of all 1200 pupils — Method: a sample is random when every member of the population has the same chance of being chosen and nobody, including the pupils themselves, can influence who ends up in it; test each method against that. Working: drawing names from a list of all 1200 pupils gives each pupil the same chance, 60 out of 1200, whatever their year group, class or opinion, so the method is random. Answer: drawing 60 names at random from a list of all 1200 pupils. The distractors: asking the pupils who volunteer is self-selection, and the pupils with the strongest views volunteer first, so they decide the sample; asking the pupils nearest the door is convenience sampling, which reaches only those who happen to be in one place at one time; asking two Year 10 classes samples a cluster, so every pupil in the other year groups has no chance of being chosen at all.
- (d) 1100 — Method: to combine two samples of different sizes, add the faulty counts together and add the sample sizes together before scaling up, rather than treating the two samples separately. Working: the combined sample found 34 + 21 = 55 scratched cases out of 100 + 50 = 150 cases checked, a proportion of 55 ÷ 150. Applying that proportion to the week's production of 3,000 gives an estimate of 55 ÷ 150 × 3000 = 1100 scratched cases. Averaging the two shifts' proportions instead of combining their totals, (34 ÷ 100 + 21 ÷ 50) ÷ 2 = 0.38, gives 0.38 × 3000 = 1140 — this treats the two samples as equally weighted even though Shift A checked twice as many cases as Shift B. Using only Shift A's sample, 34 ÷ 100 × 3000 = 1020, ignores Shift B's cases completely. Using only Shift B's sample, 21 ÷ 50 × 3000 = 1260, ignores Shift A's cases completely. When two samples are different sizes, combine their totals before finding the proportion — do not average the two proportions, and do not use only one shift's sample.
- (b) All 2,000 members of the sports centre. — Method: in a survey, the population is the whole group the survey is trying to find out about, and the sample is the smaller group actually asked. Working: this survey wants to know what the sports centre's members think, so the population is every one of the 2,000 members — whether or not they were personally asked. Saying the population is the 100 members who were asked names the sample, not the population; the sample is drawn FROM the population, so it is smaller than it, not the same as it. Saying the population is everybody who lives in Ipswich widens the group far beyond who the survey is actually about — plenty of Ipswich residents are not members of the sports centre at all, so they are outside this survey altogether. Saying the population is the members who say they are unhappy confuses the population with a result of the survey: whether a member turns out to be happy or unhappy is something the survey finds out, not part of the definition of who is being studied. The population is always the whole group the question is about, before any sampling or any results come in.
- (a) No, 150 pupils are a tenth of the school, chosen at random — Method: judge a sample on two things, whether every member of the population had the same chance of being chosen, and whether the sample is large enough to carry a pattern. Working: the 150 pupils were drawn from the register of every pupil in the school, so no year group or set is shut out and no pupil chooses to take part; and 150 ÷ 1,500 = 0.1, so one pupil in ten has been asked. A random sample of that share is ample for an estimate of how long the school's pupils spend on homework. Answer: no, because 150 pupils are a tenth of the school and were chosen at random. The distractors: saying a random sample always gives the exact school figure reaches the same verdict for a reason that is false, since a second random sample of 150 would give a slightly different mean; saying 150 pupils cannot be picked at random from 1,500 treats randomness as something only a whole population can have, when drawing names from the register is exactly how a random sample is taken; saying that only asking all 1,500 could show anything rejects sampling altogether, which would leave no way to study any population too large to count.
- (a) 150 bulbs — Method: assume the proportion faulty in a random sample is the proportion faulty in the whole day's output, and scale the sample up to the population. Working: the sample of 80 has to be scaled up to 4,000 bulbs, and 4,000 ÷ 80 = 50, so the day's output is 50 sample-sized batches. Each batch is expected to contain the same 3 faulty bulbs, so the estimate is 3 × 50 = 150. Answer: 150 bulbs, and it is an estimate, because another sample of 80 would probably contain a different number of faulty bulbs. The distractors: 50 bulbs is the scale factor 4,000 ÷ 80 written down as though it were the answer, so it reports how many batches there are rather than how many faulty bulbs; 120 bulbs comes from reading 3 out of 80 as 3%, then taking 0.03 × 4,000 = 120, but 3 out of 80 is 3.75%; 240 bulbs comes from 3 × 80 = 240, multiplying the faulty bulbs by the size of the sample instead of by the scale factor, which uses the 80 twice and the 4,000 not at all.
- (d) Only internet users reach the website; others are excluded. — Method: a sample is biased when it systematically leaves out part of the population, or systematically over-represents another part. Working: anyone without internet access, or who does not visit the council's website, has NO chance of being included — the sample is drawn only from internet-using residents, which is not the whole town. Saying too many people might respond because the survey is free confuses bias with sample size — bias is about who CAN be reached, not how many respond. Saying people might lie describes a different problem, response honesty, not who was sampled in the first place. Saying online surveys cannot be anonymous is not a reason connected to bias at all. A sample is biased when part of the population has no chance of being included, whatever the reason for that.
- (c) Because every pupil has an equal chance of being picked — Method: whether a sample represents its population is decided by the selection method, not by the size of the sample, so ask whether the method gives every member of the population the same chance of being chosen. Working: the names are drawn at random from a list of all 10,000 pupils, so each pupil has the same chance, 500 out of 10,000, of being drawn, and no group of pupils is more likely to appear than any other; that is what keeps bias out of the sample. Answer: because every pupil has an equal chance of being picked. The distractors: the reply about 5% treats the sampling fraction as the test of fairness, but a badly chosen 5% is still biased and a well chosen 1% is not; the reply about 500 being large enough makes size the test instead, which is the same mistake in another form, since a large sample drawn from one school would still misrepresent the city; the reply about the most willing pupils describes self-selection, which hands the choice of who is in the sample to the pupils who feel most strongly about the question.
- (d) 648 kg — Method: to estimate a total from a sample, multiply the sample's mean by the number of items in the whole population, then check the units the question asks for. Working: 32.4 g × 20,000 = 648,000 g. Converting to kilograms, 648,000 ÷ 1,000 = 648 kg. This is only an estimate, not an exact total, because it assumes every one of the 20,000 packets has exactly the sample mean mass, when in reality individual packets vary above and below it. Giving 1.62 kg multiplies the mean by 50, the SAMPLE size, instead of by 20,000, the number of packets actually made that day — this finds the total mass of the 50 sampled packets, not the day's production. Giving 32.4 kg treats the sample mean itself, in grams, as if it already were the day's total mass in kilograms, skipping the scaling up altogether. Giving 648,000 kg correctly scales the mean up to the whole day's production but never converts the answer from grams to kilograms, leaving it 1,000 times too large. Always scale a sample's mean up by the SIZE OF THE WHOLE POPULATION, and always finish by checking the units the question asks for.
- (d) Full-time workers, as most are at work at that time — Method: a sample is biased when the method of contact makes part of the population much less likely to be reached, so test each group against where its members actually are between 10 am and 2 pm on a weekday, and test each stated reason against the facts. Working: those hours are the middle of the working day, so people in full-time employment are at work and not beside a landline telephone, while people who are retired and people who are unemployed are far more likely to be at home and are reached at the usual rate; the method therefore collects far fewer replies from full-time workers than their share of the adult population the council is consulting. Answer: full-time workers, as most are at work at that time. The distractors: the reply naming retired people rests on the false claim that most retired people are at work in the daytime, when in fact a daytime call reaches them more easily than anyone; the reply naming unemployed people rests on the false claim that they are out during the day, when they too are among the easiest people to reach by a daytime call; the reply naming children rests on the false claim that children are at home at 11 am on a Tuesday, when they are at school and so are not reached by the call at all, and school-age children are in any case not the adults whose views the council is collecting.
Build your own mix at the worksheet builder.