Printable · GCSE Foundation · ages 14-16
Empirical samples and sample size worksheet — GCSE Foundation
Fifteen questions on "empirical samples and sample size" — DfE statement P5. Print it, or print three versions so neighbours cannot copy by letter; the key gives the letter for each version.
Non-calculatoronly 12 unique questions available
Empirical samples and sample size worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- 1.Four pupils each flip the same coin a number of times and record the results. Sam flips it 40 times and gets 24 heads. Priti flips it 25 times and gets 10 heads. Leo flips it 20 times and gets 9 heads. Fatima flips it 15 times and gets 6 heads. Combine all four pupils' results to estimate the probability that the coin lands on heads.
- 2.A biased spinner is spun 40 times and lands on red 16 times. It is then spun a further 60 times and lands on red 21 times. Work out the best estimate of the probability that the spinner lands on red, using the results of all 100 spins together.
- 3.Ellie drops a bottle top and records whether it lands open end up. In her first 20 drops it lands open end up 13 times. She carries on, and after 200 drops in total it has landed open end up 84 times. Work out the best estimate of the probability that the bottle top lands open end up.
- 4.A drawing pin is dropped many times and lands either point up or point down. The relative frequency of landing point up is recorded as the experiment goes on: after 50 drops it is 0.720, after 200 drops it is 0.665, and after 1000 drops it is 0.638. The pin is to be dropped a further 2000 times. Work out the best estimate of the number of times it will land point up.
- 5.A machine makes 4000 light bulbs a day and runs 5 days a week. Two inspectors test bulbs from this machine. Inspector A tests 40 bulbs and finds 4 faulty. Inspector B tests 500 bulbs and finds 30 faulty. Using the better of the two estimates, work out how many faulty bulbs the machine is expected to make in one week.
- 6.A six-sided dice is rolled 12 times. It lands on a six 4 times. Erin says that the dice must be biased. Which of these is the best comment on Erin's statement?
- 7.Priya and Ben each want to estimate the probability that a coin lands on heads. Priya flips the coin 50 times and works out her relative frequency. Ben flips the same coin 500 times and works out his relative frequency. Whose relative frequency is more likely to be close to the true probability? Give a reason for your answer.
- 8.A fair coin is flipped again and again. After the first 10 flips there have been 7 heads. After 1000 flips there have been 528 heads. Which statement best describes what these results show?
- 9.A market stall sells umbrellas. Over the last 250 days, it rained on 70 of them. Using this as an estimate of the probability of rain, work out how many rainy days would be expected in the next 365 days.
- 10.A biased spinner is spun 200 times. It lands on red 70 times, on blue 50 times, and on green 80 times. Using these results, work out the expected number of times the spinner does NOT land on red, in 500 spins of the same spinner.
- 11.In a class experiment a fair six-sided dice is rolled 60 times and lands on a six 14 times. The whole school then rolls the same fair dice 3000 times. Work out the best estimate of the number of sixes the school should expect.
- 12.A student is estimating the probability that a spinner lands on red. In her first 85 spins it landed on red 34 times. She then spins it 165 more times, and in those it lands on red 58 times. Work out the best estimate of the probability of red from all 250 spins. Give your answer as a decimal, correct to 3 decimal places.
Answer key
- (b) 0.49 — Adding all four pupils' flips gives 40 + 25 + 20 + 15 = 100 flips in total, and adding their heads gives 24 + 10 + 9 + 6 = 49 heads in total, so the combined estimate is 49/100 = 0.49. Writing 0.46 is wrong because it averages the four pupils' individual rates (0.60, 0.40, 0.45 and 0.40) as if they all came from the same number of flips, which they do not — this ignores that Sam and Priti flipped far more times than Leo and Fatima. Writing 0.60 is wrong because it only uses Sam's own result (24/40 = 0.60), ignoring the other three pupils. Writing 0.40 is wrong because it only combines Priti and Fatima's results (16/40 = 0.40), leaving out Sam and Leo entirely. The combined estimate from all four pupils is 0.49.
- (c) 0.37 — Method: pool the two runs into one combined set of results, then find the relative frequency of red across all of the spins together. Working: total reds = 16 + 21 = 37. Total spins = 40 + 60 = 100. Relative frequency = 37 ÷ 100 = 0.37. Answer: 0.37. Watch out: writing down 0.40 uses only the first run, 16 ÷ 40, and throws away the extra evidence from the second 60 spins. Writing down 0.35 uses only the second run, 21 ÷ 60, and throws away the first run instead. And writing down 0.375 averages the two runs' separate rates, (0.40 + 0.35) ÷ 2, which treats a run of 40 spins and a run of 60 spins as equally weighted, when pooling the actual counts gives the larger run its fair share of influence.
- (d) 0.420 — Method: use the relative frequency worked out from the larger number of trials as the best estimate of the probability, since a bigger sample tends to sit closer to the true probability. Working: 84 out of 200 drops land open end up, so the relative frequency from the total is 84 ÷ 200 = 0.420. Answer: 0.420. Watch out: writing down 0.650 comes from 13 ÷ 20, using only the first, much smaller sample instead of the total. Writing down 0.535 comes from averaging 0.650 and 0.420, treating the 20-drop run and the 200-drop run as equally reliable instead of using the larger sample on its own. And writing down 0.580 comes from 1 − 0.420, working out the probability that the bottle top lands the other way up instead of open end up.
- (a) 1276 — Method: an unbiased relative frequency tends towards the theoretical probability as the number of trials increases, so use the record resting on the most trials, then multiply by the number of new trials. Working: the three records rest on 50, 200 and 1000 drops, so the most reliable is the one after 1000 drops, namely 0.638, and the run is indeed settling as the trials increase. The expected number of point up landings in 2000 further drops is 2000 × 0.638 = 1276. Answer: about 1276 times. The distractors: 1440 uses the earliest record, which rests on only 50 drops, giving 2000 × 0.720 = 1440; 1330 uses the middle record, treating 200 drops as a safe compromise when 1000 drops is better still, giving 2000 × 0.665 = 1330; 1348 comes from averaging the three records, since 0.720 + 0.665 + 0.638 = 2.023 and 2.023 ÷ 3 = 0.674, then 2000 × 0.674 = 1348, which gives the 50 drop record the same weight as the 1000 drop record.
- (a) 1200 — Method: take the estimate from the larger sample, because an unbiased relative frequency tends towards the true probability as the sample grows, then multiply by the number of bulbs made in a week. Working: Inspector B tested 500 bulbs, far more than Inspector A's 40, so use B's relative frequency: 30 ÷ 500 = 0.06. A week's production is 4000 × 5 = 20000 bulbs. The expected number of faulty bulbs is 20000 × 0.06 = 1200. Answer: about 1200 faulty bulbs a week. The distractors: 2000 uses Inspector A's estimate, 4 ÷ 40 = 0.1, giving 20000 × 0.1 = 2000, and so rests on a sample of only 40 bulbs; 1600 comes from averaging the two estimates of 0.1 and 0.06 to get 0.08, and 20000 × 0.08 = 1600, which gives the small sample equal weight with the large one; 240 uses the right estimate but stops at a single day, 4000 × 0.06 = 240.
- (b) She is wrong; 12 rolls is too few to judge — Method: compare the result with what is expected, then ask whether the experiment is long enough for a difference to mean anything. Working: if the dice were fair the expected number of sixes in 12 rolls is 12 × 1 ÷ 6 = 2, so 4 sixes is 2 above what was expected. But over only 12 rolls a result like this turns up often by chance: the relative frequency here is 4/12, which is 1/3, and over so few trials a relative frequency can sit well away from 1/6 with no bias at all. Answer: Erin is wrong, because 12 rolls is far too few to decide; she should roll the dice many more times and see whether the relative frequency settles near 1/6. The distractors: saying she is right because 4 beats the expected 2 uses the correct expected value but treats any difference as proof, which so short an experiment cannot give; saying a fair dice gives each score twice in 12 rolls treats an expected value as a guaranteed one; saying that 4 sixes in 12 rolls cancels down to 1 in 6 mis-cancels the fraction, because 4/12 is 1/3, which is twice 1/6, so that comment reaches the right verdict from arithmetic that is wrong.
- (a) Ben, because a larger sample is closer to the theory — Method: a relative frequency is an estimate of a probability, and for an unbiased experiment that estimate tends towards the theoretical value as the sample grows. Working: Priya's estimate rests on 50 results, so a few unexpected heads move it a long way; one extra head shifts her relative frequency by 1 ÷ 50 = 0.02. Ben's estimate rests on 500 results, where one extra head shifts his relative frequency by only 1 ÷ 500 = 0.002. The larger sample therefore swings far less around the true value. Answer: Ben's relative frequency is the one more likely to be close, because a larger unbiased sample tends closer to the theoretical probability. The distractors: saying a small sample is less affected by luck reverses the result, since it is the small sample that swings most; saying every flip is a separate random event is true of the flips themselves but says nothing about the estimates, and is often used to argue wrongly that the number of trials does not matter; saying 500 flips must give exactly 250 heads confuses an expected value with a guaranteed one, and 500 flips very rarely give exactly 250 heads.
- (b) The relative frequency is settling near 0.5 — Method: turn each result into a relative frequency before comparing them, because it is the relative frequency, and not the difference between the two counts, that tends towards the theoretical probability. Working: after 10 flips the relative frequency of a head is 7 ÷ 10 = 0.7, which is a long way from 0.5. After 1000 flips it is 528 ÷ 1000 = 0.528, which is much closer to 0.5. Meanwhile the gap between the two counts has grown rather than shrunk: it was 7 − 3 = 4 after 10 flips and is 528 − 472 = 56 after 1000 flips. Answer: the relative frequency is settling near 0.5, which is what an unbiased experiment does as the sample grows. The distractors: saying the counts are levelling out is the usual form of this idea and the figures contradict it, since the gap went from 4 to 56; saying the coin is biased treats 28 extra heads in 1000 flips as proof, when 0.528 sits close to 0.5 and a fair coin gives results like this often; saying the next flip is more likely to be a tail is the gambler's fallacy, since each flip stays at 1/2 whatever came before.
- (c) 102 — Method: turn the past record into a relative frequency, then use it as an estimate of the probability of rain and multiply by the number of days being predicted for. Working: relative frequency of rain = 70 ÷ 250 = 0.28. Expected rainy days in 365 days = 365 × 0.28 = 102.2, which rounds to about 102 days. Answer: about 102 days. Watch out: writing down 48 swaps which number is the sample and which is the target, working out 70 ÷ 365 × 250 instead of 70 ÷ 250 × 365. Writing down 70 just repeats the original count of rainy days without scaling it up to the new, longer period at all. And writing down 110 comes from rounding the relative frequency to 0.3 before multiplying, 365 × 0.3 = 109.5, when 70 ÷ 250 is exactly 0.28 and needs no rounding at all.
- (c) 325 — Method: first find the relative frequency of NOT landing on red from the 200 spins, then scale that up to 500 spins. Working: non-red results = 50 + 80 = 130, out of 200 spins, so P(not red) = 130 ÷ 200 = 0.65. Expected non-red results in 500 spins = 500 × 0.65 = 325. Answer: 325. Watch out: writing down 175 finds the expected number of RED results instead, 70 ÷ 200 × 500 = 175, answering the opposite of what was asked. Writing down 250 assumes landing red or not landing red must be a fair 50-50 split, but the spinner is biased and the actual results do not split evenly. And writing down 130 stops after finding how many of the 200 spins were non-red and forgets to scale that figure up to the 500 spins asked for.
- (b) 500 — Method: when a dice is known to be fair, the theoretical probability is the best thing to work from, and the more trials there are the closer the results tend to it. Working: for a fair dice the probability of a six is 1/6, so the expected number of sixes in 3000 rolls is 3000 × 1 ÷ 6 = 500. The class experiment gave a relative frequency of 14/60, but 60 trials is far too few to overturn a known theoretical value, and the school's 3000 rolls will tend towards 1/6 in any case. Answer: about 500 sixes. The distractors: 700 comes from using the class relative frequency instead of the theory, 3000 × 14 ÷ 60 = 700; 600 comes from splitting the difference between the two, since 1/6 is about 0.167 and 14/60 is about 0.233, whose mean is 0.2, and 3000 × 0.2 = 600; 2500 uses 5/6 instead of 1/6 and counts the rolls expected not to be a six.
- (a) 0.368 — Combining both samples, the spinner landed on red 34 + 58 = 92 times out of a total of 85 + 165 = 250 spins, so the best estimate of the probability is 92/250 = 0.368. Writing 0.400 is wrong because it uses only the first sample, 34/85 = 0.400, ignoring the extra 165 spins recorded afterwards. Writing 0.352 is wrong because it uses only the second sample, 58/165 = 0.352 (to 3 decimal places), ignoring the first 85 spins. Writing 0.376 is wrong because it averages the two separate estimates, (0.400 + 0.352) ÷ 2 = 0.376, instead of combining the actual numbers of reds and spins across both samples. The best estimate of the probability that the spinner lands on red, using all 250 spins, is 0.368.
Build your own mix at the worksheet builder.