Printable · GCSE Higher · ages 14-16
Empirical samples and sample size worksheet — GCSE Higher
Fifteen questions on "empirical samples and sample size" — DfE statement P5. Print it, or print three versions so neighbours cannot copy by letter; the key gives the letter for each version.
Calculatoronly 9 unique questions available
Answer key: Empirical samples and sample size worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- (a) 1200 — Method: take the estimate from the larger sample, because an unbiased relative frequency tends towards the true probability as the sample grows, then multiply by the number of bulbs made in a week. Working: Inspector B tested 500 bulbs, far more than Inspector A's 40, so use B's relative frequency: 30 ÷ 500 = 0.06. A week's production is 4000 × 5 = 20000 bulbs. The expected number of faulty bulbs is 20000 × 0.06 = 1200. Answer: about 1200 faulty bulbs a week. The distractors: 2000 uses Inspector A's estimate, 4 ÷ 40 = 0.1, giving 20000 × 0.1 = 2000, and so rests on a sample of only 40 bulbs; 1600 comes from averaging the two estimates of 0.1 and 0.06 to get 0.08, and 20000 × 0.08 = 1600, which gives the small sample equal weight with the large one; 240 uses the right estimate but stops at a single day, 4000 × 0.06 = 240.
- (a) 1276 — Method: an unbiased relative frequency tends towards the theoretical probability as the number of trials increases, so use the record resting on the most trials, then multiply by the number of new trials. Working: the three records rest on 50, 200 and 1000 drops, so the most reliable is the one after 1000 drops, namely 0.638, and the run is indeed settling as the trials increase. The expected number of point up landings in 2000 further drops is 2000 × 0.638 = 1276. Answer: about 1276 times. The distractors: 1440 uses the earliest record, which rests on only 50 drops, giving 2000 × 0.720 = 1440; 1330 uses the middle record, treating 200 drops as a safe compromise when 1000 drops is better still, giving 2000 × 0.665 = 1330; 1348 comes from averaging the three records, since 0.720 + 0.665 + 0.638 = 2.023 and 2.023 ÷ 3 = 0.674, then 2000 × 0.674 = 1348, which gives the 50 drop record the same weight as the 1000 drop record.
- (b) 500 — Method: when a dice is known to be fair, the theoretical probability is the best thing to work from, and the more trials there are the closer the results tend to it. Working: for a fair dice the probability of a six is 1/6, so the expected number of sixes in 3000 rolls is 3000 × 1 ÷ 6 = 500. The class experiment gave a relative frequency of 14/60, but 60 trials is far too few to overturn a known theoretical value, and the school's 3000 rolls will tend towards 1/6 in any case. Answer: about 500 sixes. The distractors: 700 comes from using the class relative frequency instead of the theory, 3000 × 14 ÷ 60 = 700; 600 comes from splitting the difference between the two, since 1/6 is about 0.167 and 14/60 is about 0.233, whose mean is 0.2, and 3000 × 0.2 = 600; 2500 uses 5/6 instead of 1/6 and counts the rolls expected not to be a six.
- (c) 102 — Method: turn the past record into a relative frequency, then use it as an estimate of the probability of rain and multiply by the number of days being predicted for. Working: relative frequency of rain = 70 ÷ 250 = 0.28. Expected rainy days in 365 days = 365 × 0.28 = 102.2, which rounds to about 102 days. Answer: about 102 days. Watch out: writing down 48 swaps which number is the sample and which is the target, working out 70 ÷ 365 × 250 instead of 70 ÷ 250 × 365. Writing down 70 just repeats the original count of rainy days without scaling it up to the new, longer period at all. And writing down 110 comes from rounding the relative frequency to 0.3 before multiplying, 365 × 0.3 = 109.5, when 70 ÷ 250 is exactly 0.28 and needs no rounding at all.
- (c) 0.37 — Method: pool the two runs into one combined set of results, then find the relative frequency of red across all of the spins together. Working: total reds = 16 + 21 = 37. Total spins = 40 + 60 = 100. Relative frequency = 37 ÷ 100 = 0.37. Answer: 0.37. Watch out: writing down 0.40 uses only the first run, 16 ÷ 40, and throws away the extra evidence from the second 60 spins. Writing down 0.35 uses only the second run, 21 ÷ 60, and throws away the first run instead. And writing down 0.375 averages the two runs' separate rates, (0.40 + 0.35) ÷ 2, which treats a run of 40 spins and a run of 60 spins as equally weighted, when pooling the actual counts gives the larger run its fair share of influence.
- (b) The relative frequency is settling near 0.5 — Method: turn each result into a relative frequency before comparing them, because it is the relative frequency, and not the difference between the two counts, that tends towards the theoretical probability. Working: after 10 flips the relative frequency of a head is 7 ÷ 10 = 0.7, which is a long way from 0.5. After 1000 flips it is 528 ÷ 1000 = 0.528, which is much closer to 0.5. Meanwhile the gap between the two counts has grown rather than shrunk: it was 7 − 3 = 4 after 10 flips and is 528 − 472 = 56 after 1000 flips. Answer: the relative frequency is settling near 0.5, which is what an unbiased experiment does as the sample grows. The distractors: saying the counts are levelling out is the usual form of this idea and the figures contradict it, since the gap went from 4 to 56; saying the coin is biased treats 28 extra heads in 1000 flips as proof, when 0.528 sits close to 0.5 and a fair coin gives results like this often; saying the next flip is more likely to be a tail is the gambler's fallacy, since each flip stays at 1/2 whatever came before.
- (a) 753 — The estimate from 2000 spins is the most reliable, since it comes from the largest sample size, so the best estimate of the probability is 0.251. Over a further 3000 spins, the expected number landing on green is 3000 × 0.251 = 753. Writing 1050 is wrong because 3000 × 0.350 = 1050 uses the estimate from only 20 spins, the LEAST reliable of the three. Writing 870 is wrong because 3000 × 0.290 = 870 uses the estimate from 200 spins rather than the more reliable 2000-spin estimate. Writing 750 is wrong because 3000 × 0.25 = 750 ignores the recorded data completely and simply assumes each of the 4 colours is equally likely. The best estimate is 753 expected green spins.
- (d) 3400 — Combining all three greenhouses gives 200 + 150 + 250 = 600 seeds planted in total, and 172 + 126 + 212 = 510 germinated, so the combined estimate of the germination probability is 510/600 = 0.85. Out of a new batch of 4000 seeds, the expected number to germinate is 4000 × 0.85 = 3400. Writing 3440 is wrong because it uses only Greenhouse 1's rate, 172/200 = 0.86, instead of the combined rate from all three: 4000 × 0.86 = 3440. Writing 3360 is wrong because it uses only Greenhouse 2's rate, 126/150 = 0.84: 4000 × 0.84 = 3360. Writing 510 is wrong because that is the total number that germinated in the ORIGINAL trial, not scaled up to the new batch of 4000 seeds at all. The best estimate is 3400 seeds.
- (c) 325 — Method: first find the relative frequency of NOT landing on red from the 200 spins, then scale that up to 500 spins. Working: non-red results = 50 + 80 = 130, out of 200 spins, so P(not red) = 130 ÷ 200 = 0.65. Expected non-red results in 500 spins = 500 × 0.65 = 325. Answer: 325. Watch out: writing down 175 finds the expected number of RED results instead, 70 ÷ 200 × 500 = 175, answering the opposite of what was asked. Writing down 250 assumes landing red or not landing red must be a fair 50-50 split, but the spinner is biased and the actual results do not split evenly. And writing down 130 stops after finding how many of the 200 spins were non-red and forgets to scale that figure up to the 500 spins asked for.
Build your own mix at the worksheet builder.