Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Answer key: Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- (c) Club members probably eat differently from most people — Method: a sample is biased when the group it is drawn from differs from the population in the very thing the survey is measuring, so compare the subgroup with the population on that quantity. Working: the survey measures eating habits, and people who join a sports club take more exercise than average and are known to eat differently from the town as a whole, so their replies pull the results away from the true picture for the town however many of them are asked. Answer: club members probably eat differently from most people. The distractors: the reply about the number of members treats bias as a question of size, but a large biased sample is still biased; the reply that the members were picked at random is false, since the council picked a club rather than picking residents, and it confuses bias with non-response; the reply that everyone asked lives in the town notes something true of the members but draws the false conclusion that the sample therefore covers the town, when a sample must reflect a population and not merely be taken from inside it.
- (c) 76 — Method: the frequency of each class is the area of its bar, frequency density × class width, so work out all three frequencies and add them. Working: the widths are 5, 10 and 25 kg, so the frequencies are 4 × 5 = 20, 2.6 × 10 = 26 and 1.2 × 25 = 30; the total is 20 + 26 + 30 = 76. Answer: 76 dogs. The distractors: 7.8 comes from adding the three frequency densities, 4 + 2.6 + 1.2, treating each height as though it were a count; 107 comes from using each upper class boundary as the width, giving 4 × 5, 2.6 × 15 and 1.2 × 40; 39 comes from using the first class width, 5, for every bar, which ignores the unequal intervals and gives 20 + 13 + 6.
- (a) 3/10 — There are 60 − 9 = 51 right-handed pupils. Of the 24 pupils who prefer Art, 6 are left-handed, so 24 − 6 = 18 are right-handed and prefer Art. The probability that a randomly chosen pupil is right-handed and prefers Art is 18/60, which simplifies to 3/10. Giving 2/5 is 24/60 simplified — the probability of preferring Art, ignoring the right-handed condition entirely. Giving 17/20 is 51/60 simplified — the probability of being right-handed, ignoring the Art condition entirely. Giving 1/10 is 6/60 simplified — the probability of being left-handed and preferring Art, the wrong hand condition.
- (b) 6 — Method: frequency density = frequency ÷ class width. Working: the class 12 ≤ h < 18 has width 18 − 12 = 6, so frequency density = 36 ÷ 6 = 6. Answer: the frequency density is 6 seedlings per cm. Watch which numbers you use: taking the lower bound, 12, as the width instead of 18 − 12 = 6 gives 36 ÷ 12 = 3; dividing the total number of seedlings, 90, rather than this class's frequency, 36, by the width gives 90 ÷ 6 = 15, a density that belongs to no single class; and multiplying instead of dividing gives 36 × 6 = 216, far too large a density for so narrow a class.
- (d) 4, 3, 2, 6 — Method: frequency density = frequency ÷ class width for each class in turn; do not assume the classes are all the same width. Working: the four classes have widths 10 − 0 = 10, 30 − 10 = 20, 45 − 30 = 15 and 50 − 45 = 5. Dividing each frequency by its own width gives 40 ÷ 10 = 4, 60 ÷ 20 = 3, 30 ÷ 15 = 2 and 30 ÷ 5 = 6. Answer: the frequency densities, in order, are 4, 3, 2 and 6. Watch the width of each class separately: treating the last class as if it were also 10 units wide, like the first, gives 30 ÷ 10 = 3 instead of 30 ÷ 5 = 6 — the classes here are deliberately unequal, so no width can be borrowed from another class; dividing the width by the frequency instead of the frequency by the width for the third class gives 15 ÷ 30 = 0.5 in place of 2, the formula the wrong way round; and reading the frequency column straight off the table, 40, 60, 30, 30, skips the division by width altogether and reports how many fish are in each class rather than how densely packed each bar is.
- (c) 40 minutes — Method: for grouped data, estimate the mean using the midpoint of each class — multiply each midpoint by its frequency, add the results, then divide by the total frequency. Working: the midpoints are 10, 30, 50 and 70 minutes. 10 × 5 = 50. 30 × 10 = 300. 50 × 10 = 500. 70 × 5 = 350. Σfx = 50 + 300 + 500 + 350 = 1200. Σf = 5 + 10 + 10 + 5 = 30. Estimated mean = 1200 ÷ 30 = 40 minutes. Using the upper boundary of each class instead of the midpoint — 20 × 5 = 100, 40 × 10 = 400, 60 × 10 = 600, 80 × 5 = 400 — gives a total of 1500 and an estimate of 1500 ÷ 30 = 50 minutes, too high because a boundary is not the middle of the class. Averaging the frequencies themselves, 5, 10, 10 and 5, ignores the times altogether and gives 7.5. Stopping after Σfx = 1200 without dividing by the total frequency gives a number far too large to be a time in minutes. Always find the midpoint of each class before multiplying by the frequency, and always divide by Σf at the end.
- (d) 648 kg — Method: to estimate a total from a sample, multiply the sample's mean by the number of items in the whole population, then check the units the question asks for. Working: 32.4 g × 20,000 = 648,000 g. Converting to kilograms, 648,000 ÷ 1,000 = 648 kg. This is only an estimate, not an exact total, because it assumes every one of the 20,000 packets has exactly the sample mean mass, when in reality individual packets vary above and below it. Giving 1.62 kg multiplies the mean by 50, the SAMPLE size, instead of by 20,000, the number of packets actually made that day — this finds the total mass of the 50 sampled packets, not the day's production. Giving 32.4 kg treats the sample mean itself, in grams, as if it already were the day's total mass in kilograms, skipping the scaling up altogether. Giving 648,000 kg correctly scales the mean up to the whole day's production but never converts the answer from grams to kilograms, leaving it 1,000 times too large. Always scale a sample's mean up by the SIZE OF THE WHOLE POPULATION, and always finish by checking the units the question asks for.
- (a) 1.24 — Method: for data given as a frequency table, the mean is Σfx ÷ Σf — multiply each value by its frequency, add the results, then divide by the total frequency. Working: 0 × 6 = 0. 1 × 10 = 10. 2 × 6 = 12. 3 × 3 = 9. So Σfx = 0 + 10 + 12 + 9 = 31. The total frequency is Σf = 6 + 10 + 6 + 3 = 25. Mean = 31 ÷ 25 = 1.24 siblings. Averaging the frequency column itself, (6 + 10 + 6 + 3) ÷ 4 = 6.25, mixes up the frequencies with the values they belong to. Writing down 1, the number of siblings with the highest frequency, gives the mode, not the mean. Writing down 31 stops after finding Σfx and forgets to divide by the total frequency, 25. Always divide Σfx by Σf — never stop at the top of the fraction.
- (a) x = 0 gives y = −20: a negative number sold — The y-intercept is the value the line predicts when x = 0: y = 3 × 0 − 20 = −20. A kiosk cannot sell a negative number of ice creams, so this is not a sensible estimate. The 3 in the equation is the gradient, not the intercept, so an option claiming x = 0 gives y = 3 has swapped the two numbers around — substituting x = 0 makes the 3x term equal 0, leaving −20, not 3. The danger of extrapolating to very high temperatures is a real issue with this line, but it is a different issue from the y-intercept, so it does not answer this question. And whether x = 0 could occur on a trading day is beside the point: the model still makes that prediction, and it is the prediction itself, −20, that is impossible.
- (c) No, 90 cm is far outside the heights on the graph — Method: a line of best fit describes the trend only across the stretch of data it was drawn through; predicting beyond that stretch is extrapolation, and nothing in the data supports it. Working: the heights used to draw this line run from 150 cm to 180 cm, all of them Year 10 pupils, while 90 cm is 60 cm below the shortest of them and belongs to a two-year-old child, whose build follows no trend the graph has measured. Substituting anyway gives 0.9 × 90 − 85 = −4, a mass of −4 kg, which cannot exist. Answer: no, because 90 cm is far outside the heights on the graph. The distractors: saying a line of best fit cannot be used to predict at all throws away its main purpose, since a prediction made between the plotted values is perfectly sound; saying the line passes through all 20 points misdescribes a line of best fit, which is drawn to follow the trend of the points and will normally pass through few of them; saying the equation works for any value put into it treats an equation fitted to Year 10 heights as a law of nature, and the mass of −4 kg shows what that assumption produces.
- (b) Town B — higher median and smaller IQR — Method: 'higher and more consistent' needs two comparisons — the median for typical price, and the interquartile range for spread, with a smaller interquartile range meaning more consistent. Working: Town B's median, £235,000, is higher than Town A's, £220,000. Town A's interquartile range is 310 − 180 = 130 and Town B's is 260 − 175 = 85, so Town B's interquartile range is the smaller of the two. Answer: Town B has both the higher median and the smaller interquartile range, so it is the town with higher, more consistent prices. Watch which combination of median and interquartile range each statement claims, and for which town: claiming Town A has the higher median and the smaller interquartile range gets both comparisons wrong, since Town B leads on both; claiming Town A has the higher median (still wrong) but the larger interquartile range at least reads the spread correctly, without it rescuing the false median claim; and claiming Town B has the higher median (correct) but the larger interquartile range misreads the spread — Town B's interquartile range is the smaller of the two, not the larger.
- (c) 75 — Method: the number in a class is the area of its bar, frequency density × class width, so work out the frequency of each class that lies at or above 10 minutes and add them. Working: the class 10 ≤ t < 25 is 15 minutes wide with a frequency density of 3.2, giving 3.2 × 15 = 48 members; the class 25 ≤ t < 55 is 30 minutes wide with a frequency density of 0.9, giving 0.9 × 30 = 27 members; the total charged is 48 + 27 = 75. Answer: 75 members pay the extra charge. The distractors: 4.1 comes from adding the two frequency densities, 3.2 + 0.9, as though each height were a count; 93 comes from including the class 0 ≤ t < 10 as well, 1.8 × 10 = 18 added to 48 and 27, which charges every member; 27 comes from using only the class 25 ≤ t < 55 and forgetting that 10 ≤ t < 25 is also at or above 10 minutes.
- (d) 60 — Method: on a histogram the frequency of a class is its frequency density × its class width, and the frequencies of all the classes add up to the total, so turn each labelled bar into a frequency and subtract their total from 250. Working: 10 ≤ age < 20 has width 20 − 10 = 10, so its frequency is 4.5 × 10 = 45; 20 ≤ age < 35 has width 35 − 20 = 15, so its frequency is 6 × 15 = 90; 50 ≤ age < 70 has width 70 − 50 = 20, so its frequency is 2.75 × 20 = 55. Those three come to 45 + 90 + 55 = 190, and the total is 250, so the missing frequency is 250 − 190 = 60. Answer: the class 35 ≤ age < 50 has 60 members. Watch what you do with the total and the three frequencies you have found: giving the total, 250, as the answer forgets that three bars have already accounted for some of the members; giving 190, the total of the other three classes, reports how many members are not in this class rather than how many are; and leaving one of the three out of the subtraction, for example 45 + 90 = 135 and 250 − 135 = 115, still owes the class at 50 ≤ age < 70 its 55 members.
- (b) Treat them as outliers and check them before deciding — Method: a point lying a long way from the pattern the rest of the data make is called an outlier, and an outlier is investigated before anything is done with it, because it may be an error in the data or it may be a genuine but unusual case. Working: 48 of the 50 points lie close to the rising line of best fit, so the trend is set by those 48; the two remaining points do not follow it, so they are identified as outliers and checked — a mistake in measuring or recording would be corrected, while a genuine reading would be kept and reported. Answer: treat them as outliers and check them before deciding what to do with them. The distractors: deleting them at once assumes that every point far from the line must be an error, which throws away real data; moving the line so that it passes through them assumes a line of best fit must touch particular points, when it is drawn to follow all 50; taking them as proof that there is no correlation lets two points overturn the pattern that the other 48 agree on.
- (d) 210 — 350 − 150 = 200. 200 ÷ 10 = 20, so the gradient is 20. Using the point (5, 150): 20 × 5 = 100, so 150 − 100 = 50 is the intercept, giving the line y = 20x + 50. At x = 8: 20 × 8 = 160, and 160 + 50 = 210, so the estimated number of visitors is 210. Choosing 160 stops after 20 × 8 = 160 and forgets to add the intercept of 50. Choosing 250 comes from averaging the two given y-values: 150 + 350 = 500, and 500 ÷ 2 = 250, instead of using the line's equation. Choosing 240 assumes the visitors are directly proportional to the hours of sunshine using the first point, 150 × 8 ÷ 5 = 240, which ignores that the line does not pass through the origin.
Build your own mix at the worksheet builder.