Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- 1.A school has 1200 pupils. A teacher wants to take a random sample of 60 of them. Write down which of these methods gives a random sample.
- 2.A scatter graph of the number of ice creams sold at a seaside kiosk in Bournemouth and the number of sunburn cases treated at a nearby pharmacy, recorded on the same 30 days, shows strong positive correlation. Which statement about this correlation is correct?
- 3.A cumulative frequency graph for the diameters, d mm, of 320 ball bearings is plotted from these points (upper class boundary, cumulative frequency): (10, 30), (20, 90), (30, 190), (40, 280), (50, 320). Estimate the diameter below which 90% of the ball bearings measure.
- 4.The masses, m kg, of 160 fish caught by a trawler in one day are grouped into classes of unequal width: 0 ≤ m < 10, 40 fish; 10 ≤ m < 30, 60 fish; 30 ≤ m < 45, 30 fish; 45 ≤ m < 50, 30 fish. A histogram is to be drawn from this table. Which set of frequency densities, listed in the same order as the classes above, is correct?
- 5.Two classes sat the same maths test, both marked out of 100. Class A had a mean mark of 70 and a range of 30 marks. Class B had a mean mark of 70 and a range of 10 marks. Compare the marks of the two classes.
- 6.A quality inspector weighs a random sample of 50 packets of crisps from one day's production and finds their mean mass is 32.4 g. The factory makes 20,000 packets that day. Work out an estimate for the total mass, in kg, of all the packets made that day.
- 7.A council in Leeds wants to know what local people think about letting shops stay open later in the evening. It rings landline telephone numbers between 10 am and 2 pm on a Tuesday. Write down which group is most likely to be under-represented in the sample, and give a reason for your answer.
- 8.A garden centre records the heights, in cm, of eleven seedlings. In order, the heights are 5, 9, x, 17, 20, 24, 28, 31, 35, 40, 44, where x is unknown. The interquartile range of the eleven heights is 19 cm. Work out the value of x.
- 9.The times taken, in minutes, by 30 runners in a Portsmouth fun run are grouped in this table: 0 < t ≤ 20 — 5 runners, 20 < t ≤ 40 — 10 runners, 40 < t ≤ 60 — 10 runners, 60 < t ≤ 80 — 5 runners. Work out an estimate for the mean time, in minutes.
- 10.The times, t minutes, taken by 120 runners to finish a fun run are summarised by these cumulative frequencies: t < 20, 8 runners; t < 30, 26 runners; t < 40, 74 runners; t < 50, 110 runners; t < 60, 120 runners. Work out the number of runners who took 40 minutes or longer to finish.
- 11.A histogram is drawn for the masses, m grams, of 200 letters. The bar for 0 ≤ m < 50 has a frequency density of 1.2 per gram and the bar for 50 ≤ m < 100 has a frequency density of 1.8 per gram. All the remaining letters lie in the class 100 ≤ m < 200. Work out the frequency density of the bar for 100 ≤ m < 200.
- 12.An estate agent in Leeds lists the prices of the five houses sold on one street last year: £170,000, £150,000, £580,000, £140,000 and £160,000. A newspaper wants to print one figure for a typical price on this street. Work out the mean and the median, and write down which figure the newspaper should print.
- 13.The marks scored by 11 pupils in a test are given in order: 12, 15, 18, 21, 24, 27, 30, 33, 36, 39, 42. Work out the lower quartile of these marks.
- 14.A pictogram shows the number of cars sold by a garage in Plymouth each month, where each whole car symbol represents 8 cars sold and a half symbol represents 4 cars. March shows 3 whole symbols and one half symbol. Work out how many cars were sold in March.
- 15.A histogram shows the ages, in years, of 250 members of a running club. The bar for the class 10 ≤ age < 20 has a frequency density of 4.5 members per year, the bar for 20 ≤ age < 35 has a frequency density of 6 members per year, and the bar for 50 ≤ age < 70 has a frequency density of 2.75 members per year. Work out the frequency of the remaining class, 35 ≤ age < 50.
Answer key
- (a) Drawing 60 names at random from a list of all 1200 pupils — Method: a sample is random when every member of the population has the same chance of being chosen and nobody, including the pupils themselves, can influence who ends up in it; test each method against that. Working: drawing names from a list of all 1200 pupils gives each pupil the same chance, 60 out of 1200, whatever their year group, class or opinion, so the method is random. Answer: drawing 60 names at random from a list of all 1200 pupils. The distractors: asking the pupils who volunteer is self-selection, and the pupils with the strongest views volunteer first, so they decide the sample; asking the pupils nearest the door is convenience sampling, which reaches only those who happen to be in one place at one time; asking two Year 10 classes samples a cluster, so every pupil in the other year groups has no chance of being chosen at all.
- (c) Neither causes the other; sunshine links both. — Both ice cream sales and sunburn cases tend to rise on hot, sunny days, so the amount of sunshine is a third factor linked to both — neither variable causes the other. Saying ice cream sales cause the sunburn assumes a causal link in one direction that the correlation alone cannot establish. Saying sunburn cases cause the ice cream sales assumes the reverse causal link, which is no more justified. Saying a strong correlation always means causation is the general error this question is testing: correlation, however strong, does not by itself prove that one variable causes the other.
- (c) 42 — Method: find the target cumulative frequency, 90% of the total, locate the class it falls in from the plotted points, then interpolate: lower boundary, plus the extra distance needed into the class divided by the class's frequency, times its width. Working: 90% of 320 is 0.9 × 320 = 288. The plotted points show a cumulative frequency of 280 at d = 40 and 320 at d = 50, so the class 40 ≤ d < 50 has frequency 320 − 280 = 40 and width 50 − 40 = 10, and 288 falls inside it. The extra distance needed into the class is 288 − 280 = 8, and 8 ÷ 40 × 10 = 2, so the diameter is 40 + 2 = 42. Answer: the estimated diameter is 42 mm. Watch which point and which class the interpolation actually uses: reading off d = 40, the plotted point just below the target, instead of interpolating the extra 8 ball bearings into the next 10 mm, stops one step short of the true answer; finding the diameter below which only 10% lie instead of 90% gives a target of 0.1 × 320 = 32, which falls in the class 10 ≤ d < 20 — the extra distance into that class is 32 − 30 = 2, and 2 ÷ 60 × 10 = 0.3, so this route gives 10 + 0.3 = 10.3, the bottom decile rather than the top 90%; and interpolating within the class 30 ≤ d < 40 instead of 40 ≤ d < 50, as though 288 had not yet reached a cumulative frequency of 280, treats the extra distance as 288 − 190 = 98, and 98 ÷ 90 × 10 = 10.9, giving 30 + 10.9 = 40.9, one class too early.
- (d) 4, 3, 2, 6 — Method: frequency density = frequency ÷ class width for each class in turn; do not assume the classes are all the same width. Working: the four classes have widths 10 − 0 = 10, 30 − 10 = 20, 45 − 30 = 15 and 50 − 45 = 5. Dividing each frequency by its own width gives 40 ÷ 10 = 4, 60 ÷ 20 = 3, 30 ÷ 15 = 2 and 30 ÷ 5 = 6. Answer: the frequency densities, in order, are 4, 3, 2 and 6. Watch the width of each class separately: treating the last class as if it were also 10 units wide, like the first, gives 30 ÷ 10 = 3 instead of 30 ÷ 5 = 6 — the classes here are deliberately unequal, so no width can be borrowed from another class; dividing the width by the frequency instead of the frequency by the width for the third class gives 15 ÷ 30 = 0.5 in place of 2, the formula the wrong way round; and reading the frequency column straight off the table, 40, 60, 30, 30, skips the division by width altogether and reports how many fish are in each class rather than how densely packed each bar is.
- (a) The means are equal, and Class B's marks are the more consistent because its range is smaller. — Method: comparing two distributions needs two things — a measure of average and a measure of spread — and each must be put into the context of the question. Working: both classes have a mean mark of 70, so on average the two classes scored the same; the range measures spread, and Class A's range of 30 marks is three times Class B's range of 10 marks, so Class B's marks sit closely around the mean while Class A's are far more spread out. Answer: the means are equal, and Class B's marks are the more consistent because its range is smaller. The distractors: the reply crediting Class A with more consistency reverses the meaning of the range, treating a larger range as tighter data when a larger range means more spread; the reply that Class A's mean mark is higher compares the wrong pair of figures, reading the range of 30 as an average; the reply that Class B's mean mark is higher reads the spread correctly but its claim about the means is false, since both means are 70.
- (d) 648 kg — Method: to estimate a total from a sample, multiply the sample's mean by the number of items in the whole population, then check the units the question asks for. Working: 32.4 g × 20,000 = 648,000 g. Converting to kilograms, 648,000 ÷ 1,000 = 648 kg. This is only an estimate, not an exact total, because it assumes every one of the 20,000 packets has exactly the sample mean mass, when in reality individual packets vary above and below it. Giving 1.62 kg multiplies the mean by 50, the SAMPLE size, instead of by 20,000, the number of packets actually made that day — this finds the total mass of the 50 sampled packets, not the day's production. Giving 32.4 kg treats the sample mean itself, in grams, as if it already were the day's total mass in kilograms, skipping the scaling up altogether. Giving 648,000 kg correctly scales the mean up to the whole day's production but never converts the answer from grams to kilograms, leaving it 1,000 times too large. Always scale a sample's mean up by the SIZE OF THE WHOLE POPULATION, and always finish by checking the units the question asks for.
- (d) Full-time workers, as most are at work at that time — Method: a sample is biased when the method of contact makes part of the population much less likely to be reached, so test each group against where its members actually are between 10 am and 2 pm on a weekday, and test each stated reason against the facts. Working: those hours are the middle of the working day, so people in full-time employment are at work and not beside a landline telephone, while people who are retired and people who are unemployed are far more likely to be at home and are reached at the usual rate; the method therefore collects far fewer replies from full-time workers than their share of the adult population the council is consulting. Answer: full-time workers, as most are at work at that time. The distractors: the reply naming retired people rests on the false claim that most retired people are at work in the daytime, when in fact a daytime call reaches them more easily than anyone; the reply naming unemployed people rests on the false claim that they are out during the day, when they too are among the easiest people to reach by a daytime call; the reply naming children rests on the false claim that children are at home at 11 am on a Tuesday, when they are at school and so are not reached by the call at all, and school-age children are in any case not the adults whose views the council is collecting.
- (d) 16 — Method: rearrange interquartile range = upper quartile − lower quartile to make the lower quartile the subject: lower quartile = upper quartile − interquartile range, then check the answer sits in the right place in the list. Working: there are 11 values, so 11 + 1 = 12; the upper quartile sits at position 3 × 12 ÷ 4 = 9, which is 35, and x sits at position 12 ÷ 4 = 3, which is the lower quartile. So x = 35 − 19 = 16, and 16 does sit between the 2nd value, 9, and the 4th value, 17, as it should. Answer: x = 16. Watch how you rearrange and where you count to: adding instead of subtracting, 35 + 19 = 54, treats the interquartile range as something added on rather than a gap taken away; subtracting in the wrong order, 19 − 35 = −16, finds the right two numbers but flips the sign; and counting to the 8th value instead of the 9th treats 31 as the upper quartile, giving 31 − 19 = 12, one position short of where the upper quartile actually sits.
- (c) 40 minutes — Method: for grouped data, estimate the mean using the midpoint of each class — multiply each midpoint by its frequency, add the results, then divide by the total frequency. Working: the midpoints are 10, 30, 50 and 70 minutes. 10 × 5 = 50. 30 × 10 = 300. 50 × 10 = 500. 70 × 5 = 350. Σfx = 50 + 300 + 500 + 350 = 1200. Σf = 5 + 10 + 10 + 5 = 30. Estimated mean = 1200 ÷ 30 = 40 minutes. Using the upper boundary of each class instead of the midpoint — 20 × 5 = 100, 40 × 10 = 400, 60 × 10 = 600, 80 × 5 = 400 — gives a total of 1500 and an estimate of 1500 ÷ 30 = 50 minutes, too high because a boundary is not the middle of the class. Averaging the frequencies themselves, 5, 10, 10 and 5, ignores the times altogether and gives 7.5. Stopping after Σfx = 1200 without dividing by the total frequency gives a number far too large to be a time in minutes. Always find the midpoint of each class before multiplying by the frequency, and always divide by Σf at the end.
- (c) 46 — Method: the cumulative frequency table gives the number of runners below each time; to find the number at or above a time, subtract that cumulative frequency from the total. Working: the cumulative frequency for t < 40 is 74, so 120 runners in total take away the 74 who finished in under 40 minutes: 120 − 74 = 46. Answer: 46 runners took 40 minutes or longer. Watch which boundary and which subtraction you use: reading off t < 50 instead of t < 40 and subtracting, 120 − 110 = 10, answers a different question, '50 minutes or longer'; giving 74 itself as the answer reports how many finished below 40 minutes, the opposite of what was asked; and subtracting the two nearby cumulative frequencies, 110 − 74 = 36, finds how many took between 40 and 50 minutes, not everyone from 40 minutes upward.
- (c) 0.5 per gram — Method: turn the two known bars into frequencies using area, subtract from the total to find how many letters are left, then divide that frequency by the width of the last class to get its height. Working: the first bar covers 50 g at a frequency density of 1.2, giving 1.2 × 50 = 60 letters, and the second covers 50 g at 1.8, giving 1.8 × 50 = 90 letters; together that is 60 + 90 = 150 letters, so 200 − 150 = 50 letters remain; the class 100 ≤ m < 200 is 100 g wide, so its frequency density is 50 ÷ 100 = 0.5 per gram. Answer: 0.5 per gram. The distractors: 0.25 per gram comes from dividing the remaining 50 letters by the upper class boundary, 200, instead of by the class width of 100; 2 per gram comes from dividing the class width by the frequency, 100 ÷ 50, reversing the formula; 1.4 per gram comes from subtracting only the first bar's 60 letters, leaving 140, and then dividing by 100.
- (c) The median, £160,000, as one very high price lifts the mean — Method: find both averages, then choose the one that sits closer to the bulk of the data. Working: in order the prices are 140,000, 150,000, 160,000, 170,000 and 580,000, so the median is the third of the five, £160,000. For the mean, 140,000 + 150,000 + 160,000 + 170,000 + 580,000 = 1,200,000 and 1,200,000 ÷ 5 = 240,000, so the mean is £240,000. Four of the five houses sold for £170,000 or less, so a reader told that a typical price is £240,000 would expect to pay at least £70,000 more than any of those four cost. Answer: the median, £160,000, as one very high price lifts the mean. The distractors: £580,000 is the middle value of the list as it is printed, which is the median only when the values have first been put in order; £240,000 is the mean, chosen on the ground that a median ignores three of the five prices, but a median uses all five to find which one is central and is then untroubled by how extreme the outer values are; £155,000 comes from deleting the £580,000 house and taking the mean of what is left, since 140,000 + 150,000 + 160,000 + 170,000 = 620,000 and 620,000 ÷ 4 = 155,000, but a real sale may not be thrown away merely for being large.
- (c) 18 — Method: for n ordered values, GCSE convention places the lower quartile at position (n + 1) ÷ 4, counting from the smallest value. Working: there are 11 marks, so n + 1 = 11 + 1 = 12 and 12 ÷ 4 = 3, so the lower quartile is the 3rd value in the ordered list 12, 15, 18, 21, 24, 27, 30, 33, 36, 39, 42, which is 18. Answer: the lower quartile is 18 marks. Watch the position you count to: dividing 11 ÷ 4 = 2.75 without adding 1 first, then rounding down, lands on the 2nd value, 15, not the 3rd; reaching for the middle of the whole list instead gives the median, 27, a different statistic; and averaging the 3rd and 4th values, 18 + 21 = 39 and 39 ÷ 2 = 19.5, borrows a method for an even split where it is not needed here.
- (c) 28 — Method: multiply the number of whole symbols by the value of one symbol, then add the value of any half symbol shown. Working: 3 whole symbols represent 3 × 8 = 24 cars. The half symbol represents 4 cars. Total cars sold in March = 24 + 4 = 28. Leaving out the half symbol, 3 × 8 = 24, undercounts by exactly the value of that half symbol. Treating the half symbol as if it were a full symbol, 4 × 8 = 32, overcounts because it doubles the value the half symbol is worth. Giving 3.5 reports the number of symbols shown, not the number of cars they represent — the key still needs to be applied. Always apply the key to every symbol shown, including a half symbol, rather than reading off the symbol count itself.
- (d) 60 — Method: on a histogram the frequency of a class is its frequency density × its class width, and the frequencies of all the classes add up to the total, so turn each labelled bar into a frequency and subtract their total from 250. Working: 10 ≤ age < 20 has width 20 − 10 = 10, so its frequency is 4.5 × 10 = 45; 20 ≤ age < 35 has width 35 − 20 = 15, so its frequency is 6 × 15 = 90; 50 ≤ age < 70 has width 70 − 50 = 20, so its frequency is 2.75 × 20 = 55. Those three come to 45 + 90 + 55 = 190, and the total is 250, so the missing frequency is 250 − 190 = 60. Answer: the class 35 ≤ age < 50 has 60 members. Watch what you do with the total and the three frequencies you have found: giving the total, 250, as the answer forgets that three bars have already accounted for some of the members; giving 190, the total of the other three classes, reports how many members are not in this class rather than how many are; and leaving one of the three out of the subtraction, for example 45 + 90 = 135 and 250 − 135 = 115, still owes the class at 50 ≤ age < 70 its 55 members.
Build your own mix at the worksheet builder.