Printable · GCSE Higher · ages 14-16
Statistics worksheet — GCSE Higher
Fifteen questions across the statistics statements at Higher tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Non-calculator
Statistics worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- 1.The times, t minutes, of 80 journeys are summarised by these cumulative frequencies: t < 10, 8 journeys; t < 20, 28 journeys; t < 30, 52 journeys; t < 40, 72 journeys; t < 50, 80 journeys. Estimate the interquartile range.
- 2.In a histogram of the masses, m grams, of some pebbles, the bar for the class 50 ≤ m < 80 has a frequency density of 2.4 per gram. Work out the number of pebbles in this class.
- 3.A council wants to find out what all residents of a town think about a new cycle lane. It posts a survey only on its website and asks people to fill it in online. Give a reason why this sample is likely to be biased.
- 4.A two-way table records whether each of the 30 pupils in a class passed maths and whether they passed science. 18 pupils passed maths, 12 pupils passed science and 8 pupils passed both. Work out how many pupils passed at least one of the two subjects.
- 5.The mean of 5 numbers is 8. Work out the total of the 5 numbers.
- 6.Two classes sat the same test. The 30 pupils in Class A had a mean mark of 72. The 20 pupils in Class B had a mean mark of 82. Work out the mean mark of all 50 pupils.
- 7.A frequency polygon for the mass, in kg, of 40 parcels at a delivery depot is drawn by plotting one point at the midpoint of each class, joined by straight lines: (5, 6), (15, 10), (25, 16), (35, 6), (45, 2). Every class has a width of 10 kg. Write down the modal class.
- 8.A factory makes 4,000 light bulbs a day. In a random sample of 80 of one day's bulbs, 3 were faulty. Work out an estimate for the number of faulty bulbs the factory makes in a day.
- 9.Four pairs of variables are listed below. Write down the pair that you would expect to show negative correlation.
- 10.A council in Leeds wants to know what local people think about letting shops stay open later in the evening. It rings landline telephone numbers between 10 am and 2 pm on a Tuesday. Write down which group is most likely to be under-represented in the sample, and give a reason for your answer.
- 11.A teacher records the number of pets owned by each of 25 pupils in a class; each pupil owns 0, 1, 2, 3 or 4 pets. The teacher wants to show how many pupils own each number of pets. Write down the most suitable type of chart for this data, and give a reason for your answer.
- 12.The times, t minutes, taken by 80 people to travel to work are grouped like this: 0 ≤ t < 10, 6 people; 10 ≤ t < 20, 14 people; 20 ≤ t < 30, 25 people; 30 ≤ t < 40, 20 people; 40 ≤ t < 50, 15 people. Work out the cumulative frequency for t < 30.
- 13.The mean of three numbers is 50. A fourth number, 100, is added to the set. Work out the mean of the four numbers.
- 14.A histogram is drawn for the masses, m grams, of 200 letters. The bar for 0 ≤ m < 50 has a frequency density of 1.2 per gram and the bar for 50 ≤ m < 100 has a frequency density of 1.8 per gram. All the remaining letters lie in the class 100 ≤ m < 200. Work out the frequency density of the bar for 100 ≤ m < 200.
- 15.A sports centre in Ipswich has 2,000 members. It wants to know what its members think of its opening hours, so it asks a random sample of 100 of them. Write down what the population is in this survey.
Answer key
- (d) 18 minutes — Method: the lower quartile is the 80 ÷ 4 = 20th value and the upper quartile is the 3 × 80 ÷ 4 = 60th value; locate each inside its class by linear interpolation, then subtract. Working: the 20th value lies between the running totals 8 and 28, so it is in the class 10 ≤ t < 20, which holds 20 journeys across 10 minutes, and it is the 20 − 8 = 12th of them, giving 10 + (12 ÷ 20) × 10 = 16 minutes; the 60th value lies between the running totals 52 and 72, so it is in the class 30 ≤ t < 40, which also holds 20 journeys across 10 minutes, and it is the 60 − 52 = 8th of them, giving 30 + (8 ÷ 20) × 10 = 34 minutes; subtracting, 34 − 16 = 18. Answer: an estimated interquartile range of 18 minutes. The distractors: 20 minutes comes from taking the lower boundaries of the two quartile classes, 30 − 10, which locates the classes but never the values inside them; 40 minutes comes from subtracting the two positions, 60 − 20, instead of the two times; 22 minutes comes from interpolating downwards from each upper boundary rather than upwards from each lower boundary, giving 20 − 6 = 14 and 40 − 4 = 36.
- (c) 72 — Method: on a histogram the frequency of a class is the area of its bar, so frequency = frequency density × class width. Working: the class 50 ≤ m < 80 has width 80 − 50 = 30 grams and a frequency density of 2.4 per gram, so the frequency is 2.4 × 30 = 72. Answer: 72 pebbles. The distractors: 192 comes from using the upper class boundary, 80, as the width, giving 2.4 × 80; 12.5 comes from dividing the width by the density, 30 ÷ 2.4, which reverses the area rule; 2.4 comes from reading the height of the bar as the frequency itself, the commonest mistake on histograms, where a height is a density and only an area is a count.
- (d) Only internet users reach the website; others are excluded. — Method: a sample is biased when it systematically leaves out part of the population, or systematically over-represents another part. Working: anyone without internet access, or who does not visit the council's website, has NO chance of being included — the sample is drawn only from internet-using residents, which is not the whole town. Saying too many people might respond because the survey is free confuses bias with sample size — bias is about who CAN be reached, not how many respond. Saying people might lie describes a different problem, response honesty, not who was sampled in the first place. Saying online surveys cannot be anonymous is not a reason connected to bias at all. A sample is biased when part of the population has no chance of being included, whatever the reason for that.
- (a) 22 — Method: the two subject totals overlap, because every pupil who passed both subjects has been counted once in the maths total and once again in the science total; adding the totals therefore counts those pupils twice, and the overlap has to be taken off once. Working: 18 + 12 = 30, and the 8 pupils who passed both have been counted twice in that 30, so the number who passed at least one subject is 30 − 8 = 22. Answer: 22 pupils, a count of pupils, and it is less than the 30 in the class, which leaves 8 pupils who passed neither. The distractors: 30 comes from adding the two subject totals and never removing the overlap, so it counts the 8 pupils twice; 14 comes from taking the 8 away twice, 18 + 12 − 8 − 8, removing an overlap that was only counted twice once too often; 18 comes from writing down the larger of the two subject totals on its own, which leaves out every pupil who passed science but not maths.
- (a) 40 — Method: the mean is the total divided by how many values there are, so rearranging gives total = mean × number of values. Working: the mean is 8 and there are 5 numbers, so the total is 8 × 5 = 40. Answer: 40, and checking, 40 ÷ 5 = 8, which is the mean given. The distractors: 13 comes from adding the mean and the count, 8 + 5, instead of multiplying them; 1.6 comes from dividing the mean by the count, 8 ÷ 5, which reverses the relationship; 8 comes from quoting the mean itself as the total, which is only true when there is a single number.
- (c) 76 marks — Method: a mean of means only works when the groups are the same size, so rebuild each class's total mark, add the totals and divide by the number of pupils altogether. Working: Class A scored 30 × 72 = 2160 marks and Class B scored 20 × 82 = 1640 marks, giving 2160 + 1640 = 3800 marks between 50 pupils, so the overall mean is 3800 ÷ 50 = 76 marks. Answer: 76 marks. The distractors: 77 marks comes from averaging the two class means, (72 + 82) ÷ 2, which ignores the different class sizes; 78 marks comes from attaching each mean to the other class's size, (30 × 82 + 20 × 72) ÷ 50; 3800 marks comes from stopping at the combined total and never dividing by 50.
- (c) 20 kg ≤ mass < 30 kg — The modal class is the class with the highest frequency. Reading the plotted points, the frequencies are 6, 10, 16, 6 and 2, so the highest frequency is 16, plotted at the midpoint 25. A class of width 10 centred on 25 runs from 25 − 5 = 20 to 25 + 5 = 30, so the modal class is 20 kg ≤ mass < 30 kg. Writing '25 kg' gives only the midpoint, not the class — the modal class is an interval, not a single value. '10 kg ≤ mass < 20 kg' is the class before the peak, centred on 15, which has frequency 10, not the highest. '30 kg ≤ mass < 40 kg' is the class after the peak, centred on 35, which has frequency 6, not the highest.
- (a) 150 bulbs — Method: assume the proportion faulty in a random sample is the proportion faulty in the whole day's output, and scale the sample up to the population. Working: the sample of 80 has to be scaled up to 4,000 bulbs, and 4,000 ÷ 80 = 50, so the day's output is 50 sample-sized batches. Each batch is expected to contain the same 3 faulty bulbs, so the estimate is 3 × 50 = 150. Answer: 150 bulbs, and it is an estimate, because another sample of 80 would probably contain a different number of faulty bulbs. The distractors: 50 bulbs is the scale factor 4,000 ÷ 80 written down as though it were the answer, so it reports how many batches there are rather than how many faulty bulbs; 120 bulbs comes from reading 3 out of 80 as 3%, then taking 0.03 × 4,000 = 120, but 3 out of 80 is 3.75%; 240 bulbs comes from 3 × 80 = 240, multiplying the faulty bulbs by the size of the sample instead of by the scale factor, which uses the 80 twice and the 4,000 not at all.
- (a) Minutes a candle has burned and length remaining — As a candle burns for longer, less of it remains, so these two variables move in opposite directions as one increases — that is negative correlation. A pupil's shoe size generally increases as they get older, so age and shoe size show positive correlation, not negative, since both rise together. A football team's shirt colour is not a numerical quantity linked to how many matches it wins, so shirt colour and number of wins show no correlation at all. The number of letters in a pupil's name has no real connection to their ability in maths, so that pair also shows no correlation.
- (d) Full-time workers, as most are at work at that time — Method: a sample is biased when the method of contact makes part of the population much less likely to be reached, so test each group against where its members actually are between 10 am and 2 pm on a weekday, and test each stated reason against the facts. Working: those hours are the middle of the working day, so people in full-time employment are at work and not beside a landline telephone, while people who are retired and people who are unemployed are far more likely to be at home and are reached at the usual rate; the method therefore collects far fewer replies from full-time workers than their share of the adult population the council is consulting. Answer: full-time workers, as most are at work at that time. The distractors: the reply naming retired people rests on the false claim that most retired people are at work in the daytime, when in fact a daytime call reaches them more easily than anyone; the reply naming unemployed people rests on the false claim that they are out during the day, when they too are among the easiest people to reach by a daytime call; the reply naming children rests on the false claim that children are at home at 11 am on a Tuesday, when they are at school and so are not reached by the call at all, and school-age children are in any case not the adults whose views the council is collecting.
- (a) A vertical line chart (discrete numerical data) — The number of pets is discrete numerical data — whole-number values such as 0, 1, 2, 3 or 4 — recorded for one variable, so a vertical line chart is the chart specified for this kind of data. A bar chart is used for categorical data, such as favourite colour, not numerical values counted like this. A pie chart shows proportions of a whole and does not show the frequency of each separate value. A scatter graph compares two different variables against each other, and only one variable, the number of pets, is recorded here.
- (d) 45 — Method: a cumulative frequency is a running total — it counts everybody in every class up to and including the one that ends at the value given. Working: the classes that lie wholly below 30 minutes are 0 ≤ t < 10, 10 ≤ t < 20 and 20 ≤ t < 30, with frequencies 6, 14 and 25, so the running total is 6 + 14 = 20 and then 20 + 25 = 45. Answer: 45 people took less than 30 minutes. The distractors: 25 comes from quoting the frequency of the class 20 ≤ t < 30 on its own instead of the running total; 65 comes from accumulating one class too many and including 30 ≤ t < 40, which is 45 + 20; 35 comes from accumulating from the top downwards, 15 + 20, which counts the people who took 30 minutes or more rather than fewer.
- (a) 62.5 — Method: a mean cannot be averaged with a new value — rebuild the total, add the new value to it, then divide by the new count. Working: three numbers with a mean of 50 have a total of 50 × 3 = 150; adding 100 makes the total 150 + 100 = 250; there are now 4 numbers, so the new mean is 250 ÷ 4 = 62.5. Answer: 62.5. The distractors: 75 comes from averaging the old mean with the new value, (50 + 100) ÷ 2, which ignores that three numbers pull against one; 50 comes from assuming an extra value leaves the mean unchanged; 37.5 comes from dividing the old total of 150 by the new count of 4, adding the new value to the count but not to the total.
- (c) 0.5 per gram — Method: turn the two known bars into frequencies using area, subtract from the total to find how many letters are left, then divide that frequency by the width of the last class to get its height. Working: the first bar covers 50 g at a frequency density of 1.2, giving 1.2 × 50 = 60 letters, and the second covers 50 g at 1.8, giving 1.8 × 50 = 90 letters; together that is 60 + 90 = 150 letters, so 200 − 150 = 50 letters remain; the class 100 ≤ m < 200 is 100 g wide, so its frequency density is 50 ÷ 100 = 0.5 per gram. Answer: 0.5 per gram. The distractors: 0.25 per gram comes from dividing the remaining 50 letters by the upper class boundary, 200, instead of by the class width of 100; 2 per gram comes from dividing the class width by the frequency, 100 ÷ 50, reversing the formula; 1.4 per gram comes from subtracting only the first bar's 60 letters, leaving 140, and then dividing by 100.
- (b) All 2,000 members of the sports centre. — Method: in a survey, the population is the whole group the survey is trying to find out about, and the sample is the smaller group actually asked. Working: this survey wants to know what the sports centre's members think, so the population is every one of the 2,000 members — whether or not they were personally asked. Saying the population is the 100 members who were asked names the sample, not the population; the sample is drawn FROM the population, so it is smaller than it, not the same as it. Saying the population is everybody who lives in Ipswich widens the group far beyond who the survey is actually about — plenty of Ipswich residents are not members of the sports centre at all, so they are outside this survey altogether. Saying the population is the members who say they are unhappy confuses the population with a result of the survey: whether a member turns out to be happy or unhappy is something the survey finds out, not part of the definition of who is being studied. The population is always the whole group the question is about, before any sampling or any results come in.
Build your own mix at the worksheet builder.