Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- 1.On a scatter graph of the age of a car, in years, and its value, in pounds, the points fall from left to right. Write down the type of correlation shown.
- 2.A vertical line chart shows the number of pets owned by each pupil in a class: 0 pets — 5 pupils, 1 pet — 4 pupils, 2 pets — 3 pupils, 3 pets — 3 pupils. Work out the total number of pupils in the class.
- 3.The mean mass of four parcels is 17 kg. Three of the parcels have masses 12 kg, 16 kg and 18 kg. Work out the mass of the fourth parcel.
- 4.A dual bar chart shows how many boys and how many girls are in a class. The bar for boys stands at 14 and the bar for girls stands at 16. Work out how many pupils are in the class altogether.
- 5.A school has 1,500 pupils. The head teacher takes a random sample of 150 of them from the school register and asks how long they spend on homework. Rory says the sample is too small for the result to mean anything. Is Rory right? Give a reason for your answer.
- 6.The mean of the numbers x, 20 and 30 is equal to the mean of the numbers 15 and 25. Work out the value of x.
- 7.A sports centre in Ipswich has 2,000 members. It wants to know what its members think of its opening hours, so it asks a random sample of 100 of them. Write down what the population is in this survey.
- 8.A scatter graph shows the number of years of experience, x, of 18 sales assistants and their monthly sales, y hundred pounds. The plotted points run from x = 1 to x = 12 years, and the line of best fit is y = 4x + 20. A new assistant has 25 years of experience. Use the line of best fit to estimate a value of y for this assistant, and decide whether the estimate would be reliable.y = 4x + 20
- 9.Write down the statement that correctly describes the difference between correlation and causation.
- 10.A stacked bar for one day at a café in Norwich shows total drink sales of 50 drinks, split into three parts: 22 were tea, 15 were coffee and the rest were hot chocolate. Work out the number of hot chocolates sold.
- 11.Two classes at a school in Coventry sit the same maths test, out of 20 marks. Class A has a mean mark of 14 and a range of 6. Class B has a mean mark of 14 and a range of 14. Write a sentence comparing the two classes, using the mean and the range.
- 12.A shop recorded the number of books it sold on five days: 100, 40, 70, 20, 60. Work out the range of the numbers of books sold.
- 13.A shop sold seven pairs of shoes in these sizes: 4, 4, 5, 6, 6, 6, 9. Write down the modal size.
- 14.A pupil writes this question for a school survey: “Do you agree that learning matters and that we should be set more homework?” Write down what is wrong with the survey question.
- 15.A composite (stacked) bar for a charity bake sale in Durham shows the number of cakes sold, split into three types. The bar has a total height of 80 cakes: 34 were sponge cakes, 26 were chocolate cakes and the rest were fruit cakes. Work out the percentage of the cakes sold that were fruit cakes.
Answer key
- (a) Negative correlation — As the age of the car increases, the points fall towards a lower value, so the value decreases as the age increases. This falling pattern is a negative correlation. A positive correlation would show the points rising together instead. No correlation would apply only if the points showed no pattern at all, and correlation is not the same as causation — strong causation is not a type of correlation.
- (b) 15 pupils — The frequencies are 5, 4, 3 and 3 pupils, and 5 + 4 + 3 + 3 = 15, so there are 15 pupils in the class. Choosing 6 pupils comes from adding the numbers of pets, 0 + 1 + 2 + 3 = 6, instead of the frequencies. Choosing 10 pupils comes from leaving out the '0 pets' row: 4 + 3 + 3 = 10. Choosing 18 pupils comes from counting the '3 pets' row twice: 5 + 4 + 3 + 3 + 3 = 18.
- (c) 22 kg — Method: multiply the mean by the number of parcels to rebuild the total mass, then subtract the masses that are known. Working: four parcels with a mean mass of 17 kg have a total mass of 17 × 4 = 68 kg; the three known parcels total 12 + 16 + 18 = 46 kg; so the fourth parcel has mass 68 − 46 = 22 kg. Answer: 22 kg. The distractors: 68 kg comes from stopping at the total mass of all four parcels; 17 kg comes from assuming the missing parcel must have the mean mass; 5 kg comes from multiplying the mean by 3, the number of parcels whose mass is given, leaving 51 − 46 = 5.
- (d) 30 — Method: on a dual bar chart each bar is a separate frequency, so a total for the whole class is found by combining the two frequencies the bars show. Working: the bars show 14 boys and 16 girls, and 14 + 16 = 30. Answer: 30 pupils, a count of pupils in the class. The distractors: 32 comes from doubling the taller bar, 16 + 16, as though the two bars were equal; 28 comes from doubling the shorter bar, 14 + 14, in the same way; 2 comes from finding the difference between the two bars, 16 − 14, which answers how many more girls there are rather than how many pupils there are.
- (a) No, 150 pupils are a tenth of the school, chosen at random — Method: judge a sample on two things, whether every member of the population had the same chance of being chosen, and whether the sample is large enough to carry a pattern. Working: the 150 pupils were drawn from the register of every pupil in the school, so no year group or set is shut out and no pupil chooses to take part; and 150 ÷ 1,500 = 0.1, so one pupil in ten has been asked. A random sample of that share is ample for an estimate of how long the school's pupils spend on homework. Answer: no, because 150 pupils are a tenth of the school and were chosen at random. The distractors: saying a random sample always gives the exact school figure reaches the same verdict for a reason that is false, since a second random sample of 150 would give a slightly different mean; saying 150 pupils cannot be picked at random from 1,500 treats randomness as something only a whole population can have, when drawing names from the register is exactly how a random sample is taken; saying that only asking all 1,500 could show anything rejects sampling altogether, which would leave no way to study any population too large to count.
- (c) 10 — Method: work out the mean that can be found straight away, then use total = mean × number of values on the group of three to find the missing number. Working: the mean of 15 and 25 is (15 + 25) ÷ 2 = 40 ÷ 2 = 20, so the group of three must also have a mean of 20; three numbers with a mean of 20 have a total of 20 × 3 = 60, and 20 + 30 = 50 of that total is already accounted for, so x = 60 − 50 = 10. Answer: 10, and checking, (10 + 20 + 30) ÷ 3 = 20. The distractors: 20 comes from working out the mean the two groups share and writing that down as x; −10 comes from dividing the group of three by 2 instead of by 3, which gives x + 50 = 40; 70 comes from reading the total 15 + 25 = 40 as the mean of the pair, which sets the target total at 120 and leaves x = 70.
- (b) All 2,000 members of the sports centre. — Method: in a survey, the population is the whole group the survey is trying to find out about, and the sample is the smaller group actually asked. Working: this survey wants to know what the sports centre's members think, so the population is every one of the 2,000 members — whether or not they were personally asked. Saying the population is the 100 members who were asked names the sample, not the population; the sample is drawn FROM the population, so it is smaller than it, not the same as it. Saying the population is everybody who lives in Ipswich widens the group far beyond who the survey is actually about — plenty of Ipswich residents are not members of the sports centre at all, so they are outside this survey altogether. Saying the population is the members who say they are unhappy confuses the population with a result of the survey: whether a member turns out to be happy or unhappy is something the survey finds out, not part of the definition of who is being studied. The population is always the whole group the question is about, before any sampling or any results come in.
- (b) 120, unreliable — x = 25 is outside 1 to 12 — The line of best fit is y = 4x + 20. 4 × 25 = 100, and 100 + 20 = 120, so the estimate is y = 120. But x = 25 lies far outside the plotted range of 1 to 12 years, so this is an extrapolation, and the estimate is not reliable. Reaching 100 instead of 120 comes from 4 × 25 = 100 with the intercept of 20 left out — still correctly flagged as unreliable, but the wrong value. Calling the estimate reliable simply because it was calculated correctly, giving 120, wrongly assumes that a correct calculation is automatically trustworthy, ignoring that x = 25 lies far beyond the data actually collected. Reaching 68, from 4 × 12 = 48 and 48 + 20 = 68, substitutes x = 12, the top of the plotted range, instead of the assistant's actual x = 25, and wrongly calls that reliable because 12 lies inside the range.
- (d) Correlation is a link; causation is one causing the other — Method: the two words describe different claims — one is about a pattern in the data, the other is about what produced that pattern. Working: correlation says only that two quantities tend to change together, which is something a scatter graph can display; causation says that a change in one quantity actually brings about the change in the other, which needs evidence a scatter graph cannot supply, because a third quantity may be driving both. Answer: correlation is a link between the quantities, while causation is one quantity causing the change in another. The distractors: the statement giving causation as the link and correlation as the cause simply swaps the two words over; the statement that the words mean the same thing is the classic error of reading a correlation as proof of cause; the statement that a scatter graph shows causation but not correlation reverses what a scatter graph can do, since the pattern it displays is exactly the correlation.
- (b) 13 — The total is 50, and the two known parts are 22 (tea) and 15 (coffee), so 50 − 22 − 15 = 13 hot chocolates. Choosing 28 comes from 50 − 22, subtracting only the tea and forgetting the coffee. Choosing 35 comes from 50 − 15, subtracting only the coffee and forgetting the tea. Choosing 37 comes from 22 + 15, which finds how many drinks were tea or coffee, not the number left over for hot chocolate.
- (a) Equal means; Class A is more consistent, smaller range. — Method: when two data sets share a measure of location, compare a measure of spread to say more about consistency. Working: both classes have the same mean mark, 14, so on average they performed equally well. Class A has the smaller range, 6, so its marks are more tightly grouped around 14 than Class B's marks, which vary by as much as 14. So Class A's marks were more consistent, even though neither class did better on average. Saying Class B did better because it has the bigger range confuses a wide spread with a high score — a big range describes variability, not performance. Saying Class A did better because it has the smaller range makes the same mistake in the other direction: the two classes are tied on the mean, so neither one 'did better'. Saying the classes cannot be compared because their means are equal misses the whole point of also comparing the range. Always compare both an average AND a spread before describing two data sets — either one alone tells only half the story.
- (c) 80 — Method: the range is a measure of spread and is found by subtracting the smallest value from the largest. Working: the largest number sold is 100 and the smallest is 20, so the range is 100 − 20 = 80. Answer: 80. The distractors: 100 comes from writing down the largest value and never subtracting the smallest; 60 comes from working out the median, the middle value of 20, 40, 60, 70, 100, instead of the range; 58 comes from working out the mean, 290 ÷ 5, which measures centre rather than spread.
- (d) 6 — Method: the mode, or modal value, is the value that occurs most often in the data set, and it is a value from the data rather than a count. Working: size 4 occurs twice, size 5 occurs once, size 6 occurs three times and size 9 occurs once, so the highest frequency is three and the size it belongs to is 6. Answer: 6. The distractors: 3 comes from writing down the frequency of the most common size instead of the size itself; 9 comes from picking the largest size in the list, which confuses the mode with the maximum; 4 comes from stopping at the first size that repeats rather than checking which size repeats most often.
- (d) It asks two things at once and invites agreement — Method: a survey question is faulty when a reply to it cannot be read as evidence about one single thing, so check how many claims it contains and whether its wording pushes the reader one way. Working: the question joins two separate claims, that learning matters and that more homework should be set, so a reply of yes could mean either of them or both and cannot be counted as evidence about homework; the opening words “Do you agree” also invite agreement instead of leaving the reader free to say no. Answer: it asks two things at once and invites agreement. The distractors: the reply calling it too short mistakes length for clarity, when the fault is that too much has been packed in rather than too little; the reply about long words is false, since every word in the question is an everyday one and the fault lies in what is being asked rather than in the vocabulary used to ask it; the reply that the question is fine takes a yes or no answer as proof that the question works, which is exactly what a double question defeats.
- (d) 25% — First find the number of fruit cakes: 80 − 34 − 26 = 20. Then write this as a percentage of the total: 20 ÷ 80 × 100 = 25%. Giving 20% comes from reporting the count of fruit cakes, 20, directly as a percentage, without dividing by the total of 80 first. Giving 32.5% computes the percentage of chocolate cakes instead of fruit cakes: 26 ÷ 80 × 100 = 32.5%. Giving 42.5% computes the percentage of sponge cakes instead of fruit cakes: 34 ÷ 80 × 100 = 42.5%.
Build your own mix at the worksheet builder.