Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Non-calculator
Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- 1.A shop records the colour of every car that uses its car park one Tuesday. Its frequency table reads: black 44, silver 37, blue 26, red 18, white 12. Write down the frequency of black cars.
- 2.Amelia's mean mark over four tests is 29. Her first three marks are 31, 26 and 26. Work out her fourth mark.
- 3.An estate agent in Leeds lists the prices of the five houses sold on one street last year: £170,000, £150,000, £580,000, £140,000 and £160,000. A newspaper wants to print one figure for a typical price on this street. Work out the mean and the median, and write down which figure the newspaper should print.
- 4.A netball team scored 50, 83 and 68 points in three matches. Work out the mean number of points scored.
- 5.A scatter graph has 12 plotted points. Four pupils each draw a line of best fit on the same graph and count how many of the 12 points lie above their line and how many lie below it: Amir — 10 above, 2 below. Priya — 6 above, 6 below. Kofi — 2 above, 10 below. Leah — 0 above, 12 below. Write down the name of the pupil whose line was drawn correctly, so that the points are roughly balanced above and below it.
- 6.A school posts a questionnaire about school dinners to the families of all 200 pupils. Only 30 families send theirs back, and 27 of those 30 say they are unhappy with school dinners. Give a reason why this result may not represent all 200 families.
- 7.The rainfall, in millimetres, was recorded in Cambridge on six days: 12, 5, 9, 15, 3 and 11. Work out the range of the rainfall.
- 8.A survey of 200 car owners in Bristol records the colour of each car: silver 74, black 52, blue 40 and red 34. Write down which average should be used to describe a typical car in this survey, and give a reason for your answer.
- 9.A random sample of 50 pupils at a school were asked whether they prefer sport to music. 30 of the 50 pupils said they prefer sport. The school has 500 pupils altogether. Work out an estimate for the number of the 500 pupils who prefer sport.
- 10.The mean of 5 numbers is 8. Work out the total of the 5 numbers.
- 11.Noah wrote down the numbers 9, 3, 7, 1, 5 and said, “The median is 7, because 7 is in the middle of my list.” Is Noah right? Give a reason for your answer.
- 12.A school has 1,000 pupils. A student wants to estimate how many of them walk to school, so she asks 8 pupils in her own class. She says her sample is large enough to give a reliable estimate for the whole school. Is she right? Give a reason for your answer.
- 13.A charity shop in Bath holds 2,000 books. Volunteer A checks a random sample of 50 books and finds 35 paperbacks. Volunteer B checks a different random sample of 50 books and finds 31 paperbacks. Work out the estimate each sample gives for the whole stock, and write down what the shop should do next.
- 14.On a scatter graph of the age of a car, in years, and its value, in pounds, the points fall from left to right. Write down the type of correlation shown.
- 15.A café sold 10 sandwiches on each of four days: 10, 10, 10, 10. Work out the range of the numbers sold.
Answer key
- (d) 44 — Method: the frequency of a category in a frequency table is the number of times that category was counted, and it is read from the row for that category. Working: the rows of the table pair each colour with its count, and the row for black is paired with the count 44, so the frequency of black cars is 44. Answer: 44 cars — a frequency is a count of cars, not a colour and not a percentage. The distractors: 37 comes from reading the count paired with silver, that is from reading the wrong row of the table; 137 comes from adding every count in the table, 44 + 37 + 26 + 18 + 12, which gives the total number of cars rather than the frequency of one colour; 5 comes from counting how many different colours the table lists instead of how many cars were black.
- (c) 33 — Method: multiply the mean by the number of tests to get the total marks, then subtract the marks that are already known. Working: four tests with a mean of 29 give a total of 29 × 4 = 116 marks; the first three marks total 31 + 26 + 26 = 83; so the fourth mark is 116 − 83 = 33. Answer: 33, and checking, (31 + 26 + 26 + 33) ÷ 4 = 116 ÷ 4 = 29. The distractors: 116 comes from stopping at the total for all four tests; 29 comes from assuming the missing mark must be the mean itself; 4 comes from multiplying the mean by 3, the number of marks given, leaving 87 − 83 = 4.
- (c) The median, £160,000, as one very high price lifts the mean — Method: find both averages, then choose the one that sits closer to the bulk of the data. Working: in order the prices are 140,000, 150,000, 160,000, 170,000 and 580,000, so the median is the third of the five, £160,000. For the mean, 140,000 + 150,000 + 160,000 + 170,000 + 580,000 = 1,200,000 and 1,200,000 ÷ 5 = 240,000, so the mean is £240,000. Four of the five houses sold for £170,000 or less, so a reader told that a typical price is £240,000 would expect to pay at least £70,000 more than any of those four cost. Answer: the median, £160,000, as one very high price lifts the mean. The distractors: £580,000 is the middle value of the list as it is printed, which is the median only when the values have first been put in order; £240,000 is the mean, chosen on the ground that a median ignores three of the five prices, but a median uses all five to find which one is central and is then untroubled by how extreme the outer values are; £155,000 comes from deleting the £580,000 house and taking the mean of what is left, since 140,000 + 150,000 + 160,000 + 170,000 = 620,000 and 620,000 ÷ 4 = 155,000, but a real sale may not be thrown away merely for being large.
- (c) 67 — Method: the mean is the total of the values divided by how many values there are, so add first and divide second. Working: the total is 50 + 83 + 68 = 201 points and three matches were played, so the mean is 201 ÷ 3 = 67 points. Answer: 67. The distractors: 68 comes from writing down the median, the middle value of 50, 68, 83, instead of the mean; 33 comes from working out the range, 83 − 50, which measures spread and not centre; 100.5 comes from dividing the total by 2 instead of by the 3 matches played.
- (a) Priya — A line of best fit should be drawn so that the plotted points are roughly balanced above and below it. Work out the difference between the two counts for each pupil: Amir 10 − 2 = 8; Kofi 10 − 2 = 8 (10 below and 2 above); Leah 12 − 0 = 12; Priya 6 − 6 = 0. Priya's line has the smallest difference, an exact balance of 6 above and 6 below, so her line is drawn correctly. Amir's line has 10 of the 12 points above it, so it is drawn too low. Kofi's line has 10 of the 12 points below it, so it is drawn too high. Leah's line has every single point below it, so it is not a line of best fit at all.
- (d) Only families with strong feelings bothered to reply. — Method: a survey has non-response bias when only some of the people asked actually reply, and those who do are not a typical cross-section of everyone who was asked. Working: only 30 of the 200 families sent back their questionnaire, and 27 of those 30 — the great majority — said they were unhappy. Families who feel strongly about an issue, particularly those with a complaint, are far more likely to make the effort to reply than families who are simply satisfied and see no need to say anything, so the 30 replies over-represent unhappy families. Saying the families who replied were picked at random by the school gets the sampling the wrong way round: nobody picked them — they picked themselves by deciding to reply, and that is precisely why they are not a typical cross-section of all 200. Saying postal surveys always have low response rates restates that the response was low without explaining why a low response rate, on its own, makes a result unrepresentative — it is the reason FOR the low response, not the low response itself, that causes the bias here. Saying that the 27 unhappy replies show most families are unhappy is exactly the mistake the question is warning against: it treats the loudest 30 replies as if they stood for the other 170 who never sent theirs back. A low response rate is a warning sign only because the people who bother to reply are rarely typical of everyone who was asked.
- (a) 12 mm — Method: the range is the highest value minus the lowest value. Working: the highest rainfall is 15 mm and the lowest is 3 mm, so the range is 15 − 3 = 12 mm. Giving 15 mm alone states the highest value, not the range. Giving 3 mm alone states the lowest value, not the range. Sorting the six values, 3, 5, 9, 11, 12 and 15, and averaging the middle two, (9 + 11) ÷ 2 = 10 mm, finds the median, a completely different statistic. The range always needs BOTH the highest and the lowest value — never just one of them.
- (d) The mode, because colours cannot be added or ordered — Method: an average can only be used on data that supports the operation it needs. A mean needs the values to be added and divided, a median needs them to be placed in order, and a range needs one value to be taken away from another; a mode needs only counting, so it is the average available when the data are categories rather than numbers. Working: the data collected here are colours, silver, black, blue and red. The numbers 74, 52, 40 and 34 count the cars of each colour, they do not measure them, and 74 + 52 + 40 + 34 = 200 simply returns the size of the survey. No colour can be added to another, and there is no order that puts blue before red, so of the four averages only the one found by counting survives. Answer: the mode, because colours cannot be added or ordered, and the mode is silver. The distractors: the mean is said to use all 200 colours, and a mean of the four frequencies, 200 ÷ 4 = 50, is a number of cars rather than a colour, so it describes nothing about a typical car; the median is said to put the colours in order, but ordering the frequencies 34, 40, 52, 74 orders the counts, not the colours, and gives 46, again a number of cars; the range is not an average at all, and 74 − 34 = 40 measures the gap between the commonest and rarest counts, which is a measure of spread.
- (d) 300 pupils — Method: an estimate for a whole population is made by finding the proportion in the sample and applying that same proportion to the population. Working: in the sample 30 of the 50 pupils prefer sport, a proportion of 30 ÷ 50 = 0.6, and applying that proportion to the school gives 0.6 × 500 = 300 pupils. Answer: 300 pupils, and it is only an estimate, because a different random sample of 50 would give a slightly different figure. The distractors: 200 pupils comes from scaling up the 20 pupils in the sample who did not prefer sport, 20 × 10, which answers the opposite question; 150 pupils comes from reading 30 out of 50 as 30% and taking 30% of 500; 60 pupils comes from working out the proportion correctly as 60% and then writing the 60 down as a number of pupils instead of applying it to the 500.
- (a) 40 — Method: the mean is the total divided by how many values there are, so rearranging gives total = mean × number of values. Working: the mean is 8 and there are 5 numbers, so the total is 8 × 5 = 40. Answer: 40, and checking, 40 ÷ 5 = 8, which is the mean given. The distractors: 13 comes from adding the mean and the count, 8 + 5, instead of multiplying them; 1.6 comes from dividing the mean by the count, 8 ÷ 5, which reverses the relationship; 8 comes from quoting the mean itself as the total, which is only true when there is a single number.
- (b) No — in order the numbers are 1, 3, 5, 7, 9, so the median is 5. — Method: the median is the middle value of the data in order of size, so the data must be sorted before any position is read off. Working: Noah's list 9, 3, 7, 1, 5 is not in order; sorted it becomes 1, 3, 5, 7, 9, and with 5 values the middle position is the third, which now holds 5 rather than 7. Noah has read the third value of the unsorted list. Answer: no — in order the numbers are 1, 3, 5, 7, 9, so the median is 5. The distractors: the reply giving 3 as the median sorts the data correctly but then reads the value in the second place instead of the third; the reply that 7 is the third number he wrote accepts a position in the unsorted list, which is exactly the mistake the question is about; the reply using the mean claims a value of 7 for it, but the mean is 25 ÷ 5 = 5, so that reasoning is false as well.
- (c) No — 8 from one class is too small to represent the school. — Method: judge reliability by asking whether the sample is both large enough, and spread across the population, relative to what it is meant to represent. Working: 8 pupils is a tiny fraction of the school's 1,000 pupils, and all 8 come from a single class rather than a range of year groups, so the sample is both too small and too narrow to represent the whole school reliably. She is not right. Saying any sample size gives an equally reliable estimate ignores that reliability generally improves with a larger, more representative sample. Saying the method is unreliable because it was not done online is not a reason connected to sample size or representativeness at all. Saying 8 is reliable because it is more than half her class compares the sample to the wrong population — the school has 1,000 pupils, not one class. Always judge a sample's size against the population it is meant to represent, not against a smaller group within it.
- (a) 1,400 and 1,240, so combine the samples for one estimate — Method: scale each sample up to the whole stock, then use the fact that a larger sample gives a more reliable estimate than a smaller one. Working: the first sample gives 35 ÷ 50 = 0.7 and 0.7 × 2,000 = 1,400 paperbacks; the second gives 31 ÷ 50 = 0.62 and 0.62 × 2,000 = 1,240 paperbacks. Two random samples of the same size are expected to differ a little, so neither estimate is wrong. Putting the two together gives 35 + 31 = 66 paperbacks in 100 books, and 66 ÷ 100 = 0.66 with 0.66 × 2,000 = 1,320, an estimate resting on twice as many books as either volunteer checked. Answer: 1,400 and 1,240, so combine the samples for one estimate. The distractors: keeping 1,400 because it is larger picks an estimate by its size, when both samples held 50 books and neither has a stronger claim; saying a volunteer must have miscounted assumes two random samples ought to agree exactly, which is precisely what random sampling does not promise; 1,750 and 1,550 come from 35 × 50 = 1,750 and 31 × 50 = 1,550, multiplying each count by the size of the sample instead of scaling by 2,000 ÷ 50.
- (a) Negative correlation — As the age of the car increases, the points fall towards a lower value, so the value decreases as the age increases. This falling pattern is a negative correlation. A positive correlation would show the points rising together instead. No correlation would apply only if the points showed no pattern at all, and correlation is not the same as causation — strong causation is not a type of correlation.
- (d) 0 — Method: the range is the largest value minus the smallest value, whatever those two values turn out to be. Working: every value is 10, so the largest value is 10 and the smallest value is 10 as well, and the range is 10 − 10 = 0. Answer: 0 — a range of nothing says the data do not vary at all. The distractors: 10 comes from writing down the repeated value itself instead of the difference between the extremes; 20 comes from adding the largest and the smallest, 10 + 10, instead of subtracting; 40 comes from adding all four values, which gives the total sold and not a measure of spread.
Build your own mix at the worksheet builder.