Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- 1.A straight line is drawn on a scatter graph to show the trend of the points. Write down the name given to this line.
- 2.In a scatter graph of the age, in years, and the wingspan, in cm, of 20 birds of the same species, all the points lie close to a rising line of best fit except one, which lies a long way below the line. That bird was later found to have a damaged wing. Give a reason why this point should not be used when drawing the line of best fit.
- 3.On a scatter graph of the age of a car, in years, and its value, in pounds, the points fall from left to right. Write down the type of correlation shown.
- 4.In a vertical line chart of test results, the tallest line stands for 40 pupils. The same results are to be shown in a pie chart of all 100 pupils. Work out the angle at the centre of the sector that will stand for those 40 pupils.
- 5.A sports shop in Cardiff sold 40 pairs of football boots last month: 4 pairs of size 6, 5 pairs of size 7, 8 pairs of size 8, 13 pairs of size 9 and 10 pairs of size 10. The manager will order 40 pairs for next month and wants as many pairs as possible to be in a size customers will buy. Work out the mean size and the modal size, and write down which of the two he should use.
- 6.A stem-and-leaf diagram, described in words, shows the ages of 9 people at a family party. The stem is the tens digit: stem 1 has leaves 4 and 8; stem 2 has leaves 0, 3, 5 and 9; stem 3 has leaves 1 and 6; stem 4 has leaf 2. Work out the median age.
- 7.A sports centre in Ipswich has 2,000 members. It wants to know what its members think of its opening hours, so it asks a random sample of 100 of them. Write down what the population is in this survey.
- 8.A school posts a questionnaire about school dinners to the families of all 200 pupils. Only 30 families send theirs back, and 27 of those 30 say they are unhappy with school dinners. Give a reason why this result may not represent all 200 families.
- 9.A bar chart shows the highest temperature recorded on each of five days. The temperatures were Monday 20 °C, Tuesday 22 °C, Wednesday 18 °C, Thursday 25 °C and Friday 23 °C. Write down the day on which the temperature was highest.
- 10.Five pupils spent these numbers of minutes on their homework: 50, 65, 55, 90, 60. Work out the median time.
- 11.A line graph shows the temperature in a greenhouse at the end of each of six hours. The six readings were 10 °C, 14 °C, 18 °C, 22 °C, 26 °C and 30 °C. Describe the trend shown by the line graph.
- 12.The seven members of Team A took 20, 21, 22, 23, 24, 25 and 40 seconds to finish a task. The seven members of Team B took 20, 30, 32, 34, 36, 38 and 40 seconds. Tomás says that because the two teams have the same range, their times are spread out in the same way. Is Tomás right? Give a reason for your answer.
- 13.The age, in years, and the score in a reaction test are recorded for eight members of a sports club: (14, 92), (18, 88), (23, 85), (27, 80), (31, 78), (36, 74), (42, 70), (49, 65). The eight pairs are plotted on a scatter graph. Describe the correlation between age and score.
- 14.In a random sample of 40 pupils at a school, 6 are left-handed. The school has 900 pupils. Work out an estimate for the number of left-handed pupils in the school.
- 15.A bar chart is drawn for a set of categorical data. Write down what the height of each bar represents.
Answer key
- (c) A line of best fit — Method: the straight line drawn on a scatter graph is named from the job it does — it is chosen so that it follows the whole set of points as closely as possible. Working: the line passes through the middle of the points, with roughly as many points above it as below it, and it need not pass through any of the plotted points at all; the name given to the straight line chosen in that way is a line of best fit. Answer: a line of best fit. The distractors: a line of symmetry comes from confusing a trend with symmetry, which is a property of a shape rather than of a set of data; a horizontal line through the mean comes from thinking the trend is shown by an average, when a horizontal line would say that the vertical quantity does not change and so show no correlation; a line joining the first and last points comes from thinking the line must join the two extreme points, which lets two points decide a trend that all of the points should share in.
- (d) An outlier from the damaged wing, not the trend. — That bird's point lies a long way from the rising trend followed by every other bird, and its low wingspan is explained by the damaged wing rather than by its age — it is an outlier caused by an unusual factor, not part of the general relationship between age and wingspan, so it should not be used when drawing the line of best fit. Saying every point must be used ignores that an outlier caused by a separate, identifiable factor can rightly be set aside. Saying it shows no correlation ignores that the other 19 points do show a clear rising trend; one outlier does not remove that. Saying it proves the line is inaccurate confuses one unusual bird with a fault in the line itself, when the line correctly describes the trend followed by the rest of the data.
- (a) Negative correlation — As the age of the car increases, the points fall towards a lower value, so the value decreases as the age increases. This falling pattern is a negative correlation. A positive correlation would show the points rising together instead. No correlation would apply only if the points showed no pattern at all, and correlation is not the same as causation — strong causation is not a type of correlation.
- (a) 144° — Method: a pie chart shares the 360° at its centre between the categories in proportion to their frequencies, so the angle of a sector is that category's fraction of the total multiplied by 360°. Working: the sector stands for 40 pupils out of 100, which is the fraction 40/100; one pupil is worth 360 ÷ 100 = 3.6°, so 40 pupils are worth 40 × 3.6° = 144°. Answer: 144°, and the answer is an angle in degrees, not a number of pupils. The distractors: 40° comes from sharing out 100 instead of 360, so the percentage is written straight down as a number of degrees; 216° comes from working out the angle for the other 60 pupils, 60 × 3.6°, which is the rest of the pie chart; 180° comes from assuming that the tallest line must stand for half of the pupils and so take half of the chart.
- (a) The modal size, 9, bought by more customers than any other — Method: work out both averages from the frequencies, then choose the one the shop can act on. Working: for the mean, multiply each size by the number of pairs sold at it and add: 6 × 4 + 7 × 5 + 8 × 8 + 9 × 13 + 10 × 10 = 340, and 340 ÷ 40 = 8.5, so the mean size is 8.5. The largest frequency is 13, which belongs to size 9, so the modal size is 9. The mean 8.5 is a size no customer in the record asked for, so 40 pairs of it would sit unsold, while 13 of the 40 customers wanted size 9, more than wanted any other size. Answer: the modal size, 9, bought by more customers than any other. The distractors: the mean size 8.5 does take account of all 40 pairs, but a mean of sizes is a summary figure and not a size the month's customers were buying; the mean size 8 comes from averaging the five sizes on sale, 6 + 7 + 8 + 9 + 10 = 40 and 40 ÷ 5 = 8, which ignores how many pairs were sold at each size and so treats the 4 pairs of size 6 as equal in weight to the 13 pairs of size 9; the range 4 comes from 10 − 6 and measures spread, so it says how wide a set of sizes the shop must stock, not which size to stock most of.
- (d) 25 — In order, the nine ages are 14, 18, 20, 23, 25, 29, 31, 36 and 42, and with 9 values the median is the 5th one, which is 25. Choosing 23 takes the 4th value instead of the 5th. Choosing 29 takes the 6th value instead of the 5th. Choosing 26 comes from averaging the 4th and 6th values, 23 + 29 = 52, and 52 ÷ 2 = 26, a method that is only needed when there is an even number of values.
- (b) All 2,000 members of the sports centre. — Method: in a survey, the population is the whole group the survey is trying to find out about, and the sample is the smaller group actually asked. Working: this survey wants to know what the sports centre's members think, so the population is every one of the 2,000 members — whether or not they were personally asked. Saying the population is the 100 members who were asked names the sample, not the population; the sample is drawn FROM the population, so it is smaller than it, not the same as it. Saying the population is everybody who lives in Ipswich widens the group far beyond who the survey is actually about — plenty of Ipswich residents are not members of the sports centre at all, so they are outside this survey altogether. Saying the population is the members who say they are unhappy confuses the population with a result of the survey: whether a member turns out to be happy or unhappy is something the survey finds out, not part of the definition of who is being studied. The population is always the whole group the question is about, before any sampling or any results come in.
- (d) Only families with strong feelings bothered to reply. — Method: a survey has non-response bias when only some of the people asked actually reply, and those who do are not a typical cross-section of everyone who was asked. Working: only 30 of the 200 families sent back their questionnaire, and 27 of those 30 — the great majority — said they were unhappy. Families who feel strongly about an issue, particularly those with a complaint, are far more likely to make the effort to reply than families who are simply satisfied and see no need to say anything, so the 30 replies over-represent unhappy families. Saying the families who replied were picked at random by the school gets the sampling the wrong way round: nobody picked them — they picked themselves by deciding to reply, and that is precisely why they are not a typical cross-section of all 200. Saying postal surveys always have low response rates restates that the response was low without explaining why a low response rate, on its own, makes a result unrepresentative — it is the reason FOR the low response, not the low response itself, that causes the bias here. Saying that the 27 unhappy replies show most families are unhappy is exactly the mistake the question is warning against: it treats the loudest 30 replies as if they stood for the other 170 who never sent theirs back. A low response rate is a warning sign only because the people who bother to reply are rarely typical of everyone who was asked.
- (c) Thursday — Method: on a bar chart the tallest bar belongs to the greatest value, so compare the five temperatures and then read off the day that the greatest one belongs to. Working: the temperatures are 20 °C, 22 °C, 18 °C, 25 °C and 23 °C; in order of size these are 18, 20, 22, 23 and 25, so the greatest temperature is 25 °C, and the day recorded with 25 °C is Thursday. Answer: Thursday — the answer is a day, not a temperature. The distractors: Friday comes from stopping at the final temperature listed instead of comparing all five; Wednesday comes from picking out the shortest bar, 18 °C, and so answering for the lowest temperature rather than the highest; Monday comes from writing down the first day in the chart without comparing any of the temperatures at all.
- (c) 60 minutes — Method: the median is the middle value once the data have been put in order of size, so the list must be sorted before any position is read. Working: in order the times are 50, 55, 60, 65, 90 minutes; there are 5 values, so the middle position is the third and the time sitting there is 60 minutes. Answer: 60 minutes. The distractors: 64 minutes comes from working out the mean, 320 ÷ 5, instead of the median; 70 minutes comes from taking the time halfway between the shortest and the longest, (50 + 90) ÷ 2; 40 minutes comes from working out the range, 90 − 50, which measures spread rather than centre.
- (a) The temperature is rising steadily — Method: the trend of a line graph is the overall direction of the readings as time goes on, found by comparing each reading with the one before it. Working: from 10 °C to 14 °C is a rise of 4 °C, and the same comparison from 14 °C to 18 °C, from 18 °C to 22 °C, from 22 °C to 26 °C and from 26 °C to 30 °C gives a rise of 4 °C every time; every reading is greater than the one before it and none of them falls. Answer: the temperature is rising steadily — steadily because the rise is the same size each hour. The distractors: falling steadily comes from reading the six values from right to left, which reverses the direction of time; stays the same comes from noticing that the step of 4 °C is the same each hour and describing the step as constant instead of the temperature; rises and then falls comes from assuming that a line graph has to turn at some point rather than reading the values that are actually given.
- (d) No, the range uses only the fastest and slowest time — Method: check what the range is built from, then look at what it leaves out. Working: both teams have a fastest time of 20 seconds and a slowest of 40 seconds, so both ranges are 40 − 20 = 20 seconds and Tomás has that part right. But the range is calculated from those two values alone. Six of Team A's seven times lie between 20 and 25 seconds, with a single time far out at 40; Team B is the other way round, with six of its seven times at 30 seconds or more and a single time far out at 20. So Team A bunches at the fast end and Team B at the slow end. The two patterns are quite different, and the range cannot see the difference because the five middle times never enter the calculation. Answer: no, because the range uses only the fastest and slowest time. The distractors: comparing the means answers a different question, since a mean measures position rather than spread, and two sets with the same spread can have different means; saying that equal ranges mean equal spread is the very assumption that fails here; saying that seven times each forces the spreads to match confuses the size of a data set with how its values are arranged inside it.
- (c) Strong negative correlation — Method: correlation is described by two things — the direction the points take as the graph is read from left to right, and how closely the points lie to a single straight line. Working: reading the pairs in order of age, the ages rise 14, 18, 23, 27, 31, 36, 42, 49 while the scores fall 92, 88, 85, 80, 78, 74, 70, 65; the score falls at every single step, with no reversal anywhere, so the points fall from left to right and lie close to a straight line. Answer: strong negative correlation — negative for the falling direction, strong because every point follows the pattern. The distractors: strong positive correlation comes from noticing a clear pattern and calling any clear pattern positive, without checking the direction; weak negative correlation comes from reading the direction correctly but judging points that do not lie exactly on a straight line to be only loosely related, when these eight fall without a single exception; no correlation comes from reading a falling trend as though it showed no relationship at all, when a falling trend is itself a relationship.
- (b) 135 — Method: use the sample to find the PROPORTION of left-handed pupils, then apply that same proportion to the whole school population. Working: in the sample, 6 out of 40 pupils are left-handed, a proportion of 6 ÷ 40 = 0.15. Applying that proportion to the school's 900 pupils gives an estimate of 0.15 × 900 = 135 pupils. Giving 6 simply repeats the number of left-handed pupils IN THE SAMPLE, without scaling up to the whole school at all. Multiplying the population by the number of left-handed pupils in the sample without first dividing by the sample size, 900 × 6 = 5400, badly overestimates — that is more pupils than the whole school has. Dividing the population by the sample size but forgetting to multiply by the number of left-handed pupils found, 900 ÷ 40 = 22.5, finds the scale factor but stops one step short of using it. Always find the proportion in the sample first, then scale that same proportion up to the population.
- (c) The frequency of that category — Method: a bar chart for categorical data has one bar for each category, and the vertical scale on which the bars are measured is a count. Working: a bar drawn twice as tall as another tells you that twice as many items of data fell into its category, so the height measures how many items of data belong to that one category, which is exactly what a frequency is. Answer: the height of each bar is the frequency of that category — a count of items of data. The distractors: the number of different categories comes from reading the vertical scale as though it counted the bars, which is shown along the horizontal axis instead; the total of all the data comes from treating one bar as though it stood for the whole data set rather than for one category; the mean of all the data comes from confusing a bar chart with a measure of average, which no single bar can show.
Build your own mix at the worksheet builder.