Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- 1.A scatter graph shows the mass, x kg, of a parcel and the cost, y pounds, of posting it. The line of best fit is y = 1.5x + 2. Work out the estimated cost of posting a parcel with a mass of 6 kg, using the line of best fit.y = 1.5x + 2
- 2.A town council wants to find out about the eating habits of the people who live in the town. It asks only the members of a local sports club. Give a reason why this sample is biased.
- 3.A survey of 25 pupils in Derby records how many siblings each has: 0 siblings — 6 pupils, 1 sibling — 10 pupils, 2 siblings — 6 pupils, 3 siblings — 3 pupils. Calculate the mean number of siblings.
- 4.A scatter graph plots the shoe size and the spelling test score of 25 pupils. The points are scattered with no pattern across the graph. Write down the type of correlation shown.
- 5.A teacher records the number of pets owned by each of 25 pupils in a class; each pupil owns 0, 1, 2, 3 or 4 pets. The teacher wants to show how many pupils own each number of pets. Write down the most suitable type of chart for this data, and give a reason for your answer.
- 6.A school posts a questionnaire about school dinners to the families of all 200 pupils. Only 30 families send theirs back, and 27 of those 30 say they are unhappy with school dinners. Give a reason why this result may not represent all 200 families.
- 7.A two-way table records whether each of 40 people at a gym in Sheffield prefers weights or cardio, and whether they are male or female. 24 of the 40 people are female. 15 of the males prefer cardio. 10 of the females prefer weights. Work out the total number of people who prefer cardio.
- 8.Four pairs of variables are listed below. Write down the pair that you would expect to show negative correlation.
- 9.A sports centre in Newcastle asks 60 members which activity they prefer: swimming 15 members, football 25 members, badminton 12 members and other activities 8 members. Work out the angle at the centre of the pie chart sector that represents badminton.
- 10.A gym in Cardiff has a scatter graph of the number of training sessions, x, attended by each member and the weight lost, y kg. The line of best fit is y = 0.4x + 1. Members want to lose at least 9 kg. Using the line of best fit, work out the least whole number of sessions needed to reach this target.y = 0.4x + 1
- 11.A scatter graph shows the number of hours of sunshine, x, and the number of visitors, y, at an outdoor swimming pool in Torquay on each of 15 days. The line of best fit passes through the points (5, 150) and (15, 350). Work out the estimated number of visitors on a day with 8 hours of sunshine, using the line of best fit.
- 12.A scatter graph shows the number of days, x, that each of 16 tomato plants was watered and its height, y cm. The line of best fit has equation y = 1.5x + 4. Write down what the 1.5 in this equation tells you about the plants.y = 1.5x + 4
- 13.A sports club has 300 women members and 200 men members, 500 members in total. The club committee wants to choose a sample of 40 members that is in proportion to the club's membership. Work out how many women should be in the sample.
- 14.A scatter graph shows the number of guests, x, at a wedding and the length of buffet table needed, y metres. The line of best fit is y = 0.5x + 2. Write down what the 2 in this equation tells you about the buffet table.y = 0.5x + 2
- 15.Ben and Chloe each sat five maths tests. Ben's marks were 62, 64, 65, 66 and 68. Chloe's marks were 40, 52, 65, 78 and 90. Both pupils have a mean mark of 65. Their teacher says the mean on its own does not describe the two sets of marks well. Give a reason why the teacher is right.
Answer key
- (b) £11.00 — 1.5 × 6 = 9, and 9 + 2 = 11, so the estimated cost is £11.00. Choosing £9.00 stops after 1.5 × 6 = 9 and forgets to add the £2. Choosing £12.00 adds the mass and the constant first and then multiplies: 6 + 2 = 8, and 8 × 1.5 = 12.00. Choosing £13.50 swaps the gradient and the intercept, using y = 2x + 1.5 instead: 2 × 6 = 12, and 12 + 1.5 = 13.50.
- (c) Club members probably eat differently from most people — Method: a sample is biased when the group it is drawn from differs from the population in the very thing the survey is measuring, so compare the subgroup with the population on that quantity. Working: the survey measures eating habits, and people who join a sports club take more exercise than average and are known to eat differently from the town as a whole, so their replies pull the results away from the true picture for the town however many of them are asked. Answer: club members probably eat differently from most people. The distractors: the reply about the number of members treats bias as a question of size, but a large biased sample is still biased; the reply that the members were picked at random is false, since the council picked a club rather than picking residents, and it confuses bias with non-response; the reply that everyone asked lives in the town notes something true of the members but draws the false conclusion that the sample therefore covers the town, when a sample must reflect a population and not merely be taken from inside it.
- (a) 1.24 — Method: for data given as a frequency table, the mean is Σfx ÷ Σf — multiply each value by its frequency, add the results, then divide by the total frequency. Working: 0 × 6 = 0. 1 × 10 = 10. 2 × 6 = 12. 3 × 3 = 9. So Σfx = 0 + 10 + 12 + 9 = 31. The total frequency is Σf = 6 + 10 + 6 + 3 = 25. Mean = 31 ÷ 25 = 1.24 siblings. Averaging the frequency column itself, (6 + 10 + 6 + 3) ÷ 4 = 6.25, mixes up the frequencies with the values they belong to. Writing down 1, the number of siblings with the highest frequency, gives the mode, not the mean. Writing down 31 stops after finding Σfx and forgets to divide by the total frequency, 25. Always divide Σfx by Σf — never stop at the top of the fraction.
- (a) No correlation — Shoe size has no real relationship with spelling ability, and the points here are scattered with no rising or falling trend, so this is no correlation. A positive correlation would show the points rising together, and a negative correlation would show them falling as one increases; neither pattern is present here. Strong correlation is not correct either, since strength only applies once a positive or negative trend exists, and there isn't one.
- (a) A vertical line chart (discrete numerical data) — The number of pets is discrete numerical data — whole-number values such as 0, 1, 2, 3 or 4 — recorded for one variable, so a vertical line chart is the chart specified for this kind of data. A bar chart is used for categorical data, such as favourite colour, not numerical values counted like this. A pie chart shows proportions of a whole and does not show the frequency of each separate value. A scatter graph compares two different variables against each other, and only one variable, the number of pets, is recorded here.
- (d) Only families with strong feelings bothered to reply. — Method: a survey has non-response bias when only some of the people asked actually reply, and those who do are not a typical cross-section of everyone who was asked. Working: only 30 of the 200 families sent back their questionnaire, and 27 of those 30 — the great majority — said they were unhappy. Families who feel strongly about an issue, particularly those with a complaint, are far more likely to make the effort to reply than families who are simply satisfied and see no need to say anything, so the 30 replies over-represent unhappy families. Saying the families who replied were picked at random by the school gets the sampling the wrong way round: nobody picked them — they picked themselves by deciding to reply, and that is precisely why they are not a typical cross-section of all 200. Saying postal surveys always have low response rates restates that the response was low without explaining why a low response rate, on its own, makes a result unrepresentative — it is the reason FOR the low response, not the low response itself, that causes the bias here. Saying that the 27 unhappy replies show most families are unhappy is exactly the mistake the question is warning against: it treats the loudest 30 replies as if they stood for the other 170 who never sent theirs back. A low response rate is a warning sign only because the people who bother to reply are rarely typical of everyone who was asked.
- (b) 29 — There are 40 − 24 = 16 males, and 15 of them prefer cardio, so 16 − 15 = 1 male prefers weights. There are 24 females, and 10 prefer weights, so 24 − 10 = 14 females prefer cardio. Altogether, 15 + 14 = 29 people prefer cardio. Choosing 15 only counts the males who prefer cardio and forgets the females. Choosing 11 adds the two weights figures, 1 + 10 = 11, instead of the two cardio figures. Choosing 30 comes from 40 − 10, subtracting only the number of females who prefer weights from the grand total, rather than finding both cardio sub-totals separately.
- (a) Minutes a candle has burned and length remaining — As a candle burns for longer, less of it remains, so these two variables move in opposite directions as one increases — that is negative correlation. A pupil's shoe size generally increases as they get older, so age and shoe size show positive correlation, not negative, since both rise together. A football team's shirt colour is not a numerical quantity linked to how many matches it wins, so shirt colour and number of wins show no correlation at all. The number of letters in a pupil's name has no real connection to their ability in maths, so that pair also shows no correlation.
- (c) 72° — The angle is 12 ÷ 60 × 360 = 72°. Choosing 150° divides football's frequency of 25 instead of badminton's 12: 25 ÷ 60 × 360 = 150. Choosing 20° finds badminton as a percentage of the members, 12 ÷ 60 × 100 = 20, rather than an angle in degrees. Choosing 90° uses 48, the total of the other three activities, as the total instead of the full 60 members: 12 ÷ 48 × 360 = 90.
- (d) 20 — 0.4x + 1 = 9, so 0.4x = 9 − 1 = 8, and 8 ÷ 0.4 = 20, so 20 sessions are needed. Choosing 23 divides 9 by 0.4 without first subtracting the 1: 9 ÷ 0.4 = 22.5, rounded up to 23. Choosing 25 subtracts the wrong way, adding the 1 instead of taking it away: 9 + 1 = 10, and 10 ÷ 0.4 = 25. Choosing 2 misplaces the decimal point in the gradient, dividing by 4 instead of by 0.4: 8 ÷ 4 = 2.
- (d) 210 — 350 − 150 = 200. 200 ÷ 10 = 20, so the gradient is 20. Using the point (5, 150): 20 × 5 = 100, so 150 − 100 = 50 is the intercept, giving the line y = 20x + 50. At x = 8: 20 × 8 = 160, and 160 + 50 = 210, so the estimated number of visitors is 210. Choosing 160 stops after 20 × 8 = 160 and forgets to add the intercept of 50. Choosing 250 comes from averaging the two given y-values: 150 + 350 = 500, and 500 ÷ 2 = 250, instead of using the line's equation. Choosing 240 assumes the visitors are directly proportional to the hours of sunshine using the first point, 150 × 8 ÷ 5 = 240, which ignores that the line does not pass through the origin.
- (b) On average a plant grew 1.5 cm taller for each extra day — Method: in the equation of a line, the number multiplying x is the gradient, and a gradient states the change in y produced by an increase of 1 in x, read in the units of the two axes. Working: here x is measured in days and y in centimetres, so the gradient 1.5 carries the units centimetres per day. Testing it on the line, 5 days gives 1.5 × 5 + 4 = 11.5 cm and 6 days gives 1.5 × 6 + 4 = 13 cm, a rise of 1.5 cm for the one extra day. Answer: on average a plant grew 1.5 cm taller for each extra day of watering. The distractors: 1.5 cm as the height before any watering is the value of y when x is 0, which is the other number in the equation, 4 cm, so this swaps the gradient and the intercept; 1.5 cm as the gap between the tallest and the shortest plant reads the gradient as a range, when a range is a difference between two of the 16 plants and a gradient is a rate; 1.5 days for each extra centimetre inverts the rate, dividing days by centimetres instead of centimetres by days, and the line gives 1 cm of growth in two thirds of a day.
- (d) 24 — Method: for a sample in proportion to the population, apply the same fraction that each group makes up of the whole population to the size of the sample. Working: women make up 300 out of the 500 members, a fraction of 300 ÷ 500 = 0.6. Applying that fraction to the sample of 40 gives 0.6 × 40 = 24 women. Splitting the sample evenly, 40 ÷ 2 = 20, ignores that the club has more women than men and treats the two groups as equal in size, which they are not. Misreading the sample size as 50 instead of 40, then applying the 3:2 ratio of women to men, 3 ÷ 5 × 50 = 30, uses the right ratio but the wrong sample total. Working out the number of MEN instead of women, 200 ÷ 500 × 40 = 16, answers a different question — how many men, not how many women, belong in the sample. Always apply each group's own share of the population to the sample size, and check which group the question is actually asking about.
- (b) At 0 guests, the model predicts 2 m of table — The y-intercept of a line of best fit y = mx + c is the value of y when x = 0. Here y = 0.5 × 0 + 2 = 2, so the line predicts a table length of 2 m when there are 0 guests. The 2 m does not grow as more guests arrive — that role belongs to the gradient, 0.5 — so an option saying each extra guest adds 2 m has swapped the two numbers around. The 2 is a length in metres, not a number of guests, so an option requiring 2 guests before set-up has misread its units. And the table length does change with x, since it is 0.5x + 2 and not a fixed value, so an option claiming the table is always 2 m ignores the 0.5x term completely.
- (a) Chloe's marks are far more spread out than Ben's — Method: a mean reports where a set of values sits, and two sets can sit in the same place while behaving quite differently, so a measure of spread has to be worked out as well. Working: Ben's marks add to 62 + 64 + 65 + 66 + 68 = 325 and 325 ÷ 5 = 65; Chloe's add to 40 + 52 + 65 + 78 + 90 = 325 and 325 ÷ 5 = 65, so the two means agree, as the question says. The ranges do not: Ben's is 68 − 62 = 6 marks, while Chloe's is 90 − 40 = 50 marks. Ben's five marks all sit within 3 marks of 65; Chloe's lowest is 25 marks below it and her highest 25 marks above it. Answer: Chloe's marks are far more spread out than Ben's, which is exactly what the mean cannot show. The distractors: saying Ben's marks are more spread out comes from subtracting in the order the values are written, 62 − 68 = −6 against 40 − 90 = −50, and then reading −6 as the larger spread; saying Chloe scored far more marks in total assumes a wider set of marks must add to more, when both totals are 325; saying the two sets vary by the same amount assumes that equal means force equal spread, when the two ranges are 6 and 50.
Build your own mix at the worksheet builder.