Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- 1.A scatter graph shows the height, x cm, and the mass, y kg, of 20 pupils in Year 10. The heights on the graph run from 150 cm to 180 cm, and the line of best fit is y = 0.9x − 85. Nadia puts x = 90 into this equation to estimate the mass of a two-year-old child who is 90 cm tall. Is her estimate reliable? Give a reason for your answer.y = 0.9x − 85
- 2.A straight line is drawn on a scatter graph to show the trend of the points. Write down the name given to this line.
- 3.The times taken, in minutes, by 30 runners in a Portsmouth fun run are grouped in this table: 0 < t ≤ 20 — 5 runners, 20 < t ≤ 40 — 10 runners, 40 < t ≤ 60 — 10 runners, 60 < t ≤ 80 — 5 runners. Work out an estimate for the mean time, in minutes.
- 4.The seven members of Team A took 20, 21, 22, 23, 24, 25 and 40 seconds to finish a task. The seven members of Team B took 20, 30, 32, 34, 36, 38 and 40 seconds. Tomás says that because the two teams have the same range, their times are spread out in the same way. Is Tomás right? Give a reason for your answer.
- 5.Priya wants to find out the favourite sport of the 900 pupils at her school. She asks 5 pupils, chosen at random. Give a reason why her sample may not give a reliable result.
- 6.Write down what a scatter graph is used to show.
- 7.A scatter graph shows the number of years of experience, x, of 18 sales assistants and their monthly sales, y hundred pounds. The plotted points run from x = 1 to x = 12 years, and the line of best fit is y = 4x + 20. A new assistant has 25 years of experience. Use the line of best fit to estimate a value of y for this assistant, and decide whether the estimate would be reliable.y = 4x + 20
- 8.A town council wants to find out about the eating habits of the people who live in the town. It asks only the members of a local sports club. Give a reason why this sample is biased.
- 9.A survey of 200 car owners in Bristol records the colour of each car: silver 74, black 52, blue 40 and red 34. Write down which average should be used to describe a typical car in this survey, and give a reason for your answer.
- 10.A company found that the higher the price it charged for a product, the lower the satisfaction score its customers gave. The price and the satisfaction score are plotted on a scatter graph. Describe the line of best fit that would be drawn on that graph.
- 11.A scatter graph of the number of hours, x, that pupils revised against their test score, y, has the line of best fit y = 2.5x + 15. Amelia wants a score of at least 80. Work out the least whole number of hours of revision the line of best fit suggests she needs.y = 2.5x + 15
- 12.A scatter graph plots the shoe size and the spelling test score of 25 pupils. The points are scattered with no pattern across the graph. Write down the type of correlation shown.
- 13.A charity shop in Leicester made £2,400 in one month. A pie chart shows how this was spent: repairs took up an angle of 90°, wages took up an angle of 150°, and the rest was other costs. Work out how much was spent on wages.
- 14.The midday temperature in Leeds was recorded on each day of one week: 22 °C, 24 °C, 23 °C, 25 °C, 26 °C, 21 °C, 24 °C. Work out the mean midday temperature, giving your answer to 2 decimal places.
- 15.A scatter graph shows the mass, x kg, of a parcel and the cost, y pounds, of posting it. The line of best fit is y = 1.5x + 2. Work out the estimated cost of posting a parcel with a mass of 6 kg, using the line of best fit.y = 1.5x + 2
Answer key
- (c) No, 90 cm is far outside the heights on the graph — Method: a line of best fit describes the trend only across the stretch of data it was drawn through; predicting beyond that stretch is extrapolation, and nothing in the data supports it. Working: the heights used to draw this line run from 150 cm to 180 cm, all of them Year 10 pupils, while 90 cm is 60 cm below the shortest of them and belongs to a two-year-old child, whose build follows no trend the graph has measured. Substituting anyway gives 0.9 × 90 − 85 = −4, a mass of −4 kg, which cannot exist. Answer: no, because 90 cm is far outside the heights on the graph. The distractors: saying a line of best fit cannot be used to predict at all throws away its main purpose, since a prediction made between the plotted values is perfectly sound; saying the line passes through all 20 points misdescribes a line of best fit, which is drawn to follow the trend of the points and will normally pass through few of them; saying the equation works for any value put into it treats an equation fitted to Year 10 heights as a law of nature, and the mass of −4 kg shows what that assumption produces.
- (c) A line of best fit — Method: the straight line drawn on a scatter graph is named from the job it does — it is chosen so that it follows the whole set of points as closely as possible. Working: the line passes through the middle of the points, with roughly as many points above it as below it, and it need not pass through any of the plotted points at all; the name given to the straight line chosen in that way is a line of best fit. Answer: a line of best fit. The distractors: a line of symmetry comes from confusing a trend with symmetry, which is a property of a shape rather than of a set of data; a horizontal line through the mean comes from thinking the trend is shown by an average, when a horizontal line would say that the vertical quantity does not change and so show no correlation; a line joining the first and last points comes from thinking the line must join the two extreme points, which lets two points decide a trend that all of the points should share in.
- (c) 40 minutes — Method: for grouped data, estimate the mean using the midpoint of each class — multiply each midpoint by its frequency, add the results, then divide by the total frequency. Working: the midpoints are 10, 30, 50 and 70 minutes. 10 × 5 = 50. 30 × 10 = 300. 50 × 10 = 500. 70 × 5 = 350. Σfx = 50 + 300 + 500 + 350 = 1200. Σf = 5 + 10 + 10 + 5 = 30. Estimated mean = 1200 ÷ 30 = 40 minutes. Using the upper boundary of each class instead of the midpoint — 20 × 5 = 100, 40 × 10 = 400, 60 × 10 = 600, 80 × 5 = 400 — gives a total of 1500 and an estimate of 1500 ÷ 30 = 50 minutes, too high because a boundary is not the middle of the class. Averaging the frequencies themselves, 5, 10, 10 and 5, ignores the times altogether and gives 7.5. Stopping after Σfx = 1200 without dividing by the total frequency gives a number far too large to be a time in minutes. Always find the midpoint of each class before multiplying by the frequency, and always divide by Σf at the end.
- (d) No, the range uses only the fastest and slowest time — Method: check what the range is built from, then look at what it leaves out. Working: both teams have a fastest time of 20 seconds and a slowest of 40 seconds, so both ranges are 40 − 20 = 20 seconds and Tomás has that part right. But the range is calculated from those two values alone. Six of Team A's seven times lie between 20 and 25 seconds, with a single time far out at 40; Team B is the other way round, with six of its seven times at 30 seconds or more and a single time far out at 20. So Team A bunches at the fast end and Team B at the slow end. The two patterns are quite different, and the range cannot see the difference because the five middle times never enter the calculation. Answer: no, because the range uses only the fastest and slowest time. The distractors: comparing the means answers a different question, since a mean measures position rather than spread, and two sets with the same spread can have different means; saying that equal ranges mean equal spread is the very assumption that fails here; saying that seven times each forces the spreads to match confuses the size of a data set with how its values are arranged inside it.
- (a) 5 pupils are far too few to represent 900 pupils — Method: a sample can only support a claim about a population if it is chosen fairly and if it is large enough for the pattern in it to be more than chance. Working: Priya's method of choosing is fair, because the 5 pupils were picked at random, so every pupil had the same chance of being asked. The difficulty is the size: 900 ÷ 5 = 180, so each pupil she asks stands for 180 pupils. If two of the five happen to play in the same netball team, netball takes 40% of her sample on the strength of two answers, and a second sample of 5 could easily give a different favourite sport. Answer: 5 pupils are far too few to represent 900 pupils. The distractors: saying the pupils were not chosen at random contradicts the question, which states that they were; saying the 5 may each name a different sport describes what often happens in a small sample, but disagreement is not the fault, since 5 pupils who all named the same sport would be just as weak a basis for a claim about 900; saying a sample must hold at least half of the population is an invented rule, and a properly chosen sample of a few hundred can describe a population of many thousands.
- (a) The relationship between two variables — Method: what a diagram shows is decided by what has to be known before a single mark can be plotted on it. Working: every point on a scatter graph is plotted from a pair of measurements taken from the same person or object, one read on the horizontal axis and one on the vertical axis; having two measurements for each point is what makes it possible to look for a pattern between them, and the pattern between two variables is what the graph displays. Answer: a scatter graph shows the relationship between two variables. The distractors: the frequency of each single value is what a bar chart or a vertical line chart shows, and it needs only one list of values; how a total is shared between categories is what a pie chart shows; how one quantity changes over time is what a time series line graph shows, in which one of the two axes is always time.
- (b) 120, unreliable — x = 25 is outside 1 to 12 — The line of best fit is y = 4x + 20. 4 × 25 = 100, and 100 + 20 = 120, so the estimate is y = 120. But x = 25 lies far outside the plotted range of 1 to 12 years, so this is an extrapolation, and the estimate is not reliable. Reaching 100 instead of 120 comes from 4 × 25 = 100 with the intercept of 20 left out — still correctly flagged as unreliable, but the wrong value. Calling the estimate reliable simply because it was calculated correctly, giving 120, wrongly assumes that a correct calculation is automatically trustworthy, ignoring that x = 25 lies far beyond the data actually collected. Reaching 68, from 4 × 12 = 48 and 48 + 20 = 68, substitutes x = 12, the top of the plotted range, instead of the assistant's actual x = 25, and wrongly calls that reliable because 12 lies inside the range.
- (c) Club members probably eat differently from most people — Method: a sample is biased when the group it is drawn from differs from the population in the very thing the survey is measuring, so compare the subgroup with the population on that quantity. Working: the survey measures eating habits, and people who join a sports club take more exercise than average and are known to eat differently from the town as a whole, so their replies pull the results away from the true picture for the town however many of them are asked. Answer: club members probably eat differently from most people. The distractors: the reply about the number of members treats bias as a question of size, but a large biased sample is still biased; the reply that the members were picked at random is false, since the council picked a club rather than picking residents, and it confuses bias with non-response; the reply that everyone asked lives in the town notes something true of the members but draws the false conclusion that the sample therefore covers the town, when a sample must reflect a population and not merely be taken from inside it.
- (d) The mode, because colours cannot be added or ordered — Method: an average can only be used on data that supports the operation it needs. A mean needs the values to be added and divided, a median needs them to be placed in order, and a range needs one value to be taken away from another; a mode needs only counting, so it is the average available when the data are categories rather than numbers. Working: the data collected here are colours, silver, black, blue and red. The numbers 74, 52, 40 and 34 count the cars of each colour, they do not measure them, and 74 + 52 + 40 + 34 = 200 simply returns the size of the survey. No colour can be added to another, and there is no order that puts blue before red, so of the four averages only the one found by counting survives. Answer: the mode, because colours cannot be added or ordered, and the mode is silver. The distractors: the mean is said to use all 200 colours, and a mean of the four frequencies, 200 ÷ 4 = 50, is a number of cars rather than a colour, so it describes nothing about a typical car; the median is said to put the colours in order, but ordering the frequencies 34, 40, 52, 74 orders the counts, not the colours, and gives 46, again a number of cars; the range is not an average at all, and 74 − 34 = 40 measures the gap between the commonest and rarest counts, which is a measure of spread.
- (a) A straight line sloping down from left to right — Method: a line of best fit is a straight line drawn to follow the trend of the points, so its slope is decided by the direction of the relationship between the two quantities. Working: as the price rises, the satisfaction score falls, so the points start high on the left of the graph and finish low on the right; the straight line that follows them therefore slopes downwards as the graph is read from left to right, which is the line of a negative correlation. Answer: a straight line sloping down from left to right. The distractors: a line sloping up comes from reading a falling relationship as a rising one; a horizontal line comes from expecting no correlation, since a horizontal line says the satisfaction score does not change as the price changes; a curve passing through every point comes from thinking a line of best fit has to touch all of the plotted points, when it is a single straight line drawn through the middle of them.
- (a) 26 — Method: a line of best fit lets one quantity be predicted from the other, so the score is substituted into the equation of the line and the resulting inequality is solved for the number of hours. Working: a score of at least 80 means 2.5x + 15 ≥ 80; taking 15 from both sides gives 2.5x ≥ 65, and dividing both sides by 2.5 gives x ≥ 26, so the least whole number of hours is 26. Checking, 2.5 × 26 + 15 = 80, which does reach the target. Answer: 26 hours — and this is only an estimate, because a line of best fit predicts a trend rather than an individual result, and a prediction made outside the range of hours the pupils actually revised for would be an extrapolation and less reliable still. The distractors: 27 comes from reaching 26 and then rounding up again, although 26 hours already gives a score of exactly 80; 32 comes from 80 ÷ 2.5, which ignores the 15 in the equation of the line; 38 comes from (80 + 15) ÷ 2.5, that is from adding the 15 instead of subtracting it when rearranging.
- (a) No correlation — Shoe size has no real relationship with spelling ability, and the points here are scattered with no rising or falling trend, so this is no correlation. A positive correlation would show the points rising together, and a negative correlation would show them falling as one increases; neither pattern is present here. Strong correlation is not correct either, since strength only applies once a positive or negative trend exists, and there isn't one.
- (b) £1,000 — Wages take up 150° out of 360°, so the amount spent on wages is 150 ÷ 360 × 2400 = £1,000. Choosing £600 uses the repairs angle, 90°, instead of the wages angle: 90 ÷ 360 × 2400 = 600. Choosing £3,600 treats the angle in degrees as if it were a percentage, 150 ÷ 100 × 2400 = 3600, instead of dividing by 360°. Choosing £800 uses the angle for the 'other costs' sector, 360 − 90 − 150 = 120°, instead of the wages sector: 120 ÷ 360 × 2400 = 800.
- (d) 23.57 °C — Method: add all seven temperatures, divide by the number of readings and round only at the end. Working: 22 + 24 + 23 + 25 + 26 + 21 + 24 = 165, and 165 ÷ 7 = 23.5714…, which rounds to 23.57 to 2 decimal places. Answer: 23.57 °C. The distractors: 24 °C comes from writing down the mode, the only temperature recorded twice, instead of the mean; 27.5 °C comes from dividing the total by 6 instead of by the 7 days recorded; 5 °C comes from working out the range, 26 − 21, which is a measure of spread and not an average.
- (b) £11.00 — 1.5 × 6 = 9, and 9 + 2 = 11, so the estimated cost is £11.00. Choosing £9.00 stops after 1.5 × 6 = 9 and forgets to add the £2. Choosing £12.00 adds the mass and the constant first and then multiplies: 6 + 2 = 8, and 8 × 1.5 = 12.00. Choosing £13.50 swaps the gradient and the intercept, using y = 2x + 1.5 instead: 2 × 6 = 12, and 12 + 1.5 = 13.50.
Build your own mix at the worksheet builder.