Printable · GCSE Foundation · ages 14-16
Scatter graphs, correlation and lines of best fit worksheet — GCSE Foundation
Fifteen questions on "scatter graphs, correlation and lines of best fit" — DfE statement S6. Print it, or print three versions so neighbours cannot copy by letter; the key gives the letter for each version.
Calculator
Scatter graphs, correlation and lines of best fit worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- 1.A gym in Cardiff has a scatter graph of the number of training sessions, x, attended by each member and the weight lost, y kg. The line of best fit is y = 0.4x + 1. Members want to lose at least 9 kg. Using the line of best fit, work out the least whole number of sessions needed to reach this target.y = 0.4x + 1
- 2.A scatter graph shows the number of years of experience, x, of 18 sales assistants and their monthly sales, y hundred pounds. The plotted points run from x = 1 to x = 12 years, and the line of best fit is y = 4x + 20. A new assistant has 25 years of experience. Use the line of best fit to estimate a value of y for this assistant, and decide whether the estimate would be reliable.y = 4x + 20
- 3.A scatter graph shows the number of guests, x, at a wedding and the length of buffet table needed, y metres. The line of best fit is y = 0.5x + 2. Write down what the 2 in this equation tells you about the buffet table.y = 0.5x + 2
- 4.A café owner in Brighton records the midday temperature, x °C, and the number of hot chocolates sold, y, on 12 days. The temperatures recorded run from 4 °C to 18 °C, and the line of best fit is y = −3x + 74. The forecast for tomorrow gives a midday temperature of 12 °C. Work out the number the line of best fit predicts, and write down how much confidence the owner can have in it.y = -3x + 74
- 5.In a scatter graph of the age, in years, and the wingspan, in cm, of 20 birds of the same species, all the points lie close to a rising line of best fit except one, which lies a long way below the line. That bird was later found to have a damaged wing. Give a reason why this point should not be used when drawing the line of best fit.
- 6.A scatter graph plots the number of years, x, that 20 employees have worked at a company against their salary, y. All the plotted points lie between x = 1 and x = 15. Write down the word used to describe an estimate for y made using a value of x that lies between 1 and 15.
- 7.A straight line is drawn on a scatter graph to show the trend of the points. Write down the name given to this line.
- 8.A scatter graph has 50 points. Most of them lie close to a rising line of best fit, but two of them lie a long way from that line. Write down how those two points should be treated.
- 9.A scatter graph of the number of ice creams sold at a seaside kiosk in Bournemouth and the number of sunburn cases treated at a nearby pharmacy, recorded on the same 30 days, shows strong positive correlation. Which statement about this correlation is correct?
- 10.A scatter graph of taxi journeys in Bristol shows the distance, x miles, and the fare, y pounds. The line of best fit is y = 2x + 3.50. Work out the estimated fare for a journey of 6 miles, using the line of best fit.y = 2x + 3.5
- 11.A scatter graph plots the shoe size and the spelling test score of 25 pupils. The points are scattered with no pattern across the graph. Write down the type of correlation shown.
- 12.Write down the statement that correctly describes the difference between correlation and causation.
- 13.The age, in years, and the score in a reaction test are recorded for eight members of a sports club: (14, 92), (18, 88), (23, 85), (27, 80), (31, 78), (36, 74), (42, 70), (49, 65). The eight pairs are plotted on a scatter graph. Describe the correlation between age and score.
- 14.A scatter graph shows the height, x cm, and the mass, y kg, of 20 pupils in Year 10. The heights on the graph run from 150 cm to 180 cm, and the line of best fit is y = 0.9x − 85. Nadia puts x = 90 into this equation to estimate the mass of a two-year-old child who is 90 cm tall. Is her estimate reliable? Give a reason for your answer.y = 0.9x − 85
- 15.Four pairs of variables are listed below. Write down the pair that you would expect to show no correlation.
Answer key
- (d) 20 — 0.4x + 1 = 9, so 0.4x = 9 − 1 = 8, and 8 ÷ 0.4 = 20, so 20 sessions are needed. Choosing 23 divides 9 by 0.4 without first subtracting the 1: 9 ÷ 0.4 = 22.5, rounded up to 23. Choosing 25 subtracts the wrong way, adding the 1 instead of taking it away: 9 + 1 = 10, and 10 ÷ 0.4 = 25. Choosing 2 misplaces the decimal point in the gradient, dividing by 4 instead of by 0.4: 8 ÷ 4 = 2.
- (b) 120, unreliable — x = 25 is outside 1 to 12 — The line of best fit is y = 4x + 20. 4 × 25 = 100, and 100 + 20 = 120, so the estimate is y = 120. But x = 25 lies far outside the plotted range of 1 to 12 years, so this is an extrapolation, and the estimate is not reliable. Reaching 100 instead of 120 comes from 4 × 25 = 100 with the intercept of 20 left out — still correctly flagged as unreliable, but the wrong value. Calling the estimate reliable simply because it was calculated correctly, giving 120, wrongly assumes that a correct calculation is automatically trustworthy, ignoring that x = 25 lies far beyond the data actually collected. Reaching 68, from 4 × 12 = 48 and 48 + 20 = 68, substitutes x = 12, the top of the plotted range, instead of the assistant's actual x = 25, and wrongly calls that reliable because 12 lies inside the range.
- (b) At 0 guests, the model predicts 2 m of table — The y-intercept of a line of best fit y = mx + c is the value of y when x = 0. Here y = 0.5 × 0 + 2 = 2, so the line predicts a table length of 2 m when there are 0 guests. The 2 m does not grow as more guests arrive — that role belongs to the gradient, 0.5 — so an option saying each extra guest adds 2 m has swapped the two numbers around. The 2 is a length in metres, not a number of guests, so an option requiring 2 guests before set-up has misread its units. And the table length does change with x, since it is 0.5x + 2 and not a fixed value, so an option claiming the table is always 2 m ignores the 0.5x term completely.
- (d) 38, and fairly confident, as 12 °C is inside the range — Method: substitute the forecast temperature into the equation of the line of best fit, then judge the prediction by where that temperature sits among the data the line was drawn from. Working: putting x = 12 into y = −3x + 74 gives −3 × 12 + 74 = 38, so the line predicts 38 hot chocolates. The recorded temperatures run from 4 °C to 18 °C, and 12 °C lies inside that interval, so this is interpolation, the safer kind of prediction. Answer: 38, and fairly confident, as 12 °C is inside the range; the owner should still expect the true figure to differ a little, since the points only lie near the line and not on it. The distractors: being completely certain treats a line of best fit as a rule that fixes each day's sales, when it describes a trend that individual days depart from; saying 12 °C is outside the range misreads the interval 4 °C to 18 °C, and the wrong warning would be attached to a sound prediction; 110 comes from −3 × 12 being taken as +36, giving 36 + 74 = 110, which loses the negative gradient and so predicts that a warm day sells more hot chocolate than a cold one.
- (d) An outlier from the damaged wing, not the trend. — That bird's point lies a long way from the rising trend followed by every other bird, and its low wingspan is explained by the damaged wing rather than by its age — it is an outlier caused by an unusual factor, not part of the general relationship between age and wingspan, so it should not be used when drawing the line of best fit. Saying every point must be used ignores that an outlier caused by a separate, identifiable factor can rightly be set aside. Saying it shows no correlation ignores that the other 19 points do show a clear rising trend; one outlier does not remove that. Saying it proves the line is inaccurate confuses one unusual bird with a fault in the line itself, when the line correctly describes the trend followed by the rest of the data.
- (b) Interpolation — The salary is being estimated for a value of x between 1 and 15, which is inside the range of x-values that were actually plotted, so this is interpolation. Extrapolation would apply if the estimate used a value of x below 1 or above 15, outside the plotted range. Correlation describes the relationship between the two variables, not the reliability of an estimate, and causation describes one variable actually causing a change in the other, which is a different idea altogether — neither is the word being asked for here.
- (c) A line of best fit — Method: the straight line drawn on a scatter graph is named from the job it does — it is chosen so that it follows the whole set of points as closely as possible. Working: the line passes through the middle of the points, with roughly as many points above it as below it, and it need not pass through any of the plotted points at all; the name given to the straight line chosen in that way is a line of best fit. Answer: a line of best fit. The distractors: a line of symmetry comes from confusing a trend with symmetry, which is a property of a shape rather than of a set of data; a horizontal line through the mean comes from thinking the trend is shown by an average, when a horizontal line would say that the vertical quantity does not change and so show no correlation; a line joining the first and last points comes from thinking the line must join the two extreme points, which lets two points decide a trend that all of the points should share in.
- (b) Treat them as outliers and check them before deciding — Method: a point lying a long way from the pattern the rest of the data make is called an outlier, and an outlier is investigated before anything is done with it, because it may be an error in the data or it may be a genuine but unusual case. Working: 48 of the 50 points lie close to the rising line of best fit, so the trend is set by those 48; the two remaining points do not follow it, so they are identified as outliers and checked — a mistake in measuring or recording would be corrected, while a genuine reading would be kept and reported. Answer: treat them as outliers and check them before deciding what to do with them. The distractors: deleting them at once assumes that every point far from the line must be an error, which throws away real data; moving the line so that it passes through them assumes a line of best fit must touch particular points, when it is drawn to follow all 50; taking them as proof that there is no correlation lets two points overturn the pattern that the other 48 agree on.
- (c) Neither causes the other; sunshine links both. — Both ice cream sales and sunburn cases tend to rise on hot, sunny days, so the amount of sunshine is a third factor linked to both — neither variable causes the other. Saying ice cream sales cause the sunburn assumes a causal link in one direction that the correlation alone cannot establish. Saying sunburn cases cause the ice cream sales assumes the reverse causal link, which is no more justified. Saying a strong correlation always means causation is the general error this question is testing: correlation, however strong, does not by itself prove that one variable causes the other.
- (d) £15.50 — 2 × 6 = 12, and 12 + 3.50 = 15.50, so the estimated fare is £15.50. Choosing £12.00 stops after 2 × 6 = 12 and forgets to add the £3.50. Choosing £19.00 adds the distance and the constant first and then multiplies: 6 + 3.50 = 9.50, and 9.50 × 2 = 19.00, applying the ×2 to the whole sum instead of only to the distance. Choosing £13.00 multiplies only the constant term by 2 instead of the distance: 2 × 3.50 = 7, and 7 + 6 = 13.00.
- (a) No correlation — Shoe size has no real relationship with spelling ability, and the points here are scattered with no rising or falling trend, so this is no correlation. A positive correlation would show the points rising together, and a negative correlation would show them falling as one increases; neither pattern is present here. Strong correlation is not correct either, since strength only applies once a positive or negative trend exists, and there isn't one.
- (d) Correlation is a link; causation is one causing the other — Method: the two words describe different claims — one is about a pattern in the data, the other is about what produced that pattern. Working: correlation says only that two quantities tend to change together, which is something a scatter graph can display; causation says that a change in one quantity actually brings about the change in the other, which needs evidence a scatter graph cannot supply, because a third quantity may be driving both. Answer: correlation is a link between the quantities, while causation is one quantity causing the change in another. The distractors: the statement giving causation as the link and correlation as the cause simply swaps the two words over; the statement that the words mean the same thing is the classic error of reading a correlation as proof of cause; the statement that a scatter graph shows causation but not correlation reverses what a scatter graph can do, since the pattern it displays is exactly the correlation.
- (c) Strong negative correlation — Method: correlation is described by two things — the direction the points take as the graph is read from left to right, and how closely the points lie to a single straight line. Working: reading the pairs in order of age, the ages rise 14, 18, 23, 27, 31, 36, 42, 49 while the scores fall 92, 88, 85, 80, 78, 74, 70, 65; the score falls at every single step, with no reversal anywhere, so the points fall from left to right and lie close to a straight line. Answer: strong negative correlation — negative for the falling direction, strong because every point follows the pattern. The distractors: strong positive correlation comes from noticing a clear pattern and calling any clear pattern positive, without checking the direction; weak negative correlation comes from reading the direction correctly but judging points that do not lie exactly on a straight line to be only loosely related, when these eight fall without a single exception; no correlation comes from reading a falling trend as though it showed no relationship at all, when a falling trend is itself a relationship.
- (c) No, 90 cm is far outside the heights on the graph — Method: a line of best fit describes the trend only across the stretch of data it was drawn through; predicting beyond that stretch is extrapolation, and nothing in the data supports it. Working: the heights used to draw this line run from 150 cm to 180 cm, all of them Year 10 pupils, while 90 cm is 60 cm below the shortest of them and belongs to a two-year-old child, whose build follows no trend the graph has measured. Substituting anyway gives 0.9 × 90 − 85 = −4, a mass of −4 kg, which cannot exist. Answer: no, because 90 cm is far outside the heights on the graph. The distractors: saying a line of best fit cannot be used to predict at all throws away its main purpose, since a prediction made between the plotted values is perfectly sound; saying the line passes through all 20 points misdescribes a line of best fit, which is drawn to follow the trend of the points and will normally pass through few of them; saying the equation works for any value put into it treats an equation fitted to Year 10 heights as a law of nature, and the mass of −4 kg shows what that assumption produces.
- (b) A person's shoe size and their favourite colour — A person's shoe size is not linked to which colour they prefer, so these two show no correlation. The other three pairs are all genuinely correlated: distance travelled and fuel used rise together, which is positive correlation; hours of revision and test score generally rise together, which is also positive correlation; and as outdoor temperature rises, fewer woolly hats are sold, which is negative correlation. Negative correlation is still a real relationship between two variables — it is not the same thing as no relationship at all, so the temperature and hats pair is not the answer to this question.
Build your own mix at the worksheet builder.