Printable · GCSE Foundation · ages 14-16
Scatter graphs, correlation and lines of best fit worksheet — GCSE Foundation
Fifteen questions on "scatter graphs, correlation and lines of best fit" — DfE statement S6. Print it, or print three versions so neighbours cannot copy by letter; the key gives the letter for each version.
Answer key: Scatter graphs, correlation and lines of best fit worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- (d) An outlier from the damaged wing, not the trend. — That bird's point lies a long way from the rising trend followed by every other bird, and its low wingspan is explained by the damaged wing rather than by its age — it is an outlier caused by an unusual factor, not part of the general relationship between age and wingspan, so it should not be used when drawing the line of best fit. Saying every point must be used ignores that an outlier caused by a separate, identifiable factor can rightly be set aside. Saying it shows no correlation ignores that the other 19 points do show a clear rising trend; one outlier does not remove that. Saying it proves the line is inaccurate confuses one unusual bird with a fault in the line itself, when the line correctly describes the trend followed by the rest of the data.
- (b) 120, unreliable — x = 25 is outside 1 to 12 — The line of best fit is y = 4x + 20. 4 × 25 = 100, and 100 + 20 = 120, so the estimate is y = 120. But x = 25 lies far outside the plotted range of 1 to 12 years, so this is an extrapolation, and the estimate is not reliable. Reaching 100 instead of 120 comes from 4 × 25 = 100 with the intercept of 20 left out — still correctly flagged as unreliable, but the wrong value. Calling the estimate reliable simply because it was calculated correctly, giving 120, wrongly assumes that a correct calculation is automatically trustworthy, ignoring that x = 25 lies far beyond the data actually collected. Reaching 68, from 4 × 12 = 48 and 48 + 20 = 68, substitutes x = 12, the top of the plotted range, instead of the assistant's actual x = 25, and wrongly calls that reliable because 12 lies inside the range.
- (a) x = 0 gives y = −20: a negative number sold — The y-intercept is the value the line predicts when x = 0: y = 3 × 0 − 20 = −20. A kiosk cannot sell a negative number of ice creams, so this is not a sensible estimate. The 3 in the equation is the gradient, not the intercept, so an option claiming x = 0 gives y = 3 has swapped the two numbers around — substituting x = 0 makes the 3x term equal 0, leaving −20, not 3. The danger of extrapolating to very high temperatures is a real issue with this line, but it is a different issue from the y-intercept, so it does not answer this question. And whether x = 0 could occur on a trading day is beside the point: the model still makes that prediction, and it is the prediction itself, −20, that is impossible.
- (b) £11.00 — 1.5 × 6 = 9, and 9 + 2 = 11, so the estimated cost is £11.00. Choosing £9.00 stops after 1.5 × 6 = 9 and forgets to add the £2. Choosing £12.00 adds the mass and the constant first and then multiplies: 6 + 2 = 8, and 8 × 1.5 = 12.00. Choosing £13.50 swaps the gradient and the intercept, using y = 2x + 1.5 instead: 2 × 6 = 12, and 12 + 1.5 = 13.50.
- (d) Correlation is a link; causation is one causing the other — Method: the two words describe different claims — one is about a pattern in the data, the other is about what produced that pattern. Working: correlation says only that two quantities tend to change together, which is something a scatter graph can display; causation says that a change in one quantity actually brings about the change in the other, which needs evidence a scatter graph cannot supply, because a third quantity may be driving both. Answer: correlation is a link between the quantities, while causation is one quantity causing the change in another. The distractors: the statement giving causation as the link and correlation as the cause simply swaps the two words over; the statement that the words mean the same thing is the classic error of reading a correlation as proof of cause; the statement that a scatter graph shows causation but not correlation reverses what a scatter graph can do, since the pattern it displays is exactly the correlation.
- (d) 210 — 350 − 150 = 200. 200 ÷ 10 = 20, so the gradient is 20. Using the point (5, 150): 20 × 5 = 100, so 150 − 100 = 50 is the intercept, giving the line y = 20x + 50. At x = 8: 20 × 8 = 160, and 160 + 50 = 210, so the estimated number of visitors is 210. Choosing 160 stops after 20 × 8 = 160 and forgets to add the intercept of 50. Choosing 250 comes from averaging the two given y-values: 150 + 350 = 500, and 500 ÷ 2 = 250, instead of using the line's equation. Choosing 240 assumes the visitors are directly proportional to the hours of sunshine using the first point, 150 × 8 ÷ 5 = 240, which ignores that the line does not pass through the origin.
- (b) Positive correlation — Method: the type of correlation is named from the direction the points take as the scatter graph is read from left to right. Working: the points rise from left to right, so as the arm span read on the horizontal axis increases, the height read on the vertical axis increases as well; two quantities that increase together show positive correlation. Answer: positive correlation. The distractors: negative correlation comes from naming the direction the wrong way round, since a negative correlation needs the points to fall as the graph is read from left to right; no correlation comes from treating points that are spread out rather than sitting exactly on a line as though they showed no relationship; direct proportion comes from confusing a rising trend with proportion, which would additionally need the line through the points to pass through the origin and would mean doubling one quantity doubles the other.
- (a) Negative correlation — As the age of the car increases, the points fall towards a lower value, so the value decreases as the age increases. This falling pattern is a negative correlation. A positive correlation would show the points rising together instead. No correlation would apply only if the points showed no pattern at all, and correlation is not the same as causation — strong causation is not a type of correlation.
- (a) No correlation — Shoe size has no real relationship with spelling ability, and the points here are scattered with no rising or falling trend, so this is no correlation. A positive correlation would show the points rising together, and a negative correlation would show them falling as one increases; neither pattern is present here. Strong correlation is not correct either, since strength only applies once a positive or negative trend exists, and there isn't one.
- (b) A person's shoe size and their favourite colour — A person's shoe size is not linked to which colour they prefer, so these two show no correlation. The other three pairs are all genuinely correlated: distance travelled and fuel used rise together, which is positive correlation; hours of revision and test score generally rise together, which is also positive correlation; and as outdoor temperature rises, fewer woolly hats are sold, which is negative correlation. Negative correlation is still a real relationship between two variables — it is not the same thing as no relationship at all, so the temperature and hats pair is not the answer to this question.
- (a) No, the size of the fire affects both of the quantities — Method: correlation says that two quantities change together; a claim that one of them produces the other is a further claim, and it needs evidence that a scatter graph on its own cannot give. Working: the graph does show strong positive correlation, so more engines did go with greater damage. But neither quantity was set by the researchers: both were decided by how large the fire was. A large blaze brings many appliances and also destroys a great deal, while a small one brings few and destroys little, so a third quantity is driving both of the recorded ones. Answer: no, because the size of the fire affects both of the quantities. The distractors: saying the correlation is negative contradicts the graph, which shows the two quantities rising together, and reaching the right verdict from a false reading of the data is not the reason the mark is for; saying that strong positive correlation shows one quantity causes the other is the assumption the question exists to test, and no strength of correlation can establish cause; saying the points lie close to the line of best fit describes how strong the correlation is, and strength and cause are different matters entirely.
- (a) 23 minutes — Method: the equation of a line of best fit converts a value of x into a predicted value of y, so substitute the known number of pages for x and evaluate. Working: x is the number of pages, so put x = 10 into y = 2x + 3. Multiplication is carried out before addition, so 2 × 10 + 3 = 23. Answer: 23 minutes, and it is a prediction of the trend rather than a promise about any one chapter. The distractors: 20 minutes comes from working out 2 × 10 and stopping there, leaving out the 3 that the line adds; 26 minutes comes from reading the equation as y = 2(x + 3), adding first and then doubling, so 2 × 13 = 26; 13 minutes comes from adding 10 and 3 and never using the gradient at all, which treats the 2 as though it were not there.
- (a) A straight line sloping down from left to right — Method: a line of best fit is a straight line drawn to follow the trend of the points, so its slope is decided by the direction of the relationship between the two quantities. Working: as the price rises, the satisfaction score falls, so the points start high on the left of the graph and finish low on the right; the straight line that follows them therefore slopes downwards as the graph is read from left to right, which is the line of a negative correlation. Answer: a straight line sloping down from left to right. The distractors: a line sloping up comes from reading a falling relationship as a rising one; a horizontal line comes from expecting no correlation, since a horizontal line says the satisfaction score does not change as the price changes; a curve passing through every point comes from thinking a line of best fit has to touch all of the plotted points, when it is a single straight line drawn through the middle of them.
- (c) Neither causes the other; sunshine links both. — Both ice cream sales and sunburn cases tend to rise on hot, sunny days, so the amount of sunshine is a third factor linked to both — neither variable causes the other. Saying ice cream sales cause the sunburn assumes a causal link in one direction that the correlation alone cannot establish. Saying sunburn cases cause the ice cream sales assumes the reverse causal link, which is no more justified. Saying a strong correlation always means causation is the general error this question is testing: correlation, however strong, does not by itself prove that one variable causes the other.
- (a) Priya — A line of best fit should be drawn so that the plotted points are roughly balanced above and below it. Work out the difference between the two counts for each pupil: Amir 10 − 2 = 8; Kofi 10 − 2 = 8 (10 below and 2 above); Leah 12 − 0 = 12; Priya 6 − 6 = 0. Priya's line has the smallest difference, an exact balance of 6 above and 6 below, so her line is drawn correctly. Amir's line has 10 of the 12 points above it, so it is drawn too low. Kofi's line has 10 of the 12 points below it, so it is drawn too high. Leah's line has every single point below it, so it is not a line of best fit at all.
Build your own mix at the worksheet builder.