Printable · GCSE Foundation · ages 14-16
Scatter graphs, correlation and lines of best fit worksheet — GCSE Foundation
Fifteen questions on "scatter graphs, correlation and lines of best fit" — DfE statement S6. Print it, or print three versions so neighbours cannot copy by letter; the key gives the letter for each version.
Calculator
Answer key: Scatter graphs, correlation and lines of best fit worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- (b) Interpolation — The salary is being estimated for a value of x between 1 and 15, which is inside the range of x-values that were actually plotted, so this is interpolation. Extrapolation would apply if the estimate used a value of x below 1 or above 15, outside the plotted range. Correlation describes the relationship between the two variables, not the reliability of an estimate, and causation describes one variable actually causing a change in the other, which is a different idea altogether — neither is the word being asked for here.
- (b) A person's shoe size and their favourite colour — A person's shoe size is not linked to which colour they prefer, so these two show no correlation. The other three pairs are all genuinely correlated: distance travelled and fuel used rise together, which is positive correlation; hours of revision and test score generally rise together, which is also positive correlation; and as outdoor temperature rises, fewer woolly hats are sold, which is negative correlation. Negative correlation is still a real relationship between two variables — it is not the same thing as no relationship at all, so the temperature and hats pair is not the answer to this question.
- (c) Neither causes the other; sunshine links both. — Both ice cream sales and sunburn cases tend to rise on hot, sunny days, so the amount of sunshine is a third factor linked to both — neither variable causes the other. Saying ice cream sales cause the sunburn assumes a causal link in one direction that the correlation alone cannot establish. Saying sunburn cases cause the ice cream sales assumes the reverse causal link, which is no more justified. Saying a strong correlation always means causation is the general error this question is testing: correlation, however strong, does not by itself prove that one variable causes the other.
- (a) x = 0 gives y = −20: a negative number sold — The y-intercept is the value the line predicts when x = 0: y = 3 × 0 − 20 = −20. A kiosk cannot sell a negative number of ice creams, so this is not a sensible estimate. The 3 in the equation is the gradient, not the intercept, so an option claiming x = 0 gives y = 3 has swapped the two numbers around — substituting x = 0 makes the 3x term equal 0, leaving −20, not 3. The danger of extrapolating to very high temperatures is a real issue with this line, but it is a different issue from the y-intercept, so it does not answer this question. And whether x = 0 could occur on a trading day is beside the point: the model still makes that prediction, and it is the prediction itself, −20, that is impossible.
- (b) 120, unreliable — x = 25 is outside 1 to 12 — The line of best fit is y = 4x + 20. 4 × 25 = 100, and 100 + 20 = 120, so the estimate is y = 120. But x = 25 lies far outside the plotted range of 1 to 12 years, so this is an extrapolation, and the estimate is not reliable. Reaching 100 instead of 120 comes from 4 × 25 = 100 with the intercept of 20 left out — still correctly flagged as unreliable, but the wrong value. Calling the estimate reliable simply because it was calculated correctly, giving 120, wrongly assumes that a correct calculation is automatically trustworthy, ignoring that x = 25 lies far beyond the data actually collected. Reaching 68, from 4 × 12 = 48 and 48 + 20 = 68, substitutes x = 12, the top of the plotted range, instead of the assistant's actual x = 25, and wrongly calls that reliable because 12 lies inside the range.
- (a) No correlation — Shoe size has no real relationship with spelling ability, and the points here are scattered with no rising or falling trend, so this is no correlation. A positive correlation would show the points rising together, and a negative correlation would show them falling as one increases; neither pattern is present here. Strong correlation is not correct either, since strength only applies once a positive or negative trend exists, and there isn't one.
- (d) 38, and fairly confident, as 12 °C is inside the range — Method: substitute the forecast temperature into the equation of the line of best fit, then judge the prediction by where that temperature sits among the data the line was drawn from. Working: putting x = 12 into y = −3x + 74 gives −3 × 12 + 74 = 38, so the line predicts 38 hot chocolates. The recorded temperatures run from 4 °C to 18 °C, and 12 °C lies inside that interval, so this is interpolation, the safer kind of prediction. Answer: 38, and fairly confident, as 12 °C is inside the range; the owner should still expect the true figure to differ a little, since the points only lie near the line and not on it. The distractors: being completely certain treats a line of best fit as a rule that fixes each day's sales, when it describes a trend that individual days depart from; saying 12 °C is outside the range misreads the interval 4 °C to 18 °C, and the wrong warning would be attached to a sound prediction; 110 comes from −3 × 12 being taken as +36, giving 36 + 74 = 110, which loses the negative gradient and so predicts that a warm day sells more hot chocolate than a cold one.
- (b) Treat them as outliers and check them before deciding — Method: a point lying a long way from the pattern the rest of the data make is called an outlier, and an outlier is investigated before anything is done with it, because it may be an error in the data or it may be a genuine but unusual case. Working: 48 of the 50 points lie close to the rising line of best fit, so the trend is set by those 48; the two remaining points do not follow it, so they are identified as outliers and checked — a mistake in measuring or recording would be corrected, while a genuine reading would be kept and reported. Answer: treat them as outliers and check them before deciding what to do with them. The distractors: deleting them at once assumes that every point far from the line must be an error, which throws away real data; moving the line so that it passes through them assumes a line of best fit must touch particular points, when it is drawn to follow all 50; taking them as proof that there is no correlation lets two points overturn the pattern that the other 48 agree on.
- (d) 210 — 350 − 150 = 200. 200 ÷ 10 = 20, so the gradient is 20. Using the point (5, 150): 20 × 5 = 100, so 150 − 100 = 50 is the intercept, giving the line y = 20x + 50. At x = 8: 20 × 8 = 160, and 160 + 50 = 210, so the estimated number of visitors is 210. Choosing 160 stops after 20 × 8 = 160 and forgets to add the intercept of 50. Choosing 250 comes from averaging the two given y-values: 150 + 350 = 500, and 500 ÷ 2 = 250, instead of using the line's equation. Choosing 240 assumes the visitors are directly proportional to the hours of sunshine using the first point, 150 × 8 ÷ 5 = 240, which ignores that the line does not pass through the origin.
- (a) Minutes a candle has burned and length remaining — As a candle burns for longer, less of it remains, so these two variables move in opposite directions as one increases — that is negative correlation. A pupil's shoe size generally increases as they get older, so age and shoe size show positive correlation, not negative, since both rise together. A football team's shirt colour is not a numerical quantity linked to how many matches it wins, so shirt colour and number of wins show no correlation at all. The number of letters in a pupil's name has no real connection to their ability in maths, so that pair also shows no correlation.
- (d) £15.50 — 2 × 6 = 12, and 12 + 3.50 = 15.50, so the estimated fare is £15.50. Choosing £12.00 stops after 2 × 6 = 12 and forgets to add the £3.50. Choosing £19.00 adds the distance and the constant first and then multiplies: 6 + 3.50 = 9.50, and 9.50 × 2 = 19.00, applying the ×2 to the whole sum instead of only to the distance. Choosing £13.00 multiplies only the constant term by 2 instead of the distance: 2 × 3.50 = 7, and 7 + 6 = 13.00.
- (c) A line of best fit — Method: the straight line drawn on a scatter graph is named from the job it does — it is chosen so that it follows the whole set of points as closely as possible. Working: the line passes through the middle of the points, with roughly as many points above it as below it, and it need not pass through any of the plotted points at all; the name given to the straight line chosen in that way is a line of best fit. Answer: a line of best fit. The distractors: a line of symmetry comes from confusing a trend with symmetry, which is a property of a shape rather than of a set of data; a horizontal line through the mean comes from thinking the trend is shown by an average, when a horizontal line would say that the vertical quantity does not change and so show no correlation; a line joining the first and last points comes from thinking the line must join the two extreme points, which lets two points decide a trend that all of the points should share in.
- (a) No, the size of the fire affects both of the quantities — Method: correlation says that two quantities change together; a claim that one of them produces the other is a further claim, and it needs evidence that a scatter graph on its own cannot give. Working: the graph does show strong positive correlation, so more engines did go with greater damage. But neither quantity was set by the researchers: both were decided by how large the fire was. A large blaze brings many appliances and also destroys a great deal, while a small one brings few and destroys little, so a third quantity is driving both of the recorded ones. Answer: no, because the size of the fire affects both of the quantities. The distractors: saying the correlation is negative contradicts the graph, which shows the two quantities rising together, and reaching the right verdict from a false reading of the data is not the reason the mark is for; saying that strong positive correlation shows one quantity causes the other is the assumption the question exists to test, and no strength of correlation can establish cause; saying the points lie close to the line of best fit describes how strong the correlation is, and strength and cause are different matters entirely.
- (a) A straight line sloping down from left to right — Method: a line of best fit is a straight line drawn to follow the trend of the points, so its slope is decided by the direction of the relationship between the two quantities. Working: as the price rises, the satisfaction score falls, so the points start high on the left of the graph and finish low on the right; the straight line that follows them therefore slopes downwards as the graph is read from left to right, which is the line of a negative correlation. Answer: a straight line sloping down from left to right. The distractors: a line sloping up comes from reading a falling relationship as a rising one; a horizontal line comes from expecting no correlation, since a horizontal line says the satisfaction score does not change as the price changes; a curve passing through every point comes from thinking a line of best fit has to touch all of the plotted points, when it is a single straight line drawn through the middle of them.
- (b) At 0 guests, the model predicts 2 m of table — The y-intercept of a line of best fit y = mx + c is the value of y when x = 0. Here y = 0.5 × 0 + 2 = 2, so the line predicts a table length of 2 m when there are 0 guests. The 2 m does not grow as more guests arrive — that role belongs to the gradient, 0.5 — so an option saying each extra guest adds 2 m has swapped the two numbers around. The 2 is a length in metres, not a number of guests, so an option requiring 2 guests before set-up has misread its units. And the table length does change with x, since it is 0.5x + 2 and not a fixed value, so an option claiming the table is always 2 m ignores the 0.5x term completely.
Build your own mix at the worksheet builder.