Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Non-calculator
Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- 1.Work out the median of these five numbers: 2, 4, 7, 12, 26
- 2.There are 10,000 pupils in a city. A researcher picks a sample of 500 of them by drawing names at random from a list of every pupil in the city. Give the reason why this method gives a representative sample.
- 3.A scatter graph shows the number of days, x, that each of 16 tomato plants was watered and its height, y cm. The line of best fit has equation y = 1.5x + 4. Write down what the 1.5 in this equation tells you about the plants.y = 1.5x + 4
- 4.Write down what a scatter graph is used to show.
- 5.A vet records the weights, in kilograms, of 7 dogs, in the order she weighs them: 18, 12, 50, 10, 20, 14, 16. Work out the median weight.
- 6.A pupil writes this question for a school survey: “Do you agree that learning matters and that we should be set more homework?” Write down what is wrong with the survey question.
- 7.A shop recorded the number of books it sold on five days: 100, 40, 70, 20, 60. Work out the range of the numbers of books sold.
- 8.The mean mass of four parcels is 17 kg. Three of the parcels have masses 12 kg, 16 kg and 18 kg. Work out the mass of the fourth parcel.
- 9.Harry counted the coins he found on each of six days: 5, 7, 8, 9, 11, 30. Write down the outlier.
- 10.For each of 30 fires in one county, the number of fire engines sent and the cost of the damage were recorded. The scatter graph of the data shows strong positive correlation. A newspaper prints the headline "Sending more fire engines causes more damage". Is the newspaper right? Give a reason for your answer.
- 11.The seven members of Team A took 20, 21, 22, 23, 24, 25 and 40 seconds to finish a task. The seven members of Team B took 20, 30, 32, 34, 36, 38 and 40 seconds. Tomás says that because the two teams have the same range, their times are spread out in the same way. Is Tomás right? Give a reason for your answer.
- 12.A table shows the time each of four pupils took to run 100 m: Oliver 12 seconds, Grace 15 seconds, Ethan 10 seconds, Freya 14 seconds. Write down the name of the fastest runner and the time taken.
- 13.A frequency polygon for the mass, in kg, of 40 parcels at a delivery depot is drawn by plotting one point at the midpoint of each class, joined by straight lines: (5, 6), (15, 10), (25, 16), (35, 6), (45, 2). Every class has a width of 10 kg. Write down the modal class.
- 14.A two-way table records whether each of the 30 pupils in a class passed maths and whether they passed science. 18 pupils passed maths, 12 pupils passed science and 8 pupils passed both. Work out how many pupils passed at least one of the two subjects.
- 15.Write down the statement that correctly describes the difference between correlation and causation.
Answer key
- (c) 7 — Method: the median is the middle value when the data are written in order of size, and with an odd number of values there is exactly one middle value. Working: the numbers are already in order, 2, 4, 7, 12, 26, and there are 5 of them, so the middle position is the third and the value sitting there is 7. Answer: 7, with two values below it and two above it. The distractors: 10.2 comes from working out the mean, 51 ÷ 5, instead of the median; 14 comes from taking the value halfway between the smallest and the largest, (2 + 26) ÷ 2; 24 comes from working out the range, 26 − 2, which measures spread rather than centre.
- (c) Because every pupil has an equal chance of being picked — Method: whether a sample represents its population is decided by the selection method, not by the size of the sample, so ask whether the method gives every member of the population the same chance of being chosen. Working: the names are drawn at random from a list of all 10,000 pupils, so each pupil has the same chance, 500 out of 10,000, of being drawn, and no group of pupils is more likely to appear than any other; that is what keeps bias out of the sample. Answer: because every pupil has an equal chance of being picked. The distractors: the reply about 5% treats the sampling fraction as the test of fairness, but a badly chosen 5% is still biased and a well chosen 1% is not; the reply about 500 being large enough makes size the test instead, which is the same mistake in another form, since a large sample drawn from one school would still misrepresent the city; the reply about the most willing pupils describes self-selection, which hands the choice of who is in the sample to the pupils who feel most strongly about the question.
- (b) On average a plant grew 1.5 cm taller for each extra day — Method: in the equation of a line, the number multiplying x is the gradient, and a gradient states the change in y produced by an increase of 1 in x, read in the units of the two axes. Working: here x is measured in days and y in centimetres, so the gradient 1.5 carries the units centimetres per day. Testing it on the line, 5 days gives 1.5 × 5 + 4 = 11.5 cm and 6 days gives 1.5 × 6 + 4 = 13 cm, a rise of 1.5 cm for the one extra day. Answer: on average a plant grew 1.5 cm taller for each extra day of watering. The distractors: 1.5 cm as the height before any watering is the value of y when x is 0, which is the other number in the equation, 4 cm, so this swaps the gradient and the intercept; 1.5 cm as the gap between the tallest and the shortest plant reads the gradient as a range, when a range is a difference between two of the 16 plants and a gradient is a rate; 1.5 days for each extra centimetre inverts the rate, dividing days by centimetres instead of centimetres by days, and the line gives 1 cm of growth in two thirds of a day.
- (a) The relationship between two variables — Method: what a diagram shows is decided by what has to be known before a single mark can be plotted on it. Working: every point on a scatter graph is plotted from a pair of measurements taken from the same person or object, one read on the horizontal axis and one on the vertical axis; having two measurements for each point is what makes it possible to look for a pattern between them, and the pattern between two variables is what the graph displays. Answer: a scatter graph shows the relationship between two variables. The distractors: the frequency of each single value is what a bar chart or a vertical line chart shows, and it needs only one list of values; how a total is shared between categories is what a pie chart shows; how one quantity changes over time is what a time series line graph shows, in which one of the two axes is always time.
- (c) 16 kg — Method: sort the seven weights before finding the middle value. Working: in order, the weights are 10, 12, 14, 16, 18, 20 and 50 kg. There are 7 values, so the median is the 4th one: 16 kg. Reading off the 4th weight in the order the vet recorded them, 10 kg, skips the sorting step and is not the median. Working out the mean, 140 ÷ 7 = 20 kg, finds a different average altogether. Working out the range, 50 − 10 = 40 kg, finds the spread, not the middle value. Always sort your data first — the median lives in the ordered list, not the collection order.
- (d) It asks two things at once and invites agreement — Method: a survey question is faulty when a reply to it cannot be read as evidence about one single thing, so check how many claims it contains and whether its wording pushes the reader one way. Working: the question joins two separate claims, that learning matters and that more homework should be set, so a reply of yes could mean either of them or both and cannot be counted as evidence about homework; the opening words “Do you agree” also invite agreement instead of leaving the reader free to say no. Answer: it asks two things at once and invites agreement. The distractors: the reply calling it too short mistakes length for clarity, when the fault is that too much has been packed in rather than too little; the reply about long words is false, since every word in the question is an everyday one and the fault lies in what is being asked rather than in the vocabulary used to ask it; the reply that the question is fine takes a yes or no answer as proof that the question works, which is exactly what a double question defeats.
- (c) 80 — Method: the range is a measure of spread and is found by subtracting the smallest value from the largest. Working: the largest number sold is 100 and the smallest is 20, so the range is 100 − 20 = 80. Answer: 80. The distractors: 100 comes from writing down the largest value and never subtracting the smallest; 60 comes from working out the median, the middle value of 20, 40, 60, 70, 100, instead of the range; 58 comes from working out the mean, 290 ÷ 5, which measures centre rather than spread.
- (c) 22 kg — Method: multiply the mean by the number of parcels to rebuild the total mass, then subtract the masses that are known. Working: four parcels with a mean mass of 17 kg have a total mass of 17 × 4 = 68 kg; the three known parcels total 12 + 16 + 18 = 46 kg; so the fourth parcel has mass 68 − 46 = 22 kg. Answer: 22 kg. The distractors: 68 kg comes from stopping at the total mass of all four parcels; 17 kg comes from assuming the missing parcel must have the mean mass; 5 kg comes from multiplying the mean by 3, the number of parcels whose mass is given, leaving 51 − 46 = 5.
- (a) 30 — Method: an outlier is a value that lies far away from the pattern set by the rest of the data, so compare each value with the group the others form. Working: five of the counts, 5, 7, 8, 9 and 11, lie within 6 of one another and the steps between them are 2, 1, 1 and 2; the remaining count of 30 is 19 above the nearest of them, so it is the value that does not belong to the pattern. Answer: 30. The distractors: 5 comes from picking the smallest value, on the idea that the odd one out must be at the bottom of the list; 11 comes from ordering the data and stopping one value short, taking the largest of the counts that sit close together; 8.5 comes from working out the median, (8 + 9) ÷ 2, and giving a measure of centre where a value standing apart was asked for.
- (a) No, the size of the fire affects both of the quantities — Method: correlation says that two quantities change together; a claim that one of them produces the other is a further claim, and it needs evidence that a scatter graph on its own cannot give. Working: the graph does show strong positive correlation, so more engines did go with greater damage. But neither quantity was set by the researchers: both were decided by how large the fire was. A large blaze brings many appliances and also destroys a great deal, while a small one brings few and destroys little, so a third quantity is driving both of the recorded ones. Answer: no, because the size of the fire affects both of the quantities. The distractors: saying the correlation is negative contradicts the graph, which shows the two quantities rising together, and reaching the right verdict from a false reading of the data is not the reason the mark is for; saying that strong positive correlation shows one quantity causes the other is the assumption the question exists to test, and no strength of correlation can establish cause; saying the points lie close to the line of best fit describes how strong the correlation is, and strength and cause are different matters entirely.
- (d) No, the range uses only the fastest and slowest time — Method: check what the range is built from, then look at what it leaves out. Working: both teams have a fastest time of 20 seconds and a slowest of 40 seconds, so both ranges are 40 − 20 = 20 seconds and Tomás has that part right. But the range is calculated from those two values alone. Six of Team A's seven times lie between 20 and 25 seconds, with a single time far out at 40; Team B is the other way round, with six of its seven times at 30 seconds or more and a single time far out at 20. So Team A bunches at the fast end and Team B at the slow end. The two patterns are quite different, and the range cannot see the difference because the five middle times never enter the calculation. Answer: no, because the range uses only the fastest and slowest time. The distractors: comparing the means answers a different question, since a mean measures position rather than spread, and two sets with the same spread can have different means; saying that equal ranges mean equal spread is the very assumption that fails here; saying that seven times each forces the spreads to match confuses the size of a data set with how its values are arranged inside it.
- (b) Ethan, 10 seconds — Method: over the same distance the fastest runner is the one who takes the least time, so the smallest time in the table is found first and the name is then read from the same row. Working: the four times are 12 seconds, 15 seconds, 10 seconds and 14 seconds; in order of size these are 10, 12, 14 and 15, so the least time is 10 seconds, and the row holding 10 seconds is the row for Ethan. Answer: Ethan, 10 seconds — the time is in seconds, and a smaller time means a faster runner. The distractors: Grace with 15 seconds comes from taking the largest number in the table to mean the fastest runner, which reverses the relationship between time and speed over a fixed distance; Oliver with 12 seconds comes from writing down the first row of the table without comparing the four times; Ethan with 15 seconds comes from identifying the right runner but then reading the time from a different row of the table.
- (c) 20 kg ≤ mass < 30 kg — The modal class is the class with the highest frequency. Reading the plotted points, the frequencies are 6, 10, 16, 6 and 2, so the highest frequency is 16, plotted at the midpoint 25. A class of width 10 centred on 25 runs from 25 − 5 = 20 to 25 + 5 = 30, so the modal class is 20 kg ≤ mass < 30 kg. Writing '25 kg' gives only the midpoint, not the class — the modal class is an interval, not a single value. '10 kg ≤ mass < 20 kg' is the class before the peak, centred on 15, which has frequency 10, not the highest. '30 kg ≤ mass < 40 kg' is the class after the peak, centred on 35, which has frequency 6, not the highest.
- (a) 22 — Method: the two subject totals overlap, because every pupil who passed both subjects has been counted once in the maths total and once again in the science total; adding the totals therefore counts those pupils twice, and the overlap has to be taken off once. Working: 18 + 12 = 30, and the 8 pupils who passed both have been counted twice in that 30, so the number who passed at least one subject is 30 − 8 = 22. Answer: 22 pupils, a count of pupils, and it is less than the 30 in the class, which leaves 8 pupils who passed neither. The distractors: 30 comes from adding the two subject totals and never removing the overlap, so it counts the 8 pupils twice; 14 comes from taking the 8 away twice, 18 + 12 − 8 − 8, removing an overlap that was only counted twice once too often; 18 comes from writing down the larger of the two subject totals on its own, which leaves out every pupil who passed science but not maths.
- (d) Correlation is a link; causation is one causing the other — Method: the two words describe different claims — one is about a pattern in the data, the other is about what produced that pattern. Working: correlation says only that two quantities tend to change together, which is something a scatter graph can display; causation says that a change in one quantity actually brings about the change in the other, which needs evidence a scatter graph cannot supply, because a third quantity may be driving both. Answer: correlation is a link between the quantities, while causation is one quantity causing the change in another. The distractors: the statement giving causation as the link and correlation as the cause simply swaps the two words over; the statement that the words mean the same thing is the classic error of reading a correlation as proof of cause; the statement that a scatter graph shows causation but not correlation reverses what a scatter graph can do, since the pattern it displays is exactly the correlation.
Build your own mix at the worksheet builder.