Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Calculator
Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- 1.A pupil writes this question for a school survey: “Do you agree that learning matters and that we should be set more homework?” Write down what is wrong with the survey question.
- 2.The numbers 2, 4, 6, 8, 10 and 12 have a mean of 7 and a median of 7. The value 100 is now added to the list. Which measure is changed more by adding 100, and why?
- 3.A school has 1,500 pupils. The head teacher takes a random sample of 150 of them from the school register and asks how long they spend on homework. Rory says the sample is too small for the result to mean anything. Is Rory right? Give a reason for your answer.
- 4.A two-way table records whether each of 50 pupils in Year 10 or Year 11 at a school walks to school or is driven. 28 of the 50 pupils are in Year 10. In total, 22 of the 50 pupils walk to school. Of the Year 10 pupils, 15 walk to school. Work out how many Year 11 pupils are driven to school.
- 5.Four pairs of variables are listed below. Write down the pair that you would expect to show negative correlation.
- 6.A sports centre in Newcastle asks 60 members which activity they prefer: swimming 15 members, football 25 members, badminton 12 members and other activities 8 members. Work out the angle at the centre of the pie chart sector that represents badminton.
- 7.A company makes 50,000 light bulbs a day and wants to check how long they last before they fail. Testing a bulb to find out how long it lasts destroys it. Give a reason why the company should test a sample of bulbs rather than every bulb it makes.
- 8.A factory made 3,000 phone cases last week. Shift A checked a random sample of 100 cases and found that 34 were scratched. Shift B checked a different random sample of 50 cases and found that 21 were scratched. Using the COMBINED results from both shifts, work out an estimate for the number of scratched cases made last week.
- 9.A study found that people who drink more coffee tend to concentrate better at work. A coffee company says that this shows that drinking coffee improves concentration. Give the reason why this conclusion cannot be drawn.
- 10.A scatter graph has 50 points. Most of them lie close to a rising line of best fit, but two of them lie a long way from that line. Write down how those two points should be treated.
- 11.Priya wants to find out the favourite sport of the 900 pupils at her school. She asks 5 pupils, chosen at random. Give a reason why her sample may not give a reliable result.
- 12.A company found that the higher the price it charged for a product, the lower the satisfaction score its customers gave. The price and the satisfaction score are plotted on a scatter graph. Describe the line of best fit that would be drawn on that graph.
- 13.A sports centre in Ipswich has 2,000 members. It wants to know what its members think of its opening hours, so it asks a random sample of 100 of them. Write down what the population is in this survey.
- 14.A scatter graph of the number of hours, x, that pupils revised against their test score, y, has the line of best fit y = 2.5x + 15. Amelia wants a score of at least 80. Work out the least whole number of hours of revision the line of best fit suggests she needs.y = 2.5x + 15
- 15.On a scatter graph of the age of a car, in years, and its value, in pounds, the points fall from left to right. Write down the type of correlation shown.
Answer key
- (d) It asks two things at once and invites agreement — Method: a survey question is faulty when a reply to it cannot be read as evidence about one single thing, so check how many claims it contains and whether its wording pushes the reader one way. Working: the question joins two separate claims, that learning matters and that more homework should be set, so a reply of yes could mean either of them or both and cannot be counted as evidence about homework; the opening words “Do you agree” also invite agreement instead of leaving the reader free to say no. Answer: it asks two things at once and invites agreement. The distractors: the reply calling it too short mistakes length for clarity, when the fault is that too much has been packed in rather than too little; the reply about long words is false, since every word in the question is an everyday one and the fault lies in what is being asked rather than in the vocabulary used to ask it; the reply that the question is fine takes a yes or no answer as proof that the question works, which is exactly what a double question defeats.
- (c) The mean, because every value counts towards it, so 100 pulls it from 7 up to about 20.3. — Method: work each measure out before the extra value is added and again afterwards, then compare the size of the two changes. Working: before, the six values total 42, so the mean is 42 ÷ 6 = 7, and the middle pair 6 and 8 give a median of (6 + 8) ÷ 2 = 7; after, the seven values total 142, so the mean is 142 ÷ 7 = 20.29 to 2 decimal places, while the median is now the 4th of the seven ordered values, which is 8; the mean has moved by about 13.3 and the median by 1. Answer: the mean, because every value counts towards it, so 100 pulls it from 7 up to about 20.3 — this is why the median is often preferred when a data set contains an outlier. The distractors: the reply that the mean rises by 100 adds the extra value to the mean instead of adding it to the total; the reply that the median moves to 12 takes the largest of the original values as the new middle instead of counting to the 4th of the seven values; the reply about even and odd counts quotes a rule that does not exist, since the median moved because a very large value was added, not because the count of values changed.
- (a) No, 150 pupils are a tenth of the school, chosen at random — Method: judge a sample on two things, whether every member of the population had the same chance of being chosen, and whether the sample is large enough to carry a pattern. Working: the 150 pupils were drawn from the register of every pupil in the school, so no year group or set is shut out and no pupil chooses to take part; and 150 ÷ 1,500 = 0.1, so one pupil in ten has been asked. A random sample of that share is ample for an estimate of how long the school's pupils spend on homework. Answer: no, because 150 pupils are a tenth of the school and were chosen at random. The distractors: saying a random sample always gives the exact school figure reaches the same verdict for a reason that is false, since a second random sample of 150 would give a slightly different mean; saying 150 pupils cannot be picked at random from 1,500 treats randomness as something only a whole population can have, when drawing names from the register is exactly how a random sample is taken; saying that only asking all 1,500 could show anything rejects sampling altogether, which would leave no way to study any population too large to count.
- (a) 15 — Year 11 has 50 − 28 = 22 pupils in total. Of the 22 pupils who walk in total, 15 are in Year 10, so 22 − 15 = 7 Year 11 pupils walk. Subtracting that from the Year 11 total gives 22 − 7 = 15 Year 11 pupils who are driven. Choosing 28 takes the whole school's driven total, 50 − 22 = 28, and treats it as if it were Year 11's alone, without separating the year groups. Choosing 7 correctly finds how many Year 11 pupils walk but stops there, giving that figure instead of the number who are driven. Choosing 35 comes from 50 − 15, subtracting the Year 10 walkers from the whole school total rather than working within Year 11.
- (a) Minutes a candle has burned and length remaining — As a candle burns for longer, less of it remains, so these two variables move in opposite directions as one increases — that is negative correlation. A pupil's shoe size generally increases as they get older, so age and shoe size show positive correlation, not negative, since both rise together. A football team's shirt colour is not a numerical quantity linked to how many matches it wins, so shirt colour and number of wins show no correlation at all. The number of letters in a pupil's name has no real connection to their ability in maths, so that pair also shows no correlation.
- (c) 72° — The angle is 12 ÷ 60 × 360 = 72°. Choosing 150° divides football's frequency of 25 instead of badminton's 12: 25 ÷ 60 × 360 = 150. Choosing 20° finds badminton as a percentage of the members, 12 ÷ 60 × 100 = 20, rather than an angle in degrees. Choosing 90° uses 48, the total of the other three activities, as the total instead of the full 60 members: 12 ÷ 48 × 360 = 90.
- (d) Testing destroys bulbs, so testing all leaves none to sell. — Method: testing every item in a population instead of a sample is a census — sensible only when testing does not use up or destroy what is being tested. Working: here, testing a bulb to find its lifespan destroys it, so testing all 50,000 bulbs would leave nothing left to sell — a sample lets the company estimate the typical lifespan without destroying its whole stock. Extra electricity used in testing is not the real reason a census is avoided here — it is the destruction of the product that matters. Saying a sample is always more accurate than a full census is the wrong way round: a census, if it could be carried out, gives the exact figure for the whole population — it is testing being destructive, not a lack of accuracy, that rules it out here. There is no law against testing every item a company makes — nothing in the question suggests that. When testing destroys the item being tested, sampling is necessary, not just convenient.
- (d) 1100 — Method: to combine two samples of different sizes, add the faulty counts together and add the sample sizes together before scaling up, rather than treating the two samples separately. Working: the combined sample found 34 + 21 = 55 scratched cases out of 100 + 50 = 150 cases checked, a proportion of 55 ÷ 150. Applying that proportion to the week's production of 3,000 gives an estimate of 55 ÷ 150 × 3000 = 1100 scratched cases. Averaging the two shifts' proportions instead of combining their totals, (34 ÷ 100 + 21 ÷ 50) ÷ 2 = 0.38, gives 0.38 × 3000 = 1140 — this treats the two samples as equally weighted even though Shift A checked twice as many cases as Shift B. Using only Shift A's sample, 34 ÷ 100 × 3000 = 1020, ignores Shift B's cases completely. Using only Shift B's sample, 21 ÷ 50 × 3000 = 1260, ignores Shift A's cases completely. When two samples are different sizes, combine their totals before finding the proportion — do not average the two proportions, and do not use only one shift's sample.
- (c) The data show a link only; a third factor may affect both — Method: a study of this kind measures two quantities and reports how they change together; deciding that one of them produces the other is a further claim, and it needs evidence that the measurements alone cannot give. Working: the study shows that more coffee goes with better concentration, which is a positive correlation; but a third factor that was never measured, such as how motivated someone is, could raise both the coffee drinking and the concentration, and the concentration could equally be what leads to the extra coffee. Answer: the data show a link only, because a third factor may be affecting both quantities, so no claim about cause can be made. The distractors: calling the conclusion safe because the correlation is positive treats the direction of a correlation as proof of cause, which no direction can give; calling it wrong because the correlation is negative misreads the direction of the relationship, since the study reports both quantities rising together; saying the two quantities are not linked denies the correlation the study actually found, when what fails is only the claim about cause.
- (b) Treat them as outliers and check them before deciding — Method: a point lying a long way from the pattern the rest of the data make is called an outlier, and an outlier is investigated before anything is done with it, because it may be an error in the data or it may be a genuine but unusual case. Working: 48 of the 50 points lie close to the rising line of best fit, so the trend is set by those 48; the two remaining points do not follow it, so they are identified as outliers and checked — a mistake in measuring or recording would be corrected, while a genuine reading would be kept and reported. Answer: treat them as outliers and check them before deciding what to do with them. The distractors: deleting them at once assumes that every point far from the line must be an error, which throws away real data; moving the line so that it passes through them assumes a line of best fit must touch particular points, when it is drawn to follow all 50; taking them as proof that there is no correlation lets two points overturn the pattern that the other 48 agree on.
- (a) 5 pupils are far too few to represent 900 pupils — Method: a sample can only support a claim about a population if it is chosen fairly and if it is large enough for the pattern in it to be more than chance. Working: Priya's method of choosing is fair, because the 5 pupils were picked at random, so every pupil had the same chance of being asked. The difficulty is the size: 900 ÷ 5 = 180, so each pupil she asks stands for 180 pupils. If two of the five happen to play in the same netball team, netball takes 40% of her sample on the strength of two answers, and a second sample of 5 could easily give a different favourite sport. Answer: 5 pupils are far too few to represent 900 pupils. The distractors: saying the pupils were not chosen at random contradicts the question, which states that they were; saying the 5 may each name a different sport describes what often happens in a small sample, but disagreement is not the fault, since 5 pupils who all named the same sport would be just as weak a basis for a claim about 900; saying a sample must hold at least half of the population is an invented rule, and a properly chosen sample of a few hundred can describe a population of many thousands.
- (a) A straight line sloping down from left to right — Method: a line of best fit is a straight line drawn to follow the trend of the points, so its slope is decided by the direction of the relationship between the two quantities. Working: as the price rises, the satisfaction score falls, so the points start high on the left of the graph and finish low on the right; the straight line that follows them therefore slopes downwards as the graph is read from left to right, which is the line of a negative correlation. Answer: a straight line sloping down from left to right. The distractors: a line sloping up comes from reading a falling relationship as a rising one; a horizontal line comes from expecting no correlation, since a horizontal line says the satisfaction score does not change as the price changes; a curve passing through every point comes from thinking a line of best fit has to touch all of the plotted points, when it is a single straight line drawn through the middle of them.
- (b) All 2,000 members of the sports centre. — Method: in a survey, the population is the whole group the survey is trying to find out about, and the sample is the smaller group actually asked. Working: this survey wants to know what the sports centre's members think, so the population is every one of the 2,000 members — whether or not they were personally asked. Saying the population is the 100 members who were asked names the sample, not the population; the sample is drawn FROM the population, so it is smaller than it, not the same as it. Saying the population is everybody who lives in Ipswich widens the group far beyond who the survey is actually about — plenty of Ipswich residents are not members of the sports centre at all, so they are outside this survey altogether. Saying the population is the members who say they are unhappy confuses the population with a result of the survey: whether a member turns out to be happy or unhappy is something the survey finds out, not part of the definition of who is being studied. The population is always the whole group the question is about, before any sampling or any results come in.
- (a) 26 — Method: a line of best fit lets one quantity be predicted from the other, so the score is substituted into the equation of the line and the resulting inequality is solved for the number of hours. Working: a score of at least 80 means 2.5x + 15 ≥ 80; taking 15 from both sides gives 2.5x ≥ 65, and dividing both sides by 2.5 gives x ≥ 26, so the least whole number of hours is 26. Checking, 2.5 × 26 + 15 = 80, which does reach the target. Answer: 26 hours — and this is only an estimate, because a line of best fit predicts a trend rather than an individual result, and a prediction made outside the range of hours the pupils actually revised for would be an extrapolation and less reliable still. The distractors: 27 comes from reaching 26 and then rounding up again, although 26 hours already gives a score of exactly 80; 32 comes from 80 ÷ 2.5, which ignores the 15 in the equation of the line; 38 comes from (80 + 15) ÷ 2.5, that is from adding the 15 instead of subtracting it when rearranging.
- (a) Negative correlation — As the age of the car increases, the points fall towards a lower value, so the value decreases as the age increases. This falling pattern is a negative correlation. A positive correlation would show the points rising together instead. No correlation would apply only if the points showed no pattern at all, and correlation is not the same as causation — strong causation is not a type of correlation.
Build your own mix at the worksheet builder.