Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Non-calculator
Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- 1.A scatter graph plots the shoe size and the spelling test score of 25 pupils. The points are scattered with no pattern across the graph. Write down the type of correlation shown.
- 2.A scatter graph has 12 plotted points. Four pupils each draw a line of best fit on the same graph and count how many of the 12 points lie above their line and how many lie below it: Amir — 10 above, 2 below. Priya — 6 above, 6 below. Kofi — 2 above, 10 below. Leah — 0 above, 12 below. Write down the name of the pupil whose line was drawn correctly, so that the points are roughly balanced above and below it.
- 3.The masses of eight school bags, in kilograms, are 3, 4, 4, 5, 6, 7, 8 and 11. Work out the median mass.
- 4.The marks scored by four pupils in a quiz were 8, 8, 8, 8. Work out the mean, the median and the mode of these marks.
- 5.A company makes 50,000 light bulbs a day and wants to check how long they last before they fail. Testing a bulb to find out how long it lasts destroys it. Give a reason why the company should test a sample of bulbs rather than every bulb it makes.
- 6.A charity shop in Bath holds 2,000 books. Volunteer A checks a random sample of 50 books and finds 35 paperbacks. Volunteer B checks a different random sample of 50 books and finds 31 paperbacks. Work out the estimate each sample gives for the whole stock, and write down what the shop should do next.
- 7.A gardener plants 600 daffodil bulbs. In March she digs up 20 of them, chosen at random, and finds that 17 have flowered. Write down what she can conclude about the 600 bulbs.
- 8.A school has 1,000 pupils. A student wants to estimate how many of them walk to school, so she asks 8 pupils in her own class. She says her sample is large enough to give a reliable estimate for the whole school. Is she right? Give a reason for your answer.
- 9.On a scatter graph of the arm span and the height of some pupils, the points rise from left to right. Write down the type of correlation this shows.
- 10.Two classes sat the same test. The 30 pupils in Class A had a mean mark of 72. The 20 pupils in Class B had a mean mark of 82. Work out the mean mark of all 50 pupils.
- 11.A garden centre's sales, in thousands of pounds, at the end of each quarter last year were: quarter 1 — 18, quarter 2 — 34, quarter 3 — 30, quarter 4 — 22. Work out the increase in sales from quarter 1 to the quarter with the highest sales.
- 12.A scatter graph shows the midday temperature, x °C, and the number of ice creams sold at a seaside kiosk, y. The line of best fit is y = 3x − 20. Give a reason why the y-intercept of this line of best fit is not a sensible estimate of the number of ice creams sold.y = 3x − 20
- 13.A company found that the higher the price it charged for a product, the lower the satisfaction score its customers gave. The price and the satisfaction score are plotted on a scatter graph. Describe the line of best fit that would be drawn on that graph.
- 14.Priya wants to find out the favourite sport of the 900 pupils at her school. She asks 5 pupils, chosen at random. Give a reason why her sample may not give a reliable result.
- 15.A pupil writes this question for a school survey: “Do you agree that learning matters and that we should be set more homework?” Write down what is wrong with the survey question.
Answer key
- (a) No correlation — Shoe size has no real relationship with spelling ability, and the points here are scattered with no rising or falling trend, so this is no correlation. A positive correlation would show the points rising together, and a negative correlation would show them falling as one increases; neither pattern is present here. Strong correlation is not correct either, since strength only applies once a positive or negative trend exists, and there isn't one.
- (a) Priya — A line of best fit should be drawn so that the plotted points are roughly balanced above and below it. Work out the difference between the two counts for each pupil: Amir 10 − 2 = 8; Kofi 10 − 2 = 8 (10 below and 2 above); Leah 12 − 0 = 12; Priya 6 − 6 = 0. Priya's line has the smallest difference, an exact balance of 6 above and 6 below, so her line is drawn correctly. Amir's line has 10 of the 12 points above it, so it is drawn too low. Kofi's line has 10 of the 12 points below it, so it is drawn too high. Leah's line has every single point below it, so it is not a line of best fit at all.
- (b) 5.5 kg — Method: with an even number of values the median is the mean of the two middle values, taken once the data are in order of size. Working: the eight masses are already in order and 8 ÷ 2 = 4, so the middle pair are the 4th and 5th values, 5 kg and 6 kg; the median is (5 + 6) ÷ 2 = 5.5 kg. Answer: 5.5 kg. The distractors: 5 kg comes from reading the 4th value and stopping there instead of averaging the middle pair; 8 kg comes from working out the range, 11 − 3, which measures spread rather than centre; 4 kg comes from writing down the modal mass, the only value that occurs twice, instead of the median.
- (c) mean = 8, median = 8, mode = 8 — Method: work out each measure separately — the mean is the total divided by how many values there are, the median is the middle value once the data are in order, and the mode is the value that occurs most often. Working: the total is 8 + 8 + 8 + 8 = 32 and there are 4 marks, so the mean is 32 ÷ 4 = 8; in order the marks read 8, 8, 8, 8, and the mean of the middle pair is (8 + 8) ÷ 2 = 8; the value 8 occurs 4 times and no other value occurs at all, so the mode is 8. Answer: mean = 8, median = 8, mode = 8 — when every value in a data set is the same, all three measures of central tendency take that value. The distractors: a mean of 32 comes from stopping at the total and never dividing by 4; a mode of 4 comes from writing down how many times 8 occurs instead of the value that occurs; a mean of 2 comes from dividing a single value, 8, by the 4 marks instead of dividing the total by 4.
- (d) Testing destroys bulbs, so testing all leaves none to sell. — Method: testing every item in a population instead of a sample is a census — sensible only when testing does not use up or destroy what is being tested. Working: here, testing a bulb to find its lifespan destroys it, so testing all 50,000 bulbs would leave nothing left to sell — a sample lets the company estimate the typical lifespan without destroying its whole stock. Extra electricity used in testing is not the real reason a census is avoided here — it is the destruction of the product that matters. Saying a sample is always more accurate than a full census is the wrong way round: a census, if it could be carried out, gives the exact figure for the whole population — it is testing being destructive, not a lack of accuracy, that rules it out here. There is no law against testing every item a company makes — nothing in the question suggests that. When testing destroys the item being tested, sampling is necessary, not just convenient.
- (a) 1,400 and 1,240, so combine the samples for one estimate — Method: scale each sample up to the whole stock, then use the fact that a larger sample gives a more reliable estimate than a smaller one. Working: the first sample gives 35 ÷ 50 = 0.7 and 0.7 × 2,000 = 1,400 paperbacks; the second gives 31 ÷ 50 = 0.62 and 0.62 × 2,000 = 1,240 paperbacks. Two random samples of the same size are expected to differ a little, so neither estimate is wrong. Putting the two together gives 35 + 31 = 66 paperbacks in 100 books, and 66 ÷ 100 = 0.66 with 0.66 × 2,000 = 1,320, an estimate resting on twice as many books as either volunteer checked. Answer: 1,400 and 1,240, so combine the samples for one estimate. The distractors: keeping 1,400 because it is larger picks an estimate by its size, when both samples held 50 books and neither has a stronger claim; saying a volunteer must have miscounted assumes two random samples ought to agree exactly, which is precisely what random sampling does not promise; 1,750 and 1,550 come from 35 × 50 = 1,750 and 31 × 50 = 1,550, multiplying each count by the size of the sample instead of scaling by 2,000 ÷ 50.
- (d) About 510 of the 600 bulbs are likely to have flowered — Method: the proportion found in a random sample is used as an estimate of the proportion in the whole population, and the conclusion is stated as an estimate, never as a fact about every member. Working: 17 of the 20 bulbs dug up had flowered, so the sample proportion is 17 ÷ 20 = 0.85, and applying that proportion to the whole planting gives 0.85 × 600 = 510 bulbs. A different random sample of 20 would very probably give a slightly different figure, so 510 is an estimate. Answer: about 510 of the 600 bulbs are likely to have flowered. The distractors: saying exactly 510 have flowered takes an estimate from a sample of 20 as a count of all 600, which no sample can deliver; saying exactly 17 of the 600 have flowered reports the sample count as though it were the population count, leaving the other 580 bulbs out of the answer altogether; saying about 20 have flowered uses the size of the sample as the estimate, when 20 is the number of bulbs she dug up rather than a number that flowered.
- (c) No — 8 from one class is too small to represent the school. — Method: judge reliability by asking whether the sample is both large enough, and spread across the population, relative to what it is meant to represent. Working: 8 pupils is a tiny fraction of the school's 1,000 pupils, and all 8 come from a single class rather than a range of year groups, so the sample is both too small and too narrow to represent the whole school reliably. She is not right. Saying any sample size gives an equally reliable estimate ignores that reliability generally improves with a larger, more representative sample. Saying the method is unreliable because it was not done online is not a reason connected to sample size or representativeness at all. Saying 8 is reliable because it is more than half her class compares the sample to the wrong population — the school has 1,000 pupils, not one class. Always judge a sample's size against the population it is meant to represent, not against a smaller group within it.
- (b) Positive correlation — Method: the type of correlation is named from the direction the points take as the scatter graph is read from left to right. Working: the points rise from left to right, so as the arm span read on the horizontal axis increases, the height read on the vertical axis increases as well; two quantities that increase together show positive correlation. Answer: positive correlation. The distractors: negative correlation comes from naming the direction the wrong way round, since a negative correlation needs the points to fall as the graph is read from left to right; no correlation comes from treating points that are spread out rather than sitting exactly on a line as though they showed no relationship; direct proportion comes from confusing a rising trend with proportion, which would additionally need the line through the points to pass through the origin and would mean doubling one quantity doubles the other.
- (c) 76 marks — Method: a mean of means only works when the groups are the same size, so rebuild each class's total mark, add the totals and divide by the number of pupils altogether. Working: Class A scored 30 × 72 = 2160 marks and Class B scored 20 × 82 = 1640 marks, giving 2160 + 1640 = 3800 marks between 50 pupils, so the overall mean is 3800 ÷ 50 = 76 marks. Answer: 76 marks. The distractors: 77 marks comes from averaging the two class means, (72 + 82) ÷ 2, which ignores the different class sizes; 78 marks comes from attaching each mean to the other class's size, (30 × 82 + 20 × 72) ÷ 50; 3800 marks comes from stopping at the combined total and never dividing by 50.
- (a) £16,000 — Method: first find the quarter with the highest sales figure, then subtract quarter 1's sales from it — remembering that every figure is given in THOUSANDS of pounds. Working: the highest sales figure is quarter 2, at £34,000 (34 thousand pounds). The increase from quarter 1 is £34,000 − £18,000 = £16,000. Giving £34,000 reads off the highest sales figure on its own, without subtracting quarter 1's sales — that is the highest quarter's total, not the increase. Giving £12,000 uses quarter 3's sales, 30, the SECOND-highest figure, instead of quarter 2's 34, the actual highest — 30 − 18 = 12, but quarter 3 is not the quarter with the highest sales. Giving £16 gets the subtraction right, 34 − 18 = 16, but forgets that every figure in the question is in thousands of pounds, so the increase is £16,000, not £16. Always identify the correct quarter FIRST, and always check the units the numbers are given in before writing your final answer.
- (a) x = 0 gives y = −20: a negative number sold — The y-intercept is the value the line predicts when x = 0: y = 3 × 0 − 20 = −20. A kiosk cannot sell a negative number of ice creams, so this is not a sensible estimate. The 3 in the equation is the gradient, not the intercept, so an option claiming x = 0 gives y = 3 has swapped the two numbers around — substituting x = 0 makes the 3x term equal 0, leaving −20, not 3. The danger of extrapolating to very high temperatures is a real issue with this line, but it is a different issue from the y-intercept, so it does not answer this question. And whether x = 0 could occur on a trading day is beside the point: the model still makes that prediction, and it is the prediction itself, −20, that is impossible.
- (a) A straight line sloping down from left to right — Method: a line of best fit is a straight line drawn to follow the trend of the points, so its slope is decided by the direction of the relationship between the two quantities. Working: as the price rises, the satisfaction score falls, so the points start high on the left of the graph and finish low on the right; the straight line that follows them therefore slopes downwards as the graph is read from left to right, which is the line of a negative correlation. Answer: a straight line sloping down from left to right. The distractors: a line sloping up comes from reading a falling relationship as a rising one; a horizontal line comes from expecting no correlation, since a horizontal line says the satisfaction score does not change as the price changes; a curve passing through every point comes from thinking a line of best fit has to touch all of the plotted points, when it is a single straight line drawn through the middle of them.
- (a) 5 pupils are far too few to represent 900 pupils — Method: a sample can only support a claim about a population if it is chosen fairly and if it is large enough for the pattern in it to be more than chance. Working: Priya's method of choosing is fair, because the 5 pupils were picked at random, so every pupil had the same chance of being asked. The difficulty is the size: 900 ÷ 5 = 180, so each pupil she asks stands for 180 pupils. If two of the five happen to play in the same netball team, netball takes 40% of her sample on the strength of two answers, and a second sample of 5 could easily give a different favourite sport. Answer: 5 pupils are far too few to represent 900 pupils. The distractors: saying the pupils were not chosen at random contradicts the question, which states that they were; saying the 5 may each name a different sport describes what often happens in a small sample, but disagreement is not the fault, since 5 pupils who all named the same sport would be just as weak a basis for a claim about 900; saying a sample must hold at least half of the population is an invented rule, and a properly chosen sample of a few hundred can describe a population of many thousands.
- (d) It asks two things at once and invites agreement — Method: a survey question is faulty when a reply to it cannot be read as evidence about one single thing, so check how many claims it contains and whether its wording pushes the reader one way. Working: the question joins two separate claims, that learning matters and that more homework should be set, so a reply of yes could mean either of them or both and cannot be counted as evidence about homework; the opening words “Do you agree” also invite agreement instead of leaving the reader free to say no. Answer: it asks two things at once and invites agreement. The distractors: the reply calling it too short mistakes length for clarity, when the fault is that too much has been packed in rather than too little; the reply about long words is false, since every word in the question is an everyday one and the fault lies in what is being asked rather than in the vocabulary used to ask it; the reply that the question is fine takes a yes or no answer as proof that the question works, which is exactly what a double question defeats.
Build your own mix at the worksheet builder.