Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Non-calculator
Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- 1.A sports shop in Cardiff sold 40 pairs of football boots last month: 4 pairs of size 6, 5 pairs of size 7, 8 pairs of size 8, 13 pairs of size 9 and 10 pairs of size 10. The manager will order 40 pairs for next month and wants as many pairs as possible to be in a size customers will buy. Work out the mean size and the modal size, and write down which of the two he should use.
- 2.A pie chart is divided into 10 equal sectors. 6 of those sectors stand for the people who chose purple. Altogether 120 people were asked. Work out how many of them chose purple.
- 3.Ben and Chloe each sat five maths tests. Ben's marks were 62, 64, 65, 66 and 68. Chloe's marks were 40, 52, 65, 78 and 90. Both pupils have a mean mark of 65. Their teacher says the mean on its own does not describe the two sets of marks well. Give a reason why the teacher is right.
- 4.A scatter graph has 12 plotted points. Four pupils each draw a line of best fit on the same graph and count how many of the 12 points lie above their line and how many lie below it: Amir — 10 above, 2 below. Priya — 6 above, 6 below. Kofi — 2 above, 10 below. Leah — 0 above, 12 below. Write down the name of the pupil whose line was drawn correctly, so that the points are roughly balanced above and below it.
- 5.A school has 1,500 pupils. The head teacher takes a random sample of 150 of them from the school register and asks how long they spend on homework. Rory says the sample is too small for the result to mean anything. Is Rory right? Give a reason for your answer.
- 6.The mean of the four numbers 10, 15, 20 and x is 18. Work out the value of x.
- 7.Write down the statement that correctly describes the difference between correlation and causation.
- 8.A scatter graph of the number of hours, x, that pupils revised against their test score, y, has the line of best fit y = 2.5x + 15. Amelia wants a score of at least 80. Work out the least whole number of hours of revision the line of best fit suggests she needs.y = 2.5x + 15
- 9.A council in Leeds wants to know what local people think about letting shops stay open later in the evening. It rings landline telephone numbers between 10 am and 2 pm on a Tuesday. Write down which group is most likely to be under-represented in the sample, and give a reason for your answer.
- 10.The marks scored by four pupils in a quiz were 8, 8, 8, 8. Work out the mean, the median and the mode of these marks.
- 11.On a scatter graph of the age of a car, in years, and its value, in pounds, the points fall from left to right. Write down the type of correlation shown.
- 12.A garden centre's sales, in thousands of pounds, at the end of each quarter last year were: quarter 1 — 18, quarter 2 — 34, quarter 3 — 30, quarter 4 — 22. Work out the increase in sales from quarter 1 to the quarter with the highest sales.
- 13.On a scatter graph the horizontal axis shows height in centimetres and the vertical axis shows mass in kilograms. One point is plotted at (170, 65). Write down the height and the mass of that person.
- 14.A garden centre records its total sales, in pounds, at the end of each of the twelve months of one year, to see how sales change over time. Write down the most suitable type of chart to show this time series data.
- 15.A frequency polygon for the mass, in kg, of 40 parcels at a delivery depot is drawn by plotting one point at the midpoint of each class, joined by straight lines: (5, 6), (15, 10), (25, 16), (35, 6), (45, 2). Every class has a width of 10 kg. Write down the modal class.
Answer key
- (a) The modal size, 9, bought by more customers than any other — Method: work out both averages from the frequencies, then choose the one the shop can act on. Working: for the mean, multiply each size by the number of pairs sold at it and add: 6 × 4 + 7 × 5 + 8 × 8 + 9 × 13 + 10 × 10 = 340, and 340 ÷ 40 = 8.5, so the mean size is 8.5. The largest frequency is 13, which belongs to size 9, so the modal size is 9. The mean 8.5 is a size no customer in the record asked for, so 40 pairs of it would sit unsold, while 13 of the 40 customers wanted size 9, more than wanted any other size. Answer: the modal size, 9, bought by more customers than any other. The distractors: the mean size 8.5 does take account of all 40 pairs, but a mean of sizes is a summary figure and not a size the month's customers were buying; the mean size 8 comes from averaging the five sizes on sale, 6 + 7 + 8 + 9 + 10 = 40 and 40 ÷ 5 = 8, which ignores how many pairs were sold at each size and so treats the 4 pairs of size 6 as equal in weight to the 13 pairs of size 9; the range 4 comes from 10 − 6 and measures spread, so it says how wide a set of sizes the shop must stock, not which size to stock most of.
- (c) 72 — Method: when a pie chart is divided into equal sectors, each sector stands for the same share of the people asked, so the fraction of the sectors that are shaded is also the fraction of the people. Working: 6 sectors out of 10 are purple, which is the fraction 6/10 of the whole pie chart; one tenth of the 120 people is 120 ÷ 10 = 12 people, so six tenths is 6 × 12 = 72 people. Answer: 72 people, a count of people rather than a number of sectors. The distractors: 48 comes from working with the 4 sectors that are not purple, 4 × 12, and so answering for the wrong part of the chart; 60 comes from turning the fraction 6/10 into 60% and then writing the 60 down as though it were a number of people; 6 comes from writing down the number of purple sectors instead of the number of people those sectors stand for.
- (a) Chloe's marks are far more spread out than Ben's — Method: a mean reports where a set of values sits, and two sets can sit in the same place while behaving quite differently, so a measure of spread has to be worked out as well. Working: Ben's marks add to 62 + 64 + 65 + 66 + 68 = 325 and 325 ÷ 5 = 65; Chloe's add to 40 + 52 + 65 + 78 + 90 = 325 and 325 ÷ 5 = 65, so the two means agree, as the question says. The ranges do not: Ben's is 68 − 62 = 6 marks, while Chloe's is 90 − 40 = 50 marks. Ben's five marks all sit within 3 marks of 65; Chloe's lowest is 25 marks below it and her highest 25 marks above it. Answer: Chloe's marks are far more spread out than Ben's, which is exactly what the mean cannot show. The distractors: saying Ben's marks are more spread out comes from subtracting in the order the values are written, 62 − 68 = −6 against 40 − 90 = −50, and then reading −6 as the larger spread; saying Chloe scored far more marks in total assumes a wider set of marks must add to more, when both totals are 325; saying the two sets vary by the same amount assumes that equal means force equal spread, when the two ranges are 6 and 50.
- (a) Priya — A line of best fit should be drawn so that the plotted points are roughly balanced above and below it. Work out the difference between the two counts for each pupil: Amir 10 − 2 = 8; Kofi 10 − 2 = 8 (10 below and 2 above); Leah 12 − 0 = 12; Priya 6 − 6 = 0. Priya's line has the smallest difference, an exact balance of 6 above and 6 below, so her line is drawn correctly. Amir's line has 10 of the 12 points above it, so it is drawn too low. Kofi's line has 10 of the 12 points below it, so it is drawn too high. Leah's line has every single point below it, so it is not a line of best fit at all.
- (a) No, 150 pupils are a tenth of the school, chosen at random — Method: judge a sample on two things, whether every member of the population had the same chance of being chosen, and whether the sample is large enough to carry a pattern. Working: the 150 pupils were drawn from the register of every pupil in the school, so no year group or set is shut out and no pupil chooses to take part; and 150 ÷ 1,500 = 0.1, so one pupil in ten has been asked. A random sample of that share is ample for an estimate of how long the school's pupils spend on homework. Answer: no, because 150 pupils are a tenth of the school and were chosen at random. The distractors: saying a random sample always gives the exact school figure reaches the same verdict for a reason that is false, since a second random sample of 150 would give a slightly different mean; saying 150 pupils cannot be picked at random from 1,500 treats randomness as something only a whole population can have, when drawing names from the register is exactly how a random sample is taken; saying that only asking all 1,500 could show anything rejects sampling altogether, which would leave no way to study any population too large to count.
- (a) 27 — Method: turn the mean into a total using total = mean × number of values, then subtract the numbers that are already known. Working: four numbers with a mean of 18 have a total of 18 × 4 = 72; the three known numbers give 10 + 15 + 20 = 45; so x = 72 − 45 = 27. Answer: 27, and checking, (10 + 15 + 20 + 27) ÷ 4 = 72 ÷ 4 = 18. The distractors: 72 comes from stopping at the total the four numbers must reach and never subtracting the known three; 18 comes from assuming the missing number must equal the mean; 45 comes from stopping at the total of the three known numbers.
- (d) Correlation is a link; causation is one causing the other — Method: the two words describe different claims — one is about a pattern in the data, the other is about what produced that pattern. Working: correlation says only that two quantities tend to change together, which is something a scatter graph can display; causation says that a change in one quantity actually brings about the change in the other, which needs evidence a scatter graph cannot supply, because a third quantity may be driving both. Answer: correlation is a link between the quantities, while causation is one quantity causing the change in another. The distractors: the statement giving causation as the link and correlation as the cause simply swaps the two words over; the statement that the words mean the same thing is the classic error of reading a correlation as proof of cause; the statement that a scatter graph shows causation but not correlation reverses what a scatter graph can do, since the pattern it displays is exactly the correlation.
- (a) 26 — Method: a line of best fit lets one quantity be predicted from the other, so the score is substituted into the equation of the line and the resulting inequality is solved for the number of hours. Working: a score of at least 80 means 2.5x + 15 ≥ 80; taking 15 from both sides gives 2.5x ≥ 65, and dividing both sides by 2.5 gives x ≥ 26, so the least whole number of hours is 26. Checking, 2.5 × 26 + 15 = 80, which does reach the target. Answer: 26 hours — and this is only an estimate, because a line of best fit predicts a trend rather than an individual result, and a prediction made outside the range of hours the pupils actually revised for would be an extrapolation and less reliable still. The distractors: 27 comes from reaching 26 and then rounding up again, although 26 hours already gives a score of exactly 80; 32 comes from 80 ÷ 2.5, which ignores the 15 in the equation of the line; 38 comes from (80 + 15) ÷ 2.5, that is from adding the 15 instead of subtracting it when rearranging.
- (d) Full-time workers, as most are at work at that time — Method: a sample is biased when the method of contact makes part of the population much less likely to be reached, so test each group against where its members actually are between 10 am and 2 pm on a weekday, and test each stated reason against the facts. Working: those hours are the middle of the working day, so people in full-time employment are at work and not beside a landline telephone, while people who are retired and people who are unemployed are far more likely to be at home and are reached at the usual rate; the method therefore collects far fewer replies from full-time workers than their share of the adult population the council is consulting. Answer: full-time workers, as most are at work at that time. The distractors: the reply naming retired people rests on the false claim that most retired people are at work in the daytime, when in fact a daytime call reaches them more easily than anyone; the reply naming unemployed people rests on the false claim that they are out during the day, when they too are among the easiest people to reach by a daytime call; the reply naming children rests on the false claim that children are at home at 11 am on a Tuesday, when they are at school and so are not reached by the call at all, and school-age children are in any case not the adults whose views the council is collecting.
- (c) mean = 8, median = 8, mode = 8 — Method: work out each measure separately — the mean is the total divided by how many values there are, the median is the middle value once the data are in order, and the mode is the value that occurs most often. Working: the total is 8 + 8 + 8 + 8 = 32 and there are 4 marks, so the mean is 32 ÷ 4 = 8; in order the marks read 8, 8, 8, 8, and the mean of the middle pair is (8 + 8) ÷ 2 = 8; the value 8 occurs 4 times and no other value occurs at all, so the mode is 8. Answer: mean = 8, median = 8, mode = 8 — when every value in a data set is the same, all three measures of central tendency take that value. The distractors: a mean of 32 comes from stopping at the total and never dividing by 4; a mode of 4 comes from writing down how many times 8 occurs instead of the value that occurs; a mean of 2 comes from dividing a single value, 8, by the 4 marks instead of dividing the total by 4.
- (a) Negative correlation — As the age of the car increases, the points fall towards a lower value, so the value decreases as the age increases. This falling pattern is a negative correlation. A positive correlation would show the points rising together instead. No correlation would apply only if the points showed no pattern at all, and correlation is not the same as causation — strong causation is not a type of correlation.
- (a) £16,000 — Method: first find the quarter with the highest sales figure, then subtract quarter 1's sales from it — remembering that every figure is given in THOUSANDS of pounds. Working: the highest sales figure is quarter 2, at £34,000 (34 thousand pounds). The increase from quarter 1 is £34,000 − £18,000 = £16,000. Giving £34,000 reads off the highest sales figure on its own, without subtracting quarter 1's sales — that is the highest quarter's total, not the increase. Giving £12,000 uses quarter 3's sales, 30, the SECOND-highest figure, instead of quarter 2's 34, the actual highest — 30 − 18 = 12, but quarter 3 is not the quarter with the highest sales. Giving £16 gets the subtraction right, 34 − 18 = 16, but forgets that every figure in the question is in thousands of pounds, so the increase is £16,000, not £16. Always identify the correct quarter FIRST, and always check the units the numbers are given in before writing your final answer.
- (c) Height 170 cm, mass 65 kg — Method: a point on a scatter graph is written as a pair of coordinates in which the horizontal value is written first and the vertical value second, so each value is matched to the quantity named on its own axis. Working: in (170, 65) the value 170 is the horizontal coordinate and the horizontal axis shows height in centimetres, so the height is 170 cm; the value 65 is the vertical coordinate and the vertical axis shows mass in kilograms, so the mass is 65 kg. Answer: height 170 cm, mass 65 kg, each with the unit named on its own axis. The distractors: height 65 cm and mass 170 kg come from reading the pair the wrong way round, which would describe an impossible person; height 170 cm and mass 170 kg come from reading the horizontal coordinate for both quantities and never using the second number; height 235 cm and mass 105 kg come from combining the two coordinates, 170 + 65 and 170 − 65, instead of reading them separately.
- (b) A line graph — Sales recorded at the end of each of the twelve months are time series data, and a line graph is the chart built to show how a value changes over time, with the points usually joined in order. A pie chart is for showing categorical data as shares of a whole, not a trend over time. A pictogram shows a frequency for separate categories using symbols, not a continuous trend. A scatter graph is for comparing two different variables against each other, not one variable over time.
- (c) 20 kg ≤ mass < 30 kg — The modal class is the class with the highest frequency. Reading the plotted points, the frequencies are 6, 10, 16, 6 and 2, so the highest frequency is 16, plotted at the midpoint 25. A class of width 10 centred on 25 runs from 25 − 5 = 20 to 25 + 5 = 30, so the modal class is 20 kg ≤ mass < 30 kg. Writing '25 kg' gives only the midpoint, not the class — the modal class is an interval, not a single value. '10 kg ≤ mass < 20 kg' is the class before the peak, centred on 15, which has frequency 10, not the highest. '30 kg ≤ mass < 40 kg' is the class after the peak, centred on 35, which has frequency 6, not the highest.
Build your own mix at the worksheet builder.