Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- 1.A scatter graph shows the number of guests, x, at a wedding and the length of buffet table needed, y metres. The line of best fit is y = 0.5x + 2. Write down what the 2 in this equation tells you about the buffet table.y = 0.5x + 2
- 2.A scatter graph has 50 points. Most of them lie close to a rising line of best fit, but two of them lie a long way from that line. Write down how those two points should be treated.
- 3.A scatter graph of taxi journeys in Bristol shows the distance, x miles, and the fare, y pounds. The line of best fit is y = 2x + 3.50. Work out the estimated fare for a journey of 6 miles, using the line of best fit.y = 2x + 3.5
- 4.A school has 1200 pupils. A teacher wants to take a random sample of 60 of them. Write down which of these methods gives a random sample.
- 5.A garden centre's sales, in thousands of pounds, at the end of each quarter last year were: quarter 1 — 18, quarter 2 — 34, quarter 3 — 30, quarter 4 — 22. Work out the increase in sales from quarter 1 to the quarter with the highest sales.
- 6.On a scatter graph the horizontal axis shows height in centimetres and the vertical axis shows mass in kilograms. One point is plotted at (170, 65). Write down the height and the mass of that person.
- 7.A vet records the masses, m kg, of 30 dogs at a clinic in Preston: 0 < m ≤ 10 — 11 dogs, 10 < m ≤ 20 — 5 dogs, 20 < m ≤ 30 — 5 dogs, 30 < m ≤ 40 — 9 dogs. Work out an estimate for the mean mass, in kg, using the midpoint of each class interval.
- 8.A school has 1,000 pupils. A student wants to estimate how many of them walk to school, so she asks 8 pupils in her own class. She says her sample is large enough to give a reliable estimate for the whole school. Is she right? Give a reason for your answer.
- 9.A garden centre records its total sales, in pounds, at the end of each of the twelve months of one year, to see how sales change over time. Write down the most suitable type of chart to show this time series data.
- 10.A council in Leeds wants to know what local people think about letting shops stay open later in the evening. It rings landline telephone numbers between 10 am and 2 pm on a Tuesday. Write down which group is most likely to be under-represented in the sample, and give a reason for your answer.
- 11.A stem-and-leaf diagram, described in words, shows the ages of 9 people at a family party. The stem is the tens digit: stem 1 has leaves 4 and 8; stem 2 has leaves 0, 3, 5 and 9; stem 3 has leaves 1 and 6; stem 4 has leaf 2. Work out the median age.
- 12.Seven pupils were asked how many books they had read last month. Their answers were 3, 5, 5, 7, 8, 5, 3. Write down the mode.
- 13.A dual bar chart shows the number of hours of rain recorded in Leeds and in Bristol on each of four days. Leeds: Monday 3 hours, Tuesday 5 hours, Wednesday 2 hours, Thursday 4 hours. Bristol: Monday 4 hours, Tuesday 4 hours, Wednesday 6 hours, Thursday 2 hours. Work out the greatest amount, in hours, by which Bristol's rainfall exceeded Leeds's rainfall on a single day.
- 14.A two-way table records the favourite subject, Maths or Art, of 60 pupils in Year 10, and whether each pupil is left-handed or right-handed. 9 of the 60 pupils are left-handed, and 6 of those left-handed pupils prefer Art. In total, 24 of the 60 pupils prefer Art. A pupil is chosen at random from the 60. Work out the probability that the pupil is right-handed and prefers Art.
- 15.A survey of 25 pupils in Derby records how many siblings each has: 0 siblings — 6 pupils, 1 sibling — 10 pupils, 2 siblings — 6 pupils, 3 siblings — 3 pupils. Calculate the mean number of siblings.
Answer key
- (b) At 0 guests, the model predicts 2 m of table — The y-intercept of a line of best fit y = mx + c is the value of y when x = 0. Here y = 0.5 × 0 + 2 = 2, so the line predicts a table length of 2 m when there are 0 guests. The 2 m does not grow as more guests arrive — that role belongs to the gradient, 0.5 — so an option saying each extra guest adds 2 m has swapped the two numbers around. The 2 is a length in metres, not a number of guests, so an option requiring 2 guests before set-up has misread its units. And the table length does change with x, since it is 0.5x + 2 and not a fixed value, so an option claiming the table is always 2 m ignores the 0.5x term completely.
- (b) Treat them as outliers and check them before deciding — Method: a point lying a long way from the pattern the rest of the data make is called an outlier, and an outlier is investigated before anything is done with it, because it may be an error in the data or it may be a genuine but unusual case. Working: 48 of the 50 points lie close to the rising line of best fit, so the trend is set by those 48; the two remaining points do not follow it, so they are identified as outliers and checked — a mistake in measuring or recording would be corrected, while a genuine reading would be kept and reported. Answer: treat them as outliers and check them before deciding what to do with them. The distractors: deleting them at once assumes that every point far from the line must be an error, which throws away real data; moving the line so that it passes through them assumes a line of best fit must touch particular points, when it is drawn to follow all 50; taking them as proof that there is no correlation lets two points overturn the pattern that the other 48 agree on.
- (d) £15.50 — 2 × 6 = 12, and 12 + 3.50 = 15.50, so the estimated fare is £15.50. Choosing £12.00 stops after 2 × 6 = 12 and forgets to add the £3.50. Choosing £19.00 adds the distance and the constant first and then multiplies: 6 + 3.50 = 9.50, and 9.50 × 2 = 19.00, applying the ×2 to the whole sum instead of only to the distance. Choosing £13.00 multiplies only the constant term by 2 instead of the distance: 2 × 3.50 = 7, and 7 + 6 = 13.00.
- (a) Drawing 60 names at random from a list of all 1200 pupils — Method: a sample is random when every member of the population has the same chance of being chosen and nobody, including the pupils themselves, can influence who ends up in it; test each method against that. Working: drawing names from a list of all 1200 pupils gives each pupil the same chance, 60 out of 1200, whatever their year group, class or opinion, so the method is random. Answer: drawing 60 names at random from a list of all 1200 pupils. The distractors: asking the pupils who volunteer is self-selection, and the pupils with the strongest views volunteer first, so they decide the sample; asking the pupils nearest the door is convenience sampling, which reaches only those who happen to be in one place at one time; asking two Year 10 classes samples a cluster, so every pupil in the other year groups has no chance of being chosen at all.
- (a) £16,000 — Method: first find the quarter with the highest sales figure, then subtract quarter 1's sales from it — remembering that every figure is given in THOUSANDS of pounds. Working: the highest sales figure is quarter 2, at £34,000 (34 thousand pounds). The increase from quarter 1 is £34,000 − £18,000 = £16,000. Giving £34,000 reads off the highest sales figure on its own, without subtracting quarter 1's sales — that is the highest quarter's total, not the increase. Giving £12,000 uses quarter 3's sales, 30, the SECOND-highest figure, instead of quarter 2's 34, the actual highest — 30 − 18 = 12, but quarter 3 is not the quarter with the highest sales. Giving £16 gets the subtraction right, 34 − 18 = 16, but forgets that every figure in the question is in thousands of pounds, so the increase is £16,000, not £16. Always identify the correct quarter FIRST, and always check the units the numbers are given in before writing your final answer.
- (c) Height 170 cm, mass 65 kg — Method: a point on a scatter graph is written as a pair of coordinates in which the horizontal value is written first and the vertical value second, so each value is matched to the quantity named on its own axis. Working: in (170, 65) the value 170 is the horizontal coordinate and the horizontal axis shows height in centimetres, so the height is 170 cm; the value 65 is the vertical coordinate and the vertical axis shows mass in kilograms, so the mass is 65 kg. Answer: height 170 cm, mass 65 kg, each with the unit named on its own axis. The distractors: height 65 cm and mass 170 kg come from reading the pair the wrong way round, which would describe an impossible person; height 170 cm and mass 170 kg come from reading the horizontal coordinate for both quantities and never using the second number; height 235 cm and mass 105 kg come from combining the two coordinates, 170 + 65 and 170 − 65, instead of reading them separately.
- (d) 19 kg — Method: estimate the mean of grouped data by multiplying each class's midpoint by its frequency, adding the four totals, then dividing by the total frequency. Working: the midpoints are 5, 15, 25 and 35 kg. The weighted totals are 11 × 5 = 55, 5 × 15 = 75, 5 × 25 = 125 and 9 × 35 = 315, which add to 570. Dividing by the 30 dogs gives an estimate of 570 ÷ 30 = 19 kg. Giving 5 kg reads off the midpoint of the modal class, 0 < m ≤ 10, the class with the most dogs — but the class with the most dogs is not where the mean falls, and neither is a substitute for actually calculating it. Giving 20 kg averages the four midpoints, (5 + 15 + 25 + 35) ÷ 4, treating every class as equally likely and ignoring that far more dogs are in the lightest and heaviest classes than in the middle two. Giving 570 kg stops after finding the correct weighted total and forgets the final division by the 30 dogs. Always weight each midpoint by its own frequency, and always finish by dividing by the total frequency, not the number of classes.
- (c) No — 8 from one class is too small to represent the school. — Method: judge reliability by asking whether the sample is both large enough, and spread across the population, relative to what it is meant to represent. Working: 8 pupils is a tiny fraction of the school's 1,000 pupils, and all 8 come from a single class rather than a range of year groups, so the sample is both too small and too narrow to represent the whole school reliably. She is not right. Saying any sample size gives an equally reliable estimate ignores that reliability generally improves with a larger, more representative sample. Saying the method is unreliable because it was not done online is not a reason connected to sample size or representativeness at all. Saying 8 is reliable because it is more than half her class compares the sample to the wrong population — the school has 1,000 pupils, not one class. Always judge a sample's size against the population it is meant to represent, not against a smaller group within it.
- (b) A line graph — Sales recorded at the end of each of the twelve months are time series data, and a line graph is the chart built to show how a value changes over time, with the points usually joined in order. A pie chart is for showing categorical data as shares of a whole, not a trend over time. A pictogram shows a frequency for separate categories using symbols, not a continuous trend. A scatter graph is for comparing two different variables against each other, not one variable over time.
- (d) Full-time workers, as most are at work at that time — Method: a sample is biased when the method of contact makes part of the population much less likely to be reached, so test each group against where its members actually are between 10 am and 2 pm on a weekday, and test each stated reason against the facts. Working: those hours are the middle of the working day, so people in full-time employment are at work and not beside a landline telephone, while people who are retired and people who are unemployed are far more likely to be at home and are reached at the usual rate; the method therefore collects far fewer replies from full-time workers than their share of the adult population the council is consulting. Answer: full-time workers, as most are at work at that time. The distractors: the reply naming retired people rests on the false claim that most retired people are at work in the daytime, when in fact a daytime call reaches them more easily than anyone; the reply naming unemployed people rests on the false claim that they are out during the day, when they too are among the easiest people to reach by a daytime call; the reply naming children rests on the false claim that children are at home at 11 am on a Tuesday, when they are at school and so are not reached by the call at all, and school-age children are in any case not the adults whose views the council is collecting.
- (d) 25 — In order, the nine ages are 14, 18, 20, 23, 25, 29, 31, 36 and 42, and with 9 values the median is the 5th one, which is 25. Choosing 23 takes the 4th value instead of the 5th. Choosing 29 takes the 6th value instead of the 5th. Choosing 26 comes from averaging the 4th and 6th values, 23 + 29 = 52, and 52 ÷ 2 = 26, a method that is only needed when there is an even number of values.
- (c) 5 — Method: the mode is the value that occurs most often, so count how many times each different value appears and compare the counts. Working: 3 appears twice, 5 appears three times, 7 appears once and 8 appears once, so the highest frequency is three and the value carrying it is 5. Answer: 5. The distractors: 3 comes from writing down the frequency of the most common answer instead of the answer itself; 7 comes from taking the middle number of the list as it was written, which applies the median without ordering the data and without answering the question asked; 8 comes from picking the largest value, which confuses the mode with the maximum.
- (b) 4 — The difference, Bristol minus Leeds, on each day is: Monday 4 − 3 = 1, Tuesday 4 − 5 = −1, Wednesday 6 − 2 = 4, Thursday 2 − 4 = −2. The greatest amount by which Bristol exceeded Leeds is 4 hours, on Wednesday. Choosing 1 takes Monday's smaller positive difference instead of the greatest one. Choosing 2 takes the size of Thursday's difference, but that is the amount by which Leeds exceeded Bristol, the opposite direction to the one asked for. Choosing 6 takes Bristol's raw figure on Wednesday without subtracting Leeds's 2 hours first.
- (a) 3/10 — There are 60 − 9 = 51 right-handed pupils. Of the 24 pupils who prefer Art, 6 are left-handed, so 24 − 6 = 18 are right-handed and prefer Art. The probability that a randomly chosen pupil is right-handed and prefers Art is 18/60, which simplifies to 3/10. Giving 2/5 is 24/60 simplified — the probability of preferring Art, ignoring the right-handed condition entirely. Giving 17/20 is 51/60 simplified — the probability of being right-handed, ignoring the Art condition entirely. Giving 1/10 is 6/60 simplified — the probability of being left-handed and preferring Art, the wrong hand condition.
- (a) 1.24 — Method: for data given as a frequency table, the mean is Σfx ÷ Σf — multiply each value by its frequency, add the results, then divide by the total frequency. Working: 0 × 6 = 0. 1 × 10 = 10. 2 × 6 = 12. 3 × 3 = 9. So Σfx = 0 + 10 + 12 + 9 = 31. The total frequency is Σf = 6 + 10 + 6 + 3 = 25. Mean = 31 ÷ 25 = 1.24 siblings. Averaging the frequency column itself, (6 + 10 + 6 + 3) ÷ 4 = 6.25, mixes up the frequencies with the values they belong to. Writing down 1, the number of siblings with the highest frequency, gives the mode, not the mean. Writing down 31 stops after finding Σfx and forgets to divide by the total frequency, 25. Always divide Σfx by Σf — never stop at the top of the fraction.
Build your own mix at the worksheet builder.