Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Non-calculator
Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- 1.Priya wants to find out the favourite sport of the 900 pupils at her school. She asks 5 pupils, chosen at random. Give a reason why her sample may not give a reliable result.
- 2.A quality inspector weighs a random sample of 50 packets of crisps from one day's production and finds their mean mass is 32.4 g. The factory makes 20,000 packets that day. Work out an estimate for the total mass, in kg, of all the packets made that day.
- 3.A school posts a questionnaire about school dinners to the families of all 200 pupils. Only 30 families send theirs back, and 27 of those 30 say they are unhappy with school dinners. Give a reason why this result may not represent all 200 families.
- 4.A bar chart is drawn for a set of categorical data. Write down what the height of each bar represents.
- 5.A scatter graph of the number of hours, x, that pupils revised against their test score, y, has the line of best fit y = 2.5x + 15. Amelia wants a score of at least 80. Work out the least whole number of hours of revision the line of best fit suggests she needs.y = 2.5x + 15
- 6.A scatter graph shows the number of pages, x, in a chapter and the time, y minutes, a pupil took to read it, for 15 chapters. The line of best fit has equation y = 2x + 3. Use the line of best fit to estimate the time taken to read a chapter of 10 pages.y = 2x + 3
- 7.The mean of the four numbers 10, 15, 20 and x is 18. Work out the value of x.
- 8.A garden centre's sales, in thousands of pounds, at the end of each quarter last year were: quarter 1 — 18, quarter 2 — 34, quarter 3 — 30, quarter 4 — 22. Work out the increase in sales from quarter 1 to the quarter with the highest sales.
- 9.There are 10,000 pupils in a city. A researcher picks a sample of 500 of them by drawing names at random from a list of every pupil in the city. Give the reason why this method gives a representative sample.
- 10.A town council wants to find out about the eating habits of the people who live in the town. It asks only the members of a local sports club. Give a reason why this sample is biased.
- 11.Isla wants to find out the favourite television programme of the pupils at her school. She asks only her own close group of friends. Write down what is wrong with her sample.
- 12.A survey of 25 pupils in Derby records how many siblings each has: 0 siblings — 6 pupils, 1 sibling — 10 pupils, 2 siblings — 6 pupils, 3 siblings — 3 pupils. Calculate the mean number of siblings.
- 13.In a scatter graph of the age, in years, and the wingspan, in cm, of 20 birds of the same species, all the points lie close to a rising line of best fit except one, which lies a long way below the line. That bird was later found to have a damaged wing. Give a reason why this point should not be used when drawing the line of best fit.
- 14.The mean of the numbers x, 20 and 30 is equal to the mean of the numbers 15 and 25. Work out the value of x.
- 15.A frequency polygon for the mass, in kg, of 40 parcels at a delivery depot is drawn by plotting one point at the midpoint of each class, joined by straight lines: (5, 6), (15, 10), (25, 16), (35, 6), (45, 2). Every class has a width of 10 kg. Write down the modal class.
Answer key
- (a) 5 pupils are far too few to represent 900 pupils — Method: a sample can only support a claim about a population if it is chosen fairly and if it is large enough for the pattern in it to be more than chance. Working: Priya's method of choosing is fair, because the 5 pupils were picked at random, so every pupil had the same chance of being asked. The difficulty is the size: 900 ÷ 5 = 180, so each pupil she asks stands for 180 pupils. If two of the five happen to play in the same netball team, netball takes 40% of her sample on the strength of two answers, and a second sample of 5 could easily give a different favourite sport. Answer: 5 pupils are far too few to represent 900 pupils. The distractors: saying the pupils were not chosen at random contradicts the question, which states that they were; saying the 5 may each name a different sport describes what often happens in a small sample, but disagreement is not the fault, since 5 pupils who all named the same sport would be just as weak a basis for a claim about 900; saying a sample must hold at least half of the population is an invented rule, and a properly chosen sample of a few hundred can describe a population of many thousands.
- (d) 648 kg — Method: to estimate a total from a sample, multiply the sample's mean by the number of items in the whole population, then check the units the question asks for. Working: 32.4 g × 20,000 = 648,000 g. Converting to kilograms, 648,000 ÷ 1,000 = 648 kg. This is only an estimate, not an exact total, because it assumes every one of the 20,000 packets has exactly the sample mean mass, when in reality individual packets vary above and below it. Giving 1.62 kg multiplies the mean by 50, the SAMPLE size, instead of by 20,000, the number of packets actually made that day — this finds the total mass of the 50 sampled packets, not the day's production. Giving 32.4 kg treats the sample mean itself, in grams, as if it already were the day's total mass in kilograms, skipping the scaling up altogether. Giving 648,000 kg correctly scales the mean up to the whole day's production but never converts the answer from grams to kilograms, leaving it 1,000 times too large. Always scale a sample's mean up by the SIZE OF THE WHOLE POPULATION, and always finish by checking the units the question asks for.
- (d) Only families with strong feelings bothered to reply. — Method: a survey has non-response bias when only some of the people asked actually reply, and those who do are not a typical cross-section of everyone who was asked. Working: only 30 of the 200 families sent back their questionnaire, and 27 of those 30 — the great majority — said they were unhappy. Families who feel strongly about an issue, particularly those with a complaint, are far more likely to make the effort to reply than families who are simply satisfied and see no need to say anything, so the 30 replies over-represent unhappy families. Saying the families who replied were picked at random by the school gets the sampling the wrong way round: nobody picked them — they picked themselves by deciding to reply, and that is precisely why they are not a typical cross-section of all 200. Saying postal surveys always have low response rates restates that the response was low without explaining why a low response rate, on its own, makes a result unrepresentative — it is the reason FOR the low response, not the low response itself, that causes the bias here. Saying that the 27 unhappy replies show most families are unhappy is exactly the mistake the question is warning against: it treats the loudest 30 replies as if they stood for the other 170 who never sent theirs back. A low response rate is a warning sign only because the people who bother to reply are rarely typical of everyone who was asked.
- (c) The frequency of that category — Method: a bar chart for categorical data has one bar for each category, and the vertical scale on which the bars are measured is a count. Working: a bar drawn twice as tall as another tells you that twice as many items of data fell into its category, so the height measures how many items of data belong to that one category, which is exactly what a frequency is. Answer: the height of each bar is the frequency of that category — a count of items of data. The distractors: the number of different categories comes from reading the vertical scale as though it counted the bars, which is shown along the horizontal axis instead; the total of all the data comes from treating one bar as though it stood for the whole data set rather than for one category; the mean of all the data comes from confusing a bar chart with a measure of average, which no single bar can show.
- (a) 26 — Method: a line of best fit lets one quantity be predicted from the other, so the score is substituted into the equation of the line and the resulting inequality is solved for the number of hours. Working: a score of at least 80 means 2.5x + 15 ≥ 80; taking 15 from both sides gives 2.5x ≥ 65, and dividing both sides by 2.5 gives x ≥ 26, so the least whole number of hours is 26. Checking, 2.5 × 26 + 15 = 80, which does reach the target. Answer: 26 hours — and this is only an estimate, because a line of best fit predicts a trend rather than an individual result, and a prediction made outside the range of hours the pupils actually revised for would be an extrapolation and less reliable still. The distractors: 27 comes from reaching 26 and then rounding up again, although 26 hours already gives a score of exactly 80; 32 comes from 80 ÷ 2.5, which ignores the 15 in the equation of the line; 38 comes from (80 + 15) ÷ 2.5, that is from adding the 15 instead of subtracting it when rearranging.
- (a) 23 minutes — Method: the equation of a line of best fit converts a value of x into a predicted value of y, so substitute the known number of pages for x and evaluate. Working: x is the number of pages, so put x = 10 into y = 2x + 3. Multiplication is carried out before addition, so 2 × 10 + 3 = 23. Answer: 23 minutes, and it is a prediction of the trend rather than a promise about any one chapter. The distractors: 20 minutes comes from working out 2 × 10 and stopping there, leaving out the 3 that the line adds; 26 minutes comes from reading the equation as y = 2(x + 3), adding first and then doubling, so 2 × 13 = 26; 13 minutes comes from adding 10 and 3 and never using the gradient at all, which treats the 2 as though it were not there.
- (a) 27 — Method: turn the mean into a total using total = mean × number of values, then subtract the numbers that are already known. Working: four numbers with a mean of 18 have a total of 18 × 4 = 72; the three known numbers give 10 + 15 + 20 = 45; so x = 72 − 45 = 27. Answer: 27, and checking, (10 + 15 + 20 + 27) ÷ 4 = 72 ÷ 4 = 18. The distractors: 72 comes from stopping at the total the four numbers must reach and never subtracting the known three; 18 comes from assuming the missing number must equal the mean; 45 comes from stopping at the total of the three known numbers.
- (a) £16,000 — Method: first find the quarter with the highest sales figure, then subtract quarter 1's sales from it — remembering that every figure is given in THOUSANDS of pounds. Working: the highest sales figure is quarter 2, at £34,000 (34 thousand pounds). The increase from quarter 1 is £34,000 − £18,000 = £16,000. Giving £34,000 reads off the highest sales figure on its own, without subtracting quarter 1's sales — that is the highest quarter's total, not the increase. Giving £12,000 uses quarter 3's sales, 30, the SECOND-highest figure, instead of quarter 2's 34, the actual highest — 30 − 18 = 12, but quarter 3 is not the quarter with the highest sales. Giving £16 gets the subtraction right, 34 − 18 = 16, but forgets that every figure in the question is in thousands of pounds, so the increase is £16,000, not £16. Always identify the correct quarter FIRST, and always check the units the numbers are given in before writing your final answer.
- (c) Because every pupil has an equal chance of being picked — Method: whether a sample represents its population is decided by the selection method, not by the size of the sample, so ask whether the method gives every member of the population the same chance of being chosen. Working: the names are drawn at random from a list of all 10,000 pupils, so each pupil has the same chance, 500 out of 10,000, of being drawn, and no group of pupils is more likely to appear than any other; that is what keeps bias out of the sample. Answer: because every pupil has an equal chance of being picked. The distractors: the reply about 5% treats the sampling fraction as the test of fairness, but a badly chosen 5% is still biased and a well chosen 1% is not; the reply about 500 being large enough makes size the test instead, which is the same mistake in another form, since a large sample drawn from one school would still misrepresent the city; the reply about the most willing pupils describes self-selection, which hands the choice of who is in the sample to the pupils who feel most strongly about the question.
- (c) Club members probably eat differently from most people — Method: a sample is biased when the group it is drawn from differs from the population in the very thing the survey is measuring, so compare the subgroup with the population on that quantity. Working: the survey measures eating habits, and people who join a sports club take more exercise than average and are known to eat differently from the town as a whole, so their replies pull the results away from the true picture for the town however many of them are asked. Answer: club members probably eat differently from most people. The distractors: the reply about the number of members treats bias as a question of size, but a large biased sample is still biased; the reply that the members were picked at random is false, since the council picked a club rather than picking residents, and it confuses bias with non-response; the reply that everyone asked lives in the town notes something true of the members but draws the false conclusion that the sample therefore covers the town, when a sample must reflect a population and not merely be taken from inside it.
- (c) It is not representative, as she picked her own friends — Method: judge a sample by asking whether the pupils in it were chosen in a way that gives the whole school a fair chance of being heard. Working: Isla's friends are a group she formed herself, and friends tend to share tastes, so their favourite programme is likely to match hers rather than the school's, and pupils in other year groups and other friendship groups had no chance at all of being asked; the fault lies in how the pupils were selected, not in how many of them there were. Answer: it is not representative, as she picked her own friends. The distractors: the reply blaming the size claims the pupils were picked at random, which is false here, and it is the common mistake of thinking a biased sample can be cured by making it bigger; the reply calling the sample too large is false in the other direction, as a survey is never spoilt by collecting more replies; the reply that the sample is fine treats attending the school as enough, which would make any group of pupils in the building a fair sample.
- (a) 1.24 — Method: for data given as a frequency table, the mean is Σfx ÷ Σf — multiply each value by its frequency, add the results, then divide by the total frequency. Working: 0 × 6 = 0. 1 × 10 = 10. 2 × 6 = 12. 3 × 3 = 9. So Σfx = 0 + 10 + 12 + 9 = 31. The total frequency is Σf = 6 + 10 + 6 + 3 = 25. Mean = 31 ÷ 25 = 1.24 siblings. Averaging the frequency column itself, (6 + 10 + 6 + 3) ÷ 4 = 6.25, mixes up the frequencies with the values they belong to. Writing down 1, the number of siblings with the highest frequency, gives the mode, not the mean. Writing down 31 stops after finding Σfx and forgets to divide by the total frequency, 25. Always divide Σfx by Σf — never stop at the top of the fraction.
- (d) An outlier from the damaged wing, not the trend. — That bird's point lies a long way from the rising trend followed by every other bird, and its low wingspan is explained by the damaged wing rather than by its age — it is an outlier caused by an unusual factor, not part of the general relationship between age and wingspan, so it should not be used when drawing the line of best fit. Saying every point must be used ignores that an outlier caused by a separate, identifiable factor can rightly be set aside. Saying it shows no correlation ignores that the other 19 points do show a clear rising trend; one outlier does not remove that. Saying it proves the line is inaccurate confuses one unusual bird with a fault in the line itself, when the line correctly describes the trend followed by the rest of the data.
- (c) 10 — Method: work out the mean that can be found straight away, then use total = mean × number of values on the group of three to find the missing number. Working: the mean of 15 and 25 is (15 + 25) ÷ 2 = 40 ÷ 2 = 20, so the group of three must also have a mean of 20; three numbers with a mean of 20 have a total of 20 × 3 = 60, and 20 + 30 = 50 of that total is already accounted for, so x = 60 − 50 = 10. Answer: 10, and checking, (10 + 20 + 30) ÷ 3 = 20. The distractors: 20 comes from working out the mean the two groups share and writing that down as x; −10 comes from dividing the group of three by 2 instead of by 3, which gives x + 50 = 40; 70 comes from reading the total 15 + 25 = 40 as the mean of the pair, which sets the target total at 120 and leaves x = 70.
- (c) 20 kg ≤ mass < 30 kg — The modal class is the class with the highest frequency. Reading the plotted points, the frequencies are 6, 10, 16, 6 and 2, so the highest frequency is 16, plotted at the midpoint 25. A class of width 10 centred on 25 runs from 25 − 5 = 20 to 25 + 5 = 30, so the modal class is 20 kg ≤ mass < 30 kg. Writing '25 kg' gives only the midpoint, not the class — the modal class is an interval, not a single value. '10 kg ≤ mass < 20 kg' is the class before the peak, centred on 15, which has frequency 10, not the highest. '30 kg ≤ mass < 40 kg' is the class after the peak, centred on 35, which has frequency 6, not the highest.
Build your own mix at the worksheet builder.