Printable · GCSE Foundation · ages 14-16
Statistics worksheet — GCSE Foundation
Fifteen questions across the statistics statements at Foundation tier. Choose the non-calculator filter to rehearse Paper 1, which counts for a third of the marks.
Statistics worksheet — GCSE Foundation
MathsUKwww.geekhero.co.uk
- 1.A two-way table records the favourite subject, Maths or Art, of 60 pupils in Year 10, and whether each pupil is left-handed or right-handed. 9 of the 60 pupils are left-handed, and 6 of those left-handed pupils prefer Art. In total, 24 of the 60 pupils prefer Art. A pupil is chosen at random from the 60. Work out the probability that the pupil is right-handed and prefers Art.
- 2.A study found that people who drink more coffee tend to concentrate better at work. A coffee company says that this shows that drinking coffee improves concentration. Give the reason why this conclusion cannot be drawn.
- 3.A charity shop in Leicester made £2,400 in one month. A pie chart shows how this was spent: repairs took up an angle of 90°, wages took up an angle of 150°, and the rest was other costs. Work out how much was spent on wages.
- 4.A shop sold seven pairs of shoes in these sizes: 4, 4, 5, 6, 6, 6, 9. Write down the modal size.
- 5.On a scatter graph of the arm span and the height of some pupils, the points rise from left to right. Write down the type of correlation this shows.
- 6.A factory makes 4,000 light bulbs a day. In a random sample of 80 of one day's bulbs, 3 were faulty. Work out an estimate for the number of faulty bulbs the factory makes in a day.
- 7.Five pupils spent these numbers of minutes on their homework: 50, 65, 55, 90, 60. Work out the median time.
- 8.A vet records the masses, m kg, of 30 dogs at a clinic in Preston: 0 < m ≤ 10 — 11 dogs, 10 < m ≤ 20 — 5 dogs, 20 < m ≤ 30 — 5 dogs, 30 < m ≤ 40 — 9 dogs. Work out an estimate for the mean mass, in kg, using the midpoint of each class interval.
- 9.A company found that the higher the price it charged for a product, the lower the satisfaction score its customers gave. The price and the satisfaction score are plotted on a scatter graph. Describe the line of best fit that would be drawn on that graph.
- 10.A scatter graph has 12 plotted points. Four pupils each draw a line of best fit on the same graph and count how many of the 12 points lie above their line and how many lie below it: Amir — 10 above, 2 below. Priya — 6 above, 6 below. Kofi — 2 above, 10 below. Leah — 0 above, 12 below. Write down the name of the pupil whose line was drawn correctly, so that the points are roughly balanced above and below it.
- 11.A straight line is drawn on a scatter graph to show the trend of the points. Write down the name given to this line.
- 12.A shop recorded the number of books it sold on five days: 100, 40, 70, 20, 60. Work out the range of the numbers of books sold.
- 13.The mean of 5 numbers is 8. Work out the total of the 5 numbers.
- 14.A scatter graph plots the number of years, x, that 20 employees have worked at a company against their salary, y. All the plotted points lie between x = 1 and x = 15. Write down the word used to describe an estimate for y made using a value of x that lies between 1 and 15.
- 15.A recruitment agency in Manchester compares the weekly pay, in pounds, of seven employees at two branches. Branch A: £480, £495, £500, £505, £510, £515, £1,200. Branch B: £480, £490, £500, £510, £520, £530, £540. An advert for the agency claims 'Branch A pays more on average.' Decide whether this claim is fairly supported, using an appropriate average, and choose the correct conclusion.
Answer key
- (a) 3/10 — There are 60 − 9 = 51 right-handed pupils. Of the 24 pupils who prefer Art, 6 are left-handed, so 24 − 6 = 18 are right-handed and prefer Art. The probability that a randomly chosen pupil is right-handed and prefers Art is 18/60, which simplifies to 3/10. Giving 2/5 is 24/60 simplified — the probability of preferring Art, ignoring the right-handed condition entirely. Giving 17/20 is 51/60 simplified — the probability of being right-handed, ignoring the Art condition entirely. Giving 1/10 is 6/60 simplified — the probability of being left-handed and preferring Art, the wrong hand condition.
- (c) The data show a link only; a third factor may affect both — Method: a study of this kind measures two quantities and reports how they change together; deciding that one of them produces the other is a further claim, and it needs evidence that the measurements alone cannot give. Working: the study shows that more coffee goes with better concentration, which is a positive correlation; but a third factor that was never measured, such as how motivated someone is, could raise both the coffee drinking and the concentration, and the concentration could equally be what leads to the extra coffee. Answer: the data show a link only, because a third factor may be affecting both quantities, so no claim about cause can be made. The distractors: calling the conclusion safe because the correlation is positive treats the direction of a correlation as proof of cause, which no direction can give; calling it wrong because the correlation is negative misreads the direction of the relationship, since the study reports both quantities rising together; saying the two quantities are not linked denies the correlation the study actually found, when what fails is only the claim about cause.
- (b) £1,000 — Wages take up 150° out of 360°, so the amount spent on wages is 150 ÷ 360 × 2400 = £1,000. Choosing £600 uses the repairs angle, 90°, instead of the wages angle: 90 ÷ 360 × 2400 = 600. Choosing £3,600 treats the angle in degrees as if it were a percentage, 150 ÷ 100 × 2400 = 3600, instead of dividing by 360°. Choosing £800 uses the angle for the 'other costs' sector, 360 − 90 − 150 = 120°, instead of the wages sector: 120 ÷ 360 × 2400 = 800.
- (d) 6 — Method: the mode, or modal value, is the value that occurs most often in the data set, and it is a value from the data rather than a count. Working: size 4 occurs twice, size 5 occurs once, size 6 occurs three times and size 9 occurs once, so the highest frequency is three and the size it belongs to is 6. Answer: 6. The distractors: 3 comes from writing down the frequency of the most common size instead of the size itself; 9 comes from picking the largest size in the list, which confuses the mode with the maximum; 4 comes from stopping at the first size that repeats rather than checking which size repeats most often.
- (b) Positive correlation — Method: the type of correlation is named from the direction the points take as the scatter graph is read from left to right. Working: the points rise from left to right, so as the arm span read on the horizontal axis increases, the height read on the vertical axis increases as well; two quantities that increase together show positive correlation. Answer: positive correlation. The distractors: negative correlation comes from naming the direction the wrong way round, since a negative correlation needs the points to fall as the graph is read from left to right; no correlation comes from treating points that are spread out rather than sitting exactly on a line as though they showed no relationship; direct proportion comes from confusing a rising trend with proportion, which would additionally need the line through the points to pass through the origin and would mean doubling one quantity doubles the other.
- (a) 150 bulbs — Method: assume the proportion faulty in a random sample is the proportion faulty in the whole day's output, and scale the sample up to the population. Working: the sample of 80 has to be scaled up to 4,000 bulbs, and 4,000 ÷ 80 = 50, so the day's output is 50 sample-sized batches. Each batch is expected to contain the same 3 faulty bulbs, so the estimate is 3 × 50 = 150. Answer: 150 bulbs, and it is an estimate, because another sample of 80 would probably contain a different number of faulty bulbs. The distractors: 50 bulbs is the scale factor 4,000 ÷ 80 written down as though it were the answer, so it reports how many batches there are rather than how many faulty bulbs; 120 bulbs comes from reading 3 out of 80 as 3%, then taking 0.03 × 4,000 = 120, but 3 out of 80 is 3.75%; 240 bulbs comes from 3 × 80 = 240, multiplying the faulty bulbs by the size of the sample instead of by the scale factor, which uses the 80 twice and the 4,000 not at all.
- (c) 60 minutes — Method: the median is the middle value once the data have been put in order of size, so the list must be sorted before any position is read. Working: in order the times are 50, 55, 60, 65, 90 minutes; there are 5 values, so the middle position is the third and the time sitting there is 60 minutes. Answer: 60 minutes. The distractors: 64 minutes comes from working out the mean, 320 ÷ 5, instead of the median; 70 minutes comes from taking the time halfway between the shortest and the longest, (50 + 90) ÷ 2; 40 minutes comes from working out the range, 90 − 50, which measures spread rather than centre.
- (d) 19 kg — Method: estimate the mean of grouped data by multiplying each class's midpoint by its frequency, adding the four totals, then dividing by the total frequency. Working: the midpoints are 5, 15, 25 and 35 kg. The weighted totals are 11 × 5 = 55, 5 × 15 = 75, 5 × 25 = 125 and 9 × 35 = 315, which add to 570. Dividing by the 30 dogs gives an estimate of 570 ÷ 30 = 19 kg. Giving 5 kg reads off the midpoint of the modal class, 0 < m ≤ 10, the class with the most dogs — but the class with the most dogs is not where the mean falls, and neither is a substitute for actually calculating it. Giving 20 kg averages the four midpoints, (5 + 15 + 25 + 35) ÷ 4, treating every class as equally likely and ignoring that far more dogs are in the lightest and heaviest classes than in the middle two. Giving 570 kg stops after finding the correct weighted total and forgets the final division by the 30 dogs. Always weight each midpoint by its own frequency, and always finish by dividing by the total frequency, not the number of classes.
- (a) A straight line sloping down from left to right — Method: a line of best fit is a straight line drawn to follow the trend of the points, so its slope is decided by the direction of the relationship between the two quantities. Working: as the price rises, the satisfaction score falls, so the points start high on the left of the graph and finish low on the right; the straight line that follows them therefore slopes downwards as the graph is read from left to right, which is the line of a negative correlation. Answer: a straight line sloping down from left to right. The distractors: a line sloping up comes from reading a falling relationship as a rising one; a horizontal line comes from expecting no correlation, since a horizontal line says the satisfaction score does not change as the price changes; a curve passing through every point comes from thinking a line of best fit has to touch all of the plotted points, when it is a single straight line drawn through the middle of them.
- (a) Priya — A line of best fit should be drawn so that the plotted points are roughly balanced above and below it. Work out the difference between the two counts for each pupil: Amir 10 − 2 = 8; Kofi 10 − 2 = 8 (10 below and 2 above); Leah 12 − 0 = 12; Priya 6 − 6 = 0. Priya's line has the smallest difference, an exact balance of 6 above and 6 below, so her line is drawn correctly. Amir's line has 10 of the 12 points above it, so it is drawn too low. Kofi's line has 10 of the 12 points below it, so it is drawn too high. Leah's line has every single point below it, so it is not a line of best fit at all.
- (c) A line of best fit — Method: the straight line drawn on a scatter graph is named from the job it does — it is chosen so that it follows the whole set of points as closely as possible. Working: the line passes through the middle of the points, with roughly as many points above it as below it, and it need not pass through any of the plotted points at all; the name given to the straight line chosen in that way is a line of best fit. Answer: a line of best fit. The distractors: a line of symmetry comes from confusing a trend with symmetry, which is a property of a shape rather than of a set of data; a horizontal line through the mean comes from thinking the trend is shown by an average, when a horizontal line would say that the vertical quantity does not change and so show no correlation; a line joining the first and last points comes from thinking the line must join the two extreme points, which lets two points decide a trend that all of the points should share in.
- (c) 80 — Method: the range is a measure of spread and is found by subtracting the smallest value from the largest. Working: the largest number sold is 100 and the smallest is 20, so the range is 100 − 20 = 80. Answer: 80. The distractors: 100 comes from writing down the largest value and never subtracting the smallest; 60 comes from working out the median, the middle value of 20, 40, 60, 70, 100, instead of the range; 58 comes from working out the mean, 290 ÷ 5, which measures centre rather than spread.
- (a) 40 — Method: the mean is the total divided by how many values there are, so rearranging gives total = mean × number of values. Working: the mean is 8 and there are 5 numbers, so the total is 8 × 5 = 40. Answer: 40, and checking, 40 ÷ 5 = 8, which is the mean given. The distractors: 13 comes from adding the mean and the count, 8 + 5, instead of multiplying them; 1.6 comes from dividing the mean by the count, 8 ÷ 5, which reverses the relationship; 8 comes from quoting the mean itself as the total, which is only true when there is a single number.
- (b) Interpolation — The salary is being estimated for a value of x between 1 and 15, which is inside the range of x-values that were actually plotted, so this is interpolation. Extrapolation would apply if the estimate used a value of x below 1 or above 15, outside the plotted range. Correlation describes the relationship between the two variables, not the reliability of an estimate, and causation describes one variable actually causing a change in the other, which is a different idea altogether — neither is the word being asked for here.
- (c) No — median £505 at A vs £510 at B. — Branch A's seven wages in order are £480, £495, £500, £505, £510, £515 and £1,200, so the median, the 4th value, is £505. Branch B's in order are £480, £490, £500, £510, £520, £530 and £540, so the median is £510. Since £505 is lower than £510, the median wage is not higher at Branch A, so the claim is not fairly supported. Choosing 'Yes — mean £600.71 at A vs £510 at B' uses the mean: 480 + 495 + 500 + 505 + 510 + 515 + 1200 = 4205, and 4205 ÷ 7 = 600.71, a figure pulled upward by the £1,200 outlier that does not represent a typical wage. Choosing 'Yes — median £515 at A vs £510 at B' miscounts the middle position, taking the 6th wage, £515, instead of the correct 4th value, £505. Choosing 'Yes — highest wage £1,200 at A vs £540 at B' compares the highest wage at each branch rather than a measure of the typical, or average, wage.
Build your own mix at the worksheet builder.