Printable · GCSE Higher · ages 14-16
Histograms and cumulative frequency graphs worksheet — GCSE Higher
Fifteen questions on "histograms and cumulative frequency graphs" — DfE statement S3. Print it, or print three versions so neighbours cannot copy by letter; the key gives the letter for each version.
Higher only
Answer key: Histograms and cumulative frequency graphs worksheet — GCSE Higher
MathsUKwww.geekhero.co.uk
- (c) 145 g — Method: find the position of the median from the total frequency, locate the class that contains it, then use linear interpolation inside that class, assuming the apples in it are spread evenly. Working: the median is the 100 ÷ 2 = 50th apple; the running totals are 10, then 10 + 30 = 40, then 40 + 40 = 80, so the 50th apple lies in the class 140 ≤ m < 160; it is the 50 − 40 = 10th of the 40 apples in that class, and the class is 20 g wide, so the median is 140 + (10 ÷ 40) × 20 = 140 + 5 = 145. Answer: an estimated median of 145 g. The distractors: 150 g comes from giving the midpoint of the class that contains the median instead of interpolating inside it; 140 g comes from stopping at the lower boundary of that class, which locates the class but not the value; 155 g comes from measuring the 5 g step down from the upper boundary, 160 − 5, instead of up from the lower boundary.
- (d) 4, 3, 2, 6 — Method: frequency density = frequency ÷ class width for each class in turn; do not assume the classes are all the same width. Working: the four classes have widths 10 − 0 = 10, 30 − 10 = 20, 45 − 30 = 15 and 50 − 45 = 5. Dividing each frequency by its own width gives 40 ÷ 10 = 4, 60 ÷ 20 = 3, 30 ÷ 15 = 2 and 30 ÷ 5 = 6. Answer: the frequency densities, in order, are 4, 3, 2 and 6. Watch the width of each class separately: treating the last class as if it were also 10 units wide, like the first, gives 30 ÷ 10 = 3 instead of 30 ÷ 5 = 6 — the classes here are deliberately unequal, so no width can be borrowed from another class; dividing the width by the frequency instead of the frequency by the width for the third class gives 15 ÷ 30 = 0.5 in place of 2, the formula the wrong way round; and reading the frequency column straight off the table, 40, 60, 30, 30, skips the division by width altogether and reports how many fish are in each class rather than how densely packed each bar is.
- (c) The class with times from 10 up to 20 — Method: to find the median class from a histogram, first turn each bar's frequency density into a frequency using density × class width, build up the cumulative frequency, and find the first class whose cumulative frequency reaches or passes n ÷ 2. Working: the four classes have widths 10, 10, 20 and 20, so their frequencies are 5 × 10 = 50, 2 × 10 = 20, 1.5 × 20 = 30 and 1 × 20 = 20, which add to the 120 visitors stated. The median sits at position 120 ÷ 2 = 60. The cumulative frequency is 50 after the first class and 50 + 20 = 70 after the second, so the 60th visitor is reached during the second class. Answer: the median lies in the class 10 ≤ t < 20. Watch which class each shortcut lands on: the tallest bar belongs to the first class, with the highest frequency density, 5 — but the tallest bar shows where visitors are packed most densely, not where the middle visitor falls, and picking it lands one class too early, at 0 ≤ t < 10; taking half of the total TIME span instead of half of the total NUMBER of visitors, 60 minutes ÷ 2 = 30 minutes, lands in the class 20 ≤ t < 40, confusing a value on the horizontal axis with a position in the data; and using the full 120 visitors as the target position, rather than 120 ÷ 2 = 60, reaches all the way to the last class, 40 ≤ t < 60, treating the whole data set's size as though it were the position of a single middle value.
- (c) 72 — Method: on a histogram the frequency of a class is the area of its bar, so frequency = frequency density × class width. Working: the class 50 ≤ m < 80 has width 80 − 50 = 30 grams and a frequency density of 2.4 per gram, so the frequency is 2.4 × 30 = 72. Answer: 72 pebbles. The distractors: 192 comes from using the upper class boundary, 80, as the width, giving 2.4 × 80; 12.5 comes from dividing the width by the density, 30 ÷ 2.4, which reverses the area rule; 2.4 comes from reading the height of the bar as the frequency itself, the commonest mistake on histograms, where a height is a density and only an area is a count.
- (c) 0.5 per gram — Method: turn the two known bars into frequencies using area, subtract from the total to find how many letters are left, then divide that frequency by the width of the last class to get its height. Working: the first bar covers 50 g at a frequency density of 1.2, giving 1.2 × 50 = 60 letters, and the second covers 50 g at 1.8, giving 1.8 × 50 = 90 letters; together that is 60 + 90 = 150 letters, so 200 − 150 = 50 letters remain; the class 100 ≤ m < 200 is 100 g wide, so its frequency density is 50 ÷ 100 = 0.5 per gram. Answer: 0.5 per gram. The distractors: 0.25 per gram comes from dividing the remaining 50 letters by the upper class boundary, 200, instead of by the class width of 100; 2 per gram comes from dividing the class width by the frequency, 100 ÷ 50, reversing the formula; 1.4 per gram comes from subtracting only the first bar's 60 letters, leaving 140, and then dividing by 100.
- (a) 49 — Method: add the frequencies of the classes that lie wholly below 35 seconds, then use linear interpolation for the class that 35 cuts through, assuming the calls in that class are spread evenly across it. Working: below 30 seconds there are 15 + 24 = 39 calls; the value 35 lies in the class 30 ≤ t < 50, which is 20 seconds wide and holds 40 calls, and 35 is 35 − 30 = 5 seconds into it, so the estimated share is (5 ÷ 20) × 40 = 10 calls; the estimate is 39 + 10 = 49. Answer: about 49 calls met the target. The distractors: 79 comes from adding the whole of the class 30 ≤ t < 50, 39 + 40, and so counting calls of up to 50 seconds as being under 35; 39 comes from stopping at the class boundary 30 and ignoring the part class altogether; 69 comes from measuring the part of the class from 35 up to 50 instead of from 30 up to 35, giving (15 ÷ 20) × 40 = 30 and then 39 + 30.
- (c) 70 — Method: estimate the median from the cumulative frequency table by interpolation: find its position, n ÷ 2, locate the class it falls in, then add the fraction of the way through that class (adjusted for the cumulative frequency reached before it) to the class's lower boundary. Working: there are 180 riders, so the median is at position 180 ÷ 2 = 90. Before the class 60 ≤ d < 80 the cumulative frequency is 60, and by the end of it, 120, so the 90th rider falls in this class; its frequency is 120 − 60 = 60 and its width is 80 − 60 = 20. The extra distance needed into the class is 90 − 60 = 30, and 30 ÷ 60 × 20 = 10, so the median is 60 + 10 = 70. Answer: the estimated median distance is 70 km. Watch which numbers the interpolation actually uses: reading off just the class's lower boundary, 60, ignores how far into the class the 90th rider falls; using the target position, 90, as the extra distance instead of subtracting the 60 riders already counted before the class gives 90 ÷ 60 × 20 = 30, so 60 + 30 = 90, overshooting by treating the whole position as if none of it had already been counted; and using the total number of riders, 180, instead of half of it as the target position lands in the very last class, giving an estimate of 130 km — further than any rider is known to have ridden by that point in the table.
- (b) Q was faster on average and more consistent — Method: compare the medians for the average and the interquartile ranges for the spread, remembering that a shorter time is faster and a smaller interquartile range means more consistent. Working: the median for class Q is 35 seconds against 38 seconds for class P, so class Q was faster on average; the interquartile range for class P is 46 − 24 = 22 seconds and for class Q it is 44 − 30 = 14 seconds, so class Q's times are more tightly grouped. Answer: class Q was faster on average and more consistent. The distractors: calling Q slower comes from comparing the lower quartiles, 30 against 24, as though a quartile were the average; calling Q less consistent comes from using the gap between the median and the upper quartile as the spread, 44 − 35 = 9 against 46 − 38 = 8, instead of the full interquartile range; the statement that Q was both slower and less consistent comes from making both of those mistakes together.
- (c) 46 — Method: the cumulative frequency table gives the number of runners below each time; to find the number at or above a time, subtract that cumulative frequency from the total. Working: the cumulative frequency for t < 40 is 74, so 120 runners in total take away the 74 who finished in under 40 minutes: 120 − 74 = 46. Answer: 46 runners took 40 minutes or longer. Watch which boundary and which subtraction you use: reading off t < 50 instead of t < 40 and subtracting, 120 − 110 = 10, answers a different question, '50 minutes or longer'; giving 74 itself as the answer reports how many finished below 40 minutes, the opposite of what was asked; and subtracting the two nearby cumulative frequencies, 110 − 74 = 36, finds how many took between 40 and 50 minutes, not everyone from 40 minutes upward.
- (a) 31.5 — Method: to estimate the mean from a histogram, first turn each bar into a frequency (frequency density × class width), then use mean = Σ(frequency × midpoint) ÷ Σfrequency, with the midpoint standing in for every value in that class. Working: the four classes have widths 20, 10, 20 and 20, so their frequencies are 1 × 20 = 20, 3 × 10 = 30, 2 × 20 = 40 and 0.5 × 20 = 10, which do add to the 100 vehicles stated. Their midpoints are 10, 25, 40 and 60, so Σfx = 20 × 10 + 30 × 25 + 40 × 40 + 10 × 60 = 200 + 750 + 1600 + 600 = 3150, and the mean is 3150 ÷ 100 = 31.5. Answer: the estimated mean speed is 31.5 mph. Watch which numbers you treat as the frequencies and which as the values: using the frequency densities themselves as the frequencies, without multiplying by the class widths first, gives 1 × 10 + 3 × 25 + 2 × 40 + 0.5 × 60 = 195 spread over 1 + 3 + 2 + 0.5 = 6.5, and 195 ÷ 6.5 = 30, a mean built from the wrong 'frequencies' altogether; averaging the four midpoints on their own, (10 + 25 + 40 + 60) ÷ 4 = 33.75, ignores how many vehicles are actually in each class; and using each class's lower boundary in place of its midpoint, 20 × 0 + 30 × 20 + 40 × 30 + 10 × 50 = 2300 and 2300 ÷ 100 = 23, systematically underestimates every class by roughly half its width.
- (d) 60 — Method: on a histogram the frequency of a class is its frequency density × its class width, and the frequencies of all the classes add up to the total, so turn each labelled bar into a frequency and subtract their total from 250. Working: 10 ≤ age < 20 has width 20 − 10 = 10, so its frequency is 4.5 × 10 = 45; 20 ≤ age < 35 has width 35 − 20 = 15, so its frequency is 6 × 15 = 90; 50 ≤ age < 70 has width 70 − 50 = 20, so its frequency is 2.75 × 20 = 55. Those three come to 45 + 90 + 55 = 190, and the total is 250, so the missing frequency is 250 − 190 = 60. Answer: the class 35 ≤ age < 50 has 60 members. Watch what you do with the total and the three frequencies you have found: giving the total, 250, as the answer forgets that three bars have already accounted for some of the members; giving 190, the total of the other three classes, reports how many members are not in this class rather than how many are; and leaving one of the three out of the subtraction, for example 45 + 90 = 135 and 250 − 135 = 115, still owes the class at 50 ≤ age < 70 its 55 members.
- (d) 25 — Method: the median is estimated at position n ÷ 2 in the cumulative frequency table, then interpolated across the class it falls in: lower boundary, plus the fraction of the way through the class, times the class width. Working: there are 80 sacks, so the median sits at position 80 ÷ 2 = 40. Before the class 20 ≤ m < 30 the cumulative frequency is 22, and by the end of it, it is 58, so this class holds the 40th sack; its frequency is 58 − 22 = 36 and its width is 30 − 20 = 10. The extra distance needed into the class is 40 − 22 = 18, and 18 ÷ 36 × 10 = 5, so the median is 20 + 5 = 25. Answer: the estimated median mass is 25 kg. Watch which numbers the interpolation uses: reading off just the lower boundary of the median class, 20, ignores how far into that class the 40th sack actually falls; treating n ÷ 2 = 40 itself as the median mass mistakes a position in the list for a mass in kilograms; and using the target position, 40, as the extra distance into the class instead of subtracting the sacks already counted changes the calculation to 20 + 40 ÷ 36 × 10. That comes to 20 + 11.1 = 31.1, overshooting the class because it never subtracts the 22 sacks already counted before it.
- (c) 20 ≤ h < 40 — Method: with 80 values the lower quartile is the 80 ÷ 4 = 20th value in order, so build a running total until it first reaches 20. Working: the running totals are 14, then 14 + 22 = 36, then 52, then 80; the 20th plant is past 14 but not past 36, so it lies in the second class. Answer: the lower quartile lies in the class 20 ≤ h < 40. The distractors: 0 ≤ h < 20 comes from believing that the bottom quarter of the data must all sit in the first class, when that class holds only 14 of the 80 plants; 40 ≤ h < 50 comes from using the position 80 ÷ 2 = 40 and so locating the median rather than the lower quartile; 50 ≤ h < 80 comes from counting 20 plants down from the tallest instead of up from the shortest, which locates the upper quartile at the 60th plant.
- (c) A histogram, with frequency density up the vertical axis — Method: decide which diagram makes area stand for frequency, which is the property the question asks for. Working: on a histogram the vertical axis is frequency density, so the area of a bar is frequency density × class width, and that product is the frequency; this is exactly what is wanted, and it is what allows classes of unequal width to be shown fairly. Answer: a histogram, with frequency density up the vertical axis. The distractors: a bar chart plots frequency as the height, so with unequal widths a wide class would cover far more area than a narrow class holding the same number of batteries, and area would measure nothing; a cumulative frequency diagram plots running totals against upper class boundaries, so a point on it gives how many lie below a value rather than how many lie in a class; a pie chart shows each class as a share of the whole 300 and loses the class widths entirely, so no area on it is tied to a scale of hours.
- (c) 90 — Method: the height of a bar is its frequency density, so twice as tall means twice the frequency density — not twice the frequency, because the two classes have different widths. Then frequency = frequency density × class width. Working: the first bar has frequency density 3 per cm, so the second has frequency density 2 × 3 = 6 per cm; the class 30 ≤ x < 45 is 45 − 30 = 15 cm wide, so its frequency is 6 × 15 = 90. Answer: 90 rods. The distractors: 120 comes from doubling the first bar's frequency instead of its height — the first class holds 3 × 20 = 60 rods, and doubling that ignores the fact that the second class is narrower; 45 comes from using the first bar's frequency density, 3, for the second bar, 3 × 15, and so never using the information that it is twice as tall; 6 comes from stopping at the frequency density of the taller bar and quoting a height as though it were a count.
Build your own mix at the worksheet builder.