Statistics
Core Revision Module
Revision & Practice Book
Interactive revision notes with exam tips and worked examples for this chapter.
Practice & Resources
3 toolsChapter overview
A summary of this Mathematics chapter — open a section to read it. The full notes, worked examples and practice questions are in the study modules above.
What is Statistics about?
If you remember nothing else: continuous data goes into intervals, not categories; the median needs ordered data; a stem-and-leaf diagram needs ordered leaves and a key. Extended adds three more: a grouped mean is an estimate; cumulative frequency is plotted at upper class boundaries; and in a histogram the area of a bar is its frequency.
The classification is not a labelling exercise for its own sake — it decides what is legal downstream. Categorical data can have a mode but never a mean. Discrete data can be listed value by value. Continuous data is collected into intervals as soon as the values are many and varied, and it is that grouped continuous data which leads to cumulative-frequency curves and histograms. A short list of measurements can still be handled value by value, with an exact mean and median — see WE 9.2a in 9.2. Get this line wrong and every later step inherits the error.
A statistical conclusion is a sentence with evidence attached, not a number. “Class A did better” states no evidence at all. “Class A has the higher median (\(6\) against \(5\)), so its typical mark is higher” states all of it. The pattern to write, every single time, is: which measure, which way round, what that means here.
The four measures are not interchangeable. The mean uses every value, which is its strength and its weakness — one extreme observation drags it. The median ignores everything except position, so extremes cannot move it. The mode is the only average available for categorical data. The range uses nothing but the two extremes, so it is the measure most easily distorted by a single unusual value. Choosing the right one is the statistics.
Charts differ in what they make easy to see, and in how much of the original data they throw away. A bar chart makes it easy to compare categories against each other. A pie chart makes it easy to see each category's share of the whole — and hard to compare two similar slices. A pictogram makes the count approachable and the precision poor. A stem-and-leaf diagram is the only one of the five that keeps every original value, which is why an exact median, mode and range can be read straight off it. Choosing the diagram is choosing which question the reader will find easy to answer.
Direction and strength are two separate statements. “Positive” says which way the trend runs; “strong” says how closely the points follow it. The syllabus names the direction explicitly — positive, negative or zero — so that is the part it asks for by name; the strength costs one extra word and makes the description sharper, so give both. A relationship can be weak and positive, or strong and negative, and the two words are chosen independently. Then — every time — add the third sentence: association is not causation.
Key ideas to remember
- DATA — Define the variable and its type · Arrange it in a table or diagram · Take the measure asked for · Answer in context, saying whether the figure is exact or an estimate.
What you need to be able to do
- I can decide whether a variable is categorical, discrete or continuous, and justify the choice.
- I can build a tally table from raw data and check the frequency total against the number of observations.
- I can choose grouped class intervals that are equal in width where sensible and that never overlap.
- I can complete a two-way table and verify the grand total from both the row totals and the column totals.
- I read the title, key, units and axis scale of a chart before extracting any value from it.
- I can draw an inference from a table or a statistical diagram and state it in context.
- I can compare two data sets using an average and the range, and then interpret both in context.
- I can name the specific limitation that weakens a stated conclusion — sample size, bias, missing data, measurement, wording, scale or extrapolation.
- I can explain why correlation does not establish causation, using the data in front of me.
- I can find the mean, median, mode and range of a list, ordering the list first.
- I can find the mean, median, mode and range from an ungrouped frequency table using \(\dfrac{\sum fx}{\sum f}\), without confusing the value row with the frequency row.
- I can say what each of the four measures is good for and how each one fails.
- I can draw and read bar charts, including composite and dual bar charts with a legend and a shared scale.
- I can calculate a pie-chart sector angle, do the reverse calculation from angle to frequency, and check the angles total \(360^\circ\).
- I can apply a pictogram key to partial symbols as well as whole ones.
- I can draw a stem-and-leaf diagram with ordered leaves and a key, and read the mode, median and range straight off it.
- I can draw and read a simple frequency distribution, and explain why a category bar chart has separated bars.
- I can plot paired data accurately as crosses and describe the correlation as positive, negative or zero, and by strength.
- I can draw one ruled line of best fit by inspection, across the whole data set, with a rough balance of points either side of it.
- I can use the line to interpolate by reading from it, and can explain why extrapolating it is unreliable.
- I can identify an outlier and say what should be done about it.
- I can compare two data sets using an average together with any measure of spread, choosing the interquartile range rather than the range when extreme values would distort the comparison.
- I can find the lower quartile, upper quartile and interquartile range of individual data from an ordered list, and say what the IQR measures.
- I can estimate the mean of grouped discrete or grouped continuous data from class midpoints, and state that the answer is an estimate.
- I can identify the modal class from a grouped frequency distribution and explain why I do not quote an exact mode from grouped data.
- I can build a cumulative-frequency column as a running total and plot it at upper class boundaries.
- I start the curve at the lower boundary of the first class with cumulative frequency \(0\), mark the points clearly and join them with a smooth increasing curve.
- I can estimate and interpret the median, \(Q_1\), \(Q_3\) and any percentile from the diagram, and calculate the interquartile range.
- I can read the curve backwards to estimate how many observations lie above or below a stated value.
- I can compare two distributions using median and IQR together.
- I can calculate frequency density and use it as the bar height.
- I can recover a frequency from a histogram using \(\text{frequency}=\text{frequency density}\times\text{class width}\).
- I can explain why bar area represents frequency and why the tallest bar need not be the largest class.
- I label the vertical axis “frequency density” and draw adjacent bars touching.
- I can find true class boundaries for values recorded to the nearest whole unit.
Common mistakes to avoid
- “The median of \(7,3,9,4,6\) is \(9\), because it is the middle one written down.” Why it is wrong The median is the middle of the ordered list. Unordered, the middle position means nothing. Write instead Order first: \(3,4,6,7,9\). Median \(=6\). Section 9.3
- “Class A is better than Class B because its mean is higher.” Why it is wrong One measure of centre is half an answer, and “better” is not a statistical word. Two sets can share a mean and be nothing alike. Write instead “The two classes have the same median (\(6\) each), so their typical mark is the same. Class A has the smaller range (\(4\) against \(10\)), so its marks are more consistent.” Section 9.2
- A stem-and-leaf diagram drawn with the leaves in the order the raw data happened to arrive, and no key. Why it is wrong The syllabus note is explicit: stem-and-leaf diagrams have ordered data with a key. Unordered leaves make the median unreadable, and without a key the digits have no size — a stem of \(3\) and a leaf of \(7\) could be \(37\), \(3.7\) or \(370\). Write instead Sort each row of leaves ascending, then write the key underneath: “Key: \(3\;|\;7\) means \(37\) marks”. Section 9.4
- “There is strong positive correlation, so more revision causes higher marks.” Why it is wrong A scatter diagram shows that two variables move together. It cannot rule out a third factor driving both, or the causation running the other way. Write instead “There is strong positive correlation: students who revised longer tended to score higher. This does not prove that revision caused the higher marks.” Section 9.5
- “The mean of the grouped table is \(25.25\) minutes.” Extended E9.3 Why it is wrong The exact values inside each class are unknown. The midpoint is a stand-in, so the result is an estimate of the mean, not the mean. Write instead “An estimate of the mean is \(25.25\) minutes.” One word. Section 9.3
- “The modal class is \(14\).” Extended E9.3 Why it is wrong A modal class is an interval. \(14\) is the frequency of that interval, not the class itself. Write instead “The modal class is \(20<t\le30\).” Section 9.3
- “Plot each cumulative frequency against the midpoint of its class.” Extended E9.6 Why it is wrong A cumulative frequency counts everything up to and including the class. That total is only reached at the upper boundary, not halfway through. Write instead Plot \((2,6)\), \((4,20)\), \((6,42)\)… and start at \((0,0)\), the lower boundary of the first class. Section 9.6
- “The classes have different widths, but the frequencies are \(20,15,18,12\), so those are the bar heights.” Extended E9.7 Why it is wrong With unequal widths, using frequency as height makes a wide class look enormous. It is the area that must equal the frequency. Write instead Divide each frequency by its class width to get frequency density: \(2.0,\;3.0,\;1.2,\;0.4\). Section 9.7
- Core C9.3 Pets \(0,1,2,3,4\) with frequencies \(11,14,9,4,2\). “The mode is \(14\).” Why it convinces \(14\) is the largest number in the table, and “mode = most” is half-remembered correctly. The error \(14\) is how often, not what. It lives in the frequency row, which is a count of the data, not the data. Correct The largest frequency is \(14\), which belongs to \(x=1\). The mode is \(1\) pet.
- Core C9.3 Same table. “The range is \(14-2=12\).” Why it convinces It is largest minus smallest — of the wrong row. The error Range is the spread of the values, so it is computed from the top row only. Correct Values run \(0\) to \(4\), so the range is \(4-0=4\) pets.
- Core C9.3 Times \(9,9,10,11,12,14,15,16\). “\(n=8\), so the median position is \(\frac{8+1}{2}=4.5\). The median is \(4.5\).” Why it convinces The formula was applied correctly. Only the last step is wrong. The error \(\frac{n+1}{2}\) gives a position in the ordered list, not a value. Position \(4.5\) means “between the \(4\)th and \(5\)th”. Correct The \(4\)th and \(5\)th values are \(11\) and \(12\), so the median is \(\frac{11+12}{2}=11.5\) s. A median of \(4.5\) would also fail the sense check: it lies outside the data, which begins at \(9\).
- Extended E9.3 Grouped homework times: \(\sum fx = 1010\) across five classes. “Estimated mean \(=\frac{1010}{5}=202\) minutes.” Why it convinces There genuinely are five classes, and dividing by “how many things” is the right instinct. The error The divisor is \(\sum f\) — how many students — not how many classes. Correct \(\frac{1010}{40}=\mathbf{25.25}\) minutes. And \(202\) minutes fails the boundary check instantly: the largest class ends at \(50\) minutes, so no mean can exceed \(50\).
- Extended E9.3 Egg masses in classes \(40\)–\(50\), \(50\)–\(55\), \(55\)–\(60\), \(60\)–\(80\). “Midpoints \(45, 55, 65, 75\).” Why it convinces The first midpoint is right, and midpoints usually do step evenly — when the widths are equal. The error These widths are \(10,5,5,20\). A pattern was assumed instead of four separate calculations. Correct \(\frac{40+50}{2}=45\); \(\frac{50+55}{2}=52.5\); \(\frac{55+60}{2}=57.5\); \(\frac{60+80}{2}=70\). Compute every midpoint from its own two boundaries.
- Core C9.4 \(18\) of \(60\) students choose music. “Sector angle \(=\frac{18}{60}\times100=30^\circ\).” Why it convinces \(30\) is a perfectly plausible-looking angle, and \(\times100\) is the reflex from percentage questions. The error A pie chart is a full turn, so the scaling factor is \(360^\circ\). \(\times100\) produces a percentage wearing a degree symbol. Correct \(\frac{18}{60}\times360=\mathbf{108^\circ}\). The check catches it too: with \(\times100\) the five sectors would total \(100^\circ\), not \(360^\circ\).
- Extended E9.6 Cumulative frequency curve, \(N=80\). “The median is where the curve reaches \(40\) on the mass axis.” Why it convinces The number \(\frac N2 = 40\) is right. Only the axis is wrong. The error \(\frac N2\) is a count, so it belongs on the cumulative frequency axis. On a mass axis running \(0\) to \(12\) kg, \(40\) does not even exist. Correct Go across from \(40\) on the vertical axis to the curve, then down to the mass axis: median \(\approx5.8\) kg.
- Extended E9.6 Same curve. “\(Q_1\) is at \(\frac{12}{4}=3\) kg, because the masses run from \(0\) to \(12\).” Why it convinces It uses a quarter of something, and produces a mass, which is the right kind of answer. The error \(Q_1\) is the quarter-way point through the data, not through the range of the axis. Those coincide only if the data is perfectly uniform, which it never is. Correct Read across at \(\frac N4=20\): \(Q_1=4.0\) kg. The two answers differ by a whole kilogram, and the wrong method would give the same \(3\) kg for any distribution on that axis — which is the tell that it cannot be right.
- Extended E9.7 A class recorded as \(10\)–\(19\) to the nearest unit, frequency \(25\). “Width \(=19-10=9\), so density \(=\frac{25}{9}=2.78\).” Why it convinces Subtracting the two printed numbers is exactly what you do when the boundaries are printed. The error These are rounded values, so the class really occupies \(9.5\le x<19.5\). Correct Width \(=19.5-9.5=10\), so density \(=\frac{25}{10}=\mathbf{2.5}\). Check by counting the labels \(10,11,\ldots,19\): ten of them, so the width is \(10\).
- Extended E9.7 A class of width \(10\) contains \(20\) patients. “Frequency density \(=\frac{10}{20}=0.5\).” Why it convinces Both numbers are correct and the division looks symmetrical. Nothing in the answer flags an error. The error The fraction is upside down. Density means observations per unit, so frequency is on top. Correct \(\frac{20}{10}=\mathbf{2.0}\). Always reverse-check: \(2.0\times10=20\) ✓, whereas \(0.5\times10=5\), which is not the frequency you were given.
- Core C9.4 Marks \(34,21,45,33,28,41,37,22,30,46,35,29,38,42,33\), written up as “\(2\;|\;1\;8\;2\;9\) / \(3\;|\;4\;3\;7\;0\;5\;8\;3\) / \(4\;|\;5\;1\;6\;2\)”, and the median given as \(30\). Why it convinces Every value is present, in the right row, and \(30\) is the \(8\)th leaf as written. The diagram looks finished. The error The leaves are in arrival order, not ascending order. Counting to the \(8\)th leaf only finds the median if the leaves are sorted — and there is no key, so the digits have no size either. Correct Sort each row: \(2\;|\;1\;2\;8\;9\) / \(3\;|\;0\;3\;3\;4\;5\;7\;8\) / \(4\;|\;1\;2\;5\;6\), and add “Key: \(2\;|\;1\) means \(21\) marks”. The \(8\)th leaf is now \(4\) on row \(3\), so the median is \(34\) marks. Recount the leaves after sorting: \(4+7+4=15\) ✓
- Extended E9.3 Ordered marks \(12,15,18,21,23,26,28,31,35,40,44\). “\(n=11\), so \(Q_1\) is at \(\frac{11}{4}=2.75\), giving \(Q_1\approx17\).” Why it convinces A quarter of something has been taken, and the answer sits plausibly inside the data. The error The quartile position uses \(n+1\), not \(n\) — the same \(+1\) that appears in \(\frac{n+1}{2}\) for the median, and for the same reason: you are counting gaps between values, not values. Correct \(\frac{11+1}{4}=3\), so \(Q_1\) is the \(3\)rd value, \(\mathbf{18}\). Likewise \(Q_3\) is at \(\frac{3\times12}{4}=9\), the \(9\)th value, \(35\), so \(\text{IQR}=17\). Check the order: \(18<26<35\) ✓
Examiner tips
- Do the arithmetic on paper even on a calculator paper. Topic 9 is arithmetically light but bookkeeping-heavy: a mean is a sum of twelve products, a cumulative frequency column is five additions in a chain, a histogram is four divisions. Every one of those is a place to slip a digit. Write the \(fx\) column out; write the running totals out. The calculator — on Paper 3 for Core candidates, Paper 4 for Extended — is for the final division, not for holding the table in its memory.
- Write the totals row and column even when the question does not print one. A two-way table without totals is a table you cannot check. Adding them costs two additions and converts a guess into a verified answer — and on a “complete the table” question the totals are the working that shows how each entry was found.
- Never write “better”, “worse” or “good”. They are judgements, not statistics, and they name no evidence. Replace them with what you actually measured: higher median, smaller range, more consistent, more spread out, typically larger. And always attach the numbers — “the higher median” without the two values is only half a statement.
- The median position formula gives a position, not a value. For \(n=9\), \(\frac{9+1}{2}=5\) means “the \(5\)th value in the ordered list”, not “the median is \(5\)”. For \(n=8\), \(\frac{8+1}{2}=4.5\) means “halfway between the \(4\)th and \(5\)th values”, so you average those two. Writing \(4.5\) as the answer reports a position where a value belongs — and the sense check catches it at once, since \(4.5\) need not even lie inside the data.
- One method, used consistently, with the positions written down. Some textbooks find quartiles instead by taking the median of the lower half and the median of the upper half of the ordered data. For \(n=11\) above that gives the same \(18\) and \(35\); for the eight times it gives \(9.5\) and \(14.5\), an IQR of \(5\) rather than \(5.5\). The two conventions agree for every odd \(n\) — try them on the eleven marks above, or on any list with an odd number of values — and differ slightly only when \(n\) is even. Pick one, use it every time, and always show the positions you used — the working is what makes your answer readable.
- A grouped mean has a built-in sanity check: it must lie between the smallest lower boundary and the largest upper boundary. In WE 9.3c that is \(40\) to \(80\) g, and \(57.6\) sits comfortably inside. If your answer escapes that interval you have almost certainly divided by \(\sum fx\) instead of \(\sum f\), or lost a digit in one of the products. The check takes two seconds and catches the worst errors.
- Three marks are commonly lost on one diagram, and all three are avoidable. Leaves not ordered. No key. A leaf column that does not line up, so the rows cannot be compared by length. Rule the vertical line between stem and leaf, space the leaves evenly, and write the key underneath as soon as you have drawn the first row — not at the end, when you may forget.
- When you are asked to draw, draw completely. A bar chart with correct heights but no axis label, no scale label and no title is an incomplete answer — the labelling is part of what was asked for, not decoration. Get into the habit of writing the title first, before you plot anything — it takes four seconds and it also forces you to reread what the data actually is.
- Supporting insight—not required syllabus content: a useful anchor, not a rule. A line of best fit will usually pass close to the point \((\bar x,\bar y)\) — the mean of the \(x\) values against the mean of the \(y\) values. Plotting that point lightly gives you somewhere sensible to pivot the ruler. Cambridge accepts a range of lines, so this is a way of landing inside that range quickly, not a way of finding “the” line. There is no unique correct line of best fit at this level.
- Write the three heights down before you touch the ruler. For \(N=80\): “\(Q_1\) at \(20\), median at \(40\), \(Q_3\) at \(60\)”. Doing this as a separate line of working means that if your read-off drifts, the method is still visibly correct. It also guards against a very easy error — reading the median at \(N/2\) of the mass axis instead of the cumulative frequency axis.
- Two columns, then draw. Whatever the question gives you, add a class width column and a frequency density column to the table before touching the graph paper — and if the data was recorded to the nearest unit, add a true boundaries column first. Those columns are the method. Trying to compute densities in your head while drawing is how the widths \(10, 5, 15, 30\) turn into \(10, 5, 15, 15\).
- If you only remember one sentence from this clinic: densities never add up, areas always do. Any time you find yourself summing a frequency-density column, stop — convert each one back to a frequency first.
- Write the working out, not just the answer. The points below are split into method (M), accuracy (A) and statement (B) lines so you can see which part of an answer each one is testing. A correct \(fx\) column with a slipped final division still shows a correct method; a bare correct number shows none. These are this resource’s own marking points, written to structure your self-marking — they are not an official mark scheme.
Frequently asked questions
Why does the mean sometimes look “wrong”, like \(1.3\) pets or \(2.4\) children?
Because a mean summarises a group, not a member of it. No household owns \(1.3\) pets, and none needs to. The mean is the value each household would own if the total were shared out equally, which is a statement about the total and the count — both of which are whole numbers. Only the mode is guaranteed to be a value that actually occurred.
When should I use the median rather than the mean?
When the data contains an extreme value or is clearly lopsided. The mean uses every value, so a single unusual observation drags it; the median depends only on position, so it does not move. In Figure 9.5, changing one value from \(10\) to \(20\) moved the mean from \(6\) to \(8\) and left the median at \(5\). If a question asks which average “best represents” the data and there is an obvious outlier, the median is almost always the expected answer — and say why: “because it is not affected by the extreme value”.
Is \(\frac{n+1}{2}\) or \(\frac N2\) the right median rule? Extended E9.6
Both, in different situations — and a Core candidate only ever meets the first. For a list of individual values you are counting positions, so the median is at position \(\frac{n+1}{2}\). That is the Core rule, and it is the only one C9.3 needs. For a cumulative frequency curve — Extended only — you are reading a continuous model at its halfway height, so you read across at \(\frac N2\). At the sample sizes these questions use, the graphical difference between \(\frac N2\) and \(\frac{N+1}{2}\) is invisible anyway; what matters is knowing which situation you are in, so the two methods do not get crossed.
Why is the modal class not just “the tallest bar”? Extended E9.3 Extended E9.7
On a histogram with unequal class widths, the tallest bar is the one where data is most densely packed, which is not the same as the one containing the most observations. The modal class is defined by the greatest frequency, which on a histogram is the greatest area. In Figure 9.11 the tallest bar holds \(15\) patients while a shorter, wider bar holds \(20\). When the class widths are all equal the two do coincide — which is why the confusion survives.
Why is cumulative frequency plotted at the upper boundary and not the midpoint? Extended E9.6
Because a cumulative frequency answers “how many are at or below this value?”. Halfway through a class you have not counted all of it yet, so the running total is not correct there. The total is only reached at the top of the class. Plotting at midpoints shifts the entire curve to the left and makes every median, quartile and percentile too small.
Do I need to know the formula for the median of grouped data? Extended E9.6
No, and only Extended candidates meet the question at all. E9.6 asks for the median of grouped data to be estimated from a cumulative frequency diagram, by reading across at \(\frac N2\) and down. An algebraic interpolation formula is not required anywhere in 0580, and a question that supplies a grid is asking you to use the grid. (Interpolating between two plotted points arithmetically will give a very similar answer and is a legitimate private check — but the graphical reading, with visible reading lines, is what the marks are for.)
How exact does a reading from a graph have to be?
Reading a curve by eye is inherently approximate, so an answer of this kind is judged against a band of acceptable values rather than one exact number. What protects you is showing the construction: rule the horizontal line at \(\frac N2\), rule the vertical line down, and leave both on the diagram. A reading with visible construction lines shows the method behind it; a bare number shows nothing but itself.
Can I draw a line of best fit through the origin?
Only if the trend genuinely goes there. Forcing the line through \((0,0)\) is a common error: it is a decision about the data, not a tidying-up convention. The line should follow the overall trend with a rough balance of points above and below along its whole length — and if that line happens to miss the origin, it misses the origin. Equally, do not force it through the first point, the last point, or any particular point.
Should I delete an outlier?
No — investigate it. An outlier may be a recording error, a genuine unusual case, or a sign that another factor is at work. Removing a point because it spoils the trend is not a statistical decision, it is the opposite of one. Only a reason found outside the data — a known measurement fault, for instance — justifies excluding it, and you would say so.
Is strong correlation ever enough to prove cause?
No, and the strength of the correlation makes no difference. A scatter diagram shows that two variables move together. It cannot rule out a third factor driving both, and it cannot tell you which direction any causation runs. The standard sentence is: “This shows association only; it does not prove that one variable causes the other.”
What is the difference between a bar chart and a histogram, in one line? Extended E9.7
A bar chart has separated bars of equal width whose height is the frequency, and it is for categorical or discrete data. A histogram has touching bars whose widths are the class widths, whose height is the frequency density, and whose area is the frequency — and it is for grouped continuous data. The three differences always travel together; see Figure 9.12.
Why does a stem-and-leaf diagram need a key when a bar chart does not? Core C9.4
Because its numbers have no axis to give them size. A bar chart carries a labelled vertical scale, so a bar of height \(7\) is unambiguous. A stem-and-leaf diagram carries only digits: the pair \(3\;|\;7\) could be \(37\), \(3.7\) or \(370\), in any units at all. The key — “\(3\;|\;7\) means \(37\) marks” — is what turns the digits back into data, which is why the syllabus note makes it compulsory rather than optional. The same note requires the leaves to be ordered, for a separate reason: the median is found by counting along the leaves, and counting an unordered row finds the wrong value.
Do I need standard deviation or box plots for 0580?
No, on either route. Standard deviation, variance, box-and-whisker plots, correlation coefficients, regression equations, the normal distribution and hypothesis testing are all outside Cambridge IGCSE Mathematics 0580. The measures of spread the syllabus does require differ by route: Core needs the range, and that is the whole of C9.3's spread requirement. Extended adds the interquartile range — for individual data in E9.3, and read from a cumulative frequency diagram in E9.6. Using an out-of-syllabus method will not gain marks, and using an Extended measure on a Core paper will not either, because the question will not have asked for it.
Every chapter note, MCQ explanation, and structured mark scheme is rigorously vetted by Cambridge curriculum specialists.

