Representation of Data
Revision chapter for Cambridge International AS and A Level Mathematics 9709, Probability and Statistics 1 (Paper 5), syllabus section 5.1 Representation of data, written to the 2028-2030 syllabus (version 1, identical in teaching content to 2026-2027). It covers all five learning outcomes, 5.1.1 to 5.1.5. Choosing a representation is taught as sentences a student can write: a stem-and-leaf diagram keeps every raw value and gives the exact median, quartiles and mode by counting; a box-and-whisker plot shows five values and is the tool for comparing distributions on one scale; a histogram shows the shape of grouped data; a cumulative frequency graph gives medians, quartiles, percentiles and proportions; grouping always turns a statistic into an estimate. Stem-and-leaf diagrams carry a key with the unit and ordered leaves, and the back-to-back form has leaves increasing outwards on both sides. Box plots have whiskers to the minimum and maximum with no outlier rule. Histograms with unequal widths use frequency density, frequency divided by class width, so the area of a bar is its frequency and the modal class is the class of greatest density; true class boundaries depend on whether the data are ages, rounded measurements or counts. Cumulative frequencies are plotted at upper class boundaries and read at n/2, n/4, 3n/4 and kn/100. The measures of centre (mean, median, mode) and of spread (range, interquartile range, standard deviation) are compared: the median and interquartile range resist an extreme value, the mean and standard deviation use every value. A comparison of two data sets names centre and spread, each in context. The mean and standard deviation are calculated in MF19's form dividing by n, from raw values, grouped data with mid-interval values, the totals of x and x squared, coded totals of x - a and (x - a) squared, and two combined data sets, never by averaging standard deviations. Three fictional data sets run through the chapter: puzzle times A and B (medians 28 and 33 minutes, interquartile ranges 14 and 13, A has mean 27.9 and standard deviation 9.33 minutes) and homework times C for 70 students (densities 0.8, 1.4, 1.2, 0.6, 0.2; estimated mean 35.3 and standard deviation 22.5 minutes; median about 31 minutes from the graph). Eight recomputed worked examples, six computed figures, four sketches, drills on choosing, class boundaries, histograms, graph readings and totals, an MF19 card, a mistake clinic, seventeen retrieval questions, Paper 5-style structured questions with mark allocations, a mastery checklist and a spaced-review plan. The n - 1 variance, outlier rules, skewness coefficients, correlation and interpolation formulae are excluded.Show moreShow less
Revision notes
Interactive notes with exam tips and worked examples.
Study path
Chapter overview
A summary of this Mathematics chapter — open a section to read it. The full notes, worked examples and practice questions are in the study modules above.
What is Representation of Data about?
Statistics starts with a list of numbers. This chapter gives four ways of showing a data set — the stem-and-leaf diagram, the box-and-whisker plot, the histogram and the cumulative frequency graph — and two families of numbers that summarise it: the centre (mean, median, mode) and the spread (range, interquartile range, standard deviation). The skill is choosing: which diagram suits the data, which measure survives an extreme value, and how to compare two data sets in a sentence that names centre and spread in context. It ends with the calculation the rest of Paper 5 leans on: the mean and standard deviation from raw values, grouped data, the totals \(\Sigma x\) and \(\Sigma x^2\), coded totals, and two data sets combined — always dividing by \(n\).
Key ideas to remember
- Unequal widths: height is frequency density and the frequency is the area. Cumulative frequency at the upper boundary, read at \(n/2\). Standard deviation: the mean of the squares minus the square of the mean, divided by \(n\), never by \(n - 1\).
- Frequency density = frequency ÷ width, and area is frequency; cumulative frequency at the upper boundary, read at \(n/2\); \(\sigma = \sqrt{\dfrac{\Sigma x^2}{n} - \bar{x}^2}\), dividing by \(n\), coded totals leave \(\sigma\) alone, and two sets combine by adding \(n\), \(\Sigma x\) and \(\Sigma x^2\).
What you need to be able to do
- 5.1.1 I can — select a suitable way of presenting raw statistical data, and discuss advantages and/or disadvantages that particular representations may have
- 5.1.2 I can draw — draw and interpret stem-and-leaf diagrams, box-and-whisker plots, histograms and cumulative frequency graphs
- 5.1.3 I can understand — understand and use different measures of central tendency (mean, median, mode) and variation (range, interquartile range, standard deviation)
- 5.1.4 I can use — use a cumulative frequency graph
- 5.1.5 I can calculate — calculate and use the mean and standard deviation of a set of data (including grouped data) either from the data itself or from given totals Σx and Σx², or coded totals Σ(x − a) and Σ(x − a)², and use such totals in solving problems which may involve up to two data sets
Why Representation of Data matters
Accuracy for this chapter. Give means and standard deviations to 3 significant figures with the unit of the data. Keep the variance unrounded through the working (86.9956, 508.490, 28.75, 25.4) and take the square root once: rounding data set A's mean to 27.9 before squaring gives \(867.2667 - 778.41 = 88.8567\) and \(\sigma = 9.43\), not 9.33. Write \(n\), \(\Sigma x\) and \(\Sigma x^2\) before the formula so the method is visible; an unsupported calculator answer earns nothing. State every graph reading, and every statistic from grouped data, as an estimate.
Common mistakes to avoid
- “With unequal class widths, the bar height is the frequency.” Correct Unequal widths mean the height is frequency density = frequency ÷ class width, and the frequency is the area of the bar. Drawing heights as frequencies makes a wide class look far bigger than it is (5.1.2).
- “The standard deviation divides by \(n - 1\).” Correct In Paper 5 the standard deviation divides by \(n\), as MF19 prints it: \(\sigma = \sqrt{\dfrac{\Sigma x^2}{n} - \bar{x}^2}\). The version dividing by \(n - 1\) is a different quantity that belongs to Paper 6 (5.1.5).
- “The modal class is the class with the biggest frequency.” Correct With unequal widths the modal class has the greatest frequency density: in data set C it is \(10 \le t < 20\) (density 1.4), not \(20 \le t < 40\) (frequency 24, density 1.2).
- “Cumulative frequencies go at the middle of each class.” Correct A cumulative frequency counts everything up to the end of a class, so it is plotted at the upper class boundary, starting from (lower boundary of the first class, 0) (5.1.4).
- “Ages 10–19 run from 9.5 to 19.5.” Correct Ages are completed years: 10–19 runs from 10 to 20. Only rounded measurements and counts take the half-unit boundaries.
- “The combined standard deviation is the average of the two.” Correct Add the \(n\), \(\Sigma x\) and \(\Sigma x^2\) of the two sets and recompute. The combined spread includes the gap between the two means, which an average of standard deviations ignores.
- “Class B took longer.” Correct A comparison needs two statements, each with numbers and context: centre (“class B took longer on average: median 33 minutes against 28”) and spread (“the IQRs, 13 and 14 minutes, show a similar spread”).
- A stem-and-leaf diagram drawn with no key. Repair “2 | 1 means 21 minutes” is part of the diagram; without it 2 | 1 could be 2.1 or 210.
- Leaves written in the order the data came. Repair Order the leaves: the median and quartiles are found by counting an ordered list.
- Back-to-back leaves on the left increasing towards the stem. Repair On the left the leaves increase outwards: the smallest is next to the stem, as on the right.
- Bar heights equal to the frequencies, with unequal widths. Repair Height = frequency ÷ class width; the frequency is the area.
- “The modal class is \(20 \le t < 40\) because it has the biggest frequency.” Repair With unequal widths the modal class has the greatest frequency density: \(10 \le t < 20\) (1.4 > 1.2).
- “10–19 years” drawn from 9.5 to 19.5. Repair Ages are completed years: 10 to 20. Only rounded measurements and counts take the half-unit boundaries.
- Cumulative frequencies plotted at the class midpoints. Repair Plot at the upper boundary: the cumulative frequency counts everything up to that boundary.
- “The median of 70 grouped values is read at 35.5.” Repair For grouped data read at \(\tfrac{n}{2} = 35\); the \(\tfrac{n + 1}{2}\) rule is for a small ordered list.
- Whiskers cut off at an “outlier limit”. Repair 9709 draws whiskers to the minimum and maximum; there is no outlier rule.
- “Variance \(= 262.5 - 15.5\).” Repair Subtract the square of the mean: \(262.5 - 15.5^2 = 22.25\).
- Dividing by \(n - 1\). Repair Paper 5's standard deviation divides by \(n\), as MF19 prints it.
- “\(\Sigma(x - 50) = -18\) for 12 values, so the mean is \(-1.5\).” Repair \(-1.5\) is the coded mean; \(\bar{x} = 50 - 1.5 = 48.5\). The standard deviation, though, is not adjusted by 50.
- “Combined \(\sigma = \tfrac{1}{2}(4.72 + 5)\).” Repair Add the \(n\), \(\Sigma x\) and \(\Sigma x^2\) of the two sets and recompute: 5.04, not 4.86 (worked example 7).
Examiner tips
- Read the command word before you decide how much to write. This syllabus uses eleven: calculate, describe, determine, evaluate, explain, identify, justify, show (that), sketch, state and verify. Show that and verify give you the answer and mark the route to it, so every step must be visible and the argument must run forwards from what is given, never backwards from the result. Sketch means a simple freehand drawing showing the key features, taking care over proportions; it is not a plot. Determine means establish with certainty; justify means support a case with evidence or argument. Find, solve, express and hence are ordinary question wording; hence means the previous part is the intended route.
- Interleave with the chapters that use this one. Chapter 26 uses the same “mean of the squares minus the square of the mean” as worked example 6: re-answer it there. Chapter 27 uses a data set's mean and standard deviation as its starting point: re-answer worked example 5 before starting it. Recalling a method inside a new problem is worth more than another pass over this chapter on its own, and the rest of Paper 5 uses this chapter’s totals throughout.
How Representation of Data is examined
- Chapter 23 · Probability & Statistics 1 · How it is assessed
- Cambridge International AS & A Level Mathematics 9709 has six components, and a candidate takes two of them for the AS Level and four for the A Level. This chapter is Probability & Statistics 1 content, examined in Paper 5. Paper 5 (Probability & Statistics 1) is 40% of an AS Level that includes it and 20% of the A Level, for which it is compulsory. Its questions use no algebraic methods beyond the Paper 1 content, and it is the foundation for Paper 6. Every paper is a written examination of compulsory structured questions, answered on the question paper, with MF19 (the list of formulae and statistical tables) supplied. Examinations are available in the June and November series, and in March in India.
- Across the whole qualification the assessment objectives are weighted AO1 55% (knowledge and understanding: concepts, terminology, notation and accurate manipulative technique) and AO2 45% (application and communication: choosing the procedure, combining techniques to solve problems, and presenting the work clearly and logically) at AS Level, and AO1 52%, AO2 48% at A Level. AS candidates are graded a–e; A Level candidates A*–E.
- A grouped table turned into a histogram with frequency density or a cumulative frequency graph, then read for a frequency, a median, quartiles or a proportion; a stem-and-leaf diagram or two box plots, followed by a comparison of the two data sets; and a calculation of a mean and standard deviation from totals, coded totals or two combined sets, sometimes with a “show that” for a recovered total. The reasoning is carried by the boundaries you choose, the cumulative frequency you read at, and the totals you add.
- MF19's Summary statistics gives the mean and the standard deviation for ungrouped and grouped data, both dividing by n. Frequency density, the plotting and reading rules, the coded-total and combining rules, and recovering totals from a mean and standard deviation must all be known. See the MF19 card.
- Heights drawn as frequencies; boundaries of 9.5 for ages; cumulative frequencies at midpoints; a median read at (n + 1)/2 from grouped data; a mean rounded before it is squared; a standard deviation adjusted by the coding constant. Keep the variance unrounded, give means and standard deviations to 3 significant figures with the unit, call graph readings estimates, and write the totals before the formula: an unsupported calculator answer earns nothing.
Syllabus reference and sources
Written against: Cambridge International AS & A Level Mathematics (9709). Syllabus for 2028, 2029 and 2030 (version 1, September 2025). Topic 23: Representation of Data.
Written by: Academiq Edu Instructor Panel
Source documents
- Cambridge International AS & A Level Mathematics 9709
- Section 5 of the same syllabus, “List of formulae and statistical tables (MF19)”
- Section 4 of the same syllabus, “Details of the assessment”
All educational content, structured explanations, diagrams, worked examples, and pedagogical materials contained within this chapter revision note are the exclusive intellectual property of Academiq Edu. Unauthorized reproduction, distribution, resale, or extraction of this content without prior written permission is strictly prohibited under international copyright laws. Cambridge Assessment International Education (CAIE) is a registered trademark of Cambridge University Press & Assessment. This revision guide is independently authored by the Academiq Edu Instructor Panel for educational purposes and is not affiliated with or endorsed by Cambridge Assessment International Education.
Verified content
Every chapter note, MCQ explanation and structured mark scheme is checked by Cambridge curriculum specialists.