Sampling and estimation
Revision chapter for Cambridge International AS and A Level Mathematics 9709, section 6.4 Sampling and estimation, examined in Paper 6 (Probability and Statistics 2) on the A Level route through Papers 5 and 6, for the 2028 to 2030 syllabus. It teaches all eight learning outcomes. Population and sample, parameter and statistic, and why a sample must be random: every member has an equal chance of selection, chosen independently, which avoids bias. How to explain in context why a given sampling method is unsatisfactory by naming who has no chance, or an unequal chance, of being chosen, and how to use random numbers to choose a simple random sample, ignoring out-of-range values and repeats. The sample mean as a random variable with E(X bar) = mu and Var(X bar) = sigma squared over n, derived from the variance of a sum of independent observations. The fact that the sample mean is exactly normal when the population is normal, for every sample size, and the Central Limit Theorem used informally: for a large sample the sample mean is approximately normal whatever the population, with n of at least 30 labelled as a working convention rather than a syllabus rule. Unbiased estimates of the population mean and variance from raw, summarised and coded data using the MF19 forms of s squared, which divide by n minus 1. Confidence intervals for a population mean when the population is normal with known variance or the sample is large, with the critical values 1.645, 1.960, 2.326 and 2.576, the width of an interval, the smallest sample size for a required width rounded up, and the correct interpretation: about 95 percent of intervals built this way from repeated samples contain the fixed parameter. An approximate confidence interval for a population proportion from a large sample. Includes eight worked examples recomputed with MF19 table values, four computed figures, method cards, drills, a mistake clinic, retrieval practice with answers and exam-style structured questions with marking points.Show moreShow less
Revision notes
Interactive notes with exam tips and worked examples.
Study path
Chapter overview
A summary of this Mathematics chapter — open a section to read it. The full notes, worked examples and practice questions are in the study modules above.
What is Sampling and estimation about?
Until now every statistics question told you \(\mu\) and \(\sigma\) and asked about one observation. This chapter turns that round: the population mean is unknown, so you take a sample and use it to estimate. Two ideas make the estimate honest. First, the sample must be random — every member of the population has an equal chance of being chosen, independently — or the way it was chosen biases the answer. Second, the sample mean is itself a random variable \(\bar{X}\): another sample would give another \(\bar{x}\). It has mean \(\mu\) and variance \(\dfrac{\sigma^2}{n}\); it is exactly normal when the population is normal, and approximately normal for a large sample whatever the population (the Central Limit Theorem). From that one fact come the chapter's two products: unbiased estimates of \(\mu\) and \(\sigma^2\) (the variance estimate divides by \(n - 1\)), and a confidence interval, \(\bar{x} \pm z\dfrac{\sigma}{\sqrt{n}}\), that says how precisely \(\mu\) has been pinned down. Eight outcomes, one syllabus section (6.4), Paper 6 only.
Key ideas to remember
- The sample mean has standard deviation \(\sigma/\sqrt{n}\), not \(\sigma/n\); it is exactly normal for a normal population and approximately normal (CLT) for a large sample; the variance estimate divides by \(n - 1\); and it is the interval that varies from sample to sample, never \(\mu\).
- “X̄ has standard deviation σ/√n; exact if X is normal, approximate by the CLT if n is large; s2 divides by n − 1; the interval varies, μ does not.” If those four come back instantly on day 30, the chapter has stuck.
What you need to be able to do
- 6.4.1 I can understand — understand the distinction between a sample and a population, and appreciate the necessity for randomness in choosing samples
- 6.4.2 I can explain — explain in simple terms why a given sampling method may be unsatisfactory
- 6.4.3 I can recognise — recognise that a sample mean can be regarded as a random variable, and use the facts that E(X̄) = μ and that Var(X̄) = σ²/n
- 6.4.4 I can use — use the fact that X̄ has a normal distribution if X has a normal distribution
- 6.4.5 I can use — use the Central Limit Theorem where appropriate
- 6.4.6 I can calculate — calculate unbiased estimates of the population mean and variance from a sample, using either raw or summarised data
- 6.4.7 I can determine — determine and interpret a confidence interval for a population mean in cases where the population is normally distributed with known variance or where a large sample is used
- 6.4.8 I can determine — determine, from a large sample, an approximate confidence interval for a population proportion
Why Sampling and estimation matters
Accuracy for this chapter. Keep intermediate values unrounded and write at least four significant figures of each: \(\dfrac{(\Sigma x)^2}{n} = 25\,603.6\), \(\dfrac{\sigma}{\sqrt{n}} = 0.19365\), \(\dfrac{s}{\sqrt{n}} = 0.49769\). Give \(z\) to 3 d.p. and read \(\Phi\) to 4 d.p. with the ADD column for the third decimal as written, never rounded: Φ(1.033) = 0.8485 + 0.0007 = 0.8492. Quote critical values as 1.645, 1.960, 2.326, 2.576 exactly. Give probabilities and interval ends to 3 s.f., and round a required sample size up. Show the standardisation or the interval formula with numbers substituted before the result: a bare calculator answer earns nothing.
Common mistakes to avoid
- “The unbiased estimate of the variance is \(\dfrac{\Sigma(x - \bar{x})^2}{n}\).” Correct \(s^2\) divides by \(n - 1\). \(s^2 = \dfrac{\Sigma(x - \bar{x})^2}{n - 1} = \dfrac{1}{n - 1}\left\{\Sigma x^2 - \dfrac{(\Sigma x)^2}{n}\right\}\), both forms in MF19. The divide-by-\(n\) formula gives the Paper 5 variance of the data; it is on average too small as an estimate of \(\sigma^2\). On a calculator, the key that divides by \(n\) is the wrong key here.
- “There is a 95% probability that \(\mu\) lies between 51.0 and 54.2.” Correct The interval varies; \(\mu\) does not. \(\mu\) is a fixed number, so a particular interval either contains it or does not. What 95% describes is the method: if many samples were taken and an interval calculated from each in the same way, about 95% of those intervals would contain \(\mu\).
- “The standard deviation of \(\bar{X}\) is \(\dfrac{\sigma}{n}\).” Correct Only the variance is divided by \(n\): \(\operatorname{Var}(\bar{X}) = \dfrac{\sigma^2}{n}\), so the standard deviation is \(\dfrac{\sigma}{\sqrt{n}}\). For \(\sigma = 6\), \(n = 9\) it is 2, not \(\dfrac{2}{3}\).
- “\(n \ge 61.47\), so the smallest sample is 61.” Correct Round a required sample size up. 61 gives a width of 2.008, which is too wide; 62 is the smallest whole number that meets the condition. Rounding to the nearest is wrong even when the decimal part is small.
- “The population is normal, so by the Central Limit Theorem \(\bar{X}\) is normal.” Correct A normal population makes \(\bar{X}\) exactly normal for every \(n\); no theorem about large samples is needed. The CLT is for a population that is not normal, or not known to be, and a large sample, and it gives only an approximation.
- “A 99% interval uses \(z = 2.326\).” Correct 99% in the middle leaves 0.5% in each tail, so \(P(Z \le z) = 0.995\) and \(z = 2.576\). 2.326 is the 98% value. Write the table: 90% 1.645, 95% 1.960, 98% 2.326, 99% 2.576.
- “\(\operatorname{Var}(\bar{X}) = \sigma^2\).” Repair \(\operatorname{Var}(\bar{X}) = \dfrac{\sigma^2}{n}\): the mean of \(n\) observations varies less than one observation does.
- “The standard deviation of \(\bar{X}\) is \(\dfrac{\sigma}{n}\).” Repair It is \(\dfrac{\sigma}{\sqrt{n}}\); only the variance is divided by \(n\).
- “The population is normal, so we need the CLT.” Repair \(\bar{X}\) is then exactly normal for any \(n\). The CLT is for a population that is not normal, or not known to be, with \(n\) large.
- “By the CLT, the sample data are normally distributed.” Repair The CLT is about the distribution of \(\bar{X}\), not of the individual observations; a large sample from a skewed population is still skewed.
- “The sample has \(n = 12\), so by the CLT \(\bar{X}\) is normal” (population skewed). Repair 12 is not large, and the population is skewed with no distribution given, so the distribution of \(\bar{X}\) is not known. Only \(E(\bar{X})\) and \(\operatorname{Var}(\bar{X})\) can be stated.
- “\(s^2 = \dfrac{\Sigma(x - \bar{x})^2}{n}\).” Repair The unbiased estimate divides by \(n - 1\).
- “\(s^2 = \dfrac{\Sigma x^2 - \bar{x}^2}{n - 1}\).” Repair It is \(\dfrac{1}{n - 1}\left\{\Sigma x^2 - \dfrac{(\Sigma x)^2}{n}\right\}\); \(\dfrac{(\Sigma x)^2}{n}\) is \(n\bar{x}^2\), not \(\bar{x}^2\).
- “Asking the first 30 people at a gym is random because the interviewer did not choose them.” Repair Random means every member of the population has an equal chance; people who do not use the gym, or come at other times, have none.
- “The random number 047 appears twice, so student 47 is chosen twice.” Repair Ignore repeats; the sample is of \(n\) different members.
- “There is a 95% probability that \(\mu\) lies in \((51.0, 54.2)\).” Repair \(\mu\) is fixed; about 95% of intervals constructed this way from repeated samples would contain \(\mu\).
- “A 99% interval uses \(z = 2.326\).” Repair 99% needs a central area of 0.99, so \(P(Z \le z) = 0.995\) and \(z = 2.576\); 2.326 is for 98%.
- “A 95% interval uses \(z = 1.645\).” Repair 1.645 leaves 5% in one tail, which gives a 90% interval; 95% uses 1.960.
- “\(n \ge 61.47\), so \(n = 61\).” Repair Round up: \(n = 62\). 61 gives a width of 2.008, over the limit.
- “The proportion interval is \(\hat{p} \pm z\dfrac{\hat{p}(1 - \hat{p})}{n}\).” Repair Take the square root: \(\hat{p} \pm z\sqrt{\dfrac{\hat{p}(1 - \hat{p})}{n}}\).
Examiner tips
- Read the command word before you decide how much to write. This syllabus uses eleven: calculate, describe, determine, evaluate, explain, identify, justify, show (that), sketch, state and verify. Show that and verify give you the answer and mark the route to it, so every step must be visible and the argument must run forwards from what is given, never backwards from the result. Sketch means a simple freehand drawing showing the key features, taking care over proportions; it is not a plot. Determine means establish with certainty; justify means support a case with evidence or argument. Find, solve, express and hence are ordinary question wording; hence means the previous part is the intended route.
- Interleave with the chapters that use this one. Chapter 32 (hypothesis tests) uses the distribution of X̄ from this chapter in every test of a population mean: when you reach it, re-answer worked example 3 and ask how unlikely a sample mean of 3 minutes would be if the claimed mean were 3.2. Recalling a method inside a new problem is worth more than another pass over this chapter on its own: later chapters use these methods without re-teaching them, and the syllabus says an individual examination question may involve ideas and methods from more than one section of the content for that paper, so nothing here is ever finished with.
How Sampling and estimation is examined
- Chapter 31 · Probability & Statistics 2 · How it is assessed
- Cambridge International AS & A Level Mathematics 9709 has six components, and a candidate takes two of them for the AS Level and four for the A Level. This chapter is Probability & Statistics 2 content, examined in Paper 6. Paper 6 (Probability & Statistics 2) is offered only as part of the A Level, where it is 20%. It assumes the whole of the Paper 5 content and the calculus of Paper 3. Every paper is a written examination of compulsory structured questions, answered on the question paper, with MF19 (the list of formulae and statistical tables) supplied. Examinations are available in the June and November series, and in March in India.
- Across the whole qualification the assessment objectives are weighted AO1 55% (knowledge and understanding: concepts, terminology, notation and accurate manipulative technique) and AO2 45% (application and communication: choosing the procedure, combining techniques to solve problems, and presenting the work clearly and logically) at AS Level, and AO1 52%, AO2 48% at A Level. AS candidates are graded a–e; A Level candidates A*–E.
- A short verbal part (explain why a method of sampling is unsatisfactory, or what a confidence interval means) sitting beside a calculation: a probability about X̄ with the distribution stated and justified; estimates from summarised totals followed by an interval from them; a sample size for a required width. Parts can lean on earlier Paper 6 sections, such as a Poisson population averaged by the CLT, and chapter 32 uses the same distribution of X̄ for its tests.
- Printed: the unbiased estimators x̄ and both forms of s2, the line “Central Limit Theorem: X̄ ~ N(μ, σ2/n)”, the approximate distribution of a sample proportion, and the normal tables with the critical values. To know: E(X̄) and Var(X̄) as results for any population, exact normality for a normal population, every interval formula, the width, and the interpretation. The MF19 card has the detail.
- Dividing by n instead of n − 1; standardising with σ or σ/n instead of σ/√n; the wrong critical value for the level; a sample size rounded down; an interval read as a probability statement about μ; and a probability about X̄ with no line saying whether it is exact or by the CLT. Keep intermediate values unrounded, read the table with the ADD column, and give answers to 3 significant figures.
Syllabus reference and sources
Written against: Cambridge International AS & A Level Mathematics (9709). Syllabus for 2028, 2029 and 2030 (version 1, September 2025). Chapter 31: Sampling and estimation.
Written by: Academiq Edu Instructor Panel
Source documents
- Cambridge International AS & A Level Mathematics 9709
- Section 5 of the same syllabus, “List of formulae and statistical tables (MF19)”
- Section 4 of the same syllabus, “Details of the assessment”
All educational content, structured explanations, diagrams, worked examples, and pedagogical materials contained within this chapter revision note are the exclusive intellectual property of Academiq Edu. Unauthorized reproduction, distribution, resale, or extraction of this content without prior written permission is strictly prohibited under international copyright laws. Cambridge Assessment International Education (CAIE) is a registered trademark of Cambridge University Press & Assessment. This revision guide is independently authored by the Academiq Edu Instructor Panel for educational purposes and is not affiliated with or endorsed by Cambridge Assessment International Education.
Verified content
Every chapter note, MCQ explanation and structured mark scheme is checked by Cambridge curriculum specialists.