Probability
Cambridge O Level Mathematics (Syllabus D) 4024 Topic 8 revision chapter covering the whole of Probability for the 2025-2027 syllabus, version 2. It teaches all three official subtopics in order. Introduction to probability covers the probability scale from 0 to 1 with 0 meaning impossible and 1 meaning certain, the notation P(A) for the probability of an event and P(A prime) for the probability that the event does not occur, the calculation of a single-event probability as the number of favourable outcomes divided by the total number of equally likely outcomes, the requirement to confirm that outcomes really are equally likely before counting, the complement rule that the probability of an event not occurring is one minus the probability that it occurs, and the extraction of probabilities from lists, two-way tables, graphs and Venn diagrams, with answers given as a fraction, a decimal or a percentage and fractional answers simplified. Relative and expected frequencies covers relative frequency as the observed frequency divided by the number of trials, its status as an experimental estimate rather than a guarantee, the tendency of the estimate to settle as the number of trials grows while conditions stay unchanged, expected frequency as the number of trials multiplied by the probability, the use of a probability to estimate an expected value from a population, and the version 2 vocabulary of fair, bias and random, including why a short surprising run is not by itself proof of bias and why randomness does not require every outcome to appear equally often in a small sample. Probability of combined events covers systematic listing, sample-space diagrams and two-way outcome tables in which the thirty-six ordered outcomes of two dice are equally likely while the eleven totals are not, Venn diagrams filled from the intersection outwards with universal, intersection, union, only and neither regions checked against the total, the addition rule with the intersection subtracted once and its mutually exclusive special case, and tree diagrams with outcomes written at the ends of the branches and probabilities beside the branches, multiplying along a single route, adding alternative successful routes, restoring the original composition when sampling with replacement and updating both numerator and denominator when sampling without replacement. Independence is taught as supporting understanding and carefully separated from mutual exclusivity. A visible S-P-A-C-E routine, a model-selection decision map, fully worked examples, a fair-bias-random clinic, a replacement comparison, a mistake clinic, retrieval practice with answers, a mixed exam-style challenge set and a spaced-review plan support both first-pass learning and last-week revision.Show moreShow less
Core Revision Module
Revision & Practice Book
Interactive revision notes with exam tips and worked examples for this chapter.
Practice & Resources
3 toolsChapter overview
A summary of this Mathematics (Syllabus D) chapter — open a section to read it. The full notes, worked examples and practice questions are in the study modules above.
What is Probability about?
The probability of an event \(A\), written \(P(A)\), is a number from 0 to 1: \(P(A)=0\) means the event is impossible and \(P(A)=1\) means it is certain. When every outcome is equally likely, \(P(A)=\dfrac{\text{favourable outcomes}}{\text{total outcomes}}\), and the probability that \(A\) does not happen is \(P(A')=1-P(A)\). Two-stage experiments are modelled with a sample-space table, a Venn diagram or a tree diagram, multiplying along a single route and adding separate routes that both succeed, and the answer is given as a simplified fraction, a decimal or a percentage.
Probability measures how likely an event is, on a scale from 0 to 1. Every question in this topic is answered the same way: say exactly what the event is, build a model of all the possible outcomes — a list, a table, a Venn diagram or a tree — attach a probability to each part of that model, then combine those parts using the right operation. Multiply along a single route; add separate routes that both succeed. Finish by checking the answer lies between 0 and 1.
Key ideas to remember
- Define the event, model every outcome, assign the probabilities, multiply along a route and add across routes, then check \(0\leq P\leq1\).
- Equally likely? Replaced or not? Do they overlap? Does one change the other? Observed or theoretical? Five questions, and the model is decided.
- Every one of these eight is caught by the same two checks: does the answer lie between 0 and 1, and do my exhaustive outcomes total exactly 1?
- If you remember nothing else: define the event, model every outcome, multiply along, add across, and check the total is 1.
What you need to be able to do
- 8.1 — place an event on the probability scale from 0 to 1, and say what \(P(A)=0\) and \(P(A)=1\) mean.
- 8.1 — read and write the notation \(P(A)\) and \(P(A')\), and give an answer as a fraction, a decimal or a percentage, with fractions simplified.
- 8.1 — calculate the probability of a single event by counting favourable outcomes out of equally likely outcomes, having first checked that the outcomes really are equally likely.
- 8.1 — use \(P(A')=1-P(A)\) in both directions, including finding \(P(A)\) when you are given the probability that \(A\) does not happen.
- 8.1 — extract a probability from a list, a two-way table, a graph or a Venn diagram.
- 8.2 — calculate a relative frequency from experimental results and use it as an estimate of a probability.
- 8.2 — explain why a longer run of trials usually gives a more stable estimate, provided the conditions do not change.
- 8.2 — calculate an expected frequency as \(n\times P(A)\), including estimating an expected value from a population, and say clearly why it is a prediction rather than a guarantee.
- 8.2 — use the terms fair, bias and random precisely, and judge fairness by comparing observed with expected frequencies while taking the number of trials into account.
- 8.3 — list the outcomes of a two-stage experiment systematically, and build a sample-space diagram or two-way outcome table without omission or duplication.
- 8.3 — complete a two-set Venn diagram from the intersection outwards, and read \(P(A\cap B)\), \(P(A\cup B)\), \(P(A')\), an "only" region and the "neither" region from it.
- 8.3 — use \(P(A\cup B)=P(A)+P(B)-P(A\cap B)\), and recognise the mutually exclusive case where the intersection is empty.
- 8.3 — draw a tree diagram with outcomes at the ends of the branches and probabilities beside them, multiply along a route, and add alternative successful routes.
- 8.3 — handle selection with replacement, where the composition is restored, and without replacement, where both the numerator and the denominator change.
- Across the topic — check that every probability you write lies between 0 and 1, that the probabilities out of any one node total 1, and that a set of exhaustive outcomes totals 1.
Why Probability matters
A revision habit worth keeping past the exam. Whenever you meet a percentage in the news — a weather forecast, a medical risk, a poll — ask the two questions this chapter has been drilling: is that number theoretical or observed, and how many trials is it based on? Those two questions separate a reliable figure from a meaningless one, and they are the same two the examiner is testing.
Key terms in Probability
- Relative and Expected Frequencies
- Relative frequency is the observed frequency of an outcome divided by the number of trials, and it is used as an experimental estimate of a probability when the theoretical probability is unknown or when a device may not be fair. It is an estimate rather than a guarantee: repeating the experiment can give a different value, and a longer run of trials usually gives a more stable estimate provided the conditions do not change. Expected frequency is the number of trials multiplied by the probability of the outcome, and the same relationship is used to estimate an expected value from a population by multiplying the population size by the probability. An expected frequency is a prediction made by the model rather than a promise about what must occur. A device is fair when its intended outcomes have the appropriate equal theoretical probabilities and biased when they do not, while an outcome is random when an individual result cannot be predicted with certainty in advance even though its probabilities may be known. Fairness is judged by comparing observed with expected frequencies while taking the number of trials into account, because a short surprising run is not by itself proof of bias.
- Introduction to Probability
- Probability is a number from 0 to 1 that measures how likely an event is, where 0 means the event is impossible, 1 means it is certain, and values nearer 1 mean greater likelihood. The probability of event A is written P(A) and the probability that A does not occur is written P(A prime). When every possible outcome is equally likely, P(A) is the number of favourable outcomes divided by the total number of possible outcomes, and the result may be given as a fraction, a decimal or a percentage, with fractional answers simplified. Because an event either occurs or does not occur, the two probabilities total 1, giving the complement rule P(A prime) equals 1 minus P(A). The outcome information needed may come from a systematic list, a two-way table, a graph such as a bar chart, or a Venn diagram, and the counting method is valid only when the listed outcomes are genuinely equally likely.
- Probability of Combined Events
- A combined event involves more than one thing happening, either two features of the same object or two stages of an experiment, and its probability is calculated from a model of all the possible outcomes. Sample-space diagrams and two-way outcome tables list every ordered outcome of a two-part experiment, so that two fair dice give thirty-six equally likely ordered outcomes even though they produce only eleven possible totals, which are not equally likely. Venn diagrams organise two overlapping sets and are completed from the intersection outwards, giving the intersection, the two only regions, the neither region and the union, all of which must total the universal set. The addition rule states that the probability of A union B is the probability of A plus the probability of B minus the probability of A intersection B, so that the overlap is not counted twice; when the events are mutually exclusive the intersection has probability zero and the rule reduces to simple addition. Tree diagrams show successive stages with outcomes written at the ends of the branches and probabilities beside them, probabilities are multiplied along a single complete route and the products of alternative successful routes are added, and the second-stage probabilities are unchanged with replacement but updated in both numerator and denominator without replacement.
Common mistakes to avoid
- 1. Counting outcomes that are not equally likely \(P(A)=\dfrac{\text{favourable}}{\text{total}}\) is a statement about equally likely outcomes only. Two dice have 36 equally likely ordered outcomes but only 11 possible totals, and those totals run from probability \(\frac{1}{36}\) up to \(\frac{6}{36}\). Answering "the total can be 2 to 12, so \(P(\text{total}=7)=\frac{1}{11}\)" is an expensive error, because it invalidates every answer built on top of it. Fix Before dividing, ask: is every outcome in my list as likely as every other? If not, go back down to the ordered outcomes underneath them.
- 2. Leaving the second denominator unchanged without replacement If a counter is not put back, the bag is smaller. Both the numerator and the denominator of the second-stage probability may change, and which numerator changes depends on which branch you are on. Writing \(\frac{3}{5}\times\frac{2}{5}\) for a without-replacement question is not a small slip — it answers a different question. Fix Write the composition of the bag beside every node: "3R 2B" at the start, "2R 2B" after a red, "3R 1B" after a blue.
- 3. Adding overlapping events without subtracting the overlap \(P(A\cup B)=P(A)+P(B)\) is true only when \(A\) and \(B\) cannot happen together. If they can, everything in the intersection has been counted twice, and \(P(A\cap B)\) must be subtracted once. Adding blindly is what produces answers greater than 1. Fix Ask "can both happen at once?" before adding. If the answer is yes, draw the Venn diagram.
- 4. Confusing mutually exclusive with independent These are opposite kinds of statement. Mutually exclusive says the two events cannot both happen. Independent says one happening does not change the chance of the other. Two mutually exclusive events with non-zero probability are as far from independent as it is possible to be: if one happens, the other's probability drops straight to zero. Fix Exclusive is a question about the sample space — do the regions overlap? Independent is a question about influence — does knowing one change the other?
- 5. Treating an experimental result as a certainty A relative frequency is an estimate. An expected frequency is a prediction from a model. Neither is a promise. "The expected number of sixes in 300 rolls is 50, so there will be 50 sixes" is wrong, and so is "the spinner landed on blue 84 times out of 240, so \(P(\text{blue})=0.35\)" stated as a fact rather than as an estimate. Fix Use the words the syllabus uses: estimate for a relative frequency, expected for a model prediction. Write \(\approx\), not \(=\), when the number came out of an experiment.
Examiner tips
- What version 2 changed here. The February 2024 revision made exactly one change to Topic 8: the guidance for 8.2.2 was updated to include the term random. That is why this chapter treats fair, bias and random as three separate ideas with three separate definitions rather than as loose synonyms — see the fair, bias and random clinic inside section 8.2.
- Why the model is worth drawing. Where a multi-part question opens by asking you to complete a Venn diagram or a tree, the later parts are read off what you built. A modelling error in part (a) then propagates through the rest — and, conversely, the completed diagram is itself the communicated method the syllabus asks for. Draw it even when the question does not explicitly ask you to.
- Two phrases worth translating on sight. "At least one" means "one or more" — its complement is "none", which is usually a single route. "Exactly one" means "one and not the other" — on a two-stage tree that is two routes, added. Mistaking one for the other changes the answer, not just the working.
- The complement is a shortcut, not just a definition. Whenever an event is awkward to count directly — "at least one", "not all the same", "more than two" — count the opposite instead and subtract from 1. The opposite is usually a single simple case.
- Three sentences that sound right and are wrong. "It's random, so all the outcomes are equally likely." No — random means unpredictable in advance, not equally likely. A biased spinner is perfectly random. "Each colour came up a different number of times, so the spinner is biased." No — exact equality would be astonishing, not reassuring. Compare with the expected frequency and think about how many trials there were. "I got 7 heads in 10 tosses, so the coin is biased." No — 10 trials tells you almost nothing. Ten thousand would tell you a great deal.
- Do not cancel too early on Paper 1. Leaving \(\frac35\times\frac24\) as \(\frac{6}{20}\) until every route has been calculated keeps all four products over the same denominator, so adding them is trivial and the total-to-1 check is instant. Simplify once, at the end.
- The habit all three share. Each solution finishes by checking a set of exhaustive probabilities against 1 — eight outcomes of \(\frac18\), four spinner probabilities, or same-colour against different-colour. That check costs one line and catches almost every modelling error this topic can produce.
- Marking yourself honestly. Give yourself the method marks only if the working was on your page before you looked — the completed tree, the filled Venn, the products before they were added. A correct final answer with no visible model would not earn full marks on a question that asks for working.
How Probability is examined
- Both components can ask about every part of Topic 8. There is no "probability paper" and no part of this topic that is safe to skip because it belongs to the other one.
- No probability formula is printed on the paper. The list of formulas on page 2 covers areas, volumes, the quadratic formula and the trigonometry rules. There is nothing for probability, so \(P(A')=1-P(A)\) and \(P(A\cup B)=P(A)+P(B)-P(A\cap B)\) have to be known.
- Every probability lies between 0 and 1. An answer outside that range is certainly wrong, and it is the fastest check you own.
- Answers should be given in their simplest form unless the question says otherwise, so \(\frac{6}{20}\) should be written \(\frac{3}{10}\).
- Do not mix fractions and decimals inside one number. The syllabus's mathematical conventions state that a combination such as \(\frac{0.3}{4}\) is not acceptable as a final answer.
- Show the model, not just the answer. Where a question asks for working, full marks need the method communicated — a completed tree with its branch probabilities, or the products before they are added, earns method marks even if the final arithmetic slips.
Frequently asked questions
What is probability and what does the scale from 0 to 1 mean?
Probability is a number from 0 to 1 that measures how likely an event is. \(P(A)=0\) means the event is impossible, \(P(A)=1\) means it is certain, and values nearer 1 mean greater likelihood. The probability that \(A\) does not happen is \(P(A')=1-P(A)\). Every probability you write must lie between 0 and 1, and a set of exhaustive outcomes must total exactly 1.
What is the difference between relative frequency and theoretical probability?
Relative frequency is what an experiment did: the observed frequency of an outcome divided by the number of trials, used as an estimate of a probability. Theoretical probability is what the model says. A longer run of trials usually gives a more stable estimate, provided the conditions do not change, and a short run that looks surprising is not by itself proof of bias. Expected frequency, \(n\times P(A)\), is a prediction from the model, not a promise.
When do you multiply probabilities and when do you add them?
Multiply along a single route through a tree diagram, where both stages must happen; add separate routes that each succeed. On a tree the outcomes sit at the ends of the branches with probabilities beside them, and the probabilities out of any one node total 1. For events that can overlap, \(P(A\cup B)=P(A)+P(B)-P(A\cap B)\), so the intersection is subtracted once.
Why can you not simply add \(P(A)\) and \(P(B)\) for overlapping events?
Because everything in the intersection has been counted twice. \(P(A\cup B)=P(A)+P(B)\) is true only when \(A\) and \(B\) are mutually exclusive and cannot happen together. If they can, subtract \(P(A\cap B)\) once. Adding blindly is what produces answers greater than 1, which the final check — does the answer lie between 0 and 1 — will catch.
What changes in a tree diagram when an item is not replaced?
The bag is smaller at the second stage, so the denominator falls by one, and the numerator may also change depending on which branch you are on. With replacement, the composition is restored and the second-stage probabilities are the same as the first. Writing \(\frac{3}{5}\times\frac{2}{5}\) for two draws without replacement leaves the second denominator unchanged and is wrong; it should be \(\frac{3}{5}\times\frac{2}{4}\).
What is the difference between mutually exclusive and independent events?
They are opposite kinds of statement. Mutually exclusive means the two events cannot both happen, so the intersection is empty and \(P(A\cup B)=P(A)+P(B)\). Independent means one happening does not change the chance of the other, so the second-stage probability on a tree is unchanged. Two mutually exclusive events with non-zero probability are as far from independent as possible, because one happening makes the other impossible.
Why is counting outcomes only valid when they are equally likely?
Because \(P(A)=\dfrac{\text{favourable outcomes}}{\text{total outcomes}}\) is a statement about equally likely outcomes only. Two dice have 36 equally likely ordered outcomes but only 11 possible totals, and those totals are not equally likely: a total of 2 has probability \(\frac{1}{36}\) while a total of 7 has \(\frac{6}{36}\). Build the sample-space table of ordered outcomes first, then count from it.
Syllabus reference and sources
Written against: Cambridge O Level Mathematics – Syllabus D (4024) 2025–2027 Syllabus (Subject Content, Topic 8: Probability).
Written by: Academiq Edu Instructor Panel
Source documents
All educational content, structured explanations, diagrams, worked examples, and pedagogical materials contained within this chapter revision note are the exclusive intellectual property of Academiq Edu. Unauthorized reproduction, distribution, resale, or extraction of this content without prior written permission is strictly prohibited under international copyright laws. Cambridge Assessment International Education (CAIE) is a registered trademark of Cambridge University Press & Assessment. This revision guide is independently authored by the Academiq Edu Instructor Panel for educational purposes and is not affiliated with or endorsed by Cambridge Assessment International Education.
Every chapter note, MCQ explanation, and structured mark scheme is rigorously vetted by Cambridge curriculum specialists.

