Independent revision built for the Bristol first- and second-year Economics, Accounting & Finance syllabi. Not affiliated with or endorsed by the University of Bristol.

Probability

Sample spaces and events, set operations and Venn diagrams, the rules of probability, conditional probability and independence, the law of total probability, Bayes’ rule, and probability trees

What is probability?

Foundations

Probability is the mathematics of uncertainty — a number between 0 and 1 measuring how likely an event is. It is foundational to economics and finance, where almost every quantity of interest (a return, a price, a policy outcome) is uncertain. There are three complementary ways to say what a probability means.

Classical

When outcomes are equally likely, P(A) = (favourable outcomes)/(possible outcomes). A fair die gives P(even) = 3/6 = 1/2.

Frequentist (empirical)

The long-run proportion of times an event occurs over many identical trials. Objective, data-driven — the basis of statistical inference.

Subjective (personal)

A degree of belief about an event, updated as evidence arrives via Bayes' rule. Useful where trials cannot be repeated.

Three views, one calculus. Whichever interpretation you adopt, the rules of probability are identical. The interpretation tells you where the numbers come from; the axioms tell you how to combine them.

Sample spaces, outcomes and events

Core concept

An experiment is any process with an uncertain result. Each possible result is an outcome; the set of all outcomes is the sample space S; and an event is any subset of S — a collection of outcomes we care about.

Figure 1 — the sample space of two coin flips, with an event highlighted

Flip a coin twice and S = {HH, HT, TH, TT}. The event "at least one head" is the subset A = {HH, HT, TH}. When outcomes are equally likely, the classical rule gives the probability directly:

P(A) = (number of favourable outcomes) / (number of possible outcomes) = 3/4
Everything starts with S. Write down the sample space before computing anything. Most early mistakes come from miscounting the outcomes or forgetting one — get S right and the probabilities follow.

Set operations and Venn diagrams

Core concept

Because events are sets, we combine them with set operations, and Venn diagrams make the combinations visible.

Figure 2 — union (either) versus intersection (both)
Union A ∪ B
All outcomes in A or B (or both).
Intersection A ∩ B
Outcomes in both A and B at once.
Complement Aᶜ
All outcomes in S that are not in A.
Disjoint
A and B share no outcomes: A ∩ B = ∅ (the empty set).
Figure 3 — the complement of A, and two disjoint events

A collection A₁, A₂, …, Aₙ is a partition of S if the events are mutually disjoint and together they fill S. Partitions are the key to the law of total probability later.

The rules of probability

Core technique

A few axioms and rules govern every probability calculation. The axioms fix the scale; the rules tell you how to combine events.

0 ≤ P(A) ≤ 1;  P(S) = 1;  P(∅) = 0
Complement rule
P(Aᶜ) = 1 − P(A). Often the fastest route: to find "at least one", subtract the probability of "none".
Addition (disjoint)
If A and B are disjoint, P(A ∪ B) = P(A) + P(B). E.g. a die: P(1 or 2) = 1/6 + 1/6 = 1/3.
Multiplication
For a sequence of independent events, multiply: P(A ∩ B) = P(A)P(B).
Inclusion–exclusion
In general P(A ∪ B) = P(A) + P(B) − P(A ∩ B) — subtract the overlap.
Figure 4 — why the overlap is subtracted once

With P(A) = 0.5, P(B) = 0.4 and P(A ∩ B) = 0.2, the union is 0.5 + 0.4 − 0.2 = 0.7. Adding P(A) and P(B) alone would count the overlap twice.

The overlap is the trap. P(A ∪ B) = P(A) + P(B) holds only when the events are disjoint. If they can happen together, you must subtract P(A ∩ B) — otherwise the shared region is double-counted.

Conditional probability

Core technique

New information changes probabilities. The conditional probability of A given that B has occurred rescales A to B's smaller world:

P(A | B) = P(A ∩ B) / P(B),  provided P(B) ≠ 0
Figure 5 — conditioning shrinks the universe to B

Joint-probability tables make this concrete. Suppose students are cross-classified by whether they passed and whether they attended tutorials:

AttendedSkippedTotal
Pass0.400.200.60
Fail0.100.300.40
Total0.500.501.00

Then P(Pass | Attended) = P(Pass ∩ Attended)/P(Attended) = 0.40 / 0.50 = 0.8, while P(Pass | Skipped) = 0.20 / 0.50 = 0.4. Attendance is associated with a higher pass rate.

Condition = divide by the new total. Conditioning on B restricts attention to the row or column for B, and rescales so those entries sum to 1. The denominator is always P(the thing you conditioned on).

Independence

Core concept

Two events are independent when knowing one tells you nothing about the other — the conditional probability equals the unconditional one:

A, B independent  ⟺  P(A | B) = P(A)  ⟺  P(A ∩ B) = P(A) P(B)

The multiplication rule P(A ∩ B) = P(A)P(B) is exactly the test. In the tutorials table, P(Pass) · P(Attended) = 0.60 × 0.50 = 0.30, but the actual joint probability P(Pass ∩ Attended) = 0.40. Since 0.30 ≠ 0.40, passing and attending are not independent — they are positively associated. By contrast, two fair dice are independent: P(both six) = 1/6 × 1/6 = 1/36.

Independent ≠ disjoint. Disjoint events cannot both occur, so learning one happened rules the other out — that is the opposite of independence. Never conflate the two: disjoint means P(A ∩ B) = 0; independent means P(A ∩ B) = P(A)P(B).

The law of total probability and Bayes' rule

Core technique

When a sample space is partitioned into cases A₁, …, Aₙ, the probability of any event B is a weighted average across the cases — the law of total probability:

P(B) = Σᵢ P(B | Aᵢ) P(Aᵢ)

Bayes' rule then reverses a conditional probability — turning P(B | A) into P(A | B) — which is how we update beliefs from evidence:

P(A | B) = P(B | A) P(A) / P(B)
Figure 6 — a partition as a tree: total probability, then reverse it with Bayes

Suppose a component comes from three suppliers with shares 50%, 30%, 20% and defect rates 2%, 4%, 5%. Total probability gives P(defective) = 0.5(0.02) + 0.3(0.04) + 0.2(0.05) = 0.032. Bayes then reverses it: given a defective part, P(from Supplier 3) = 0.2(0.05) / 0.032 = 0.3125. As a second example, if 1/3 of applicants are strong and a strong applicant passes an assessment with probability 3/4 versus 1/4 for others, then given a pass, P(strong) = (3/4 · 1/3)/(5/12) = 3/5.

Do not confuse P(A | B) with P(B | A). "Probability of defect given Supplier 3" and "probability of Supplier 3 given a defect" are different numbers. Bayes' rule is precisely the machinery for converting one into the other — and the base rate P(A) matters enormously.

Probability trees

Applied tool

A probability tree organises sequential or multi-stage experiments. Each branch carries a probability; you multiply along a path and add across paths that give the same outcome.

Figure 7 — drawing two balls without replacement

A bag holds 3 red and 2 blue balls; draw two without replacement. Multiplying along the branches, P(both red) = 3/5 × 2/4 = 3/10, and P(one of each) = P(RB) + P(BR) = 3/10 + 3/10 = 3/5. The second-stage probabilities change because the first ball is not replaced — the two draws are not independent.

Multiply along, add across. A tree turns a wordy sequential problem into simple arithmetic: probabilities on a single path multiply (an intersection), and separate paths to the same event add (a union of disjoint routes).
Flashcards

Tap card to reveal answer

Quick-reference note sheet
Your score
0 / 0
Build the statement — type each figure, then check