Sample spaces and events, set operations and Venn diagrams, the rules of probability, conditional probability and independence, the law of total probability, Bayes’ rule, and probability trees
Probability is the mathematics of uncertainty — a number between 0 and 1 measuring how likely an event is. It is foundational to economics and finance, where almost every quantity of interest (a return, a price, a policy outcome) is uncertain. There are three complementary ways to say what a probability means.
When outcomes are equally likely, P(A) = (favourable outcomes)/(possible outcomes). A fair die gives P(even) = 3/6 = 1/2.
The long-run proportion of times an event occurs over many identical trials. Objective, data-driven — the basis of statistical inference.
A degree of belief about an event, updated as evidence arrives via Bayes' rule. Useful where trials cannot be repeated.
An experiment is any process with an uncertain result. Each possible result is an outcome; the set of all outcomes is the sample space S; and an event is any subset of S — a collection of outcomes we care about.
Flip a coin twice and S = {HH, HT, TH, TT}. The event "at least one head" is the subset A = {HH, HT, TH}. When outcomes are equally likely, the classical rule gives the probability directly:
Because events are sets, we combine them with set operations, and Venn diagrams make the combinations visible.
A collection A₁, A₂, …, Aₙ is a partition of S if the events are mutually disjoint and together they fill S. Partitions are the key to the law of total probability later.
A few axioms and rules govern every probability calculation. The axioms fix the scale; the rules tell you how to combine events.
P(Aᶜ) = 1 − P(A). Often the fastest route: to find "at least one", subtract the probability of "none".P(A ∪ B) = P(A) + P(B). E.g. a die: P(1 or 2) = 1/6 + 1/6 = 1/3.P(A ∩ B) = P(A)P(B).P(A ∪ B) = P(A) + P(B) − P(A ∩ B) — subtract the overlap.With P(A) = 0.5, P(B) = 0.4 and P(A ∩ B) = 0.2, the union is 0.5 + 0.4 − 0.2 = 0.7. Adding P(A) and P(B) alone would count the overlap twice.
New information changes probabilities. The conditional probability of A given that B has occurred rescales A to B's smaller world:
Joint-probability tables make this concrete. Suppose students are cross-classified by whether they passed and whether they attended tutorials:
| Attended | Skipped | Total | |
|---|---|---|---|
| Pass | 0.40 | 0.20 | 0.60 |
| Fail | 0.10 | 0.30 | 0.40 |
| Total | 0.50 | 0.50 | 1.00 |
Then P(Pass | Attended) = P(Pass ∩ Attended)/P(Attended) = 0.40 / 0.50 = 0.8, while P(Pass | Skipped) = 0.20 / 0.50 = 0.4. Attendance is associated with a higher pass rate.
Two events are independent when knowing one tells you nothing about the other — the conditional probability equals the unconditional one:
The multiplication rule P(A ∩ B) = P(A)P(B) is exactly the test. In the tutorials table, P(Pass) · P(Attended) = 0.60 × 0.50 = 0.30, but the actual joint probability P(Pass ∩ Attended) = 0.40. Since 0.30 ≠ 0.40, passing and attending are not independent — they are positively associated. By contrast, two fair dice are independent: P(both six) = 1/6 × 1/6 = 1/36.
When a sample space is partitioned into cases A₁, …, Aₙ, the probability of any event B is a weighted average across the cases — the law of total probability:
Bayes' rule then reverses a conditional probability — turning P(B | A) into P(A | B) — which is how we update beliefs from evidence:
Suppose a component comes from three suppliers with shares 50%, 30%, 20% and defect rates 2%, 4%, 5%. Total probability gives P(defective) = 0.5(0.02) + 0.3(0.04) + 0.2(0.05) = 0.032. Bayes then reverses it: given a defective part, P(from Supplier 3) = 0.2(0.05) / 0.032 = 0.3125. As a second example, if 1/3 of applicants are strong and a strong applicant passes an assessment with probability 3/4 versus 1/4 for others, then given a pass, P(strong) = (3/4 · 1/3)/(5/12) = 3/5.
A probability tree organises sequential or multi-stage experiments. Each branch carries a probability; you multiply along a path and add across paths that give the same outcome.
A bag holds 3 red and 2 blue balls; draw two without replacement. Multiplying along the branches, P(both red) = 3/5 × 2/4 = 3/10, and P(one of each) = P(RB) + P(BR) = 3/10 + 3/10 = 3/5. The second-stage probabilities change because the first ball is not replaced — the two draws are not independent.
Tap card to reveal answer