Two strands that meet in Week 2: what a causal question actually asks and why observational data cannot answer it on its own, and the probability machinery — distributions, moments, and the behaviour of an average — that every later method is built from.
Three questions, all about the same industry, all answerable with data, and all asking for something different.
Has the share of Bristol economics graduates still working in the South West five years on risen since 2010?
Does a longer apprenticeship actually raise a marine fitter’s hourly pay — or do the people who take one simply earn more anyway?
What will Cornish visitor spending be next August?
The first is a matter of counting carefully. The third asks for a number you will be judged on when it arrives. The second is the hard one, and it is where most of this unit goes: it asks what would have happened to the same person under a different choice, which is not something any dataset contains.
Underneath all three sits one idea. The data you have are treated as a realisation — one draw — from an underlying probabilistic model of how the world generates outcomes. That model is the population; your spreadsheet is a sample from it. Econometrics is the business of reasoning backwards from the sample to the model, and being honest about how far that reasoning carries.
Empirical economics has a standard shape, and it is worth seeing it whole before meeting any of the pieces.
Step 2 is economics. You decide, from theory or from plain reasoning, which things plausibly bear on the outcome. Sometimes that comes from a formal model — a utility-maximisation problem, say, whose solution is a demand equation. More often it comes from knowing the industry: a fitter’s pay depends on training, schooling and years on the tools, and you do not need a Lagrangian to say so.
Step 3 is where econometrics starts, and it demands two commitments the economic model never had to make. You must choose a functional form — linear, in this unit, almost always — and you must decide what to do about everything you left out.
The economic model says only that pay depends on training, schooling and experience, without saying how. The econometric model commits:
That is a much stronger statement, and a much more useful one: it has named quantities in it that data can pin down.
The βs are the parameters. They are fixed, unknown numbers describing the population — not things you compute, things you estimate. β1 is the one the question was about: how much an extra week of training adds to hourly pay, holding schooling and experience fixed.
Everything else that affects pay and is not on the list goes into u, the error term or disturbance. For a fitter that might be aptitude, how hard they are willing to work, whether the yard they joined happened to be well run, and the ordinary noise of wage-setting. You cannot list it all, and you certainly cannot measure it all.
Once the model is written down, a question becomes arithmetic on its parameters. “Training does nothing for pay” is H0: β1 = 0. “An extra week is worth at least 20p an hour” is β1 ≥ 0.20. Translating an argument into a restriction on parameters, and then asking whether the data can reject it, is most of what the exam will ask you to do.
Before any method, a practical question: what does one row of your data represent? There are four common answers, and which one you have determines what you are allowed to do.
Many units, one moment. A row is a fitter. The ordering of the rows carries no information whatsoever — shuffle them and nothing is lost — and if the sample was drawn at random from the population, that randomness is what makes the standard results work.
One unit, many moments. A row is a year. Now the ordering is the information, and the comfortable assumption of independence is gone: this year’s average wage in a yard is obviously related to last year’s. That dependence is why time series needs its own machinery, and why this unit leaves it until much later.
Several independent cross-sections stacked up. Survey 300 fitters in 2022, survey 300 different fitters in 2024, and put the two together with a year variable to tell them apart. Nobody appears twice. It is a bigger sample, and — more usefully — it straddles whatever changed in between.
The same units, followed over time. Fitter 1 in 2022 and fitter 1 in 2024 are both in the data. This is the structurally richest of the four, because you can watch a single person change, which means the things about that person that did not change can be swept out of the problem entirely.
A second question, and a more awkward one: who decided the values of the variable you care about?
In an experiment, you did. You choose which plots get the new feed mix, which fitters get the extra twelve weeks, and you choose by a method unrelated to anything about them — a coin, in effect. In observational data — also called non-experimental, or retrospective — the subjects chose, or circumstance chose, and you merely recorded what happened. You were a passive collector.
Nearly all economic data are observational, and not by oversight. You cannot randomly assign people to leave school at sixteen. You cannot randomly set the Bank of England’s base rate to different levels in different years to see what each does. Some of it is prohibitively expensive, some is impossible, and a good deal of it would be straightforwardly wrong.
“Other relevant factors being equal.” You met the phrase in first year as a modelling convenience. It is in fact the definition of the thing you are trying to measure.
Ask what a longer apprenticeship is worth. The honest version of the question is: take this fitter, and compare what they earn having done the extra twelve weeks with what the same fitter would have earned having not done them. Not a different fitter. The same one, in a different state of the world.
Those two numbers are the potential outcomes, or counterfactual outcomes. The causal effect for this person is the difference between them. Stated this way, the ceteris paribus condition is free: everything else about the fitter is held fixed automatically, because it is the same fitter either way.
It will not always be possible to hold all else equal, and you should be suspicious of anyone who claims otherwise. The practical question in applied work is never “is this perfect?” but “have enough of the right things been held fixed for this comparison to be worth believing?” That is the question a referee asks, and it is the question an exam answer should address.
Here is the obvious thing to do, and why it does not work. Take your sample of fitters, average the pay of those who did the longer apprenticeship, average the pay of those who did not, and subtract.
The trouble is that nobody assigned the training. Fitters chose it. And the kind of person who signs up for twelve more weeks of unpaid study is, on average, also the kind of person who turns up early, takes the awkward jobs and gets promoted — none of which you measured. Those traits raise pay on their own. They also made the training more likely.
So the raw gap contains the effect of training plus the effect of whatever made people select into it. The two arrive welded together, and no amount of arithmetic on those two averages can separate them.
Aptitude and drive affect both the choice and the outcome, and appear in no dataset. This is the fitter case.
A yard that is already paying well can afford to train. Does training raise pay, or does pay fund training?
Towns with more crime hire more police, and more police may reduce crime. Both are decided at once, by each other.
A random process is one whose individual outcome you cannot predict but whose pattern over many repetitions you can. Toss a coin. Roll a die. Knock on a randomly chosen door in Falmouth and ask the household’s income. In each case the next answer is unknown; the distribution of answers is not.
Three pieces of vocabulary, which the exam will expect you to use precisely.
The set of every outcome the process could produce. For two coins: {HH, HT, TH, TT}. Also called the population.
One individual outcome in that set. HT is a sample point; so is TH, and they are different ones.
Any subset of the sample space. “Exactly one Head” is the event {HT, TH} — two sample points, one event.
A random variable is a number whose value is settled by the outcome of a chance experiment. Write x = 1 if the first coin lands Heads and x = 0 otherwise, and x is a random variable: discrete, because it takes values from a list you could write out. Household income is continuous — between any two incomes there is another — and the formulas pick up integrals where the discrete ones have sums, but nothing conceptual changes.
P(A), the probability of an event, is the proportion of times A occurs in repeated trials — its relative frequency in the long run. Toss a fair coin ten thousand times and the share of Heads will be very close to one half; that is what the number 0.5 means here.
Three rules, and everything else follows from them.
1 · Bounded. 0 ≤ P(A) ≤ 1 for every event A. Nothing is less likely than impossible or more likely than certain.
2 · Exhaustive events sum to one. If A, B, C… between them cover every outcome, then P(A + B + C + …) = 1. Something must happen.
3 · Mutually exclusive events add. If no two can occur together, P(A + B + …) = P(A) + P(B) + …
Apply them to two fair coins. “Exactly one Head” is {HT, TH}; the four sample points are equally likely and mutually exclusive, so the probability is 0.25 + 0.25 = 0.5. Students who answer 1⁄3 have counted outcomes as {two Heads, one Head, no Heads} and assumed the three are equally likely. They are not: one Head happens two ways, the others one way each.
Toss two coins. Let x = 1 if the first is Heads, y = 1 if the second is. Every probability you could want about the pair lives in one small table.
f(x, y) = P(x and y) sits in the body of the table. Here every cell is 0.25. For continuous variables the same object is integrated rather than read off:
Sum a row or a column and you get the margin — the distribution of one variable on its own, with the other summed away. f(x) = Σy f(x, y), so f(1) = 0.25 + 0.25 = 0.5. The name is literal: it is written in the margin of the table.
Conditioning means restricting attention to one row or column, then rescaling it so it sums to one again:
So P(y = 1 | x = 1) = 0.25 ÷ 0.50 = 0.5. The denominator is doing the rescaling: having been told the first coin was Heads, you are in a world half the size, and the probabilities inside it must still add to one.
Here the answer came back as 0.5 — exactly what P(y = 1) was before you were told anything. The first coin carried no information about the second. That is independence, and the formal condition is
A whole distribution is more than you usually need. Moments are the summary.
The expected value is the value repeated draws average towards:
For one coin, E(x) = 1(0.5) + 0(0.5) = 0.5. Note that the expected value is not a value you expect: a single toss never returns 0.5. It is the centre of gravity of the distribution.
The variance is the average squared distance from the mean:
For the coin: 0.5(0 − 0.5)2 + 0.5(1 − 0.5)2 = 0.25, so σx = 0.5. The squaring is not arbitrary — the raw deviations average to zero by construction, so without it the measure would be zero for every distribution. The cost is that the variance is in squared units, which is why the standard deviation, its square root, is what gets quoted.
With two variables, a third question arises: does one tend to be above its own mean when the other is?
The second form is almost always the easier one to compute. For the two coins, E(xy) = 0.25 and E(x)E(y) = 0.25, so cov(x, y) = 0.
Take a bag with three orange marbles and three white. Draw two, and do it in two different ways.
Put the first back and the bag resets: the second draw faces 3 orange out of 6 regardless of what came first, so P(2nd orange | 1st orange) = 0.5 = P(2nd orange). The draws are independent and identically distributed — iid — which is exactly the coin-toss world.
Keep the first out and it does not: an orange first draw leaves 2 orange among 5, so P(2nd orange | 1st orange) = 0.4, while a white first draw leaves 3 among 5 and gives 0.6. Conditioning now moves the answer, so the draws are dependent.
Notice what did not change. The unconditional P(2nd is orange) is 0.5 under both schemes. Dependence is not visible in the margin; it only shows up when you condition.
Essentially every result in this unit assumes an iid random sample. Yet nobody runs a survey that way. A household, once interviewed, is struck off the list for that wave — which is the keep-it-out scheme, not the put-it-back one. Strictly, then, your data fall on the dependent side.
It is tolerated because the population is enormous. Removing one household from twenty-eight million barely moves the proportions, where removing one marble from six moves them from 0.5 to 0.4. The approximation is excellent when the sample is a tiny fraction of the population, and it degrades when it is not.
And sometimes the dependence is the economics. This year’s income depends on last year’s. One yard’s order book depends on its neighbour’s. Those cases are not nuisances to be assumed away — they are modelled, with covariance and conditional expectation as the tools.
The conditional expectation is just the mean of the conditional distribution:
Back to the bag, drawing without replacement. After an orange first draw, E(y | x = 1) = 0.4. After a white one, E(y | x = 0) = 0.6.
This is almost always what an economist actually wants. Not average pay, but average pay given a level of schooling. Not average consumption, but consumption given income. A regression line is nothing more or less than an estimate of a conditional expectation function, which is why this idea is the hinge of the whole unit.
If E(y | X) is a random variable, it has an expectation of its own. Take it:
Weight each conditional mean by the probability of the condition it is conditioned on, add them up, and the unconditional mean reappears. In general:
Look at what it is saying. Neither conditional mean equals 0.5 — one is below, one above — and yet the weighted average of them is exactly 0.5. The second marble is still equally likely to be orange overall, even though every piece of information you could receive about the first draw moves the answer away from a half.
Everything so far has been about a population. Now the awkward part: you do not have one. You have a sample, drawn once, and you have to say something about the population anyway.
The population has a fixed, unknown mean µ = E(x). For a fair coin it happens to be 0.5, which is the one case where you can check your working. From an iid sample {x1…xn} you compute x̄ = (1/n)Σxi.
The estimator X̄ = (1/n)ΣXi is a formula — a recipe you could apply to any sample. Before the coins are tossed it is a random variable with a distribution of its own, and that distribution is called the sampling distribution. The estimate x̄ is the single number that falls out once the tossing is done.
This is why different groups in a lecture theatre, all tossing the same fair coin, report different averages. Nobody did anything wrong. x̄ is computed from a random sample, so x̄ is itself random, and the variation between groups is the sampling distribution revealing itself.
Two things can be proved about X̄ without knowing anything much about the population at all.
Each step is small: expectation is linear, so it passes through the sum and the constant; each Xi is drawn from the same population, so each has expectation µ; there are n of them. When E(θ̂) = θ, the estimator is unbiased.
The second step is the one doing the work, and it is only available because the draws are independent: variances of independent variables add, and the constant comes out squared. Drop independence and a covariance term appears that does not vanish. This is the first place the iid assumption pays for itself, and it will not be the last.
At n = 1 the standard deviation is 0.500, at n = 10 it is 0.158, at n = 50 it is 0.071. All three are centred on 0.5. The centre is a property of the formula; the width is what more data buys.
Draw the samples yourself. The first view builds the sampling distribution of x̄ one batch at a time and lays the Normal curve over it; the second follows a handful of single runs and watches them settle.
If var(X̄) = σ2/n, then letting n grow without bound sends the variance to zero. The distribution collapses onto a point, and the point is µ.
Written more compactly, plim X̄n = µ, and an estimator with this property is consistent. Pick any tolerance you like, however tight, and a large enough sample will land inside it with probability approaching one.
The running proportion wanders early — after five tosses it can be anywhere — and then settles. The dashed envelope is ±2σ/√t, which for a fair coin is simply ±1/√t.
Unbiasedness is about the centre of the sampling distribution at a fixed n. Consistency is about what happens to the whole distribution as n grows. An estimator can be biased in every finite sample and still consistent, if the bias shrinks away; it can be unbiased and inconsistent, if the variance refuses to. They are worth keeping apart, and the exam does separate them.
The LLN says the distribution of X̄ collapses to a point. That is almost too much information — in the limit there is no shape left to describe. Rescale so the spread neither vanishes nor explodes:
The √n is exactly what offsets the 1/√n shrinkage. And then the remarkable part:
Look at what the population was. A single toss takes two values, 0 and 1, with nothing between them — about as far from a bell as a distribution can get. Average fifty of them and the result tracks a Normal curve closely enough that the difference is hard to draw.
Equivalently, X̄ ≈ N(µ, σ2/n), and this holds without knowing the population distribution. That last clause is the one that earns the theorem its place: you need µ and σ2 to exist, and essentially nothing else.
X ~ N(µ, σ2) has density
with µ = 0, σ = 1 giving the standard Normal. Roughly 68%, 95% and 99.7% of the mass lies within one, two and three standard deviations — 68.3%, 95.4% and 99.7% to be exact.
This is why the Normal turns up relentlessly from here on. Not because economic data are Normal — wages and firm sizes are conspicuously not — but because the things econometrics computes are averages, and averages are Normal almost regardless of what they average.
For an iid sample, the sample mean has three properties, and between them they are most of why it is used:
E(X̄) = µ. Across repeated samples it is centred on the truth, at every sample size.
plim X̄ = µ. As n grows it collapses onto the truth. This is the LLN.
√n(X̄ − µ)/σ → N(0, 1). Its shape is known even when the population’s is not. This is the CLT.
Three statements about an estimator, none of which required you to know the shape of F(·) at all. For a subject that spends most of its time worrying about what it cannot observe, that is a considerable amount of free information.
Suppose a group tosses a coin and reports x̄ = 0.3. Should they doubt the coin? A gap of 0.2 proves nothing by itself: sampling variation guarantees that most groups land somewhere other than 0.5. What has to be settled is whether a gap that size is improbable enough to be worth explaining — and that cannot be settled until you know how many tosses produced it.
From ten tosses, three or fewer Heads happens 17.2% of the time with a perfectly fair coin — about one group in six, in any lecture theatre. From a hundred tosses, thirty or fewer happens 0.0039% of the time, roughly one in twenty-five thousand. Same x̄. Verdicts four thousand times apart.
Deciding where to draw the line between “unlucky” and “the coin is not fair” is hypothesis testing, and it is where Week 2 begins. Everything needed for it is now in place: a population parameter, an estimator, and a known distribution for that estimator under an assumption about the world.
Tap card to reveal answer