Independent revision built for the Bristol first- and second-year Economics, Accounting & Finance syllabi. Not affiliated with or endorsed by the University of Bristol.

What Econometrics Is, and the Probability It Runs On

Two strands that meet in Week 2: what a causal question actually asks and why observational data cannot answer it on its own, and the probability machinery — distributions, moments, and the behaviour of an average — that every later method is built from.

What Econometrics Is For

Framing the unit

Three questions, all about the same industry, all answerable with data, and all asking for something different.

Descriptive

Has the share of Bristol economics graduates still working in the South West five years on risen since 2010?

Causal

Does a longer apprenticeship actually raise a marine fitter’s hourly pay — or do the people who take one simply earn more anyway?

Forecasting

What will Cornish visitor spending be next August?

The first is a matter of counting carefully. The third asks for a number you will be judged on when it arrives. The second is the hard one, and it is where most of this unit goes: it asks what would have happened to the same person under a different choice, which is not something any dataset contains.

Underneath all three sits one idea. The data you have are treated as a realisation — one draw — from an underlying probabilistic model of how the world generates outcomes. That model is the population; your spreadsheet is a sample from it. Econometrics is the business of reasoning backwards from the sample to the model, and being honest about how far that reasoning carries.

Why this unit spends its first week on probability. If data are a draw from a random process, then anything you calculate from them — an average, a regression coefficient, a forecast — is itself random. Calculate it again on a different sample and you get a different number. Everything econometrics can tell you depends on knowing how much that number moves about, which is a question about distributions, not about economics. Hence a week of coins.

From a Question to an Estimate

The shape of empirical work

Empirical economics has a standard shape, and it is worth seeing it whole before meeting any of the pieces.

Figure 1 — the five steps, and where econometrics begins

Step 2 is economics. You decide, from theory or from plain reasoning, which things plausibly bear on the outcome. Sometimes that comes from a formal model — a utility-maximisation problem, say, whose solution is a demand equation. More often it comes from knowing the industry: a fitter’s pay depends on training, schooling and years on the tools, and you do not need a Lagrangian to say so.

Step 3 is where econometrics starts, and it demands two commitments the economic model never had to make. You must choose a functional form — linear, in this unit, almost always — and you must decide what to do about everything you left out.

Economic model, econometric model

The economic model says only that pay depends on training, schooling and experience, without saying how. The econometric model commits:

pay = β0 + β1 train + β2 educ + β3 exper + u

That is a much stronger statement, and a much more useful one: it has named quantities in it that data can pin down.

Parameters, and the Error Term

What the symbols are

The βs are the parameters. They are fixed, unknown numbers describing the population — not things you compute, things you estimate. β1 is the one the question was about: how much an extra week of training adds to hourly pay, holding schooling and experience fixed.

Everything else that affects pay and is not on the list goes into u, the error term or disturbance. For a fitter that might be aptitude, how hard they are willing to work, whether the yard they joined happened to be well run, and the ordinary noise of wage-setting. You cannot list it all, and you certainly cannot measure it all.

The error term is not a nuisance you tidy up at the end — it is the subject. Almost every technique in this unit exists because of something that might be hiding in u. Whether β1 can be estimated honestly turns entirely on how u relates to train. If the people who choose more training also happen to have more of whatever is in u, the estimate is measuring both at once and there is no way to tell them apart from the data alone. Sections 6 and 7 are about exactly this.

Hypotheses are statements about parameters

Once the model is written down, a question becomes arithmetic on its parameters. “Training does nothing for pay” is H0: β1 = 0. “An extra week is worth at least 20p an hour” is β1 ≥ 0.20. Translating an argument into a restriction on parameters, and then asking whether the data can reject it, is most of what the exam will ask you to do.

Four Shapes of Data

Structure

Before any method, a practical question: what does one row of your data represent? There are four common answers, and which one you have determines what you are allowed to do.

Figure 2 — what identifies a row, in each of the four structures

Cross-sectional

Many units, one moment. A row is a fitter. The ordering of the rows carries no information whatsoever — shuffle them and nothing is lost — and if the sample was drawn at random from the population, that randomness is what makes the standard results work.

Time series

One unit, many moments. A row is a year. Now the ordering is the information, and the comfortable assumption of independence is gone: this year’s average wage in a yard is obviously related to last year’s. That dependence is why time series needs its own machinery, and why this unit leaves it until much later.

Pooled cross-section

Several independent cross-sections stacked up. Survey 300 fitters in 2022, survey 300 different fitters in 2024, and put the two together with a year variable to tell them apart. Nobody appears twice. It is a bigger sample, and — more usefully — it straddles whatever changed in between.

Panel

The same units, followed over time. Fitter 1 in 2022 and fitter 1 in 2024 are both in the data. This is the structurally richest of the four, because you can watch a single person change, which means the things about that person that did not change can be swept out of the problem entirely.

Panel and pooled cross-section look identical in a spreadsheet. Both have a unit column and a year column; both put 2022 and 2024 rows together. The difference is whether the unit identifiers repeat. If fitter 1 appears in both years, it is a panel and you can difference within a person. If every identifier is unique, it is pooled and you cannot. Checking which you have is a thirty-second job that determines half of what follows.

Experimental and Observational Data

Where the data came from

A second question, and a more awkward one: who decided the values of the variable you care about?

In an experiment, you did. You choose which plots get the new feed mix, which fitters get the extra twelve weeks, and you choose by a method unrelated to anything about them — a coin, in effect. In observational data — also called non-experimental, or retrospective — the subjects chose, or circumstance chose, and you merely recorded what happened. You were a passive collector.

Nearly all economic data are observational, and not by oversight. You cannot randomly assign people to leave school at sixteen. You cannot randomly set the Bank of England’s base rate to different levels in different years to see what each does. Some of it is prohibitively expensive, some is impossible, and a good deal of it would be straightforwardly wrong.

This is the single fact that shapes the discipline. Statistics developed largely around designed experiments, where randomisation does the hard work for you. Econometrics is what you are left with when randomisation is unavailable: a set of methods for getting at causal effects from data that nobody randomised. Instrumental variables, differences-in-differences and regression discontinuity — the back half of this unit — are all attempts to find, or manufacture, a source of variation that behaves as if someone had tossed a coin.

Ceteris Paribus and the Counterfactual

What a causal effect is

“Other relevant factors being equal.” You met the phrase in first year as a modelling convenience. It is in fact the definition of the thing you are trying to measure.

Ask what a longer apprenticeship is worth. The honest version of the question is: take this fitter, and compare what they earn having done the extra twelve weeks with what the same fitter would have earned having not done them. Not a different fitter. The same one, in a different state of the world.

Figure 3 — the comparison the question asks for

Those two numbers are the potential outcomes, or counterfactual outcomes. The causal effect for this person is the difference between them. Stated this way, the ceteris paribus condition is free: everything else about the fitter is held fixed automatically, because it is the same fitter either way.

The fundamental problem. Exactly one of those two columns is ever observed. The fitter either did the extra weeks or did not; the other number is not missing from your data through carelessness, it never existed. Every technique in this unit is a strategy for standing in for a quantity that is unobservable in principle — which is why “just collect more data” is never the answer to a causal question.

It will not always be possible to hold all else equal, and you should be suspicious of anyone who claims otherwise. The practical question in applied work is never “is this perfect?” but “have enough of the right things been held fixed for this comparison to be worth believing?” That is the question a referee asks, and it is the question an exam answer should address.

Why Observational Data Resists a Causal Reading

The central difficulty

Here is the obvious thing to do, and why it does not work. Take your sample of fitters, average the pay of those who did the longer apprenticeship, average the pay of those who did not, and subtract.

Figure 4 — two arrows, measured as one

The trouble is that nobody assigned the training. Fitters chose it. And the kind of person who signs up for twelve more weeks of unpaid study is, on average, also the kind of person who turns up early, takes the awkward jobs and gets promoted — none of which you measured. Those traits raise pay on their own. They also made the training more likely.

So the raw gap contains the effect of training plus the effect of whatever made people select into it. The two arrive welded together, and no amount of arithmetic on those two averages can separate them.

Three ways the same problem shows up

Unobserved traits

Aptitude and drive affect both the choice and the outcome, and appear in no dataset. This is the fitter case.

Reverse causation

A yard that is already paying well can afford to train. Does training raise pay, or does pay fund training?

Simultaneity

Towns with more crime hire more police, and more police may reduce crime. Both are decided at once, by each other.

Correlation is not nothing. It is a fact about the data, and it is often the first thing worth knowing. What it is not is an answer to a ceteris paribus question, and the gap between the two is the whole reason this unit runs for eleven weeks rather than two. The rest of Week 1 builds the probability tools; from Week 2 onwards they get pointed at this problem.

Random Variables and the Sample Space

The vocabulary

A random process is one whose individual outcome you cannot predict but whose pattern over many repetitions you can. Toss a coin. Roll a die. Knock on a randomly chosen door in Falmouth and ask the household’s income. In each case the next answer is unknown; the distribution of answers is not.

Three pieces of vocabulary, which the exam will expect you to use precisely.

Sample space

The set of every outcome the process could produce. For two coins: {HH, HT, TH, TT}. Also called the population.

Sample point

One individual outcome in that set. HT is a sample point; so is TH, and they are different ones.

Event

Any subset of the sample space. “Exactly one Head” is the event {HT, TH} — two sample points, one event.

A random variable is a number whose value is settled by the outcome of a chance experiment. Write x = 1 if the first coin lands Heads and x = 0 otherwise, and x is a random variable: discrete, because it takes values from a list you could write out. Household income is continuous — between any two incomes there is another — and the formulas pick up integrals where the discrete ones have sums, but nothing conceptual changes.

Capital letters and lower-case letters are doing different jobs, and the notation is not decoration. X is the random variable — the rule, before anything has happened. x is the realisation — the number you got. The same distinction returns later as X̄ against x̄, and the entire idea of a sampling distribution depends on holding the two apart.

Probability, and Three Rules

Foundations

P(A), the probability of an event, is the proportion of times A occurs in repeated trials — its relative frequency in the long run. Toss a fair coin ten thousand times and the share of Heads will be very close to one half; that is what the number 0.5 means here.

Three rules, and everything else follows from them.

The axioms

1 · Bounded. 0 ≤ P(A) ≤ 1 for every event A. Nothing is less likely than impossible or more likely than certain.

2 · Exhaustive events sum to one. If A, B, C… between them cover every outcome, then P(A + B + C + …) = 1. Something must happen.

3 · Mutually exclusive events add. If no two can occur together, P(A + B + …) = P(A) + P(B) + …

Apply them to two fair coins. “Exactly one Head” is {HT, TH}; the four sample points are equally likely and mutually exclusive, so the probability is 0.25 + 0.25 = 0.5. Students who answer 1⁄3 have counted outcomes as {two Heads, one Head, no Heads} and assumed the three are equally likely. They are not: one Head happens two ways, the others one way each.

The Joint Table: Marginals and Conditionals

Two variables at once

Toss two coins. Let x = 1 if the first is Heads, y = 1 if the second is. Every probability you could want about the pair lives in one small table.

Figure 5 — one table, read three ways

The joint pdf

f(x, y) = P(x and y) sits in the body of the table. Here every cell is 0.25. For continuous variables the same object is integrated rather than read off:

P(a ≤ X ≤ b, c ≤ Y ≤ d) = ∫ab ∫cd f(x, y) dx dy

The marginal pdf

Sum a row or a column and you get the margin — the distribution of one variable on its own, with the other summed away. f(x) = Σy f(x, y), so f(1) = 0.25 + 0.25 = 0.5. The name is literal: it is written in the margin of the table.

The conditional pdf

Conditioning means restricting attention to one row or column, then rescaling it so it sums to one again:

f(y | x) = f(x, y) ÷ f(x)

So P(y = 1 | x = 1) = 0.25 ÷ 0.50 = 0.5. The denominator is doing the rescaling: having been told the first coin was Heads, you are in a world half the size, and the probabilities inside it must still add to one.

Independence

Here the answer came back as 0.5 — exactly what P(y = 1) was before you were told anything. The first coin carried no information about the second. That is independence, and the formal condition is

f(x, y) = f(x) · f(y)   for every pair — here 0.25 = 0.5 × 0.5 ✓
Independence is a property of the whole table, not of one cell. Checking a single pair and finding the product works is not enough; the factorisation has to hold everywhere. This matters in the exam, where a table engineered to satisfy it in one cell and fail in another is a standard trap.

Mean, Variance, Covariance

Moments

A whole distribution is more than you usually need. Moments are the summary.

The first moment

The expected value is the value repeated draws average towards:

E(x) = Σx x f(x)  (discrete)    E(x) = ∫−∞∞ x f(x) dx  (continuous)

For one coin, E(x) = 1(0.5) + 0(0.5) = 0.5. Note that the expected value is not a value you expect: a single toss never returns 0.5. It is the centre of gravity of the distribution.

The second moment

The variance is the average squared distance from the mean:

var(x) = σx2 = E[(x − E(x))2] = Σx [x − E(x)]2 f(x)
Figure 6 — the mean, and what the variance is measuring

For the coin: 0.5(0 − 0.5)2 + 0.5(1 − 0.5)2 = 0.25, so σx = 0.5. The squaring is not arbitrary — the raw deviations average to zero by construction, so without it the measure would be zero for every distribution. The cost is that the variance is in squared units, which is why the standard deviation, its square root, is what gets quoted.

Covariance

With two variables, a third question arises: does one tend to be above its own mean when the other is?

cov(x, y) = E[(x − E(x))(y − E(y))] = E(xy) − E(x)E(y)

The second form is almost always the easier one to compute. For the two coins, E(xy) = 0.25 and E(x)E(y) = 0.25, so cov(x, y) = 0.

Independent implies zero covariance. The converse does not hold. Covariance detects linear co-movement only. Two variables can be tightly, deterministically related and still have covariance zero — take X symmetric about zero and Y = X2, where knowing X tells you Y exactly, yet cov(X, Y) = 0. “Uncorrelated” and “independent” are not synonyms, and the distinction does real work when regression assumptions are stated later.

With and Without Replacement

Where iid comes from

Take a bag with three orange marbles and three white. Draw two, and do it in two different ways.

Figure 7 — the same bag, two sampling schemes

Put the first back and the bag resets: the second draw faces 3 orange out of 6 regardless of what came first, so P(2nd orange | 1st orange) = 0.5 = P(2nd orange). The draws are independent and identically distributed — iid — which is exactly the coin-toss world.

Keep the first out and it does not: an orange first draw leaves 2 orange among 5, so P(2nd orange | 1st orange) = 0.4, while a white first draw leaves 3 among 5 and gives 0.6. Conditioning now moves the answer, so the draws are dependent.

Notice what did not change. The unconditional P(2nd is orange) is 0.5 under both schemes. Dependence is not visible in the margin; it only shows up when you condition.

Put a number on it. Without replacement the joint probabilities are 0.2, 0.3, 0.3 and 0.2 rather than 0.25 each, which makes E(xy) = 0.2 and so cov(x, y) = 0.2 − 0.25 = −0.05 — a correlation of −0.2. Negative, because taking an orange one out makes the next draw less likely to be orange. The coins gave exactly zero.

Why an economist should care

Essentially every result in this unit assumes an iid random sample. Yet nobody runs a survey that way. A household, once interviewed, is struck off the list for that wave — which is the keep-it-out scheme, not the put-it-back one. Strictly, then, your data fall on the dependent side.

It is tolerated because the population is enormous. Removing one household from twenty-eight million barely moves the proportions, where removing one marble from six moves them from 0.5 to 0.4. The approximation is excellent when the sample is a tiny fraction of the population, and it degrades when it is not.

And sometimes the dependence is the economics. This year’s income depends on last year’s. One yard’s order book depends on its neighbour’s. Those cases are not nuisances to be assumed away — they are modelled, with covariance and conditional expectation as the tools.

Conditional Expectation and the Law of Iterated Expectations

The bridge to regression

The conditional expectation is just the mean of the conditional distribution:

E(y | x) = Σy y f(y | x)  (discrete)    E(y | x) = ∫−∞∞ y f(y | x) dy  (continuous)

Back to the bag, drawing without replacement. After an orange first draw, E(y | x = 1) = 0.4. After a white one, E(y | x = 0) = 0.6.

This is almost always what an economist actually wants. Not average pay, but average pay given a level of schooling. Not average consumption, but consumption given income. A regression line is nothing more or less than an estimate of a conditional expectation function, which is why this idea is the hinge of the whole unit.

E(y | X) is itself a random variable. It equals 0.4 or 0.6 depending on how the first draw lands, and the first draw is random. That is a genuinely awkward idea on first meeting, and worth sitting with: conditioning on something random leaves you holding something random.

The Law of Iterated Expectations

If E(y | X) is a random variable, it has an expectation of its own. Take it:

Figure 8 — averaging the conditional means
E[E(y | X)] = ½(0.4) + ½(0.6) = 0.5 = E(y)

Weight each conditional mean by the probability of the condition it is conditioned on, add them up, and the unconditional mean reappears. In general:

E[E(y | X)] = E(y)    (the Law of Iterated Expectations)

Look at what it is saying. Neither conditional mean equals 0.5 — one is below, one above — and yet the weighted average of them is exactly 0.5. The second marble is still equally likely to be orange overall, even though every piece of information you could receive about the first draw moves the answer away from a half.

This identity is used constantly, and usually without announcement. It is the licence to move between conditional statements — which is what a regression gives you — and unconditional ones, which is usually what the policy question asked. When a later proof writes E[E(u | X)] = E(u) = 0 in a single step, this is what it is doing.

The Sampling Distribution

The central idea

Everything so far has been about a population. Now the awkward part: you do not have one. You have a sample, drawn once, and you have to say something about the population anyway.

Figure 9 — two arrows between the population and your data

The population has a fixed, unknown mean µ = E(x). For a fair coin it happens to be 0.5, which is the one case where you can check your working. From an iid sample {x1…xn} you compute x̄ = (1/n)Σxi.

The distinction that matters

The estimator X̄ = (1/n)ΣXi is a formula — a recipe you could apply to any sample. Before the coins are tossed it is a random variable with a distribution of its own, and that distribution is called the sampling distribution. The estimate x̄ is the single number that falls out once the tossing is done.

This is why different groups in a lecture theatre, all tossing the same fair coin, report different averages. Nobody did anything wrong. x̄ is computed from a random sample, so x̄ is itself random, and the variation between groups is the sampling distribution revealing itself.

You only ever see one draw from it. That is the whole predicament. The sampling distribution describes what would happen across the thousands of samples you did not collect, and yet every statement you will ever make — a standard error, a confidence interval, a p-value — is a statement about it. The next three sections are about what can be known without looking.

Unbiased, and Narrowing as 1/√n

Two properties, derived

Two things can be proved about X̄ without knowing anything much about the population at all.

It is centred on µ

E(X̄) = E[(1/n)ΣXi] = (1/n)ΣE(Xi) = (1/n)(nµ) = µ

Each step is small: expectation is linear, so it passes through the sum and the constant; each Xi is drawn from the same population, so each has expectation µ; there are n of them. When E(θ̂) = θ, the estimator is unbiased.

Unbiased does not mean right. It is a statement about the average of the estimator across every sample you might have drawn — not about the sample you actually have. Your particular x̄ can be well off µ and the estimator still be unbiased. What unbiasedness rules out is a systematic tilt in one direction.

It gets tighter with n

var(X̄) = var[(1/n)ΣXi] = (1/n2)Σvar(Xi) = (1/n2)(nσ2) = σ2/n

The second step is the one doing the work, and it is only available because the draws are independent: variances of independent variables add, and the constant comes out squared. Drop independence and a covariance term appears that does not vanish. This is the first place the iid assumption pays for itself, and it will not be the last.

Figure 10 — the exact distribution of x̄ at three sample sizes

At n = 1 the standard deviation is 0.500, at n = 10 it is 0.158, at n = 50 it is 0.071. All three are centred on 0.5. The centre is a property of the formula; the width is what more data buys.

Try it

Draw the samples yourself. The first view builds the sampling distribution of x̄ one batch at a time and lays the Normal curve over it; the second follows a handful of single runs and watches them settle.

Interactive — the sampling distribution, built one sample at a time
Tosses per sample n
Number of samples

The Law of Large Numbers

Taking n to the limit

If var(X̄) = σ2/n, then letting n grow without bound sends the variance to zero. The distribution collapses onto a point, and the point is µ.

P(|X̄n − µ| > ε) → 0 as n → ∞, for any tolerance ε > 0

Written more compactly, plim X̄n = µ, and an estimator with this property is consistent. Pick any tolerance you like, however tight, and a large enough sample will land inside it with probability approaching one.

Figure 11 — convergence, and what it costs

The running proportion wanders early — after five tosses it can be anywhere — and then settles. The dashed envelope is ±2σ/√t, which for a fair coin is simply ±1/√t.

Consistency is a promise about the limit, and the limit is a long way off. Because the standard error falls as √n rather than n, halving it costs four times the data. Going from a standard error of 0.01 to 0.001 means going from 2,500 observations to 250,000. The theorem guarantees you arrive; it says nothing about whether the sample you can afford is anywhere near large enough, and in applied work it usually is not.

Unbiased and consistent are different claims

Unbiasedness is about the centre of the sampling distribution at a fixed n. Consistency is about what happens to the whole distribution as n grows. An estimator can be biased in every finite sample and still consistent, if the bias shrinks away; it can be unbiased and inconsistent, if the variance refuses to. They are worth keeping apart, and the exam does separate them.

The Central Limit Theorem

Why the Normal is everywhere

The LLN says the distribution of X̄ collapses to a point. That is almost too much information — in the limit there is no shape left to describe. Rescale so the spread neither vanishes nor explodes:

Z = √n · (X̄ − µ)/σ,    E(Z) = 0,  var(Z) = 1

The √n is exactly what offsets the 1/√n shrinkage. And then the remarkable part:

√n · (X̄ − µ)/σ  →D  N(0, 1)    as n → ∞
Figure 12 — fifty coins, and a bell that has no business being there

Look at what the population was. A single toss takes two values, 0 and 1, with nothing between them — about as far from a bell as a distribution can get. Average fifty of them and the result tracks a Normal curve closely enough that the difference is hard to draw.

Equivalently, X̄ ≈ N(µ, σ2/n), and this holds without knowing the population distribution. That last clause is the one that earns the theorem its place: you need µ and σ2 to exist, and essentially nothing else.

The Normal distribution

X ~ N(µ, σ2) has density

f(x) = [1 ÷ (√(2π) σ)] · e−(x−µ)2/(2σ2)

with µ = 0, σ = 1 giving the standard Normal. Roughly 68%, 95% and 99.7% of the mass lies within one, two and three standard deviations — 68.3%, 95.4% and 99.7% to be exact.

The “95% rule” is not two standard deviations. Two gives 95.4%; the figure that gives exactly 95% is 1.96. For rough mental work two is fine, and for anything you write down 1.96 is the number — which is why it appears in every confidence interval you will construct this year.

This is why the Normal turns up relentlessly from here on. Not because economic data are Normal — wages and firm sizes are conspicuously not — but because the things econometrics computes are averages, and averages are Normal almost regardless of what they average.

Three Properties, and One Last Question

Pulling it together

For an iid sample, the sample mean has three properties, and between them they are most of why it is used:

Unbiased

E(X̄) = µ. Across repeated samples it is centred on the truth, at every sample size.

Consistent

plim X̄ = µ. As n grows it collapses onto the truth. This is the LLN.

Asymptotically Normal

√n(X̄ − µ)/σ → N(0, 1). Its shape is known even when the population’s is not. This is the CLT.

Three statements about an estimator, none of which required you to know the shape of F(·) at all. For a subject that spends most of its time worrying about what it cannot observe, that is a considerable amount of free information.

So: should the group doubt the coin?

Suppose a group tosses a coin and reports x̄ = 0.3. Should they doubt the coin? A gap of 0.2 proves nothing by itself: sampling variation guarantees that most groups land somewhere other than 0.5. What has to be settled is whether a gap that size is improbable enough to be worth explaining — and that cannot be settled until you know how many tosses produced it.

Figure 13 — the same result, against two sample sizes

From ten tosses, three or fewer Heads happens 17.2% of the time with a perfectly fair coin — about one group in six, in any lecture theatre. From a hundred tosses, thirty or fewer happens 0.0039% of the time, roughly one in twenty-five thousand. Same x̄. Verdicts four thousand times apart.

A note on using the Normal here. At n = 10 the CLT approximation gives 10.3% against the exact 17.2% — wrong by nearly seven percentage points. Applying the continuity correction, treating the discrete count as covering an interval, brings it to 17.1%, which is within a twentieth of a point. The CLT is asymptotic, and ten is not a large number; when the variable is discrete and n is small, the correction is not optional.

Deciding where to draw the line between “unlucky” and “the coin is not fair” is hypothesis testing, and it is where Week 2 begins. Everything needed for it is now in place: a population parameter, an estimator, and a known distribution for that estimator under an assumption about the world.

Flashcards

Tap card to reveal answer

Quick-reference note sheet
Your score
0 / 12
Self-check — fill in the figures
Type each value, then press Check answers. Cells turn green when correct; any you miss show the right figure beside them.