Advanced Econometrics 1
Complete Course Summary
Weekly Lecture Notes, Tutorial Solutions & Step-by-Step Guides
Estimation Theory, Asymptotics, GLS & Heteroskedasticity, GMM,
Instrumental Variables, Panel Data, Maximum Likelihood, Time Series & More
Part I of the Course Summary
,Advanced Econometrics 1 — Complete Course Summary 1
Contents
1 Week 1: Econometric Tools and Linear Regression 3
1.1 Conditioning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
1.2 Regressions and Loss Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
1.3 Best Linear Prediction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
2 Week 1 (continued): Ordinary Least Squares 4
2.1 OLS Estimator and Assumptions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
2.2 Asymptotic Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
2.3 Heteroskedasticity and GLS . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
3 Week 1 (continued): Convergence, LLN, CLT & the Delta Method 6
3.1 Modes of Convergence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
3.2 Law of Large Numbers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
3.3 Central Limit Theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
3.4 Transformation Theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
3.5 The Delta Method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
4 Tutorial Week 1–2 7
5 Week 3–4: Instrumental Variables 10
5.1 Exogeneity, Endogeneity and Inconsistency of OLS . . . . . . . . . . . . . . . . . . . . . . . . 10
5.2 Instrumental Variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
5.3 Two-Stage Least Squares (2SLS) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
5.4 Indirect Least Squares and LIML . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
5.5 Testing Instrument Validity and Relevance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
6 Tutorial Week 3–4 12
7 Week 5–6: Non-Linear Models 14
7.1 Extremum Estimator Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
7.2 Nonlinear Least Squares . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
7.3 Maximum Likelihood . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
7.4 Quasi-Maximum Likelihood . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17
8 Tutorial Week 5–6 18
9 Additional Worked Exam Exercises: Weeks 1–6 19
10 Week 7–8: Generalized Method of Moments 22
10.1 From Method of Moments to GMM . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
10.2 Consistency and Asymptotic Normality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
10.3 Variance Estimation and Two-Step GMM . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23
10.4 Testing Overidentifying Restrictions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24
11 Tutorial Week 7–8 24
12 Multivariate GMM and Systems of Equations 25
12.1 Matrix Algebra: the Kronecker Product . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
12.2 The Matrix Normal Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
12.3 Multivariate (Seemingly Unrelated) Regression via GMM . . . . . . . . . . . . . . . . . . . . 25
12.4 Systems Estimation: SUR . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
12.5 Three-Stage Least Squares (3SLS) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26
,Advanced Econometrics 1 — Complete Course Summary 2
13 Week 9–10: Hypothesis Testing 26
13.1 Basic Concepts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26
13.2 Type I/II Errors, Size, and Power . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26
13.3 Quadratic Forms in Normal Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
13.4 The Wald Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
13.5 Nonlinear Hypotheses: the Delta Method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
13.6 The Likelihood Ratio (LR) Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
13.7 The Lagrange Multiplier (Score) Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
14 Tutorial Week 9–10 28
14.1 The Neyman–Pearson Lemma and UMP Tests . . . . . . . . . . . . . . . . . . . . . . . . . . 29
14.2 Likelihood-Based Tests: the Trinity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30
14.3 Asymptotic Local Power . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31
14.4 Multiple Testing and Confidence Intervals . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31
14.5 Specification (M-)Testing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32
15 Week 11: Nonparametric Density Estimation 33
15.1 From Histograms to Kernel Density Estimators . . . . . . . . . . . . . . . . . . . . . . . . . . 33
15.2 Properties of the Kernel Density Estimator . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33
15.3 Optimal Bandwidth Selection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34
15.4 Multivariate Kernel Density Estimation and Confidence Intervals . . . . . . . . . . . . . . . . 34
15.5 Nonparametric (Kernel) Regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35
16 Additional Worked Exercises: GMM, Hypothesis Testing, and Nonparametrics (Old
Exam Questions) 35
17 Tutorial Week 11 36
18 Week 12: Nonparametric Regression 37
18.1 From the Regressogram to Kernel Regression . . . . . . . . . . . . . . . . . . . . . . . . . . . 37
18.2 The Nadaraya–Watson Estimator . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37
18.3 Properties . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
19 Further Worked Exam Exercises: GMM, Bandwidth Selection, and the EDF 38
, Advanced Econometrics 1 — Complete Course Summary 3
1 Week 1: Econometric Tools and Linear Regression
1.1 Conditioning
Conditioning is important in econometrics — e.g. what is the variance today, given yesterday? Remember
that an assumption of the classical linear regression model is that X should be fixed, therefore we condition
on X.
Some important formulas
Marginal density: f (y) = f (x, y) dx
R R
or f (x) = f (x, y) dy
f (y, x) f (y, x)
Conditional density: f (y | x) = =R
f (x) f (x, y) dy
Conditional expectation: E[y | x] = y f (y | x) dy
R
Conditional variance: Var[y | x] = E (y − E[y | x])2 | x
Law of iterated expectations: E[y] = Ex Ey|x [y | x]
Marginal variance: Var(y) = E[Var(y | x)] + Var(E[y | x])
Unconditional moment conditions (GMM): E[u | z] = 0 ⇒ E[uz] = 0 ⇔ E[(y − x′ β)z] = 0
1.2 Regressions and Loss Functions
Real value: y = x′ β + ε; predictor: ŷ = x′ β̂. Residuals: e = y − ŷ. Expected loss: E[L(y − ŷ) | x].
Loss function L(e) Optimal ŷ
Squared error e2 ŷ = E[y | x]
Absolute error |e| ŷ = med(y | x)
Asymmetric absolute error αe+ + (1 − α)e− ŷ = qα (y | x)
Step loss ⊮(|e| > δ) ŷ = mode(y | x)
Proof: optimal ŷ for squared error is ŷ = E[y | x]
Define g(x) = E[y | x] and u = y − g(x). Then
L(e) = e2 = (y − ŷ)2 = (u + g(x) − ŷ)2 = u2 + 2u(g(x) − ŷ) + (g(x) − ŷ)2 .
Taking the conditional expectation:
E[(y − ŷ)2 | x] = E[u2 | x] +2(g(x) − ŷ) E[u | x] +(g(x) − ŷ)2 = σ 2 + (g(x) − ŷ)2 ,
| {z } | {z }
=σ 2 =0
using that functions of x can be taken out of the conditional expectation, and E[u | x] = 0 by definition of g.
This does not depend on the choice of ŷ except through the last term, so we minimize (g(x) − ŷ)2 , which is
minimized when
ŷ = g(x) = E[y | x]. ■
Proof: optimal ŷ for (mean) absolute error is ŷ = med(y | x)
L(e) = |e| = |y − ŷ|, so
Z ∞ Z ŷ
E[|y − ŷ| | x] = (y − ŷ)f (y | x) dy + (ŷ − y)f (y | x) dy.
ŷ −∞
Taking the derivative with respect to ŷ and setting it to zero:
∂ E[|y − ŷ| | x]
= − Pr(y ≥ ŷ | x) + Pr(y ≤ ŷ | x) = 0.
∂ ŷ
Since Pr(y ≤ ŷ | x) + Pr(y ≥ ŷ | x) = 1, this gives 2 Pr(y ≤ ŷ | x) = 1, i.e. Pr(y ≤ ŷ | x) = 21 . Hence
ŷ = med(y | x). ■
Complete Course Summary
Weekly Lecture Notes, Tutorial Solutions & Step-by-Step Guides
Estimation Theory, Asymptotics, GLS & Heteroskedasticity, GMM,
Instrumental Variables, Panel Data, Maximum Likelihood, Time Series & More
Part I of the Course Summary
,Advanced Econometrics 1 — Complete Course Summary 1
Contents
1 Week 1: Econometric Tools and Linear Regression 3
1.1 Conditioning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
1.2 Regressions and Loss Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
1.3 Best Linear Prediction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
2 Week 1 (continued): Ordinary Least Squares 4
2.1 OLS Estimator and Assumptions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
2.2 Asymptotic Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
2.3 Heteroskedasticity and GLS . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
3 Week 1 (continued): Convergence, LLN, CLT & the Delta Method 6
3.1 Modes of Convergence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
3.2 Law of Large Numbers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
3.3 Central Limit Theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
3.4 Transformation Theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
3.5 The Delta Method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
4 Tutorial Week 1–2 7
5 Week 3–4: Instrumental Variables 10
5.1 Exogeneity, Endogeneity and Inconsistency of OLS . . . . . . . . . . . . . . . . . . . . . . . . 10
5.2 Instrumental Variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
5.3 Two-Stage Least Squares (2SLS) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
5.4 Indirect Least Squares and LIML . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
5.5 Testing Instrument Validity and Relevance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
6 Tutorial Week 3–4 12
7 Week 5–6: Non-Linear Models 14
7.1 Extremum Estimator Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
7.2 Nonlinear Least Squares . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
7.3 Maximum Likelihood . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
7.4 Quasi-Maximum Likelihood . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17
8 Tutorial Week 5–6 18
9 Additional Worked Exam Exercises: Weeks 1–6 19
10 Week 7–8: Generalized Method of Moments 22
10.1 From Method of Moments to GMM . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
10.2 Consistency and Asymptotic Normality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
10.3 Variance Estimation and Two-Step GMM . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23
10.4 Testing Overidentifying Restrictions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24
11 Tutorial Week 7–8 24
12 Multivariate GMM and Systems of Equations 25
12.1 Matrix Algebra: the Kronecker Product . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
12.2 The Matrix Normal Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
12.3 Multivariate (Seemingly Unrelated) Regression via GMM . . . . . . . . . . . . . . . . . . . . 25
12.4 Systems Estimation: SUR . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
12.5 Three-Stage Least Squares (3SLS) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26
,Advanced Econometrics 1 — Complete Course Summary 2
13 Week 9–10: Hypothesis Testing 26
13.1 Basic Concepts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26
13.2 Type I/II Errors, Size, and Power . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26
13.3 Quadratic Forms in Normal Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
13.4 The Wald Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
13.5 Nonlinear Hypotheses: the Delta Method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
13.6 The Likelihood Ratio (LR) Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
13.7 The Lagrange Multiplier (Score) Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
14 Tutorial Week 9–10 28
14.1 The Neyman–Pearson Lemma and UMP Tests . . . . . . . . . . . . . . . . . . . . . . . . . . 29
14.2 Likelihood-Based Tests: the Trinity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30
14.3 Asymptotic Local Power . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31
14.4 Multiple Testing and Confidence Intervals . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31
14.5 Specification (M-)Testing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32
15 Week 11: Nonparametric Density Estimation 33
15.1 From Histograms to Kernel Density Estimators . . . . . . . . . . . . . . . . . . . . . . . . . . 33
15.2 Properties of the Kernel Density Estimator . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33
15.3 Optimal Bandwidth Selection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34
15.4 Multivariate Kernel Density Estimation and Confidence Intervals . . . . . . . . . . . . . . . . 34
15.5 Nonparametric (Kernel) Regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35
16 Additional Worked Exercises: GMM, Hypothesis Testing, and Nonparametrics (Old
Exam Questions) 35
17 Tutorial Week 11 36
18 Week 12: Nonparametric Regression 37
18.1 From the Regressogram to Kernel Regression . . . . . . . . . . . . . . . . . . . . . . . . . . . 37
18.2 The Nadaraya–Watson Estimator . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37
18.3 Properties . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
19 Further Worked Exam Exercises: GMM, Bandwidth Selection, and the EDF 38
, Advanced Econometrics 1 — Complete Course Summary 3
1 Week 1: Econometric Tools and Linear Regression
1.1 Conditioning
Conditioning is important in econometrics — e.g. what is the variance today, given yesterday? Remember
that an assumption of the classical linear regression model is that X should be fixed, therefore we condition
on X.
Some important formulas
Marginal density: f (y) = f (x, y) dx
R R
or f (x) = f (x, y) dy
f (y, x) f (y, x)
Conditional density: f (y | x) = =R
f (x) f (x, y) dy
Conditional expectation: E[y | x] = y f (y | x) dy
R
Conditional variance: Var[y | x] = E (y − E[y | x])2 | x
Law of iterated expectations: E[y] = Ex Ey|x [y | x]
Marginal variance: Var(y) = E[Var(y | x)] + Var(E[y | x])
Unconditional moment conditions (GMM): E[u | z] = 0 ⇒ E[uz] = 0 ⇔ E[(y − x′ β)z] = 0
1.2 Regressions and Loss Functions
Real value: y = x′ β + ε; predictor: ŷ = x′ β̂. Residuals: e = y − ŷ. Expected loss: E[L(y − ŷ) | x].
Loss function L(e) Optimal ŷ
Squared error e2 ŷ = E[y | x]
Absolute error |e| ŷ = med(y | x)
Asymmetric absolute error αe+ + (1 − α)e− ŷ = qα (y | x)
Step loss ⊮(|e| > δ) ŷ = mode(y | x)
Proof: optimal ŷ for squared error is ŷ = E[y | x]
Define g(x) = E[y | x] and u = y − g(x). Then
L(e) = e2 = (y − ŷ)2 = (u + g(x) − ŷ)2 = u2 + 2u(g(x) − ŷ) + (g(x) − ŷ)2 .
Taking the conditional expectation:
E[(y − ŷ)2 | x] = E[u2 | x] +2(g(x) − ŷ) E[u | x] +(g(x) − ŷ)2 = σ 2 + (g(x) − ŷ)2 ,
| {z } | {z }
=σ 2 =0
using that functions of x can be taken out of the conditional expectation, and E[u | x] = 0 by definition of g.
This does not depend on the choice of ŷ except through the last term, so we minimize (g(x) − ŷ)2 , which is
minimized when
ŷ = g(x) = E[y | x]. ■
Proof: optimal ŷ for (mean) absolute error is ŷ = med(y | x)
L(e) = |e| = |y − ŷ|, so
Z ∞ Z ŷ
E[|y − ŷ| | x] = (y − ŷ)f (y | x) dy + (ŷ − y)f (y | x) dy.
ŷ −∞
Taking the derivative with respect to ŷ and setting it to zero:
∂ E[|y − ŷ| | x]
= − Pr(y ≥ ŷ | x) + Pr(y ≤ ŷ | x) = 0.
∂ ŷ
Since Pr(y ≤ ŷ | x) + Pr(y ≥ ŷ | x) = 1, this gives 2 Pr(y ≤ ŷ | x) = 1, i.e. Pr(y ≤ ŷ | x) = 21 . Hence
ŷ = med(y | x). ■