Maximum Likelihood Estimation

The normal regression model is the linear regression model with an independent normal error

\[\begin{equation} \begin{split} y &= \bx'\bbeta + \varepsilon \\ \varepsilon &\sim N(0, \sigma^2) \end{split}\label{eq-normal-regerssion} \end{equation}\]

Normal regression does NOT require joint normality of $(y, \bx)$. All that is required is that the conditional distribution of $y$ given $\bx$ is normal (the marginal distribution of $\bx$ is unrestricted).

The likelihood is the name for the joint probability density of the data, evaluated at the observed sample, and viewed as a function of the parameters. The maximum likelihood estimator is the value which maximizes the likelihood function.

Model $\eqref{eq-normal-regerssion}$ implies that the conditional distribution of $y$ given $\bx$ takes the form

\[f(y\mid \bx ; \bbeta, \sigma^2) = \frac{1}{\sqrt{2\pi\sigma^2}}\exp\left[-\frac{(y - \bx'\bbeta)^2}{2\sigma^2}\right]\]

Under the assumption of independent observations, the likelihood function is

\[\begin{equation} \begin{split} \green{L(\bbeta, \sigma^2)} &= f(y_1, \ldots, y_n \mid \bx_1, \ldots, \bx_n ; \bbeta, \sigma^2) \\ &= \prod_{i=1}^{n} f(y_i\mid \bx_i ; \bbeta, \sigma^2) \\ &= \frac{1}{(2\pi\sigma^2)^{n/2}} \exp\left[-\frac{1}{2\sigma^2}\sum_{i=1}^{n}(y_i - \bx_i'\bbeta)^2\right] \end{split} \end{equation}\]

For convenience, it is typical to work with the natural logarithm

\[\begin{equation} \begin{split} \log L(\bbeta, \sigma^2) &= \log f(y_1, \ldots, y_n \mid \bx_1, \ldots, \bx_n ; \bbeta, \sigma^2) \\ &= -\frac{n}{2}\log(2\pi\sigma^2) - \frac{1}{2\sigma^2}\sum_{i=1}^{n}(y_i - \bx_i'\bbeta)^2 \end{split} \label{eq-log-likelihood} \end{equation}\]

The maximum likelihood estimator is the value which maximizes the log-likelihood function

\[(\hat{\bbeta}, \hat{\sigma}^2) = \arg\max_{\bbeta, \sigma^2} \log L(\bbeta, \sigma^2)\]

In most applications of maximum likelihood, the MLE must be found by numerical methods. However, in the case of the normal regression model, the MLE can be found analytically by solving $\eqref{eq-mle-beta}$ and $\eqref{eq-mle-sigma}$.

\[\begin{equation} \begin{split} 0 = \frac{\partial \log L(\bbeta, \sigma^2)}{\partial \bbeta} &= \frac{1}{\sigma^2}\sum_{i=1}^{n}(y_i - \bx_i'\bbeta)\bx_i \\ &= \frac{1}{\sigma^2}\bX'(\by - \bX\bbeta) \end{split} \label{eq-mle-beta} \end{equation}\] \[\begin{equation} \begin{split} 0 = \frac{\partial \log L(\bbeta, \sigma^2)}{\partial \sigma^2} &= -\frac{n}{2\sigma^2} + \frac{1}{2\sigma^4}\sum_{i=1}^{n}(y_i - \bx_i'\bbeta)^2 \\ &= -\frac{n}{2\sigma^2} + \frac{1}{2\sigma^4}(\by - \bX\bbeta)'(\by - \bX\bbeta) \end{split} \label{eq-mle-sigma} \end{equation}\]

Plugging in $\hat{\sigma}^2=\frac{1}{n} \sum_{i=1}^n (y_i-\bx’\hat{\bbeta})^2$ to $\eqref{eq-log-likelihood},$ we obtain

\[\log L(\hat{\bbeta}, \hat{\sigma}^2) = -\frac{n}{2}\log(2\pi\hat{\sigma}^2) - \frac{n}{2}\]

The log-likelihood is typicalled reported as a measure of fit.


AIC and BIC

Alternative measures of fit include information criteria such as AIC and BIC.

  • Akaike Information Criterion (AIC)

    \[AIC(K) = \log \left(\frac{\be'\be}{n}\right) + \frac{2K}{n}\]
  • Schwarz or Bayesian Information Criterion (BIC)

    \[BIC(K) = \log \left(\frac{\be'\be}{n}\right) + \frac{K\log(n)}{n}\]

    BIC imposes a heavier penalty for degrees of freedom lost, hence it will lean toward a simpler model than AIC.

Choose the model with the smallest AIC or BIC. AIC and BIC leaves out the constant terms in the log-likelihood function because they shift every model by the same amount and do not affect the model selection.

Alternative expressions for AIC and BIC are:

\[\begin{split} AIC(K) &= n\log \left(\frac{\text{SSR}}{n}\right) + 2K \\ BIC(K) &= n\log \left(\frac{\text{SSR}}{n}\right) + K\log(n) \end{split}\]

$\be’\be=\text{SSR}$ sum of squared residuals.

For a fixed sample size $n$, the two expressions differ only by a constant. Both will lead to the same model selection.


Ref:

  • pp 147, Chap 5 Normal Regression and Maximum Likelihood, Hansen Econometrics