Maximum Likelihood
Maximum Likelihood Estimation
The normal regression model is the linear regression model with an independent normal error
\[\begin{equation} \begin{split} y &= \bx'\bbeta + \varepsilon \\ \varepsilon &\sim N(0, \sigma^2) \end{split}\label{eq-normal-regerssion} \end{equation}\]Normal regression does NOT require joint normality of $(y, \bx)$. All that is required is that the conditional distribution of $y$ given $\bx$ is normal (the marginal distribution of $\bx$ is unrestricted).
The likelihood is the name for the joint probability density of the data, evaluated at the observed sample, and viewed as a function of the parameters. The maximum likelihood estimator is the value which maximizes the likelihood function.
Model $\eqref{eq-normal-regerssion}$ implies that the conditional distribution of $y$ given $\bx$ takes the form
\[f(y\mid \bx ; \bbeta, \sigma^2) = \frac{1}{\sqrt{2\pi\sigma^2}}\exp\left[-\frac{(y - \bx'\bbeta)^2}{2\sigma^2}\right]\]Under the assumption of independent observations, the likelihood function is
\[\begin{equation} \begin{split} \green{L(\bbeta, \sigma^2)} &= f(y_1, \ldots, y_n \mid \bx_1, \ldots, \bx_n ; \bbeta, \sigma^2) \\ &= \prod_{i=1}^{n} f(y_i\mid \bx_i ; \bbeta, \sigma^2) \\ &= \frac{1}{(2\pi\sigma^2)^{n/2}} \exp\left[-\frac{1}{2\sigma^2}\sum_{i=1}^{n}(y_i - \bx_i'\bbeta)^2\right] \end{split} \end{equation}\]For convenience, it is typical to work with the natural logarithm
\[\begin{equation} \begin{split} \log L(\bbeta, \sigma^2) &= \log f(y_1, \ldots, y_n \mid \bx_1, \ldots, \bx_n ; \bbeta, \sigma^2) \\ &= -\frac{n}{2}\log(2\pi\sigma^2) - \frac{1}{2\sigma^2}\sum_{i=1}^{n}(y_i - \bx_i'\bbeta)^2 \end{split} \label{eq-log-likelihood} \end{equation}\]The maximum likelihood estimator is the value which maximizes the log-likelihood function
\[(\hat{\bbeta}, \hat{\sigma}^2) = \arg\max_{\bbeta, \sigma^2} \log L(\bbeta, \sigma^2)\]In most applications of maximum likelihood, the MLE must be found by numerical methods. However, in the case of the normal regression model, the MLE can be found analytically by solving $\eqref{eq-mle-beta}$ and $\eqref{eq-mle-sigma}$.
\[\begin{equation} \begin{split} 0 = \frac{\partial \log L(\bbeta, \sigma^2)}{\partial \bbeta} &= \frac{1}{\sigma^2}\sum_{i=1}^{n}(y_i - \bx_i'\bbeta)\bx_i \\ &= \frac{1}{\sigma^2}\bX'(\by - \bX\bbeta) \end{split} \label{eq-mle-beta} \end{equation}\] \[\begin{equation} \begin{split} 0 = \frac{\partial \log L(\bbeta, \sigma^2)}{\partial \sigma^2} &= -\frac{n}{2\sigma^2} + \frac{1}{2\sigma^4}\sum_{i=1}^{n}(y_i - \bx_i'\bbeta)^2 \\ &= -\frac{n}{2\sigma^2} + \frac{1}{2\sigma^4}(\by - \bX\bbeta)'(\by - \bX\bbeta) \end{split} \label{eq-mle-sigma} \end{equation}\]Plugging in $\hat{\sigma}^2=\frac{1}{n} \sum_{i=1}^n (y_i-\bx’\hat{\bbeta})^2$ to $\eqref{eq-log-likelihood},$ we obtain
\[\log L(\hat{\bbeta}, \hat{\sigma}^2) = -\frac{n}{2}\log(2\pi\hat{\sigma}^2) - \frac{n}{2}\]The log-likelihood is typicalled reported as a measure of fit.
AIC and BIC
Alternative measures of fit include information criteria such as AIC and BIC.
-
Akaike Information Criterion (AIC)
\[AIC(K) = \log \left(\frac{\be'\be}{n}\right) + \frac{2K}{n}\] -
Schwarz or Bayesian Information Criterion (BIC)
\[BIC(K) = \log \left(\frac{\be'\be}{n}\right) + \frac{K\log(n)}{n}\]BIC imposes a heavier penalty for degrees of freedom lost, hence it will lean toward a simpler model than AIC.
Choose the model with the smallest AIC or BIC. AIC and BIC leaves out the constant terms in the log-likelihood function because they shift every model by the same amount and do not affect the model selection.
Alternative expressions for AIC and BIC are:
\[\begin{split} AIC(K) &= n\log \left(\frac{\text{SSR}}{n}\right) + 2K \\ BIC(K) &= n\log \left(\frac{\text{SSR}}{n}\right) + K\log(n) \end{split}\]$\be’\be=\text{SSR}$ sum of squared residuals.
For a fixed sample size $n$, the two expressions differ only by a constant. Both will lead to the same model selection.
Ref:
- pp 147, Chap 5 Normal Regression and Maximum Likelihood, Hansen Econometrics