---
title: "Variance-bias trade-off"
author: "Torben"
date: "August 20, 2018"
output: html_document
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(echo = TRUE)
```

# Variance-bias trade-off

The Root Mean Squared Error (RMSE) is a measure on the same scale of as the response of our model. 
The MSE (Mean Squared Error) also often used to assess the model performance.

When we discuss the generalisation of models, we consider *what happens on new unseen dataset* -- that is conceptual 
data. To quantify what we mean by that, we look at the *expected* MSE, $E(MSE)$, where expectation is with respect
to the average MSE when $f$ is fitted on a large number of datasets. 

Again, let $y = f(x) + \varepsilon$ as before. The $\sigma^2$ is an irreducible error that we can't get rid of - 
it is simply the nature of the phenomenon we model. However, we can attempt to learn $f$.

We let $\hat{y}_0 = \hat{f}(x_0)$, where $x_0$ is some instance of the explanatory variables and $y_0$ is the observed 
response with estimate $\hat{y}_0$. When, 
$$E[\{y_0 − \hat{f}(x_0)\}^2] = Var[\hat{f}(x_0)] + [Bias\{\hat{f}(x_0)\}]^2 + Var(\varepsilon)$$
In order to minimize the expected test error, we need to select a statistical learning method that simultaneously 
achieves low variance and low bias. Note that variance is inherently a nonnegative quantity, and squared bias is also
onnegative. Hence, we see that the expected test MSE can never lie below $\sigma^2$, the irreducible error.

## Examples 

Three different situations. The data is generated from the black curves. The fitted functions (orange: `lm`, 
blue: low-degree spline, and green: higher-order spline) has varying flexibility (degrees of freedom).

![](02_09.png)
![](02_10.png)
![](02_11.png)

## Pay a little bias to get a reduction in variance?

Why use an unbiased estimator when we have an unbiased one from OLS? One reason is that the estimator has a lower
variance. Hence, by introducing a little bias, we are able to reduce the variance of the estimated coefficients.