# Exercises

```{r }
library(tidyverse)
theme_set(theme_bw())
```

Load the `breastcancer` dataset.

Use a logistic regression model (`glm(..., family = binomial())`) and a principal component analysis to predict the `code` (`control`/`case`) using the first $m$ principal components. What should $m$ be?

Use a LASSO model to perform a similar analysis. What are the most important genes for predicting the `code`?

How does the logistic regression using principal components perform compared to the LASSO model? Use a train/test data approach. For PCA, first scale the variables (remember to use same scaling factor for train and test data).

How are the pros/cons of the types of models?
