---
output:
  pdf_document: default
  html_document: default
---

# Exercises

```{r, message=FALSE}
library(tidyverse)
theme_set(theme_bw())
```

Load the `titanic` package:

```{r}
library(titanic)
```

It contains these two datasets that must be loaded:

```{r}
data("titanic_train")
data("titanic_test")
```

Convert to `tibble`'s:

```{r}
titanic_train <- titanic_train %>% as_tibble()
titanic_test <- titanic_test %>% as_tibble()
```

In this exercise, we would like to predict survival (`Survived`) based on passenger class (`Pclass`) and `Sex`.

Consider the data type of `Pclass`...

```{r}

```

Construct a table using `titanic_train` giving the number of survivers for each value of sex.

```{r}

```

Does the sexes have the same survival probability? Try to answer by constructing just one expression based on a number of `dplyr` verbs and pipes (`%>%`).

```{r}

```

Now, do a logistic regression analysis (remember the type of `Pclass`...).

```{r}

```

For the training data, how many different predicted probabilities are there (theoretically and in practise)?

```{r}

```

For each combination of `Pclass` and `Sex`, what is the predicted survival probability?

```{r}

```

Can you construct another logistic regression model, and how are these results compared to the previous model?

```{r}

```


Now, use the model based on `titanic_train` to predict the results for `titanic_test`. If forced to convert the predicted survival probabilities to a binary response (survived/non-survived), how can this be done? What is overall chance of survival in the test dataset compared to that of the training dataset?

```{r}

```

Consider joining <https://www.kaggle.com/c/titanic> :-)
