PSTAT 131 UPDATED CORE ANSWERS AND
QUESTIONS SET A+
ML Workflow - ✔✔goal, data, learn, evaluate
✔✔Predictors are also called - ✔✔(inputs, features, covariates, independent variables)
✔✔response is also called - ✔✔(output, target, dependent variable)
✔✔Learner vs Supervisor - ✔✔machine is the learner, Y (response) is the supervisor
✔✔Supervised learning - ✔✔need predictors and response:
- prediction, estimation, model selection, inference
- regression and classification
✔✔Unsupervised Learning - ✔✔goal: to find predictors that behave similarly without
learning from a teacher, or for a Y
✔✔Supervised learning examples - ✔✔logistic regression, decision tree, k-nearest
neighbors, naive bayes, linear regression, ridge, lasso, SVM, NN
✔✔Unsupervised learning examples - ✔✔PCA, NBSCAN, hierarchical clustering, k-
means clustering, Gaussian Mixtures
✔✔Why should we ever choose a less flexible method over a flexible one? - ✔✔Better
interpretability, and sometimes even smaller error (overfitting)
✔✔flexibility vs interpretability of subset selection, lasso, least squares, trees, bagging,
boosting, SVM, NN - ✔✔subset selection/Lasso = High Interpretability, Low Flex
Least squares/trees = mid for both
bagging/boosting , SVM, NN = High flex, low interpretability
,✔✔what does it mean when a model is flexible/complex? - ✔✔specific to training data
(too flexible may result in overfitting)
✔✔Training MSE - ✔✔1/n sum{ (y_i - f_hat(x_i))^2
- average square difference between actual and predicted value
- y_i is the actual value, y_i_hat is predicted value at i-th X
✔✔Test MSE - ✔✔1/n sum{ (y_i - f_hat(x_i))^2
- average square difference between actual and predicted value
- y_i is the actual value, y_i_hat is the predicted value of the target variable for the i-th
data point in the test set.
- we want model with lowest test MSE
✔✔Overfitting - ✔✔- training MSE is low, but test MSE is high
✔✔Cross-Validation - ✔✔method for estimating test MSE and test error using only
training data
✔✔Bias-Variance Decomposition - ✔✔Expected test MSE = variance + bias^2 +
Var(epsilon)
- Var(epsilon) = irreducible error
- it is often not feasible to evaluate a model on all possible test datasets, so the concept
of expected test MSE is used to represent the average performance of a model on new,
unseen data
- Expected MSE accounts for variability in performance that may arise due to different
random samples in the test data.
- joint distribution of (X,Y) is unknown in practice
- Y = f(x) + Epsilon, where f(x) is non-random and epsilon is zero-mean noise
- MSE = E[(Y - f(x)^hat)^2]
✔✔Bias - ✔✔error that is introduced by approximating a real-life problem
ex: real relationship between response and predictors is nonlinear, but we fit a linear
model, which causes bias
,simple model = high bias + low variance
flexible model = low bias + high variance
Bias(f_hat(x_0)) = E[f_hat(x_0)] - f(x_0)
✔✔Variance - ✔✔the amount by which f ̂ model change if we estimated it using a
different training set
Variance(f_hat(x_0)) = E[((f_hat(x_0)-E(f_hat(x_0))^2]
✔✔Training Error Rate - ✔✔fraction of incorrect classifications in training set
✔✔Test error rate - ✔✔fraction of incorrect classifications on test set
- predicted class labels that the model gives for i-th test observation that doesn't match
i_th test response value
✔✔training/test MSE vs training/test error - ✔✔MSE used for regression, error used for
classification
✔✔Bayes Classifier - ✔✔- assigns each observation to the most likely class, given
predictor values
- produces lowest possible test error rate
- still makes mistakes
- bayes error rate: 1- E[maxProb(Y = j | X)]
✔✔k-nearest neighbors - ✔✔- used when we don't know conditional distribution Y|X
- good for small p (<= 4) and large n
- non-parametric method
Classification:
if k increases, variance decreases, bias^2 increases, var(e) out of control
if k decreases, variance increases, bias^2 decreases, var(e) out of control
✔✔What is the curse of dimensionality? - ✔✔High dimensionality makes clustering
hard, because having lots of dimensions means that everything is "far away" from each
other. ... use PCA
✔✔Prediction Error vs model complexity graph - ✔✔test error: U shaped
training error: \ shaped
, ✔✔k-nearest neighbors, k=1 vs k=100 graph - ✔✔k=1:
- high model complexity/flexibility
- low bias, high variance
k=100:
- less wiggly decision boundary
- low model complexity/flex
- high bias, low variance
✔✔how to choose the best k value in k nearest neighbors? - ✔✔K-fold Cross Validation
This involves splitting your training data into k subsets (folds), training the model on k−1
folds, and evaluating it on the remaining fold. Repeat this process for different k values.
✔✔Method of least squares (OLS) - ✔✔minimizing sum of squared residuals
- unbiased and have minimum variance among all unbiased linear estimators
✔✔the conditional expectation of Y is _____ in the parameters? - ✔✔Linear
✔✔How to estimate Beta_0 and Beta_1 in linear regression - ✔✔- find line that
minimizes SSR (sum of squared residuals)
- SSR = sum(y_i - beta_0 - beta_1*x_i)^2
beta_1_hat = SUM (xi- Xbar)(yi - Ybar)/ SUM (xi - xBar)^2
beta_0_hat = y_bar - beta_1_hat * Xbar
✔✔Least Squares Coefficient Estimates - ✔✔both are unbiased estimators
✔✔Variance of least squares coefficient estimates - ✔✔- Var(beta_1_hat) =
SE(beta_1_hat)^2 = σ^2/∑i=1 (xi − x )̄ ^2
- Var(beta_0_hat) = sigma^2[1/n + x ̄^2/∑ (xi−x ̄)^2]
sigma^2 = var(epsilon)
✔✔estimate of sigma^2 = sigma_hat^2 = - ✔✔SSR/n-2
✔✔95% Confidence Interval for beta_1 and beta_0 - ✔✔[beta_ _hat +- 2SE(beta_
_hat)]
✔✔Hypothesis test on coefficients - ✔✔t = beta_1_hat/ SE(beta_1_hat)
df = n-2
QUESTIONS SET A+
ML Workflow - ✔✔goal, data, learn, evaluate
✔✔Predictors are also called - ✔✔(inputs, features, covariates, independent variables)
✔✔response is also called - ✔✔(output, target, dependent variable)
✔✔Learner vs Supervisor - ✔✔machine is the learner, Y (response) is the supervisor
✔✔Supervised learning - ✔✔need predictors and response:
- prediction, estimation, model selection, inference
- regression and classification
✔✔Unsupervised Learning - ✔✔goal: to find predictors that behave similarly without
learning from a teacher, or for a Y
✔✔Supervised learning examples - ✔✔logistic regression, decision tree, k-nearest
neighbors, naive bayes, linear regression, ridge, lasso, SVM, NN
✔✔Unsupervised learning examples - ✔✔PCA, NBSCAN, hierarchical clustering, k-
means clustering, Gaussian Mixtures
✔✔Why should we ever choose a less flexible method over a flexible one? - ✔✔Better
interpretability, and sometimes even smaller error (overfitting)
✔✔flexibility vs interpretability of subset selection, lasso, least squares, trees, bagging,
boosting, SVM, NN - ✔✔subset selection/Lasso = High Interpretability, Low Flex
Least squares/trees = mid for both
bagging/boosting , SVM, NN = High flex, low interpretability
,✔✔what does it mean when a model is flexible/complex? - ✔✔specific to training data
(too flexible may result in overfitting)
✔✔Training MSE - ✔✔1/n sum{ (y_i - f_hat(x_i))^2
- average square difference between actual and predicted value
- y_i is the actual value, y_i_hat is predicted value at i-th X
✔✔Test MSE - ✔✔1/n sum{ (y_i - f_hat(x_i))^2
- average square difference between actual and predicted value
- y_i is the actual value, y_i_hat is the predicted value of the target variable for the i-th
data point in the test set.
- we want model with lowest test MSE
✔✔Overfitting - ✔✔- training MSE is low, but test MSE is high
✔✔Cross-Validation - ✔✔method for estimating test MSE and test error using only
training data
✔✔Bias-Variance Decomposition - ✔✔Expected test MSE = variance + bias^2 +
Var(epsilon)
- Var(epsilon) = irreducible error
- it is often not feasible to evaluate a model on all possible test datasets, so the concept
of expected test MSE is used to represent the average performance of a model on new,
unseen data
- Expected MSE accounts for variability in performance that may arise due to different
random samples in the test data.
- joint distribution of (X,Y) is unknown in practice
- Y = f(x) + Epsilon, where f(x) is non-random and epsilon is zero-mean noise
- MSE = E[(Y - f(x)^hat)^2]
✔✔Bias - ✔✔error that is introduced by approximating a real-life problem
ex: real relationship between response and predictors is nonlinear, but we fit a linear
model, which causes bias
,simple model = high bias + low variance
flexible model = low bias + high variance
Bias(f_hat(x_0)) = E[f_hat(x_0)] - f(x_0)
✔✔Variance - ✔✔the amount by which f ̂ model change if we estimated it using a
different training set
Variance(f_hat(x_0)) = E[((f_hat(x_0)-E(f_hat(x_0))^2]
✔✔Training Error Rate - ✔✔fraction of incorrect classifications in training set
✔✔Test error rate - ✔✔fraction of incorrect classifications on test set
- predicted class labels that the model gives for i-th test observation that doesn't match
i_th test response value
✔✔training/test MSE vs training/test error - ✔✔MSE used for regression, error used for
classification
✔✔Bayes Classifier - ✔✔- assigns each observation to the most likely class, given
predictor values
- produces lowest possible test error rate
- still makes mistakes
- bayes error rate: 1- E[maxProb(Y = j | X)]
✔✔k-nearest neighbors - ✔✔- used when we don't know conditional distribution Y|X
- good for small p (<= 4) and large n
- non-parametric method
Classification:
if k increases, variance decreases, bias^2 increases, var(e) out of control
if k decreases, variance increases, bias^2 decreases, var(e) out of control
✔✔What is the curse of dimensionality? - ✔✔High dimensionality makes clustering
hard, because having lots of dimensions means that everything is "far away" from each
other. ... use PCA
✔✔Prediction Error vs model complexity graph - ✔✔test error: U shaped
training error: \ shaped
, ✔✔k-nearest neighbors, k=1 vs k=100 graph - ✔✔k=1:
- high model complexity/flexibility
- low bias, high variance
k=100:
- less wiggly decision boundary
- low model complexity/flex
- high bias, low variance
✔✔how to choose the best k value in k nearest neighbors? - ✔✔K-fold Cross Validation
This involves splitting your training data into k subsets (folds), training the model on k−1
folds, and evaluating it on the remaining fold. Repeat this process for different k values.
✔✔Method of least squares (OLS) - ✔✔minimizing sum of squared residuals
- unbiased and have minimum variance among all unbiased linear estimators
✔✔the conditional expectation of Y is _____ in the parameters? - ✔✔Linear
✔✔How to estimate Beta_0 and Beta_1 in linear regression - ✔✔- find line that
minimizes SSR (sum of squared residuals)
- SSR = sum(y_i - beta_0 - beta_1*x_i)^2
beta_1_hat = SUM (xi- Xbar)(yi - Ybar)/ SUM (xi - xBar)^2
beta_0_hat = y_bar - beta_1_hat * Xbar
✔✔Least Squares Coefficient Estimates - ✔✔both are unbiased estimators
✔✔Variance of least squares coefficient estimates - ✔✔- Var(beta_1_hat) =
SE(beta_1_hat)^2 = σ^2/∑i=1 (xi − x )̄ ^2
- Var(beta_0_hat) = sigma^2[1/n + x ̄^2/∑ (xi−x ̄)^2]
sigma^2 = var(epsilon)
✔✔estimate of sigma^2 = sigma_hat^2 = - ✔✔SSR/n-2
✔✔95% Confidence Interval for beta_1 and beta_0 - ✔✔[beta_ _hat +- 2SE(beta_
_hat)]
✔✔Hypothesis test on coefficients - ✔✔t = beta_1_hat/ SE(beta_1_hat)
df = n-2