PSTAT 131 CORRECT EXAMS ANSWERS AND
QUESTIONS SET A+
✔✔what is the downside of best subset selection? - ✔✔Computationally unfeasible.
p=20 results in 1Billion models
use forward stepwise selection
✔✔How many models can be made with forward/backward stepwise selection? -
✔✔(p(p+1)/2) + 1
✔✔forward/backward selection uses a greedy approach? - ✔✔true
✔✔in forward/backward selection, a variable that gives the greatest additional
improvement belongs to the best model? t/f - ✔✔FALSE
✔✔both forward and backward selection require n > p for fitting model? - ✔✔no, only
backward selection requires n > p
✔✔In model selection, RSS or R^2 can not be used to select the best model? T/F -
✔✔true,
RSS/^2 are used to select candidates for best model, but not used to choose the final
best model.
✔✔C_p criterion - ✔✔1/n(SSR + 2*d*σ̂^2)
2*d*σ̂^2 = penalty
SSR = RSS of least squares fit of model
d = num predictors
σ̂^2 = estimate of variance of random error
As d gets large, SSR decreases but penalty increases
select model with lowest C_p
,✔✔AIC - ✔✔1/(nσ̂^2) (SSR + 2*d*σ̂^2)
select model with lowest AIC
defined for large class of models
✔✔BIC - ✔✔1/n(SSR + log(n)d*σ̂^2)
heavier penalty than C_p
select model with lowest BIC
✔✔Adjusted R^2 - ✔✔1 - [SSR/(n−d−1)] / [SST /(n−1)], Total sum of squares ∑ (yi −
y ̄)2
R^2 = 1 - SSR/ SST
✔✔as number of parameters increases when computing AIC/BIC/C_p, RSS...? -
✔✔decreases
✔✔Which criteria to use for model selection? - ✔✔direct estimate of test MSE: CV!
indirect approaches: AIC, BIC, C_p, Adj. R^2
✔✔Least squares performs feature selection, t/f? - ✔✔false
✔✔Regularization - ✔✔a type of model that shrinks coefficient ESTIMATES to 0
significantly reduces variance, with small increase in bias
✔✔Least squares estimates minimize ...? - ✔✔SSR
✔✔ridge regression coefficients estimates minimize - ✔✔β^R
SSR + shrinkage penalty
✔✔what is shrinkage penalty in ridge regression? - ✔✔tuning parameter * l2 norm
squared
✔✔when lambda equals zero, ridge regression will produce - ✔✔least squares
estimates
✔✔In ridge regression, as lambda approaches infinity, the coefficient estimates will
approach what value? - ✔✔0
, when lambda closer to 0, low bias high variance
when lambda closer to infinity, high bias low variance
✔✔tuning parameter must always be greater than or equal to 0, t/f? - ✔✔True
✔✔Ridge and Lasso both force coefficient to 0? T/F? - ✔✔False, Lasso forces
coefficients to 0, yielding sparse models
✔✔Lasso's shrinkage penalty is - ✔✔tuning parameter * l1 norm
✔✔l1 norm follows square shape and l2 norm^2 follows circle shape. what is S? - ✔✔s
is the radius, or distance from corner to center in square.
✔✔Both lasso and ridge perform variable selection, T/F? - ✔✔False, only Lasso
performs variable selection
✔✔Lasso is better than ridge, T/F? - ✔✔False, neither will dominate the other
✔✔least squares estimates are square equivariant, meaning? - ✔✔multiplying X_j by C
equals scaling β^LS by factor of 1/c
✔✔ridge, lasso, and least square estimates are square equivariant - ✔✔False, ridge
and lasso are NOT. so ridge/lasso should be applied after standardizing predictors
✔✔how to select best tuning parameter for ridge/lasso? - ✔✔Cross validation!
1. Choose a grid of λ values
2. For each value of λ, compute the cross-validation estimate of test MSE for ridge/lasso
3. Select the value of λ for which the cross-validation estimate of test MSE is smallest
✔✔is a decision tree better than linear model? - ✔✔higher complexity/flexibility
multiple trees are less interpretable but much better prediction accuracy.
✔✔decision tree can be applied to both regression and classification problems? -
✔✔True
✔✔decision trees are a non parametric method - ✔✔True
✔✔when building a regression tree, how do we split the predictor space into J non-
overlapping regions? - ✔✔If an observation falls into a region R_j, we predict it to be the
mean of the response for training observations in R_j
QUESTIONS SET A+
✔✔what is the downside of best subset selection? - ✔✔Computationally unfeasible.
p=20 results in 1Billion models
use forward stepwise selection
✔✔How many models can be made with forward/backward stepwise selection? -
✔✔(p(p+1)/2) + 1
✔✔forward/backward selection uses a greedy approach? - ✔✔true
✔✔in forward/backward selection, a variable that gives the greatest additional
improvement belongs to the best model? t/f - ✔✔FALSE
✔✔both forward and backward selection require n > p for fitting model? - ✔✔no, only
backward selection requires n > p
✔✔In model selection, RSS or R^2 can not be used to select the best model? T/F -
✔✔true,
RSS/^2 are used to select candidates for best model, but not used to choose the final
best model.
✔✔C_p criterion - ✔✔1/n(SSR + 2*d*σ̂^2)
2*d*σ̂^2 = penalty
SSR = RSS of least squares fit of model
d = num predictors
σ̂^2 = estimate of variance of random error
As d gets large, SSR decreases but penalty increases
select model with lowest C_p
,✔✔AIC - ✔✔1/(nσ̂^2) (SSR + 2*d*σ̂^2)
select model with lowest AIC
defined for large class of models
✔✔BIC - ✔✔1/n(SSR + log(n)d*σ̂^2)
heavier penalty than C_p
select model with lowest BIC
✔✔Adjusted R^2 - ✔✔1 - [SSR/(n−d−1)] / [SST /(n−1)], Total sum of squares ∑ (yi −
y ̄)2
R^2 = 1 - SSR/ SST
✔✔as number of parameters increases when computing AIC/BIC/C_p, RSS...? -
✔✔decreases
✔✔Which criteria to use for model selection? - ✔✔direct estimate of test MSE: CV!
indirect approaches: AIC, BIC, C_p, Adj. R^2
✔✔Least squares performs feature selection, t/f? - ✔✔false
✔✔Regularization - ✔✔a type of model that shrinks coefficient ESTIMATES to 0
significantly reduces variance, with small increase in bias
✔✔Least squares estimates minimize ...? - ✔✔SSR
✔✔ridge regression coefficients estimates minimize - ✔✔β^R
SSR + shrinkage penalty
✔✔what is shrinkage penalty in ridge regression? - ✔✔tuning parameter * l2 norm
squared
✔✔when lambda equals zero, ridge regression will produce - ✔✔least squares
estimates
✔✔In ridge regression, as lambda approaches infinity, the coefficient estimates will
approach what value? - ✔✔0
, when lambda closer to 0, low bias high variance
when lambda closer to infinity, high bias low variance
✔✔tuning parameter must always be greater than or equal to 0, t/f? - ✔✔True
✔✔Ridge and Lasso both force coefficient to 0? T/F? - ✔✔False, Lasso forces
coefficients to 0, yielding sparse models
✔✔Lasso's shrinkage penalty is - ✔✔tuning parameter * l1 norm
✔✔l1 norm follows square shape and l2 norm^2 follows circle shape. what is S? - ✔✔s
is the radius, or distance from corner to center in square.
✔✔Both lasso and ridge perform variable selection, T/F? - ✔✔False, only Lasso
performs variable selection
✔✔Lasso is better than ridge, T/F? - ✔✔False, neither will dominate the other
✔✔least squares estimates are square equivariant, meaning? - ✔✔multiplying X_j by C
equals scaling β^LS by factor of 1/c
✔✔ridge, lasso, and least square estimates are square equivariant - ✔✔False, ridge
and lasso are NOT. so ridge/lasso should be applied after standardizing predictors
✔✔how to select best tuning parameter for ridge/lasso? - ✔✔Cross validation!
1. Choose a grid of λ values
2. For each value of λ, compute the cross-validation estimate of test MSE for ridge/lasso
3. Select the value of λ for which the cross-validation estimate of test MSE is smallest
✔✔is a decision tree better than linear model? - ✔✔higher complexity/flexibility
multiple trees are less interpretable but much better prediction accuracy.
✔✔decision tree can be applied to both regression and classification problems? -
✔✔True
✔✔decision trees are a non parametric method - ✔✔True
✔✔when building a regression tree, how do we split the predictor space into J non-
overlapping regions? - ✔✔If an observation falls into a region R_j, we predict it to be the
mean of the response for training observations in R_j