PSTAT 131 COMPREHENSIVE ANSWERS AND
QUESTIONS SET A+
✔✔step function used in GAM - ✔✔break the range of X into bins
f(x) is a piecewise constant function
✔✔what is the problem with using piecewise polynomial as function in GAM -
✔✔piecewise polynomial is discontinuous so too flexible
✔✔what is the problem with using regression spline as function in GAM -
✔✔discontinuous so too flexible
✔✔What is a Natural (cubic) Spline, and one negative thing about it? - ✔✔Generalized
Additive Model allows non-linear functions for each predictor. so natural cubic spline is
another type of non-linear function that can be used in a GAM.
Con: high variance on outer range of predictors
✔✔piecewise polynomial vs regression spline vs natural spline - ✔✔constraint on lower
order derivatives of piecewise polynomial results in regression spline
constraint on outer range of regression spline results in natural spline
piecewise has high variance/flexibility, natural spline has low
✔✔how to choose number of knots in GAM spline - ✔✔Cross Validation, more knots =
more flexibile
✔✔limitations of GAMs - ✔✔no interaction terms
no two dimensional splines
✔✔Gini Index - ✔✔measures node purity: small value of G means node mostly contains
observations from a single class
,best case scenario is if Gini index = 0
✔✔Cross-Entropy - ✔✔also popular to use in addition to Gini index
measures node purity: small value of G means node mostly contains observations from
a single class
D >= 0, but takes on a small value if all proportion of training observations in the j-th
region that are from the k-th class are all near 0 or near 1.
✔✔limitation of single decision trees - ✔✔not that great prediction accuracy, not that
robust
✔✔Non-parametric methods - ✔✔k-NN regression
k-NN classification
decision tree
random forest
bagging
boosting
✔✔Given a set of B independent random variables, V1, ..., VB, each with variance σ^2,
variance of the average is - ✔✔sigma^2/B
✔✔Bagging = - ✔✔bootstrap aggregating
average/majority vote of mini models for regression
mode of mini models for classification
✔✔Out of Bag Error - ✔✔On average, around 1/3 of the observations in a bootstrap
sample are out-of-bag
✔✔Dos cross validation need to be used to estimate the test error in bootstrapping? -
✔✔No, OOB error is almost equivalent to LOOCV, but computationally cheaper
✔✔True or False, OOB MSE or error rate can be computed? - ✔✔true
✔✔true or false, it is impossible to visualize resulting model in bagging - ✔✔true
✔✔what do we look for as a result of bagging? - ✔✔the total amount that the SSR / Gini
index is decreased due to splits over that predictor, averaged over all B trees
✔✔are predictions from bagged trees independent? - ✔✔No
, ✔✔true or false, bagging decreases the bias of a single tree by aggregation? - ✔✔false,
variance
✔✔Purpose of bagging - ✔✔to make variance go down while bias stays the same
✔✔Random forest vs bagging - ✔✔When trimming the tree...
if sample only observations when doing Bootstrap => Bagging
if sampling observations and sample the variables when doing bootstrap: Random
Forest (saves time while not losing much accuracy)
✔✔Random Forest classifiers use bagged trees that split on a random subset of m < p
predictors at each branch (usually m = √p) because - ✔✔it makes the predictions from
each of the bagged tree less correlated
✔✔Random forest: variance of average of non-independent variables = - ✔✔p =
pairwise correlation
variance of each random variable = sigma^2
= p*sigma^2 + (1-p)sigma^2/B
✔✔When building trees for random forest, how many predictors should be chosen as
candidates for splitting? - ✔✔select m = sqrt(p) predictors at random as candidates for
splitting
✔✔how many predictors are in standard bagging - ✔✔m = p = original number of
predictors
✔✔True or False, in each split of bagging, the algorithm is not allowed to consider most
of the predictors? - ✔✔True, this is to de-corelate the trees and lower variance
✔✔T/F, predictors from bagged trees are independent? - ✔✔False, not independent
✔✔Boosting - ✔✔makes mini trees similar to bagging, but does not make all mini trees
at the same time
grows trees sequentially using information from previous tree
✔✔T/F, both bagging and boosting involve bootstrap sampling? - ✔✔False, only
bagging involves bootstrap sampling
✔✔if a boosting regression tree has d splits, how many terminal nodes are there? -
✔✔d+1
QUESTIONS SET A+
✔✔step function used in GAM - ✔✔break the range of X into bins
f(x) is a piecewise constant function
✔✔what is the problem with using piecewise polynomial as function in GAM -
✔✔piecewise polynomial is discontinuous so too flexible
✔✔what is the problem with using regression spline as function in GAM -
✔✔discontinuous so too flexible
✔✔What is a Natural (cubic) Spline, and one negative thing about it? - ✔✔Generalized
Additive Model allows non-linear functions for each predictor. so natural cubic spline is
another type of non-linear function that can be used in a GAM.
Con: high variance on outer range of predictors
✔✔piecewise polynomial vs regression spline vs natural spline - ✔✔constraint on lower
order derivatives of piecewise polynomial results in regression spline
constraint on outer range of regression spline results in natural spline
piecewise has high variance/flexibility, natural spline has low
✔✔how to choose number of knots in GAM spline - ✔✔Cross Validation, more knots =
more flexibile
✔✔limitations of GAMs - ✔✔no interaction terms
no two dimensional splines
✔✔Gini Index - ✔✔measures node purity: small value of G means node mostly contains
observations from a single class
,best case scenario is if Gini index = 0
✔✔Cross-Entropy - ✔✔also popular to use in addition to Gini index
measures node purity: small value of G means node mostly contains observations from
a single class
D >= 0, but takes on a small value if all proportion of training observations in the j-th
region that are from the k-th class are all near 0 or near 1.
✔✔limitation of single decision trees - ✔✔not that great prediction accuracy, not that
robust
✔✔Non-parametric methods - ✔✔k-NN regression
k-NN classification
decision tree
random forest
bagging
boosting
✔✔Given a set of B independent random variables, V1, ..., VB, each with variance σ^2,
variance of the average is - ✔✔sigma^2/B
✔✔Bagging = - ✔✔bootstrap aggregating
average/majority vote of mini models for regression
mode of mini models for classification
✔✔Out of Bag Error - ✔✔On average, around 1/3 of the observations in a bootstrap
sample are out-of-bag
✔✔Dos cross validation need to be used to estimate the test error in bootstrapping? -
✔✔No, OOB error is almost equivalent to LOOCV, but computationally cheaper
✔✔True or False, OOB MSE or error rate can be computed? - ✔✔true
✔✔true or false, it is impossible to visualize resulting model in bagging - ✔✔true
✔✔what do we look for as a result of bagging? - ✔✔the total amount that the SSR / Gini
index is decreased due to splits over that predictor, averaged over all B trees
✔✔are predictions from bagged trees independent? - ✔✔No
, ✔✔true or false, bagging decreases the bias of a single tree by aggregation? - ✔✔false,
variance
✔✔Purpose of bagging - ✔✔to make variance go down while bias stays the same
✔✔Random forest vs bagging - ✔✔When trimming the tree...
if sample only observations when doing Bootstrap => Bagging
if sampling observations and sample the variables when doing bootstrap: Random
Forest (saves time while not losing much accuracy)
✔✔Random Forest classifiers use bagged trees that split on a random subset of m < p
predictors at each branch (usually m = √p) because - ✔✔it makes the predictions from
each of the bagged tree less correlated
✔✔Random forest: variance of average of non-independent variables = - ✔✔p =
pairwise correlation
variance of each random variable = sigma^2
= p*sigma^2 + (1-p)sigma^2/B
✔✔When building trees for random forest, how many predictors should be chosen as
candidates for splitting? - ✔✔select m = sqrt(p) predictors at random as candidates for
splitting
✔✔how many predictors are in standard bagging - ✔✔m = p = original number of
predictors
✔✔True or False, in each split of bagging, the algorithm is not allowed to consider most
of the predictors? - ✔✔True, this is to de-corelate the trees and lower variance
✔✔T/F, predictors from bagged trees are independent? - ✔✔False, not independent
✔✔Boosting - ✔✔makes mini trees similar to bagging, but does not make all mini trees
at the same time
grows trees sequentially using information from previous tree
✔✔T/F, both bagging and boosting involve bootstrap sampling? - ✔✔False, only
bagging involves bootstrap sampling
✔✔if a boosting regression tree has d splits, how many terminal nodes are there? -
✔✔d+1