• Wrong document? Swap it for free
  • Written by students who passed
  • Immediately available after payment
  • Read online or as PDF
Sell
Where do you study
Your language
Document preview thumbnail
Preview 3 out of 22 pages
Exam (elaborations)

Pstat 131 Correct Exams Answers And Questions Set A.pdf

Document preview thumbnail
Preview 3 out of 22 pages

PSTAT 131 CORRECT EXAMS ANSWERS AND QUESTIONS SET A.pdf

Content preview

PSTAT 131 CORRECT EXAMS ANSWERS AND
QUESTIONS SET A+
✔✔what is the downside of best subset selection? - ✔✔Computationally unfeasible.
p=20 results in 1Billion models

use forward stepwise selection

✔✔How many models can be made with forward/backward stepwise selection? -
✔✔(p(p+1)/2) + 1

✔✔forward/backward selection uses a greedy approach? - ✔✔true

✔✔in forward/backward selection, a variable that gives the greatest additional
improvement belongs to the best model? t/f - ✔✔FALSE

✔✔both forward and backward selection require n > p for fitting model? - ✔✔no, only
backward selection requires n > p

✔✔In model selection, RSS or R^2 can not be used to select the best model? T/F -
✔✔true,
RSS/^2 are used to select candidates for best model, but not used to choose the final
best model.

✔✔C_p criterion - ✔✔1/n(SSR + 2*d*σ̂^2)

2*d*σ̂^2 = penalty
SSR = RSS of least squares fit of model
d = num predictors
σ̂^2 = estimate of variance of random error

As d gets large, SSR decreases but penalty increases

select model with lowest C_p

,✔✔AIC - ✔✔1/(nσ̂^2) (SSR + 2*d*σ̂^2)

select model with lowest AIC
defined for large class of models

✔✔BIC - ✔✔1/n(SSR + log(n)d*σ̂^2)

heavier penalty than C_p
select model with lowest BIC

✔✔Adjusted R^2 - ✔✔1 - [SSR/(n−d−1)] / [SST /(n−1)], Total sum of squares ∑ (yi −
y ̄)2

R^2 = 1 - SSR/ SST

✔✔as number of parameters increases when computing AIC/BIC/C_p, RSS...? -
✔✔decreases

✔✔Which criteria to use for model selection? - ✔✔direct estimate of test MSE: CV!

indirect approaches: AIC, BIC, C_p, Adj. R^2

✔✔Least squares performs feature selection, t/f? - ✔✔false

✔✔Regularization - ✔✔a type of model that shrinks coefficient ESTIMATES to 0

significantly reduces variance, with small increase in bias

✔✔Least squares estimates minimize ...? - ✔✔SSR

✔✔ridge regression coefficients estimates minimize - ✔✔β^R

SSR + shrinkage penalty

✔✔what is shrinkage penalty in ridge regression? - ✔✔tuning parameter * l2 norm
squared

✔✔when lambda equals zero, ridge regression will produce - ✔✔least squares
estimates

✔✔In ridge regression, as lambda approaches infinity, the coefficient estimates will
approach what value? - ✔✔0

, when lambda closer to 0, low bias high variance
when lambda closer to infinity, high bias low variance

✔✔tuning parameter must always be greater than or equal to 0, t/f? - ✔✔True

✔✔Ridge and Lasso both force coefficient to 0? T/F? - ✔✔False, Lasso forces
coefficients to 0, yielding sparse models

✔✔Lasso's shrinkage penalty is - ✔✔tuning parameter * l1 norm

✔✔l1 norm follows square shape and l2 norm^2 follows circle shape. what is S? - ✔✔s
is the radius, or distance from corner to center in square.

✔✔Both lasso and ridge perform variable selection, T/F? - ✔✔False, only Lasso
performs variable selection

✔✔Lasso is better than ridge, T/F? - ✔✔False, neither will dominate the other

✔✔least squares estimates are square equivariant, meaning? - ✔✔multiplying X_j by C
equals scaling β^LS by factor of 1/c

✔✔ridge, lasso, and least square estimates are square equivariant - ✔✔False, ridge
and lasso are NOT. so ridge/lasso should be applied after standardizing predictors

✔✔how to select best tuning parameter for ridge/lasso? - ✔✔Cross validation!

1. Choose a grid of λ values

2. For each value of λ, compute the cross-validation estimate of test MSE for ridge/lasso

3. Select the value of λ for which the cross-validation estimate of test MSE is smallest

✔✔is a decision tree better than linear model? - ✔✔higher complexity/flexibility

multiple trees are less interpretable but much better prediction accuracy.

✔✔decision tree can be applied to both regression and classification problems? -
✔✔True

✔✔decision trees are a non parametric method - ✔✔True

✔✔when building a regression tree, how do we split the predictor space into J non-
overlapping regions? - ✔✔If an observation falls into a region R_j, we predict it to be the
mean of the response for training observations in R_j

Document information

Uploaded on
September 24, 2026
Number of pages
22
Written in
2026/2027
Type
Exam (elaborations)
Contains
Questions & answers
$18.99

Wrong document? Swap it for free Within 14 days of purchase and before downloading, you can choose a different document. You can simply spend the amount again.
Written by students who passed
Immediately available after payment
Read online or as PDF

Seller avatar
Reputation scores are based on the amount of documents a seller has sold for a fee and the reviews they have received for those documents. There are three levels: Bronze, Silver and Gold. The better the reputation, the more your can rely on the quality of the sellers work.
sterlingscoress
4.0
(2)
Sold
22
Followers
2
Items
7802
Last sold
1 month ago



Why students choose Stuvia

Created by fellow students, verified by reviews

Quality you can trust: written by students who passed their tests and reviewed by others who've used these notes.

Didn't get what you expected? Choose another document

No worries! You can instantly pick a different document that better fits what you're looking for.

Pay as you like, start learning right away

No subscription, no commitments. Pay the way you're used to via credit card and download your PDF document instantly.

Student with book image

“Bought, downloaded, and aced it. It really can be that simple.”

Alisha Student

Working on your references?

Create accurate citations in APA, MLA and Harvard with our free citation generator.

Working on your references?

Frequently asked questions