PSTAT 131 REVIEW TIPS ANSWERS AND QUESTIONS
SET A+
✔✔How to estimate Beta_0 and Beta_1 in linear regression - ✔✔- find line that
minimizes SSR (sum of squared residuals)
- SSR = sum(y_i - beta_0 - beta_1*x_i)^2
beta_1_hat = SUM (xi- Xbar)(yi - Ybar)/ SUM (xi - xBar)^2
beta_0_hat = y_bar - beta_1_hat * Xbar
✔✔Least Squares Coefficient Estimates - ✔✔both are unbiased estimators
✔✔Variance of least squares coefficient estimates - ✔✔- Var(beta_1_hat) =
SE(beta_1_hat)^2 = σ^2/∑i=1 (xi − x )̄ ^2
- Var(beta_0_hat) = sigma^2[1/n + x ̄^2/∑ (xi−x ̄)^2]
sigma^2 = var(epsilon)
✔✔estimate of sigma^2 = sigma_hat^2 = - ✔✔SSR/n-2
✔✔95% Confidence Interval for beta_1 and beta_0 - ✔✔[beta_ _hat +- 2SE(beta_
_hat)]
✔✔Hypothesis test on coefficients - ✔✔t = beta_1_hat/ SE(beta_1_hat)
df = n-2
large value of |t| rejects null hypothesis
✔✔Absolute measure of lack of fit - ✔✔RSE: sqrt(SSR/n-2)
residual standard error
,✔✔R^2 - ✔✔Proportion of variability in Y that can be explained by X
1 - SSR/SST
both RSE and R^2 favor flexible methods, may overfit data
✔✔Multiple linear regression hypothesis - ✔✔H0: beta_1 = beta_2 = ... = beta_p = 0
Ha: at least one beta_j is non-zero
Use F -statistic: [(SST - SSR)/p] / [SSR(n-p-1)]
When H0 holds, numerator and denominator is about sigma^2 so F = 1
✔✔Subset selection - ✔✔compute least squares fit for all possible subsets of predictors
and choose best based on AIC, BIC, adj R^2,
impossible for large p: 2^p models in total (1 billion models when p=40)
✔✔Forward selection - ✔✔1. start with no variables in model
2. train model with each feature (add one then train, then add another then train again)
3. choose the one that performs the best based on a chosen performance metric (e.g.,
accuracy, mean squared error).
4. add
5. repeat
✔✔backward selection - ✔✔start with full model, remove one, repeat, keep repeating
until criteria met
does not work for when p > n
✔✔mixed selected - ✔✔combination of forward and backward selection
✔✔R^2 will always ____ as more predictors are added to the model - ✔✔increase
✔✔In regression, Y is quantitative. does that imply that all predictors are also
quantitative? - ✔✔No, predictors can also be qualitative. use dummy variables
(choosing the dummy variable number for each category adds bias)
dummy coding and choice of baseline are arbitrary but change the interpretation of beta
coefficients
✔✔how many dummy variables to use? - ✔✔if qualitative predictor has m levels, use m-
1 dummy variables
,✔✔How to remove additive assumption in linear model? - ✔✔Add an interaction term,
effect of X1 on Y is no longer linear
beta_0 + beta_1x_1 + beta_2x_2 + beta_3*x_1*x_2
✔✔Polynomial regression - ✔✔when X are non linear
✔✔MLR: Assumptions about random error - ✔✔- iid
- unobservable
- equal variance (sigma^2)
- normally distributed
- independent, implying y1...yn are independent
✔✔linear assumption holds if Residual Plot shows - ✔✔no distinct pattern
✔✔if residual plot shows linear plot - ✔✔heteroscedasticity
✔✔if residual plot shows upside down parabola - ✔✔nonlinear
✔✔If random errors in MLR are non-observable, how can we check them? - ✔✔hat
matrix
check residual plot for variance, constant symmetrical variation
check QQplot or Shapiro-wilk test for normality,
fix both with transformation if criteria not met
✔✔3 things that can break assumptions in MLR - ✔✔high leverage points (unusual x
value)
outliers (unusual y value given x)
influential observations (substantially changes model fit, tend to be either an outlier or a
high leverage point)
✔✔Leverages are - ✔✔diagonal elements in H of hat matrix
✔✔H_ii is ... - ✔✔leverage of x_i
and measures distance between x_i and the average of all x's in dataset
✔✔Purpose of leverage point - ✔✔large H_ii makes i'th residual have small variance,
regardless of y_i
, effect of x_i is overwhelming the effect of y_i
✔✔∑ H_ii = - ✔✔p+1
so average leverage = (p+1)/n
✔✔High leverage point is one that - ✔✔2-3 times or greatly exceeds (p+1)/n ,
✔✔how to find outliers? - ✔✔residual vs fitted plot
OR
any observations whose absolute standardized residuals ≥ 3
✔✔How to find influential observations - ✔✔either an outlier, or high leverage point or
both
= if Cook's distance is > 4/n
✔✔TRUE or FALSE? High leverage/outlier/influential point is same for all models? -
✔✔FALSE, might not be in another model
✔✔Collinearity - ✔✔if two predictors are not independent, and can't tell difference with
their individual effect on Y.
check correlation matrix
compute variance inflation factor
either drop one of the variables, or combine both,
✔✔Modern regression methods can outperform least-squares in terms of expected test
MSE, by - ✔✔having small bias but having much smaller variance
✔✔Logistic regression is - ✔✔a linear method in classification
✔✔classification methods - ✔✔k-nn, logistic regression, discriminant analysis
(LDA/QDA), predicting probability of each class
✔✔Logistic Function - ✔✔logistic function is equal to the odds function
The logistic function transforms the log-odds into probabilities :
Pr(y = 1 | x) / 1- Pr(y = 1 | x) = e^(beta_0+beta_1X)
SET A+
✔✔How to estimate Beta_0 and Beta_1 in linear regression - ✔✔- find line that
minimizes SSR (sum of squared residuals)
- SSR = sum(y_i - beta_0 - beta_1*x_i)^2
beta_1_hat = SUM (xi- Xbar)(yi - Ybar)/ SUM (xi - xBar)^2
beta_0_hat = y_bar - beta_1_hat * Xbar
✔✔Least Squares Coefficient Estimates - ✔✔both are unbiased estimators
✔✔Variance of least squares coefficient estimates - ✔✔- Var(beta_1_hat) =
SE(beta_1_hat)^2 = σ^2/∑i=1 (xi − x )̄ ^2
- Var(beta_0_hat) = sigma^2[1/n + x ̄^2/∑ (xi−x ̄)^2]
sigma^2 = var(epsilon)
✔✔estimate of sigma^2 = sigma_hat^2 = - ✔✔SSR/n-2
✔✔95% Confidence Interval for beta_1 and beta_0 - ✔✔[beta_ _hat +- 2SE(beta_
_hat)]
✔✔Hypothesis test on coefficients - ✔✔t = beta_1_hat/ SE(beta_1_hat)
df = n-2
large value of |t| rejects null hypothesis
✔✔Absolute measure of lack of fit - ✔✔RSE: sqrt(SSR/n-2)
residual standard error
,✔✔R^2 - ✔✔Proportion of variability in Y that can be explained by X
1 - SSR/SST
both RSE and R^2 favor flexible methods, may overfit data
✔✔Multiple linear regression hypothesis - ✔✔H0: beta_1 = beta_2 = ... = beta_p = 0
Ha: at least one beta_j is non-zero
Use F -statistic: [(SST - SSR)/p] / [SSR(n-p-1)]
When H0 holds, numerator and denominator is about sigma^2 so F = 1
✔✔Subset selection - ✔✔compute least squares fit for all possible subsets of predictors
and choose best based on AIC, BIC, adj R^2,
impossible for large p: 2^p models in total (1 billion models when p=40)
✔✔Forward selection - ✔✔1. start with no variables in model
2. train model with each feature (add one then train, then add another then train again)
3. choose the one that performs the best based on a chosen performance metric (e.g.,
accuracy, mean squared error).
4. add
5. repeat
✔✔backward selection - ✔✔start with full model, remove one, repeat, keep repeating
until criteria met
does not work for when p > n
✔✔mixed selected - ✔✔combination of forward and backward selection
✔✔R^2 will always ____ as more predictors are added to the model - ✔✔increase
✔✔In regression, Y is quantitative. does that imply that all predictors are also
quantitative? - ✔✔No, predictors can also be qualitative. use dummy variables
(choosing the dummy variable number for each category adds bias)
dummy coding and choice of baseline are arbitrary but change the interpretation of beta
coefficients
✔✔how many dummy variables to use? - ✔✔if qualitative predictor has m levels, use m-
1 dummy variables
,✔✔How to remove additive assumption in linear model? - ✔✔Add an interaction term,
effect of X1 on Y is no longer linear
beta_0 + beta_1x_1 + beta_2x_2 + beta_3*x_1*x_2
✔✔Polynomial regression - ✔✔when X are non linear
✔✔MLR: Assumptions about random error - ✔✔- iid
- unobservable
- equal variance (sigma^2)
- normally distributed
- independent, implying y1...yn are independent
✔✔linear assumption holds if Residual Plot shows - ✔✔no distinct pattern
✔✔if residual plot shows linear plot - ✔✔heteroscedasticity
✔✔if residual plot shows upside down parabola - ✔✔nonlinear
✔✔If random errors in MLR are non-observable, how can we check them? - ✔✔hat
matrix
check residual plot for variance, constant symmetrical variation
check QQplot or Shapiro-wilk test for normality,
fix both with transformation if criteria not met
✔✔3 things that can break assumptions in MLR - ✔✔high leverage points (unusual x
value)
outliers (unusual y value given x)
influential observations (substantially changes model fit, tend to be either an outlier or a
high leverage point)
✔✔Leverages are - ✔✔diagonal elements in H of hat matrix
✔✔H_ii is ... - ✔✔leverage of x_i
and measures distance between x_i and the average of all x's in dataset
✔✔Purpose of leverage point - ✔✔large H_ii makes i'th residual have small variance,
regardless of y_i
, effect of x_i is overwhelming the effect of y_i
✔✔∑ H_ii = - ✔✔p+1
so average leverage = (p+1)/n
✔✔High leverage point is one that - ✔✔2-3 times or greatly exceeds (p+1)/n ,
✔✔how to find outliers? - ✔✔residual vs fitted plot
OR
any observations whose absolute standardized residuals ≥ 3
✔✔How to find influential observations - ✔✔either an outlier, or high leverage point or
both
= if Cook's distance is > 4/n
✔✔TRUE or FALSE? High leverage/outlier/influential point is same for all models? -
✔✔FALSE, might not be in another model
✔✔Collinearity - ✔✔if two predictors are not independent, and can't tell difference with
their individual effect on Y.
check correlation matrix
compute variance inflation factor
either drop one of the variables, or combine both,
✔✔Modern regression methods can outperform least-squares in terms of expected test
MSE, by - ✔✔having small bias but having much smaller variance
✔✔Logistic regression is - ✔✔a linear method in classification
✔✔classification methods - ✔✔k-nn, logistic regression, discriminant analysis
(LDA/QDA), predicting probability of each class
✔✔Logistic Function - ✔✔logistic function is equal to the odds function
The logistic function transforms the log-odds into probabilities :
Pr(y = 1 | x) / 1- Pr(y = 1 | x) = e^(beta_0+beta_1X)