PSTAT 131 ACTUAL FINAL ANSWERS AND
QUESTIONS SET A+
✔✔What is Pr(Y = 1 | X) - ✔✔refers to the probability that the dependent variable Y
takes the value 1 (or the positive class) given a set of predictor variables X
probability that the result will fall in Y=1 class, given predictors.
✔✔In logistic function,
if beta_1 > 0 , then
if beta_1 < 0 , then - ✔✔X increases => Pr(y = 1| X) also increases
X increases => Pr(y = 1| X) decreases
The rate of change in Pr(Y = 1|X) per unit change in X depend on the current value of X
(not linear). same for multiple logistic regression
✔✔Confounding - ✔✔be careful when regression is only performed using 1 predictor,
when multiple may be relevant
✔✔Simpson's paradox - ✔✔a trend appears in different groups of data but disappears
or reverses when these groups are combined. In other words, the direction of an
association or the magnitude of an effect can be reversed when the data are
aggregated, even though each subgroup shows a consistent trend.
✔✔threshold probability - ✔✔after doing Pr(y =1 |x), we get a probability for
classification right?
then we set a threshold, that all probabilities greater than 0.6 or some value belongs to
1 class.
not necessary that the predicted probability has to be split at 0.5
✔✔Confusion Matrix - ✔✔to visualize TPR/FPR for binary classifier
,True positive rate (TPR) = TP / (TP + FN), a.k.a. 1 - Type II error, power, sensitivity,
recall
False positive rate (FPR) = FP / (FP + TN), a.k.a. Type I error, 1 - Specificity
✔✔ROC - ✔✔- displaying the 2 types of error, for all threshold values
- TPR on y axis FPR on x axis
✔✔AUC - ✔✔Area under curve (AUC) measures the overall performance of a classifier
Area under curve (AUC) summarized over all possible thresholds
between 0 - 1, higher the better.
✔✔Linear Discriminant Analysis (LDA) - ✔✔linear method in classification
Purpose:
1. if logistic regression coefficients are unstable
2. when n is small and distribution is close to x
3. multi class classification
4. dimensionality reduction while minimizing variance
How:
Uses Baye's theorem to make assumptions on predictors given Y.
Prior Probability: pi_k = Pr(Y = k)
Baye's theorem = prior probability * density function / SUM {pi_j * f_j(x1, .., xp)}
assume density function at Y=k is normal distribution density, if and only if (X1, ... , XP) |
Y is normally distributed
✔✔LDA vs PCA - ✔✔both reduce dimensionality by creating new axis, where first axis
is best at accounting for variation or separating categories
PCA focuses on data with most variation
LDA focuses on maximizing separability among known categories, creates a new axis
and projects data onto the new axis to separate two categories.
PCA doesn't require class labels (unsupervised)
LDA assumes normal distribution (LDA within each class, PCA within all the data)
✔✔What does each component in PCA represent? - ✔✔PC1 (the first new axis that
PCA creates) accounts for the most variation in data, PC2 does second best
, ✔✔LDA: p = 1 case - ✔✔Baye's classifier assigns the observation X to class k for which
Pr(Y=k | X=x) is the highest .
EQUIVALENT TO
classifier assigns observations to class j where δj is highest below so where
δk(x)=x⋅ Mu_k/sigma^2 − Mu^2_k/2sigma^2 +log(πk) is largest
μ1, μ2, ..., μK, π1, ..., πK, σ2 are unknown! estimate them
Estimates:
δk(x)_hat = discriminant function
Mu_k/sigma^2 _ hat = linear in X
Bayesian decision boundary follows normal smooth distribution
LDA falls normal histogram looking distribution
different mean for each k, same variance
✔✔LDA: p > 1 - ✔✔different mean for each k, same variance
bayes classifier assigns the observation X to class for which
δk(x).... is largest
but we don't know mu_k, pi_k, and δk, so estimate them using LDA estimator
SO: LDA is used to estimate those BAYES CLASSIFIER can't be used
THUS: bayesian decision boundary is better, but LDA is more feasible.
✔✔Glm vs glmnet - ✔✔Logistic (no penalty )vs penalties for ridge/lasso
✔✔When is validation set used in train/test process? - ✔✔Used in CV, same as training
data for this course
✔✔In LDA, π̂_k is.? - ✔✔n_k/n = # training observations in k_th class/total # training
observations
✔✔Quadratic Discriminant Analysis (QDA) - ✔✔different mean, different covariance
same estimating for when δk(x) is largest, where X term is quadratic in X
QUESTIONS SET A+
✔✔What is Pr(Y = 1 | X) - ✔✔refers to the probability that the dependent variable Y
takes the value 1 (or the positive class) given a set of predictor variables X
probability that the result will fall in Y=1 class, given predictors.
✔✔In logistic function,
if beta_1 > 0 , then
if beta_1 < 0 , then - ✔✔X increases => Pr(y = 1| X) also increases
X increases => Pr(y = 1| X) decreases
The rate of change in Pr(Y = 1|X) per unit change in X depend on the current value of X
(not linear). same for multiple logistic regression
✔✔Confounding - ✔✔be careful when regression is only performed using 1 predictor,
when multiple may be relevant
✔✔Simpson's paradox - ✔✔a trend appears in different groups of data but disappears
or reverses when these groups are combined. In other words, the direction of an
association or the magnitude of an effect can be reversed when the data are
aggregated, even though each subgroup shows a consistent trend.
✔✔threshold probability - ✔✔after doing Pr(y =1 |x), we get a probability for
classification right?
then we set a threshold, that all probabilities greater than 0.6 or some value belongs to
1 class.
not necessary that the predicted probability has to be split at 0.5
✔✔Confusion Matrix - ✔✔to visualize TPR/FPR for binary classifier
,True positive rate (TPR) = TP / (TP + FN), a.k.a. 1 - Type II error, power, sensitivity,
recall
False positive rate (FPR) = FP / (FP + TN), a.k.a. Type I error, 1 - Specificity
✔✔ROC - ✔✔- displaying the 2 types of error, for all threshold values
- TPR on y axis FPR on x axis
✔✔AUC - ✔✔Area under curve (AUC) measures the overall performance of a classifier
Area under curve (AUC) summarized over all possible thresholds
between 0 - 1, higher the better.
✔✔Linear Discriminant Analysis (LDA) - ✔✔linear method in classification
Purpose:
1. if logistic regression coefficients are unstable
2. when n is small and distribution is close to x
3. multi class classification
4. dimensionality reduction while minimizing variance
How:
Uses Baye's theorem to make assumptions on predictors given Y.
Prior Probability: pi_k = Pr(Y = k)
Baye's theorem = prior probability * density function / SUM {pi_j * f_j(x1, .., xp)}
assume density function at Y=k is normal distribution density, if and only if (X1, ... , XP) |
Y is normally distributed
✔✔LDA vs PCA - ✔✔both reduce dimensionality by creating new axis, where first axis
is best at accounting for variation or separating categories
PCA focuses on data with most variation
LDA focuses on maximizing separability among known categories, creates a new axis
and projects data onto the new axis to separate two categories.
PCA doesn't require class labels (unsupervised)
LDA assumes normal distribution (LDA within each class, PCA within all the data)
✔✔What does each component in PCA represent? - ✔✔PC1 (the first new axis that
PCA creates) accounts for the most variation in data, PC2 does second best
, ✔✔LDA: p = 1 case - ✔✔Baye's classifier assigns the observation X to class k for which
Pr(Y=k | X=x) is the highest .
EQUIVALENT TO
classifier assigns observations to class j where δj is highest below so where
δk(x)=x⋅ Mu_k/sigma^2 − Mu^2_k/2sigma^2 +log(πk) is largest
μ1, μ2, ..., μK, π1, ..., πK, σ2 are unknown! estimate them
Estimates:
δk(x)_hat = discriminant function
Mu_k/sigma^2 _ hat = linear in X
Bayesian decision boundary follows normal smooth distribution
LDA falls normal histogram looking distribution
different mean for each k, same variance
✔✔LDA: p > 1 - ✔✔different mean for each k, same variance
bayes classifier assigns the observation X to class for which
δk(x).... is largest
but we don't know mu_k, pi_k, and δk, so estimate them using LDA estimator
SO: LDA is used to estimate those BAYES CLASSIFIER can't be used
THUS: bayesian decision boundary is better, but LDA is more feasible.
✔✔Glm vs glmnet - ✔✔Logistic (no penalty )vs penalties for ridge/lasso
✔✔When is validation set used in train/test process? - ✔✔Used in CV, same as training
data for this course
✔✔In LDA, π̂_k is.? - ✔✔n_k/n = # training observations in k_th class/total # training
observations
✔✔Quadratic Discriminant Analysis (QDA) - ✔✔different mean, different covariance
same estimating for when δk(x) is largest, where X term is quadratic in X