ISYE 6414 STUDY EXAMS GUIDE QUESTIONS AND
ANSWERS SURE A+
✔✔Another GOF test is hypotheses testing, what is the H0 and HA? - ✔✔H0: is that the
model fits well.
HA: the alternative is that the model does not fit well.
✔✔The test statistic for the goodness of fit test is? - ✔✔sum of squared deviances
✔✔Under the null hypothesis of good fit, the test statistic's (sum of squared deviances)
distribution and DOF is...? - ✔✔Chi square with n-p-1 DF
✔✔For the G-o-F tests, do we reject the null hypotheses when the p-value is SMALL or
BIG? - ✔✔Small and we conclude the model is not a good fit
✔✔For goodness of fit test, we compare the likelihoods of the...? - ✔✔saturated model
versus the fitted model
✔✔What inference does 'testing a subset of coefficients' provide? - ✔✔provides
inferences on the predictive power of the model
✔✔What inference does 'saturated vs fitted model' provide? - ✔✔provides inferences on
the goodness of the model
✔✔Goodness of fit - ✔✔means that the model assumptions hold and fits the data well.
✔✔Predictive power - ✔✔means that the predicting variables predict the data even if
one or more of the assumptions do not hold.
✔✔What are 4 reasons why the logistic model might not be a good fit? - ✔✔1. there
may be other variables that should be included in the model and/or the relationship
, between logit of the expected probability and predictors might be multiplicative rather
than additive.
2. Initial observation outliers and leverage points are also still an issue for this model.
The model should be fitted with and without these outliers.
3. the binomial distribution isn't appropriate (overdispersion)
4. the logit function does not fit well with the data. (there could be other s shaped
functions that would work better)
✔✔How can we find that a normality transformation of the predictive variable will
improve the fit? - ✔✔Such transformation can be identified by comparing the logit of the
success rate, versus the predicted variables.
✔✔Overdispersion - ✔✔where the variability of the response variable is larger than
estimated by the model.
✔✔canonical link function - ✔✔which means that parameter estimates under logistic
regression are fully efficient and tests on those parameters are better behaved for small
samples
✔✔Simpson's paradox - ✔✔refers to reversal of an association when looking at a
marginal relationship versus a partial or conditional one. This is a situation where the
marginal relationship has a wrong sign.
✔✔marginal relationship - ✔✔Capturing the association of a predicting variable to the
response variable marginally, i.e. without consideration of other factors.
✔✔conditional relationship - ✔✔Capturing the association of a predicting variable to the
response variable, conditional of other predicting variables in the model.
✔✔Classification - ✔✔The prediction of binary responses. Classification is nothing more
than a prediction of the class of your response, y* (y star), given the predictor variable,
x* (x star). If the predicted probability is large, then classify y* as a success.
✔✔classification error rate - ✔✔the probability that the new response is equal to the
classifier.
✔✔How do we compute classification error? (2 ways) - ✔✔1. Training error 2. Cross
validation
✔✔training error - ✔✔simply use the data to fit the model then compute the classifier
from each response in the data and take the proportion of the responses we
misclassified
ANSWERS SURE A+
✔✔Another GOF test is hypotheses testing, what is the H0 and HA? - ✔✔H0: is that the
model fits well.
HA: the alternative is that the model does not fit well.
✔✔The test statistic for the goodness of fit test is? - ✔✔sum of squared deviances
✔✔Under the null hypothesis of good fit, the test statistic's (sum of squared deviances)
distribution and DOF is...? - ✔✔Chi square with n-p-1 DF
✔✔For the G-o-F tests, do we reject the null hypotheses when the p-value is SMALL or
BIG? - ✔✔Small and we conclude the model is not a good fit
✔✔For goodness of fit test, we compare the likelihoods of the...? - ✔✔saturated model
versus the fitted model
✔✔What inference does 'testing a subset of coefficients' provide? - ✔✔provides
inferences on the predictive power of the model
✔✔What inference does 'saturated vs fitted model' provide? - ✔✔provides inferences on
the goodness of the model
✔✔Goodness of fit - ✔✔means that the model assumptions hold and fits the data well.
✔✔Predictive power - ✔✔means that the predicting variables predict the data even if
one or more of the assumptions do not hold.
✔✔What are 4 reasons why the logistic model might not be a good fit? - ✔✔1. there
may be other variables that should be included in the model and/or the relationship
, between logit of the expected probability and predictors might be multiplicative rather
than additive.
2. Initial observation outliers and leverage points are also still an issue for this model.
The model should be fitted with and without these outliers.
3. the binomial distribution isn't appropriate (overdispersion)
4. the logit function does not fit well with the data. (there could be other s shaped
functions that would work better)
✔✔How can we find that a normality transformation of the predictive variable will
improve the fit? - ✔✔Such transformation can be identified by comparing the logit of the
success rate, versus the predicted variables.
✔✔Overdispersion - ✔✔where the variability of the response variable is larger than
estimated by the model.
✔✔canonical link function - ✔✔which means that parameter estimates under logistic
regression are fully efficient and tests on those parameters are better behaved for small
samples
✔✔Simpson's paradox - ✔✔refers to reversal of an association when looking at a
marginal relationship versus a partial or conditional one. This is a situation where the
marginal relationship has a wrong sign.
✔✔marginal relationship - ✔✔Capturing the association of a predicting variable to the
response variable marginally, i.e. without consideration of other factors.
✔✔conditional relationship - ✔✔Capturing the association of a predicting variable to the
response variable, conditional of other predicting variables in the model.
✔✔Classification - ✔✔The prediction of binary responses. Classification is nothing more
than a prediction of the class of your response, y* (y star), given the predictor variable,
x* (x star). If the predicted probability is large, then classify y* as a success.
✔✔classification error rate - ✔✔the probability that the new response is equal to the
classifier.
✔✔How do we compute classification error? (2 ways) - ✔✔1. Training error 2. Cross
validation
✔✔training error - ✔✔simply use the data to fit the model then compute the classifier
from each response in the data and take the proportion of the responses we
misclassified