ISYE 6414 LATEST 2026 EXAM QUESTIONS AND
SOLUTIONS RATED A+
✔✔We can use the z value to determine if a coefficient is equal to zero in logistic
regression. - ✔✔True - z value = (Beta-0)/(SE of Beta)
✔✔In testing for a subset of coefficients in logistic regression the null hypothesis is that
the coefficient is equal to zero - ✔✔True
✔✔Like standard linear regression we can use the F test to test for overall regression in
logistic regression. - ✔✔False - It's 1-pchisq(null deviance-residual deviance, DFnull-
DFresidual)
✔✔For logistic regression we can define residuals for evaluating model goodness of fit
for models with and without replication. - ✔✔False - can only be with replication under
the assumption that Yi is binary and n1 is greater than 1
✔✔The deviance residuals are the signed square root of the log-likelihood evaluated at
the saturated model - ✔✔True
✔✔From the binomial approximation with a normal distribution using the central limit
theorem, the Pearson residuals have an approximately standard chi-squared
distribution. - ✔✔False - Normal distribution
✔✔Visual Analytics for logistic regression
Normal probability plot of residuals
Residuals vs predictors
Logit of success rate vs predictors - ✔✔True
Normal probability plot of residuals - Normality
Residuals vs predictors - Linearity/Independence
Logit of success rate vs predictors - Linearity
✔✔Under the null hypothesis of good fit for logistic regression, the test statistic has a
Chi-Square distribution with n- p- 1 degrees of freedom - ✔✔True - don't forget, we want
large P values
✔✔For the testing procedure for subsets of coefficients, we compare the likelihood of a
reduced model versus a full model. This is a goodness of fit test - ✔✔False - it provides
inference of the predictive power of the model
✔✔Predictive power means that the predicting variables predict the data even if one or
more of the assumptions do not hold. - ✔✔True
, ✔✔One reason why the logistic model may not fit is the relationship between logit of the
expected probability and predictors might be multiplicative, rather than additive -
✔✔True
✔✔In logistic regression for goodness of fit, we can only use the Pearson residuals. -
✔✔False - we can use Pearson or Deviance.
✔✔An indication that a higher order non linear relationship better fits the data is that the
dummy variables are all, or nearly all, statistically significant - ✔✔True
✔✔Simpson's Paradox - the reversal of association when looking at marginal vs
conditional relationships - ✔✔True
✔✔Classification is nothing else than prediction of binary responses. - ✔✔True
✔✔We cannot use the training error rate as an estimate of the true error classification
error rate because it is biased upward. - ✔✔False - biased downward
✔✔Random sampling is computationally more expensive than the K-fold cross
validation, with no clear advantage in terms of the accuracy of the estimation
classification error rate. - ✔✔True
✔✔Leave on out cross validation is preferred - ✔✔False - K fold is preferred.
✔✔The larger K is, the larger the number of folds, the less bias the estimate of the
classification the error is but has higher variability. - ✔✔True
✔✔In Poisson regression underlying assumption is that the response variable has a
Poisson distribution, or responses could be wait times, or exponential distribution -
✔✔True
✔✔The g link function is also called the canonical link function. - ✔✔True - which means
that parameter estimates under logistic regression are fully efficient and tests on those
parameters are better behaved for small samples.
✔✔Poisson distribution, the variance is equal to the expectation. Thus, the variance is
not constant - ✔✔True
✔✔For Poisson regression we estimate the expectation of the log response variable. -
✔✔False - we estimate the log of the expectation of the response variable.
SOLUTIONS RATED A+
✔✔We can use the z value to determine if a coefficient is equal to zero in logistic
regression. - ✔✔True - z value = (Beta-0)/(SE of Beta)
✔✔In testing for a subset of coefficients in logistic regression the null hypothesis is that
the coefficient is equal to zero - ✔✔True
✔✔Like standard linear regression we can use the F test to test for overall regression in
logistic regression. - ✔✔False - It's 1-pchisq(null deviance-residual deviance, DFnull-
DFresidual)
✔✔For logistic regression we can define residuals for evaluating model goodness of fit
for models with and without replication. - ✔✔False - can only be with replication under
the assumption that Yi is binary and n1 is greater than 1
✔✔The deviance residuals are the signed square root of the log-likelihood evaluated at
the saturated model - ✔✔True
✔✔From the binomial approximation with a normal distribution using the central limit
theorem, the Pearson residuals have an approximately standard chi-squared
distribution. - ✔✔False - Normal distribution
✔✔Visual Analytics for logistic regression
Normal probability plot of residuals
Residuals vs predictors
Logit of success rate vs predictors - ✔✔True
Normal probability plot of residuals - Normality
Residuals vs predictors - Linearity/Independence
Logit of success rate vs predictors - Linearity
✔✔Under the null hypothesis of good fit for logistic regression, the test statistic has a
Chi-Square distribution with n- p- 1 degrees of freedom - ✔✔True - don't forget, we want
large P values
✔✔For the testing procedure for subsets of coefficients, we compare the likelihood of a
reduced model versus a full model. This is a goodness of fit test - ✔✔False - it provides
inference of the predictive power of the model
✔✔Predictive power means that the predicting variables predict the data even if one or
more of the assumptions do not hold. - ✔✔True
, ✔✔One reason why the logistic model may not fit is the relationship between logit of the
expected probability and predictors might be multiplicative, rather than additive -
✔✔True
✔✔In logistic regression for goodness of fit, we can only use the Pearson residuals. -
✔✔False - we can use Pearson or Deviance.
✔✔An indication that a higher order non linear relationship better fits the data is that the
dummy variables are all, or nearly all, statistically significant - ✔✔True
✔✔Simpson's Paradox - the reversal of association when looking at marginal vs
conditional relationships - ✔✔True
✔✔Classification is nothing else than prediction of binary responses. - ✔✔True
✔✔We cannot use the training error rate as an estimate of the true error classification
error rate because it is biased upward. - ✔✔False - biased downward
✔✔Random sampling is computationally more expensive than the K-fold cross
validation, with no clear advantage in terms of the accuracy of the estimation
classification error rate. - ✔✔True
✔✔Leave on out cross validation is preferred - ✔✔False - K fold is preferred.
✔✔The larger K is, the larger the number of folds, the less bias the estimate of the
classification the error is but has higher variability. - ✔✔True
✔✔In Poisson regression underlying assumption is that the response variable has a
Poisson distribution, or responses could be wait times, or exponential distribution -
✔✔True
✔✔The g link function is also called the canonical link function. - ✔✔True - which means
that parameter estimates under logistic regression are fully efficient and tests on those
parameters are better behaved for small samples.
✔✔Poisson distribution, the variance is equal to the expectation. Thus, the variance is
not constant - ✔✔True
✔✔For Poisson regression we estimate the expectation of the log response variable. -
✔✔False - we estimate the log of the expectation of the response variable.