ISYE 6414 EVALUATION TEST 2026 QUESTIONS AND
ANSWERS SURE A+
✔✔as the predicted
value x* is away from the average. - ✔✔The uncertainty in the estimated regression line
is going to be higher
✔✔confidence band - ✔✔If we take several values of x* and we construct such
confidence intervals, we get ___
✔✔prediction contains two sources of uncertainty: - ✔✔1. The new observation
(because we are predicting the response under a new setting)
2. the parameter estimation (of the response for x*)
✔✔Prediction interval is wider than then the confidence interval for the mean response.
- ✔✔This is because we have additional
uncertainty due to predicting under a new setting whereas the confidence intervals
under estimation are reflecting an average across all settings for that specific value
✔✔If the scatter plot of the residuals is not random around zero line - ✔✔relationship
between x and y may not be linear, or the variances of the error terms are
not equal, and the response data or the error terms are not independent
✔✔residuals are clustered in two separated clusters - ✔✔means that the residuals may
be correlated due to some clustering effect
✔✔independence - ✔✔residual analysis cannot be used to check for the
___ assumption
✔✔For checking normality, - ✔✔we can use the quantile plot, or normal probability plot
, ✔✔If some of the assumptions do not hold - ✔✔then we interpret that the model fit is
inadequate, but it does not mean that the regression is not useful.
✔✔to model the nonlinear relationship - ✔✔we can transform X by some nonlinear
function such as f(x) = x^a or f(x) = log(x)
✔✔If λ=0 - ✔✔we actually use
the normal logarithmic transformation
✔✔If λ=-1 - ✔✔use the inverse of y,
this is called the Box-Cox Transformation.
✔✔outliers, - ✔✔are data points
far from the majority of the data in x and/or y
✔✔leverage points - ✔✔Data points that are far from the
mean of the x's are
✔✔influential point - ✔✔A data point that is far from the mean of
the x's and/or the y's and influences the regression model fit significantly
✔✔When outliers belong in the data, - ✔✔you will have to
perform the statistical analysis with and without the outliers and inform the reader
about how an outlier influences the regression fit
✔✔To check outliers, - ✔✔a very simple approach is to use the standardized residuals
and compare the standardized residuals to the -2 and 2 band or even tighter, the -1 and
1
band.
✔✔coefficient of determination - ✔✔r^2. whether the linear model is useful to predict.
✔✔correlation coefficient - ✔✔approach to establish the linear relationship between two
variables.
✔✔the square of the correlation coefficients is actually - ✔✔R squared
✔✔To evaluate the constant variance and
the assumption of uncorrelated errors, - ✔✔we can use a scatter plot of the residuals vs
fitted values, which is the second plot
✔✔Testing for ß0 equal to zero means - ✔✔testing for statistical significance
✔✔We do not
ANSWERS SURE A+
✔✔as the predicted
value x* is away from the average. - ✔✔The uncertainty in the estimated regression line
is going to be higher
✔✔confidence band - ✔✔If we take several values of x* and we construct such
confidence intervals, we get ___
✔✔prediction contains two sources of uncertainty: - ✔✔1. The new observation
(because we are predicting the response under a new setting)
2. the parameter estimation (of the response for x*)
✔✔Prediction interval is wider than then the confidence interval for the mean response.
- ✔✔This is because we have additional
uncertainty due to predicting under a new setting whereas the confidence intervals
under estimation are reflecting an average across all settings for that specific value
✔✔If the scatter plot of the residuals is not random around zero line - ✔✔relationship
between x and y may not be linear, or the variances of the error terms are
not equal, and the response data or the error terms are not independent
✔✔residuals are clustered in two separated clusters - ✔✔means that the residuals may
be correlated due to some clustering effect
✔✔independence - ✔✔residual analysis cannot be used to check for the
___ assumption
✔✔For checking normality, - ✔✔we can use the quantile plot, or normal probability plot
, ✔✔If some of the assumptions do not hold - ✔✔then we interpret that the model fit is
inadequate, but it does not mean that the regression is not useful.
✔✔to model the nonlinear relationship - ✔✔we can transform X by some nonlinear
function such as f(x) = x^a or f(x) = log(x)
✔✔If λ=0 - ✔✔we actually use
the normal logarithmic transformation
✔✔If λ=-1 - ✔✔use the inverse of y,
this is called the Box-Cox Transformation.
✔✔outliers, - ✔✔are data points
far from the majority of the data in x and/or y
✔✔leverage points - ✔✔Data points that are far from the
mean of the x's are
✔✔influential point - ✔✔A data point that is far from the mean of
the x's and/or the y's and influences the regression model fit significantly
✔✔When outliers belong in the data, - ✔✔you will have to
perform the statistical analysis with and without the outliers and inform the reader
about how an outlier influences the regression fit
✔✔To check outliers, - ✔✔a very simple approach is to use the standardized residuals
and compare the standardized residuals to the -2 and 2 band or even tighter, the -1 and
1
band.
✔✔coefficient of determination - ✔✔r^2. whether the linear model is useful to predict.
✔✔correlation coefficient - ✔✔approach to establish the linear relationship between two
variables.
✔✔the square of the correlation coefficients is actually - ✔✔R squared
✔✔To evaluate the constant variance and
the assumption of uncorrelated errors, - ✔✔we can use a scatter plot of the residuals vs
fitted values, which is the second plot
✔✔Testing for ß0 equal to zero means - ✔✔testing for statistical significance
✔✔We do not