ISYE 6402 MIDTERM FINAL TEST 2026
QUESTIONS WITH CORRECT ANSWERS
GRADED A+
◍ When should you not use imputation?.
Answer: When more than 5% of the data is moving per factor
◍ What do we use when important data only appears in the validation or test
sets?.
Answer: cross-validation
◍ When do you use scaling?.
Answer: Data in a bounded range (e.g., neural networks, RGB values, SAT
scores, batting averages)
◍ when you have lower -pvalues.
Answer: increase the possibility of leaving out a relevant factor
◍ What does 2 < |BIC_1 - BIC_2| < 6 mean?.
Answer: smaller BIC model is "somewhat likely" better
◍ When you have higher p-values ....
Answer: increase the possibility of including irrelevant factors
◍ When is ARIMA better than exponential smoothing?.
Answer: When the data is less stable
◍ What are the parts of a box-and-whisker plot?.
Answer: The bottom and top of the box are the 25th and 75th percentile. The
middle valu is the median. The whiskers stretch up and down to reasonable
range of values (10 and 90th or 5th and 95 percentiles)
◍ What effects does randomness have on training /validation performance?.
Answer: sometimes the randomness will make the performance look worse
, than it really is, and sometimes the randomness will make the performance
look better than it really is
◍ In a simple linear regression analysis, the estimated regression equation is
given by:ŷ = 2.7x − 1.4The sum of squared residuals (SSE) is 45.2, and the
total sum of squares (SST) is 80.What is the value of the correlation
coefficient squared (ρ²)?0.5650.4350.3650.751.
Answer: 0.435Check Module 1 Topic 1.3 Lesson 9Formula: R² = 1 − (SSE /
SST)→ R² = 1 − (45.) = 0.435
◍ what's an example of heuristic optimization.
Answer: what's the best buffer size to have at each step in the process
◍ When would you use Corrected AIC and not AIC.
Answer: when you have smaller data sets
◍ What are random effects?.
Answer: They are random but look like real effects. They are different in all
data sets.
◍ What are the pros and cons of Greedy Algorithms (Forward selection,
stepwise elimination, stepwise regression).
Answer: Good for initial analysis but often don't perform as well on other
data because they fit more to random effects than you'd like and appear to
have a better fit
◍ Why are simple models better than complex ones.
Answer: less data is required; less chance of insignificant factors and easier
to interpret
◍ When do you use standardization?.
Answer: PCA or clustering
◍ What is structured data?.
Answer: Data that can be stores in a structured way
◍ Why do we need to use the same number of random numbers (seed).
Answer: Because if we wouldn't be able to accurately compare the two sets
, of replications because the replications while the same distribution are still
different
◍ what is the objective function for a time series model.
Answer: minimize prediction error
◍ Why is k-means an expectation-maximization.
Answer: finding the mean of all the points in cluster is similar to finding an
expectation.Assigning data points to cluster centers is the maximization
step. Really we are minimizing, but we could think of it as maximizing the
negative of the distance to a cluster center
◍ If we have a smaller C ....
Answer: the more sensitive the method is because S_t can get larger faster
◍ what is modularity?.
Answer: a measure of how well the graph is separated into communities or
modules that are connected a lot internally but not connected much between
each other.
◍ Whey do you focus on the first n principal components?.
Answer: Reduces the effect of randomness and earlier principal components
are likely to have higher signal to-noise ratios
◍ what is a Bernoulli distribution.
Answer: it's like a flipping coin. It can be used to model a single event and is
most useful when we put many of them together
◍ why is calculating p-values for the Wilcoxon Signed Rank Test is harder
than McNemar's test.
Answer: it is like a a normal distribution test unlike McNemar's which is
uses the binomial distribution
◍ How do you detrend data?.
Answer: Factor-by-factor. You fit a one-dimensional linear regression to the
data and subtract
◍ What does 0 < |BIC_1 - BIC_2| < 2 mean?.
, Answer: smaller BIC model is "slightly likely" better
◍ When clustering for prediction how do we choose the prediction?.
Answer: When we see a new point, we just choose whichever cluster center
is closest.
◍ what are the benefits of k-fold cross validation?.
Answer: better use of data, better estimate of model quality, and chooses
model more effectively
◍ how could we detect outliers when there are multiple dimensions?.
Answer: we could fit a model and then determine the points with a large
error
◍ When do we need a validation set?.
Answer: When we are choosing between multiple models.
◍ How do you compare two AICs?.
Answer:
◍ what is an optimal solution.
Answer: feasible solution with the best objective value
◍ Both the true errors and the model residuals in multiple linear regression
have constant variance. (T/F).
Answer: False(False. While the true errors have constant variance, the
estimated residuals do not)
◍ Do most systems exhibit the memoryless property?.
Answer: thing usually depend on the bast
◍ what is a convex optimization problem.
Answer: objective f(x) is concave (if maximizing) or convex (if
minimizing). Constraint set X is a convex set
◍ In ANOVA with k population samples, the sampling distribution of the
pooled variance is a chi-square distribution with N - 2 degrees of freedom.
(T/F).
Answer: False (False, the sampling distribution of the pooled variance is a
QUESTIONS WITH CORRECT ANSWERS
GRADED A+
◍ When should you not use imputation?.
Answer: When more than 5% of the data is moving per factor
◍ What do we use when important data only appears in the validation or test
sets?.
Answer: cross-validation
◍ When do you use scaling?.
Answer: Data in a bounded range (e.g., neural networks, RGB values, SAT
scores, batting averages)
◍ when you have lower -pvalues.
Answer: increase the possibility of leaving out a relevant factor
◍ What does 2 < |BIC_1 - BIC_2| < 6 mean?.
Answer: smaller BIC model is "somewhat likely" better
◍ When you have higher p-values ....
Answer: increase the possibility of including irrelevant factors
◍ When is ARIMA better than exponential smoothing?.
Answer: When the data is less stable
◍ What are the parts of a box-and-whisker plot?.
Answer: The bottom and top of the box are the 25th and 75th percentile. The
middle valu is the median. The whiskers stretch up and down to reasonable
range of values (10 and 90th or 5th and 95 percentiles)
◍ What effects does randomness have on training /validation performance?.
Answer: sometimes the randomness will make the performance look worse
, than it really is, and sometimes the randomness will make the performance
look better than it really is
◍ In a simple linear regression analysis, the estimated regression equation is
given by:ŷ = 2.7x − 1.4The sum of squared residuals (SSE) is 45.2, and the
total sum of squares (SST) is 80.What is the value of the correlation
coefficient squared (ρ²)?0.5650.4350.3650.751.
Answer: 0.435Check Module 1 Topic 1.3 Lesson 9Formula: R² = 1 − (SSE /
SST)→ R² = 1 − (45.) = 0.435
◍ what's an example of heuristic optimization.
Answer: what's the best buffer size to have at each step in the process
◍ When would you use Corrected AIC and not AIC.
Answer: when you have smaller data sets
◍ What are random effects?.
Answer: They are random but look like real effects. They are different in all
data sets.
◍ What are the pros and cons of Greedy Algorithms (Forward selection,
stepwise elimination, stepwise regression).
Answer: Good for initial analysis but often don't perform as well on other
data because they fit more to random effects than you'd like and appear to
have a better fit
◍ Why are simple models better than complex ones.
Answer: less data is required; less chance of insignificant factors and easier
to interpret
◍ When do you use standardization?.
Answer: PCA or clustering
◍ What is structured data?.
Answer: Data that can be stores in a structured way
◍ Why do we need to use the same number of random numbers (seed).
Answer: Because if we wouldn't be able to accurately compare the two sets
, of replications because the replications while the same distribution are still
different
◍ what is the objective function for a time series model.
Answer: minimize prediction error
◍ Why is k-means an expectation-maximization.
Answer: finding the mean of all the points in cluster is similar to finding an
expectation.Assigning data points to cluster centers is the maximization
step. Really we are minimizing, but we could think of it as maximizing the
negative of the distance to a cluster center
◍ If we have a smaller C ....
Answer: the more sensitive the method is because S_t can get larger faster
◍ what is modularity?.
Answer: a measure of how well the graph is separated into communities or
modules that are connected a lot internally but not connected much between
each other.
◍ Whey do you focus on the first n principal components?.
Answer: Reduces the effect of randomness and earlier principal components
are likely to have higher signal to-noise ratios
◍ what is a Bernoulli distribution.
Answer: it's like a flipping coin. It can be used to model a single event and is
most useful when we put many of them together
◍ why is calculating p-values for the Wilcoxon Signed Rank Test is harder
than McNemar's test.
Answer: it is like a a normal distribution test unlike McNemar's which is
uses the binomial distribution
◍ How do you detrend data?.
Answer: Factor-by-factor. You fit a one-dimensional linear regression to the
data and subtract
◍ What does 0 < |BIC_1 - BIC_2| < 2 mean?.
, Answer: smaller BIC model is "slightly likely" better
◍ When clustering for prediction how do we choose the prediction?.
Answer: When we see a new point, we just choose whichever cluster center
is closest.
◍ what are the benefits of k-fold cross validation?.
Answer: better use of data, better estimate of model quality, and chooses
model more effectively
◍ how could we detect outliers when there are multiple dimensions?.
Answer: we could fit a model and then determine the points with a large
error
◍ When do we need a validation set?.
Answer: When we are choosing between multiple models.
◍ How do you compare two AICs?.
Answer:
◍ what is an optimal solution.
Answer: feasible solution with the best objective value
◍ Both the true errors and the model residuals in multiple linear regression
have constant variance. (T/F).
Answer: False(False. While the true errors have constant variance, the
estimated residuals do not)
◍ Do most systems exhibit the memoryless property?.
Answer: thing usually depend on the bast
◍ what is a convex optimization problem.
Answer: objective f(x) is concave (if maximizing) or convex (if
minimizing). Constraint set X is a convex set
◍ In ANOVA with k population samples, the sampling distribution of the
pooled variance is a chi-square distribution with N - 2 degrees of freedom.
(T/F).
Answer: False (False, the sampling distribution of the pooled variance is a