ISYE 6402 MIDTERM ACTUAL EXAM PAPER
2026 QUESTIONS WITH ANSWERS
GRADED A+
◍ Point outlier.
Answer: A data point that is (uncommonly) far from other data points.
◍ False positive (FP).
Answer: Data point that a model incorrectly classifies as being in a certain
category. Sometimes abbreviated as "FP".
◍ Can we still make predictions with a model if the predictor and response are
highly correlated.
Answer: yes. Even though we can use it for empirical predictions, it doesn't
make sense to say that the model shows causation
◍ Setting a large value of k will ....
Answer: lead to a large model bias and lower variance.
◍ What do predictive questions ask?.
Answer: What will happen? (e.g., what will Google's stock price be?)
◍ What is C_t.
Answer: the multiplicative seasonality factor of time. It inflates or delates
the observation
◍ What does 6 < |BIC_1 - BIC_2| < 10 mean?.
Answer: smaller BIC model is "likely" better
◍ What parameter does GARCH not have the ARIMA has?.
Answer: d because GARCH doesn't deal with differences.
◍ Classification models.
Answer: CART, k-nearest-neighbor, logistic regression, random forest,
, support vector machine
◍ Q-Q plot Quantile-quantile plot.
Answer: a plot comparing the quantiles of two data sets, or one data set and
a distribution, to see whether they might have a common distribution.
◍ Why would we use two sets?.
Answer: Reason to use two different sets is because if the first set, the
training set, had unique random effects that the classifer was designed for,
we wouldn't be counting those benefits when we measure effectiveness on
the validation set.
◍ Smoothing.
Answer: Time series analysis technique to help filter out underlying
randomness/noise. Examples include moving average, exponential
smoothing, and ARIMA.
◍ Eigenvalue.
Answer: Amount by which an eigenvector gets rescaled in a linear
transformation.
◍ What is a simple linear regression?.
Answer: Linear regression with one predictor.
◍ Test set.
Answer: Estimate quality of selective model
◍ A logistic regression model can be especially useful when the response....
Answer: ...is a probability (a number between zero and one) or is binary
(either zero or one).
◍ p-value (regression).
Answer: Probability that results at least as extreme as those in the data
would be observed if the coefficient of a variable is zero.
◍ What will Regression tell you?.
Answer: How systems work (descriptive questions) and what will happen in
the future (prescriptive questions)
,◍ Additive seasonality.
Answer: Seasonal effect that is added to a baseline value.
◍ what is the initial condition for S_t?.
Answer: S_1 = x_1
◍ Wilcoxon signed rank test (one sample).
Answer: Nonparametric test for a single response, to determining whether
the median is different from a specific value.
◍ Decision.
Answer: Choice of action.
◍ what is the Area Under Curve.
Answer: probability that the model estimates a random "yes" point higher
than a random "no" point
◍ Maximum flow problem.
Answer: Network optimization model that finds the most flow that can be
sent from one specific node to another.
◍ what is the R-squared value?.
Answer: estimates how much variability the model accounts for
◍ Algorithm.
Answer: Step-by-step procedure designed to carry out a task.
◍ For ARIMA, the D parameter is used to specify ___..
Answer: , The order, or the differences of the differences of the differences
(d-times.)
◍ What do descriptive questions ask?.
Answer: What happened? (e.g., which customers are most alike)
◍ A time series outlier that seems "off the curve" is called a....
Answer: contextual outlier.
◍ The farther the wrongly classified point is from the line ___.
Answer: The bigger the mistake we've made
, ◍ Rectilinear distance.
Answer: The sum of the lengths in each dimension between two points. Also
called the Manhattan or 1-norm distance.
◍ Box and whisker plot.
Answer: Graphical representation data showing the middle range of data
(the "box"), reasonable ranges of variability ("whiskers"), and points
(possible outliers) outside those ranges.
◍ Heteroscedasticity.
Answer: When the variability of a response is different across the range of
predictor values.
◍ What are the steps of k means?.
Answer: 0. Pick k clusters within range of data.1. Assign each data point to
nearest cluster center2. Recalculate cluster centers (centroids)3. Repeat 1
and 2 until no changes
◍ What is the formula for AIC?.
Answer: AIC=2k - 2ln(L) where L* is the maximum likelihood value and K
is the number of parameters estimated.
◍ How does the exponential smoothing formula weight more recent
observations more than older ones?.
Answer: (1 - alpha) < 1
◍ Categorical data.
Answer: Data that classifies observations without quantitative meaning (for
example, colors of cars) or where quantitative amounts are categorized(for
example, "0-10, 11-20, ...").
◍ Random forest.
Answer: Machine learning model that creates many different trees and
returns their mean output. Can be used with classification trees, regression
trees, decision trees.
◍ Holt-Winters method/Winters' method.
2026 QUESTIONS WITH ANSWERS
GRADED A+
◍ Point outlier.
Answer: A data point that is (uncommonly) far from other data points.
◍ False positive (FP).
Answer: Data point that a model incorrectly classifies as being in a certain
category. Sometimes abbreviated as "FP".
◍ Can we still make predictions with a model if the predictor and response are
highly correlated.
Answer: yes. Even though we can use it for empirical predictions, it doesn't
make sense to say that the model shows causation
◍ Setting a large value of k will ....
Answer: lead to a large model bias and lower variance.
◍ What do predictive questions ask?.
Answer: What will happen? (e.g., what will Google's stock price be?)
◍ What is C_t.
Answer: the multiplicative seasonality factor of time. It inflates or delates
the observation
◍ What does 6 < |BIC_1 - BIC_2| < 10 mean?.
Answer: smaller BIC model is "likely" better
◍ What parameter does GARCH not have the ARIMA has?.
Answer: d because GARCH doesn't deal with differences.
◍ Classification models.
Answer: CART, k-nearest-neighbor, logistic regression, random forest,
, support vector machine
◍ Q-Q plot Quantile-quantile plot.
Answer: a plot comparing the quantiles of two data sets, or one data set and
a distribution, to see whether they might have a common distribution.
◍ Why would we use two sets?.
Answer: Reason to use two different sets is because if the first set, the
training set, had unique random effects that the classifer was designed for,
we wouldn't be counting those benefits when we measure effectiveness on
the validation set.
◍ Smoothing.
Answer: Time series analysis technique to help filter out underlying
randomness/noise. Examples include moving average, exponential
smoothing, and ARIMA.
◍ Eigenvalue.
Answer: Amount by which an eigenvector gets rescaled in a linear
transformation.
◍ What is a simple linear regression?.
Answer: Linear regression with one predictor.
◍ Test set.
Answer: Estimate quality of selective model
◍ A logistic regression model can be especially useful when the response....
Answer: ...is a probability (a number between zero and one) or is binary
(either zero or one).
◍ p-value (regression).
Answer: Probability that results at least as extreme as those in the data
would be observed if the coefficient of a variable is zero.
◍ What will Regression tell you?.
Answer: How systems work (descriptive questions) and what will happen in
the future (prescriptive questions)
,◍ Additive seasonality.
Answer: Seasonal effect that is added to a baseline value.
◍ what is the initial condition for S_t?.
Answer: S_1 = x_1
◍ Wilcoxon signed rank test (one sample).
Answer: Nonparametric test for a single response, to determining whether
the median is different from a specific value.
◍ Decision.
Answer: Choice of action.
◍ what is the Area Under Curve.
Answer: probability that the model estimates a random "yes" point higher
than a random "no" point
◍ Maximum flow problem.
Answer: Network optimization model that finds the most flow that can be
sent from one specific node to another.
◍ what is the R-squared value?.
Answer: estimates how much variability the model accounts for
◍ Algorithm.
Answer: Step-by-step procedure designed to carry out a task.
◍ For ARIMA, the D parameter is used to specify ___..
Answer: , The order, or the differences of the differences of the differences
(d-times.)
◍ What do descriptive questions ask?.
Answer: What happened? (e.g., which customers are most alike)
◍ A time series outlier that seems "off the curve" is called a....
Answer: contextual outlier.
◍ The farther the wrongly classified point is from the line ___.
Answer: The bigger the mistake we've made
, ◍ Rectilinear distance.
Answer: The sum of the lengths in each dimension between two points. Also
called the Manhattan or 1-norm distance.
◍ Box and whisker plot.
Answer: Graphical representation data showing the middle range of data
(the "box"), reasonable ranges of variability ("whiskers"), and points
(possible outliers) outside those ranges.
◍ Heteroscedasticity.
Answer: When the variability of a response is different across the range of
predictor values.
◍ What are the steps of k means?.
Answer: 0. Pick k clusters within range of data.1. Assign each data point to
nearest cluster center2. Recalculate cluster centers (centroids)3. Repeat 1
and 2 until no changes
◍ What is the formula for AIC?.
Answer: AIC=2k - 2ln(L) where L* is the maximum likelihood value and K
is the number of parameters estimated.
◍ How does the exponential smoothing formula weight more recent
observations more than older ones?.
Answer: (1 - alpha) < 1
◍ Categorical data.
Answer: Data that classifies observations without quantitative meaning (for
example, colors of cars) or where quantitative amounts are categorized(for
example, "0-10, 11-20, ...").
◍ Random forest.
Answer: Machine learning model that creates many different trees and
returns their mean output. Can be used with classification trees, regression
trees, decision trees.
◍ Holt-Winters method/Winters' method.