100% Correct - GT Actual Exam 2026/2027 | Complete Exam-
Style Questions with Detailed Rationales | 100% Verified | Pass
Guaranteed – A+ Graded
Classification and Regression Methods
Q1: A data scientist is building a k-Nearest Neighbors (k-NN) model to predict customer
churn based on features with vastly different scales, such as age and annual income. What is
the most critical preprocessing step before training this model?
A. Apply one-hot encoding to all numerical features to create binary indicators.
B. Standardize or normalize the features so that distance calculations are not dominated by
variables with larger magnitudes. [CORRECT]
C. Increase the value of k to ensure the model captures global patterns rather than local
noise.
D. Remove all categorical variables since k-NN cannot handle any non-numeric data types.
Correct Answer: B
Rationale: This choice is correct because k-NN relies on distance metrics like Euclidean
distance, which are highly sensitive to the scale of the input features. Standardizing ensures
each feature contributes equally to the distance calculation.
Q2: In a logistic regression model predicting loan default, the coefficient for "credit score" is
-0.05. How should this be interpreted?
A. For every one-unit increase in credit score, the probability of default decreases by exactly
5%.
B. For every one-unit increase in credit score, the log-odds of default decrease by 0.05,
holding other variables constant. [CORRECT]
C. A higher credit score causes a 0.05 increase in the odds of default.
D. The model is overfitting, as coefficients in logistic regression must always be positive.
Correct Answer: B
Rationale: This choice is correct because logistic regression coefficients represent the
change in the log-odds of the outcome for a one-unit change in the predictor. The negative
sign indicates an inverse relationship with the probability of default.
Q3: A team trains a Classification and Regression Tree (CART) model that achieves 98%
accuracy on the training data but only 65% on the validation set. Which technique should
they apply to address this issue?
A. Increase the maximum depth of the tree to capture more complex interactions.
B. Apply cost-complexity pruning to reduce the tree size and prevent overfitting. [CORRECT]
C. Switch from Gini impurity to entropy, as it inherently prevents overfitting.
D. Remove the validation set and retrain on the entire dataset to maximize learning.
Correct Answer: B
,Rationale: This choice is correct because the large gap between training and validation
performance indicates overfitting, which is common in deep, unpruned trees. Pruning
simplifies the model, improving its generalization to unseen data.
Q4: When evaluating a Random Forest model, what does the "out-of-bag" (OOB) error
estimate represent?
A. The error rate calculated on a separate, held-out test set that was never seen during
training.
B. The average error rate of predictions made for each observation using only the trees that
did not include that observation in their bootstrap sample. [CORRECT]
C. The error rate of the model when all features are randomly permuted to assess feature
importance.
D. The difference between the training error and the validation error across all folds of
cross-validation.
Correct Answer: B
Rationale: This choice is correct because the OOB error leverages the roughly one-third of
data not selected in each bootstrap sample as a built-in validation set. This provides an
unbiased estimate of model performance without needing a separate validation split.
Q5: A data scientist evaluates an SVM model and observes the following validation metrics:
Accuracy 85%, Precision 90%, Recall 60%. The business stakeholder emphasizes that missing
a positive case (e.g., a fraudulent transaction) is highly costly. Which action should the data
scientist prioritize?
A. Increase the SVM penalty parameter C to harden the margin and reduce training errors.
B. Adjust the classification threshold to favor higher recall, accepting a decrease in precision.
[CORRECT]
C. Switch to a linear kernel to simplify the decision boundary and speed up inference.
D. Reduce the training dataset size to prevent the model from memorizing noise.
Correct Answer: B
Rationale: This choice is correct because the business context explicitly prioritizes catching
positive cases, making recall the critical metric. Adjusting the decision threshold directly
trades some precision to capture more true positives, aligning with the stakeholder's cost
structure.
Q6: Which of the following scenarios is most appropriate for using logistic regression rather
than a k-NN classifier?
A. The decision boundary between classes is highly non-linear and complex.
B. The dataset is very small, and the model needs to provide interpretable coefficients for
stakeholder review. [CORRECT]
C. The features are entirely categorical with no inherent ordinal relationship.
D. The primary goal is to maximize training accuracy regardless of model complexity.
Correct Answer: B
Rationale: This choice is correct because logistic regression provides clear, interpretable
coefficients that explain the direction and magnitude of each feature's effect. It also
performs well on smaller datasets where k-NN might suffer from the curse of dimensionality
or sparse neighborhoods.
, Q7: In a CART model, what is the primary purpose of the Gini impurity metric during the
splitting process?
A. To measure the total variance of the target variable within a node.
B. To evaluate the likelihood of incorrect classification of a randomly chosen element if it
were labeled according to the node's distribution. [CORRECT]
C. To calculate the distance between the current node and the root node.
D. To penalize the model for adding too many splits, thereby controlling overfitting.
Correct Answer: B
Rationale: This choice is correct because Gini impurity quantifies the heterogeneity of a
node. The algorithm selects splits that maximize the reduction in Gini impurity, creating
child nodes that are as pure (homogeneous) as possible.
Q8: A Random Forest model indicates that "number of website visits" has a high feature
importance score, while "user age" has a near-zero score. What is the most accurate
interpretation of this result?
A. "User age" is completely irrelevant to the outcome and should be dropped from all future
models.
B. "Number of website visits" contributed significantly to reducing impurity across the trees,
whereas "user age" did not. [CORRECT]
C. The model is overfitting to "number of website visits" and requires immediate pruning.
D. "User age" is likely highly correlated with "number of website visits," causing
multicollinearity errors.
Correct Answer: B
Rationale: This choice is correct because feature importance in Random Forests measures
how much a feature contributes to decreasing impurity across all splits in all trees. A near-
zero score simply means it wasn't useful for splitting in this specific ensemble, not
necessarily that it is universally irrelevant.
Q9: In a Support Vector Machine, what is the effect of increasing the regularization
parameter C?
A. It widens the margin, allowing more misclassifications to improve generalization.
B. It narrows the margin, penalizing misclassifications more heavily and potentially leading
to overfitting. [CORRECT]
C. It automatically selects a more complex kernel function, such as the radial basis function.
D. It reduces the computational time required to train the model on large datasets.
Correct Answer: B
Rationale: This choice is correct because a higher C value places a greater penalty on
misclassified points, forcing the optimizer to find a decision boundary with a narrower
margin that correctly classifies more training points. This can lead to overfitting if the data
contains noise.
Q10: A researcher is modeling the probability of a patient responding to a new drug. Why is
linear regression inappropriate for this task?
A. Linear regression cannot handle categorical predictor variables.
B. Linear regression can produce predicted values outside the valid probability range of 0 to
1. [CORRECT]