Complete Solution Actual Exam 2026/2027 | Complete Exam-
Style Questions with Detailed Rationales | 100% Verified | Pass
Guaranteed – A+ Graded
Classification and Regression Methods
Q1: In a healthcare setting, you are predicting patient readmission using k-Nearest
Neighbors. What happens to the model if you significantly increase the value of k?
A. The model becomes more sensitive to local noise.
B. The decision boundary becomes smoother and less complex. [CORRECT]
C. The training error will always increase to 100%.
D. The model will perfectly memorize the training data.
Correct Answer: B
Rationale: The best answer is B because increasing k averages over more neighbors, which
smooths the decision boundary and reduces variance, making the model less sensitive to
local noise.
Q2: In a logistic regression model predicting loan default, the odds ratio for a specific
feature is 1.5. How should this be interpreted?
A. A one-unit increase in the feature decreases the probability of default by 50%.
B. The feature is statistically insignificant at the 5% level.
C. A one-unit increase in the feature multiplies the odds of default by 1.5. [CORRECT]
D. The model will overfit if this feature is included.
Correct Answer: C
Rationale: This choice is correct because an odds ratio greater than 1 indicates that a one-
unit increase in the predictor multiplies the odds of the outcome occurring by that factor,
holding other variables constant.
Q3: What is the primary purpose of cost-complexity pruning in Classification and Regression
Trees (CART)?
A. To increase the depth of the tree for better training accuracy.
B. To prevent overfitting by balancing tree size with classification error. [CORRECT]
C. To convert categorical variables into continuous numeric splits.
D. To ensure every leaf node contains exactly one observation.
Correct Answer: B
Rationale: This aligns with the principle that pruning removes branches that provide little
predictive power, thereby reducing model complexity and improving generalization to
unseen data.
Q4: In a Random Forest model, what does the out-of-bag (OOB) error estimate?
A. The computational time required to train the model.
B. The correlation between different decision trees in the forest.
,C. The number of features selected at each split.
D. The generalization error using observations not included in a tree's bootstrap sample.
[CORRECT]
Correct Answer: D
Rationale: The best answer is D because OOB error provides an unbiased estimate of the
model's performance on unseen data by testing each tree on the roughly one-third of data
not used to train it.
Q5: Why is the "kernel trick" used in Support Vector Machines (SVM)?
A. To reduce the number of support vectors required for the model.
B. To speed up the training process by simplifying the objective function.
C. To implicitly map data into a higher-dimensional space without computing the
coordinates. [CORRECT]
D. To force the decision boundary to be strictly linear.
Correct Answer: C
Rationale: This choice is correct because the kernel trick computes the inner products in a
higher-dimensional space directly, allowing SVMs to find non-linear decision boundaries
efficiently.
Q6: A hospital wants to predict which patients are at high risk of readmission within 30 days.
The dataset contains 50,000 patients with 200 mixed-type features, and interpretability for
clinicians is a strict requirement. Which modeling approach is most appropriate?
A. Support Vector Machine with a radial basis function kernel.
B. k-Nearest Neighbors with Euclidean distance.
C. Logistic regression with L1 regularization. [CORRECT]
D. Deep neural network with multiple hidden layers.
Correct Answer: C
Rationale: This matches the principle that logistic regression provides clear, interpretable
coefficients (odds ratios) clinicians can understand, while L1 regularization handles high
dimensionality via automatic feature selection.
Q7: Why is linear regression generally inappropriate for binary classification tasks?
A. It cannot handle categorical predictor variables.
B. It can produce predicted values outside the valid probability range of 0 to 1. [CORRECT]
C. It requires the target variable to be normally distributed.
D. It is computationally more expensive than logistic regression.
Correct Answer: B
Rationale: The best answer is B because linear regression fits an unbounded continuous line,
which can yield predictions less than 0 or greater than 1, making them invalid as
probabilities.
Q8: In CART, what does the Gini impurity measure?
A. The distance between clusters in the feature space.
B. The likelihood of an incorrect classification of a randomly chosen element. [CORRECT]
C. The variance of the target variable within a leaf node.
D. The correlation between two split variables.
Correct Answer: B
, Rationale: This choice is correct because Gini impurity quantifies how often a randomly
chosen element would be mislabeled if it were randomly labeled according to the
distribution of labels in the subset.
Q9: How does a Random Forest calculate feature importance?
A. By measuring the total decrease in node impurity (e.g., Gini) averaged over all trees.
[CORRECT]
B. By counting the number of times a feature is used as the root node.
C. By calculating the p-value of each feature in a global linear model.
D. By measuring the distance between data points in the final leaf nodes.
Correct Answer: A
Rationale: This aligns with standard Random Forest mechanics, where features that
consistently produce purer splits across the ensemble are assigned higher importance
scores.
Q10: In a linear SVM, what is the primary objective of margin maximization?
A. To minimize the number of support vectors to zero.
B. To ensure all training points are classified with 100% accuracy.
C. To improve the model's generalization by maximizing the distance to the nearest data
points of any class. [CORRECT]
D. To reduce the dimensionality of the input feature space.
Correct Answer: C
Rationale: The best answer is C because a wider margin implies greater confidence in
classifications and typically leads to better out-of-sample performance by reducing model
variance.
Q11: What is the "curse of dimensionality" in the context of k-NN?
A. The model becomes too fast to compute as dimensions increase.
B. Distance metrics become less meaningful as all points become equidistant in high-
dimensional space. [CORRECT]
C. The number of required training samples decreases exponentially.
D. The decision boundary automatically becomes smoother.
Correct Answer: B
Rationale: This choice is correct because in high dimensions, the contrast between the
nearest and farthest points diminishes, making distance-based similarity measures
unreliable.
Q12: What is the effect of applying L1 regularization (Lasso) to a logistic regression model?
A. It shrinks all coefficients proportionally but keeps them all non-zero.
B. It forces some feature coefficients to exactly zero, performing feature selection.
[CORRECT]
C. It increases the model's variance to reduce bias.
D. It guarantees the model will achieve 100% training accuracy.
Correct Answer: B
Rationale: This matches the principle that L1 regularization adds a penalty equal to the
absolute value of the magnitude of coefficients, which can drive less important feature
weights to exactly zero.