CS 559 Quiz 4: Model Selection with
correct answers and rationale
updated 2026 graded A+ new!!
Question 1: What is the primary goal of model selection in machine learning?
A) To maximize training accuracy
B) To choose the model that generalizes best to unseen data
C) To minimize the number of features
D) To increase model complexity
Correct answer B
Rationale: Model selection aims to identify the model that performs best on new, unseen data
(generalization), not just on training data. Maximizing training accuracy can lead to overfitting.
Question 2: Which of the following best describes overfitting?
A) Model performs poorly on both training and test data
B) Model performs well on training data but poorly on test data
C) Model performs well on both training and test data
D) Model performs poorly on training but well on test data
Correct answer B
Rationale: Overfitting occurs when a model learns the training data too well, including noise,
resulting in high training accuracy but poor generalization to test data.
Question 3: What does the bias-variance tradeoff describe?
A) The relationship between speed and accuracy
,B) The tradeoff between underfitting and overfitting
C) The relationship between features and samples
D) The tradeoff between training time and inference time
Correct answer B
Rationale: The bias-variance tradeoff represents the balance between underfitting (high bias) and
overfitting (high variance). Reducing one often increases the other.
Question 4: High bias in a model typically leads to:
A) Overfitting
B) Underfitting
C) Perfect fit
D) High variance
Correct answer B
Rationale: High bias means the model makes strong assumptions and is too simple to capture the
underlying pattern, leading to underfitting.
Question 5: High variance in a model typically leads to:
A) Underfitting
B) Overfitting
C) High bias
D) Perfect generalization
Correct answer B
Rationale: High variance means the model is very sensitive to training data fluctuations, capturing
noise and leading to overfitting.
Question 6: In k-fold cross-validation, what does 'k' represent?
,A) Number of features
B) Number of folds/partitions of the data
C) Number of models
D) Number of hyperparameters
Correct answer B
Rationale: In k-fold cross-validation, the data is split into k equal partitions (folds), with each fold
used once as validation while others are used for training.
Question 7: What is a common value for k in k-fold cross-validation?
A) 2
B) 5 or 10
C) 100
D) 1000
Correct answer B
Rationale: k=5 or k=10 are the most commonly used values, providing a good balance between
computational cost and reliable performance estimation.
Question 8: Leave-One-Out Cross-Validation (LOOCV) is a special case of k-fold where:
A) k = 1
B) k = 2
C) k = n (number of samples)
D) k = number of features
Correct answer C
Rationale: In LOOCV, k equals the number of samples, so each fold contains exactly one sample, and
the model is trained n times.
Question 9: Which validation method is most computationally expensive?
, A) Hold-out validation
B) 5-fold cross-validation
C) LOOCV
D) 2-fold cross-validation
Correct answer C
Rationale: LOOCV requires training the model n times (once per sample), making it the most
computationally expensive validation method.
Question 10: What is the purpose of a validation set?
A) To train the model
B) To tune hyperparameters and select models
C) To report final performance
D) To increase training data
Correct answer B
Rationale: The validation set is used to tune hyperparameters and perform model selection, while
the test set provides an unbiased final performance estimate.
Question 11: The AIC (Akaike Information Criterion) is used to:
A) Increase model complexity
B) Balance model fit and complexity
C) Maximize likelihood only
D) Minimize training error only
Correct answer B
Rationale: AIC balances goodness of fit (likelihood) against model complexity (number of
parameters), penalizing overly complex models.
correct answers and rationale
updated 2026 graded A+ new!!
Question 1: What is the primary goal of model selection in machine learning?
A) To maximize training accuracy
B) To choose the model that generalizes best to unseen data
C) To minimize the number of features
D) To increase model complexity
Correct answer B
Rationale: Model selection aims to identify the model that performs best on new, unseen data
(generalization), not just on training data. Maximizing training accuracy can lead to overfitting.
Question 2: Which of the following best describes overfitting?
A) Model performs poorly on both training and test data
B) Model performs well on training data but poorly on test data
C) Model performs well on both training and test data
D) Model performs poorly on training but well on test data
Correct answer B
Rationale: Overfitting occurs when a model learns the training data too well, including noise,
resulting in high training accuracy but poor generalization to test data.
Question 3: What does the bias-variance tradeoff describe?
A) The relationship between speed and accuracy
,B) The tradeoff between underfitting and overfitting
C) The relationship between features and samples
D) The tradeoff between training time and inference time
Correct answer B
Rationale: The bias-variance tradeoff represents the balance between underfitting (high bias) and
overfitting (high variance). Reducing one often increases the other.
Question 4: High bias in a model typically leads to:
A) Overfitting
B) Underfitting
C) Perfect fit
D) High variance
Correct answer B
Rationale: High bias means the model makes strong assumptions and is too simple to capture the
underlying pattern, leading to underfitting.
Question 5: High variance in a model typically leads to:
A) Underfitting
B) Overfitting
C) High bias
D) Perfect generalization
Correct answer B
Rationale: High variance means the model is very sensitive to training data fluctuations, capturing
noise and leading to overfitting.
Question 6: In k-fold cross-validation, what does 'k' represent?
,A) Number of features
B) Number of folds/partitions of the data
C) Number of models
D) Number of hyperparameters
Correct answer B
Rationale: In k-fold cross-validation, the data is split into k equal partitions (folds), with each fold
used once as validation while others are used for training.
Question 7: What is a common value for k in k-fold cross-validation?
A) 2
B) 5 or 10
C) 100
D) 1000
Correct answer B
Rationale: k=5 or k=10 are the most commonly used values, providing a good balance between
computational cost and reliable performance estimation.
Question 8: Leave-One-Out Cross-Validation (LOOCV) is a special case of k-fold where:
A) k = 1
B) k = 2
C) k = n (number of samples)
D) k = number of features
Correct answer C
Rationale: In LOOCV, k equals the number of samples, so each fold contains exactly one sample, and
the model is trained n times.
Question 9: Which validation method is most computationally expensive?
, A) Hold-out validation
B) 5-fold cross-validation
C) LOOCV
D) 2-fold cross-validation
Correct answer C
Rationale: LOOCV requires training the model n times (once per sample), making it the most
computationally expensive validation method.
Question 10: What is the purpose of a validation set?
A) To train the model
B) To tune hyperparameters and select models
C) To report final performance
D) To increase training data
Correct answer B
Rationale: The validation set is used to tune hyperparameters and perform model selection, while
the test set provides an unbiased final performance estimate.
Question 11: The AIC (Akaike Information Criterion) is used to:
A) Increase model complexity
B) Balance model fit and complexity
C) Maximize likelihood only
D) Minimize training error only
Correct answer B
Rationale: AIC balances goodness of fit (likelihood) against model complexity (number of
parameters), penalizing overly complex models.