CS 559 Quiz 9: Combining Models with
correct answers and rationale
updated 2026 graded A+
Section 1: Fundamentals of Combining Models (Questions 1–20)
1. What is the primary motivation behind combining multiple models in machine learning?
A) To reduce the size of the training dataset
B) To achieve better predictive performance than any single model alone
C) To eliminate the need for feature engineering
D) To guarantee zero training error
Answer: B
Rationale: The fundamental motivation for combining models (ensembles) is that a combination of
diverse models can often achieve lower generalization error than the best individual model. This
leverages the "wisdom of crowds" effect—errors made by individual models can be cancelled out
when aggregated.
2. Which of the following best describes the concept of an ensemble method?
A) A method that selects the single best model from a set of candidates
B) A method that trains one very complex model on all data
C) A method that combines the predictions of multiple base learners to produce a final prediction
D) A method that reduces the dimensionality of the feature space
Answer: C
Rationale: An ensemble method combines multiple base learners (models) and aggregates their
predictions. The aggregation can be done through voting, averaging, stacking, or other strategies to
produce a final prediction that is typically more robust and accurate.
,3. In the context of ensemble learning, what does the term "base learner" refer to?
A) The final combined model
B) An individual model that contributes to the ensemble
C) The validation set used to evaluate the ensemble
D) The loss function used for training
Answer: B
Rationale: A base learner (also called a base model or weak learner) is an individual machine learning
model that is trained as part of the ensemble. Examples include decision trees, neural networks, or
logistic regression models that are later combined.
4. Which two properties are most desirable in the base learners of an ensemble?
A) High bias and low variance
B) High accuracy and high correlation
C) Accuracy (at least better than random) and diversity
D) Low complexity and high bias
Answer: C
Rationale: For an ensemble to be effective, base learners should be individually accurate (each
better than random guessing) and diverse (they should make different errors). If all base learners
make the same mistakes, combining them provides no benefit.
5. If all base learners in an ensemble make identical predictions, what is the benefit of combining
them?
A) Significant improvement in accuracy
B) Reduction in overfitting
C) No benefit; the ensemble performs identically to any single base learner
D) Guaranteed improvement in precision
,Answer: C
Rationale: If all base learners are perfectly correlated (make identical predictions), combining them
offers no advantage. Diversity among base learners is essential for ensembles to outperform
individual models.
6. Which of the following is NOT a common ensemble method?
A) Bagging
B) Boosting
C) Stacking
D) Regularization
Answer: D
Rationale: Regularization (e.g., L1, L2) is a technique for preventing overfitting in individual models
by adding a penalty term to the loss function. It is not an ensemble method. Bagging, boosting, and
stacking are all well-known ensemble strategies.
7. The bias-variance decomposition of the expected prediction error shows that ensemble methods
primarily work by:
A) Always reducing both bias and variance simultaneously
B) Reducing variance (bagging), reducing bias (boosting), or both through different strategies
C) Increasing bias to reduce variance
D) Eliminating irreducible error
Answer: B
Rationale: Different ensemble methods target different components of the error. Bagging primarily
reduces variance by averaging over models trained on different bootstrap samples. Boosting
primarily reduces bias by sequentially focusing on hard-to-predict examples. Some methods (like
stacking) can reduce both.
8. What is the irreducible error in the bias-variance framework?
A) Error due to model complexity
, B) Error due to insufficient training data
C) Noise inherent in the data that cannot be reduced by any model
D) Error due to overfitting
Answer: C
Rationale: Irreducible error (σ²) represents the noise in the data-generating process. No matter how
good the model, this error cannot be eliminated. Ensemble methods reduce bias and variance but
cannot reduce irreducible error.
9. Which of the following statements about combining models is TRUE?
A) Combining many poor models always yields a poor ensemble
B) Combining diverse, reasonably accurate models can yield a highly accurate ensemble
C) Ensembles always outperform the best individual model
D) Ensembles only work with linear base learners
Answer: B
Rationale: The key insight is that combining diverse models that are each slightly better than random
can yield a powerful ensemble. However, combining extremely poor models (worse than random) or
combining identical models would not help.
10. In a classification setting, what is "hard voting"?
A) Each base learner outputs a class probability, and the final prediction is the weighted average
B) Each base learner casts a vote for a class label, and the majority vote wins
C) Only the most confident base learner's prediction is used
D) Base learners vote only on difficult examples
Answer: B
Rationale: Hard voting (also called majority voting) is the simplest combination strategy for
classification. Each base learner predicts a class label, and the class that receives the most votes is
selected as the final prediction.
correct answers and rationale
updated 2026 graded A+
Section 1: Fundamentals of Combining Models (Questions 1–20)
1. What is the primary motivation behind combining multiple models in machine learning?
A) To reduce the size of the training dataset
B) To achieve better predictive performance than any single model alone
C) To eliminate the need for feature engineering
D) To guarantee zero training error
Answer: B
Rationale: The fundamental motivation for combining models (ensembles) is that a combination of
diverse models can often achieve lower generalization error than the best individual model. This
leverages the "wisdom of crowds" effect—errors made by individual models can be cancelled out
when aggregated.
2. Which of the following best describes the concept of an ensemble method?
A) A method that selects the single best model from a set of candidates
B) A method that trains one very complex model on all data
C) A method that combines the predictions of multiple base learners to produce a final prediction
D) A method that reduces the dimensionality of the feature space
Answer: C
Rationale: An ensemble method combines multiple base learners (models) and aggregates their
predictions. The aggregation can be done through voting, averaging, stacking, or other strategies to
produce a final prediction that is typically more robust and accurate.
,3. In the context of ensemble learning, what does the term "base learner" refer to?
A) The final combined model
B) An individual model that contributes to the ensemble
C) The validation set used to evaluate the ensemble
D) The loss function used for training
Answer: B
Rationale: A base learner (also called a base model or weak learner) is an individual machine learning
model that is trained as part of the ensemble. Examples include decision trees, neural networks, or
logistic regression models that are later combined.
4. Which two properties are most desirable in the base learners of an ensemble?
A) High bias and low variance
B) High accuracy and high correlation
C) Accuracy (at least better than random) and diversity
D) Low complexity and high bias
Answer: C
Rationale: For an ensemble to be effective, base learners should be individually accurate (each
better than random guessing) and diverse (they should make different errors). If all base learners
make the same mistakes, combining them provides no benefit.
5. If all base learners in an ensemble make identical predictions, what is the benefit of combining
them?
A) Significant improvement in accuracy
B) Reduction in overfitting
C) No benefit; the ensemble performs identically to any single base learner
D) Guaranteed improvement in precision
,Answer: C
Rationale: If all base learners are perfectly correlated (make identical predictions), combining them
offers no advantage. Diversity among base learners is essential for ensembles to outperform
individual models.
6. Which of the following is NOT a common ensemble method?
A) Bagging
B) Boosting
C) Stacking
D) Regularization
Answer: D
Rationale: Regularization (e.g., L1, L2) is a technique for preventing overfitting in individual models
by adding a penalty term to the loss function. It is not an ensemble method. Bagging, boosting, and
stacking are all well-known ensemble strategies.
7. The bias-variance decomposition of the expected prediction error shows that ensemble methods
primarily work by:
A) Always reducing both bias and variance simultaneously
B) Reducing variance (bagging), reducing bias (boosting), or both through different strategies
C) Increasing bias to reduce variance
D) Eliminating irreducible error
Answer: B
Rationale: Different ensemble methods target different components of the error. Bagging primarily
reduces variance by averaging over models trained on different bootstrap samples. Boosting
primarily reduces bias by sequentially focusing on hard-to-predict examples. Some methods (like
stacking) can reduce both.
8. What is the irreducible error in the bias-variance framework?
A) Error due to model complexity
, B) Error due to insufficient training data
C) Noise inherent in the data that cannot be reduced by any model
D) Error due to overfitting
Answer: C
Rationale: Irreducible error (σ²) represents the noise in the data-generating process. No matter how
good the model, this error cannot be eliminated. Ensemble methods reduce bias and variance but
cannot reduce irreducible error.
9. Which of the following statements about combining models is TRUE?
A) Combining many poor models always yields a poor ensemble
B) Combining diverse, reasonably accurate models can yield a highly accurate ensemble
C) Ensembles always outperform the best individual model
D) Ensembles only work with linear base learners
Answer: B
Rationale: The key insight is that combining diverse models that are each slightly better than random
can yield a powerful ensemble. However, combining extremely poor models (worse than random) or
combining identical models would not help.
10. In a classification setting, what is "hard voting"?
A) Each base learner outputs a class probability, and the final prediction is the weighted average
B) Each base learner casts a vote for a class label, and the majority vote wins
C) Only the most confident base learner's prediction is used
D) Base learners vote only on difficult examples
Answer: B
Rationale: Hard voting (also called majority voting) is the simplest combination strategy for
classification. Each base learner predicts a class label, and the class that receives the most votes is
selected as the final prediction.