Page |1
CS 7641 Machine Learning Final Exam Practice Questions
(2026-2027 Verified Update!!!)
Instructions: This exam consists of multiple-choice questions.
Choose the best answer for each question. The questions are
organized by major topic areas based on the CS 7641
curriculum.
Section 1: Supervised Learning – Regression & Classification
(Questions)
1. In linear regression, the cost function J(θ) is typically
minimized using which method?
A) Gradient descent
B) Maximum likelihood estimation
C) Both A and B
D) Expectation-maximization
Answer: C
Rationale: Linear regression can be solved using gradient
descent (iterative) or by setting the derivative to zero using the
normal equation (which is equivalent to maximum likelihood
under Gaussian assumptions). Both approaches yield the same
optimal parameters when the problem is convex.
, Page |2
2. What is the primary difference between L1 (Lasso) and L2
(Ridge) regularization?
A) L1 penalizes the sum of squared weights; L2 penalizes the
sum of absolute weights
B) L1 penalizes the sum of absolute weights; L2 penalizes the
sum of squared weights
C) L1 is for classification; L2 is for regression
D) L1 is only used in neural networks
Answer: B
Rationale: L1 (Lasso) penalizes the sum of absolute weights,
which encourages sparsity (some weights become exactly zero).
L2 (Ridge) penalizes the sum of squared weights, which shrinks
weights but does not force them to zero.
3. What is the "curse of dimensionality" in the context of
machine learning?
A) As the number of features grows, the amount of data needed
to generalize accurately grows exponentially
B) As the number of features grows, the model becomes
simpler
C) High-dimensional data is always easier to visualize
D) More features always lead to better performance
Answer: A
Rationale: The curse of dimensionality refers to the exponential
, Page |3
growth in the amount of data required to generalize accurately
as the number of features increases.
4. What is the key advantage of lazy learning?
A) It requires less memory
B) Instead of estimating the target function once for the entire
space, it can estimate it locally and differently for each new
instance to be classified
C) It trains much faster than eager learning
D) It always produces simpler models
Answer: B
Rationale: Lazy learning defers the decision of how to
generalize beyond the training data until each new query
instance is encountered, allowing local estimation.
5. What does learning consist of in instance-based algorithms?
A) Storing the data in O(n) time
B) Building a complex decision tree
C) Training a neural network
D) Computing all possible hypotheses
Answer: A
Rationale: In instance-based (lazy) learning, learning simply
consists of storing the training data, which takes O(n) time.
6. What are the disadvantages of instance-based learning?
A) Cost of classifying a new instance is high and irrelevant
attributes can mislead distance calculations
, Page |4
B) It cannot handle large datasets
C) It always overfits
D) It requires labeled data only
Answer: A
Rationale: Disadvantages include high classification cost (O(n)
time) and the fact that if the target concept depends on only a
few of many attributes, the most similar instances may be far
apart.
7. Describe the KNN algorithm.
A) Assumes instances are points in n-dimensional space;
training stores data; classification finds k closest training
examples and takes the most common label
B) Builds a decision tree from training data
C) Trains a neural network with backpropagation
D) Uses gradient descent to minimize error
Answer: A
Rationale: KNN assumes all instances correspond to points in n-
dimensional space, nearest neighbors are defined by a distance
function, training stores each example, and classification finds
the k closest training examples and determines the most
common label.
8. In KNN, when all the training examples are considered,
what is it called?
A) A local method
CS 7641 Machine Learning Final Exam Practice Questions
(2026-2027 Verified Update!!!)
Instructions: This exam consists of multiple-choice questions.
Choose the best answer for each question. The questions are
organized by major topic areas based on the CS 7641
curriculum.
Section 1: Supervised Learning – Regression & Classification
(Questions)
1. In linear regression, the cost function J(θ) is typically
minimized using which method?
A) Gradient descent
B) Maximum likelihood estimation
C) Both A and B
D) Expectation-maximization
Answer: C
Rationale: Linear regression can be solved using gradient
descent (iterative) or by setting the derivative to zero using the
normal equation (which is equivalent to maximum likelihood
under Gaussian assumptions). Both approaches yield the same
optimal parameters when the problem is convex.
, Page |2
2. What is the primary difference between L1 (Lasso) and L2
(Ridge) regularization?
A) L1 penalizes the sum of squared weights; L2 penalizes the
sum of absolute weights
B) L1 penalizes the sum of absolute weights; L2 penalizes the
sum of squared weights
C) L1 is for classification; L2 is for regression
D) L1 is only used in neural networks
Answer: B
Rationale: L1 (Lasso) penalizes the sum of absolute weights,
which encourages sparsity (some weights become exactly zero).
L2 (Ridge) penalizes the sum of squared weights, which shrinks
weights but does not force them to zero.
3. What is the "curse of dimensionality" in the context of
machine learning?
A) As the number of features grows, the amount of data needed
to generalize accurately grows exponentially
B) As the number of features grows, the model becomes
simpler
C) High-dimensional data is always easier to visualize
D) More features always lead to better performance
Answer: A
Rationale: The curse of dimensionality refers to the exponential
, Page |3
growth in the amount of data required to generalize accurately
as the number of features increases.
4. What is the key advantage of lazy learning?
A) It requires less memory
B) Instead of estimating the target function once for the entire
space, it can estimate it locally and differently for each new
instance to be classified
C) It trains much faster than eager learning
D) It always produces simpler models
Answer: B
Rationale: Lazy learning defers the decision of how to
generalize beyond the training data until each new query
instance is encountered, allowing local estimation.
5. What does learning consist of in instance-based algorithms?
A) Storing the data in O(n) time
B) Building a complex decision tree
C) Training a neural network
D) Computing all possible hypotheses
Answer: A
Rationale: In instance-based (lazy) learning, learning simply
consists of storing the training data, which takes O(n) time.
6. What are the disadvantages of instance-based learning?
A) Cost of classifying a new instance is high and irrelevant
attributes can mislead distance calculations
, Page |4
B) It cannot handle large datasets
C) It always overfits
D) It requires labeled data only
Answer: A
Rationale: Disadvantages include high classification cost (O(n)
time) and the fact that if the target concept depends on only a
few of many attributes, the most similar instances may be far
apart.
7. Describe the KNN algorithm.
A) Assumes instances are points in n-dimensional space;
training stores data; classification finds k closest training
examples and takes the most common label
B) Builds a decision tree from training data
C) Trains a neural network with backpropagation
D) Uses gradient descent to minimize error
Answer: A
Rationale: KNN assumes all instances correspond to points in n-
dimensional space, nearest neighbors are defined by a distance
function, training stores each example, and classification finds
the k closest training examples and determines the most
common label.
8. In KNN, when all the training examples are considered,
what is it called?
A) A local method