• Wrong document? Swap it for free
  • Written by students who passed
  • Immediately available after payment
  • Read online or as PDF
Sell
Where do you study
Your language
Document preview thumbnail
Preview 4 out of 105 pages
Exam (elaborations)

CS 7641 Machine Learning Final Exam Practice Questions ( Verified Update!!!).pdf

Document preview thumbnail
Preview 4 out of 105 pages

CS 7641 Machine Learning Final Exam Practice Questions ( Verified Update!!!).pdf

Content preview

Page |1


CS 7641 Machine Learning Final Exam Practice Questions
(2026-2027 Verified Update!!!)



Instructions: This exam consists of multiple-choice questions.
Choose the best answer for each question. The questions are
organized by major topic areas based on the CS 7641
curriculum.


Section 1: Supervised Learning – Regression & Classification
(Questions)


1. In linear regression, the cost function J(θ) is typically
minimized using which method?
A) Gradient descent
B) Maximum likelihood estimation
C) Both A and B
D) Expectation-maximization
Answer: C
Rationale: Linear regression can be solved using gradient
descent (iterative) or by setting the derivative to zero using the
normal equation (which is equivalent to maximum likelihood
under Gaussian assumptions). Both approaches yield the same
optimal parameters when the problem is convex.

, Page |2


2. What is the primary difference between L1 (Lasso) and L2
(Ridge) regularization?
A) L1 penalizes the sum of squared weights; L2 penalizes the
sum of absolute weights
B) L1 penalizes the sum of absolute weights; L2 penalizes the
sum of squared weights
C) L1 is for classification; L2 is for regression
D) L1 is only used in neural networks
Answer: B
Rationale: L1 (Lasso) penalizes the sum of absolute weights,
which encourages sparsity (some weights become exactly zero).
L2 (Ridge) penalizes the sum of squared weights, which shrinks
weights but does not force them to zero.
3. What is the "curse of dimensionality" in the context of
machine learning?
A) As the number of features grows, the amount of data needed
to generalize accurately grows exponentially
B) As the number of features grows, the model becomes
simpler
C) High-dimensional data is always easier to visualize
D) More features always lead to better performance
Answer: A
Rationale: The curse of dimensionality refers to the exponential

, Page |3


growth in the amount of data required to generalize accurately
as the number of features increases.
4. What is the key advantage of lazy learning?
A) It requires less memory
B) Instead of estimating the target function once for the entire
space, it can estimate it locally and differently for each new
instance to be classified
C) It trains much faster than eager learning
D) It always produces simpler models
Answer: B
Rationale: Lazy learning defers the decision of how to
generalize beyond the training data until each new query
instance is encountered, allowing local estimation.
5. What does learning consist of in instance-based algorithms?
A) Storing the data in O(n) time
B) Building a complex decision tree
C) Training a neural network
D) Computing all possible hypotheses
Answer: A
Rationale: In instance-based (lazy) learning, learning simply
consists of storing the training data, which takes O(n) time.
6. What are the disadvantages of instance-based learning?
A) Cost of classifying a new instance is high and irrelevant
attributes can mislead distance calculations

, Page |4


B) It cannot handle large datasets
C) It always overfits
D) It requires labeled data only
Answer: A
Rationale: Disadvantages include high classification cost (O(n)
time) and the fact that if the target concept depends on only a
few of many attributes, the most similar instances may be far
apart.
7. Describe the KNN algorithm.
A) Assumes instances are points in n-dimensional space;
training stores data; classification finds k closest training
examples and takes the most common label
B) Builds a decision tree from training data
C) Trains a neural network with backpropagation
D) Uses gradient descent to minimize error
Answer: A
Rationale: KNN assumes all instances correspond to points in n-
dimensional space, nearest neighbors are defined by a distance
function, training stores each example, and classification finds
the k closest training examples and determines the most
common label.
8. In KNN, when all the training examples are considered,
what is it called?
A) A local method

Document information

Uploaded on
September 27, 2026
Number of pages
105
Written in
2026/2027
Type
Exam (elaborations)
Contains
Questions & answers
$19.49

Wrong document? Swap it for free Within 14 days of purchase and before downloading, you can choose a different document. You can simply spend the amount again.
Written by students who passed
Immediately available after payment
Read online or as PDF

Seller avatar
Reputation scores are based on the amount of documents a seller has sold for a fee and the reviews they have received for those documents. There are three levels: Bronze, Silver and Gold. The better the reputation, the more your can rely on the quality of the sellers work.
Performance
4.3
(242)
Sold
575
Followers
45
Items
19810
Last sold
12 hours ago




Why students choose Stuvia

Created by fellow students, verified by reviews

Quality you can trust: written by students who passed their tests and reviewed by others who've used these notes.

Didn't get what you expected? Choose another document

No worries! You can instantly pick a different document that better fits what you're looking for.

Pay as you like, start learning right away

No subscription, no commitments. Pay the way you're used to via credit card and download your PDF document instantly.

Student with book image

“Bought, downloaded, and aced it. It really can be that simple.”

Alisha Student

Working on your references?

Create accurate citations in APA, MLA and Harvard with our free citation generator.

Working on your references?

Frequently asked questions