• Wrong document? Swap it for free
  • Written by students who passed
  • Immediately available after payment
  • Read online or as PDF
Sell
Where do you study
Your language
Document preview thumbnail
Preview 3 out of 17 pages
Exam (elaborations)

CS 7641 Unit Quiz Problem Solutions – Fall 2025 v2 | Updated 2026/2027 | Georgia Institute of Technology | A+

Document preview thumbnail
Preview 3 out of 17 pages

CS 7641 Unit Quiz Problem Solutions – Fall 2025 v2 | Updated 2026/2027 | Georgia Institute of Technology | A+

Content preview

UL Unit Quiz
CS 7641: Machine Learning
Solutions
DO NOT DISTRIBUTE OUTSIDE OF CS7641




1 Question 1 - Clustering
Part 1. Clustering Scenario (MCMA)
You are analyzing a high-dimensional dataset (d = 100) with the following characteristics:

• There are three underlying clusters with very different sizes (one large, two small).

• The data contains moderate Gaussian noise.

• The clusters are roughly spherical in their original feature space.

You are tasked with selecting a clustering method that is robust in high-dimensional, noisy, and
imbalanced settings.
Which of the following statements are true in this scenario? Select all that apply.
Correct Answers: A, C, F

A. K-Means may be biased toward larger clusters due to its use of Euclidean distance
and sum-of-squares optimization.
True. K-Means tends to partition space such that larger clusters can dominate the centroid
updates, leading to poor performance on imbalanced datasets.

B. Single linkage clustering is generally preferred in high-dimensional noisy settings
due to its chaining property.
False. Single linkage is highly sensitive to noise and often performs poorly in high dimensions
due to chaining effects and instability.

C. GMM with EM can handle unequal cluster sizes better than K-Means by learning
separate covariance structures and mixing coefficients.
True. EM allows modeling clusters with different sizes and shapes through learned covariance
matrices and priors, making it more robust to imbalance.

D. Dimensionality has little effect on clustering because K-Means scales naturally to
high dimensions.
False. In high dimensions, distances become less informative (curse of dimensionality), and
K-Means performance degrades due to poor separation.


1

,E. Using PCA before clustering will always improve clustering results by preserving
local cluster structures.
False. PCA preserves global variance, not necessarily cluster structure, and can sometimes
collapse important separations between small or rare clusters.

F. Density-based methods like DBSCAN may fail in this scenario due to varying
cluster sizes and the curse of dimensionality.
True. DBSCAN struggles when clusters differ in density and when neighborhood definitions
become unreliable in high-dimensional spaces.

Part 2. Segmenting the Condiment Market (MCMA)
Two founders, Mr. Mayonnaise and Dr. Mustard, are competing in the premium condiments
market. They’ve collected store-level and customer-level features (e.g., weekly unit sales, price,
promo spend, shelf space, store size, region, and simple product flags like “organic”/“spicy”).
They want to discover actionable market segments using K-Means and GMMs.
Which statements are true in this context? Select all that apply.

A. Standardizing numeric features (e.g., z-scoring sales, price, and promo) is important so no single
scale-dominant variable dictates distances.

B. Because GMMs are probabilistic, they inherently handle unscaled features without issues.

C. Adding dozens of one-hot flags for every minor product variant always improves K-Means
segmentation.

D. Choosing the number of clusters k by maximizing training likelihood alone guarantees the most
useful segments.

E. Increasing k always yields more homogeneous and more actionable segments for marketing.

F. K-Means is sensitive to initialization; multiple restarts or k-means++ help avoid poor local
minima.




Part 2. Segmenting the Condiment Market (MCMA) (Answers)
Correct Answers: A, F

A. True. K-Means and GMMs operate in Euclidean feature space; scale-dominant variables (e.g.,
price vs. shelf width) can overwhelm distances and covariances. Z-scoring numeric features
makes contributions comparable, stabilizing both centroid updates (K-Means) and covariance
estimation (GMMs).

B. False. GMMs still depend on feature scales: the likelihood and fitted covariances are distorted
if one feature has much larger magnitude. Unscaled features bias component shapes and
responsibilities. Standardization (and sometimes whitening) remains important.


2

, C. False. Flooding the design with many sparse one-hot flags for minor variants can fragment
clusters and inject noise. Prefer parsimonious encodings (coarser taxonomies, target-agnostic
embeddings, or dimensionality reduction) aligned with the segmentation goal.
D. False. Maximizing in-sample fit (e.g., likelihood) often over-segments. Use multiple crite-
ria: BIC/AIC for GMMs, silhouette or stability for K-Means, and—critically—downstream
business utility (campaignability, lift, coverage).
E. False. Larger k may increase apparent homogeneity but yields tiny, hard-to-activate segments
with poor generalization and operational complexity. There is a practical optimum balancing
cohesion, separation, and activation cost.
F. True. K-Means can land in poor local minima with bad seeding. Using k-means++ and multiple
restarts improves centroid initialization, stability, and final inertia, especially with noisy retail
features.

Part 3. Comparing EM and K-Means in 1D with 3 Clusters (High Variance)
You are given the following 1D dataset of points:

D = {1, 2, 3, 8, 9, 10, 16, 17, 18}
Assume the data come from three clusters. You will compare one iteration of K-Means and
EM for Gaussian Mixture Models (GMM).
Use the following initialization for both algorithms:
• K-Means initial centroids: µ1 = 2, µ2 = 9, µ3 = 17
• EM initial parameters:
– Means: µ1 = 2, µ2 = 9, µ3 = 17
– Shared, known variance: σ 2 = 9 (i.e., standard deviation σ = 3)
– Mixing coefficients: π1 = π2 = π3 = 13
Now consider a new data point:
xnew = 7
a) (K-Means Assignment) Based on Euclidean distance, which cluster would xnew = 7 be
assigned to?
b) (EM Responsibilities) Compute the posterior probability that xnew = 7 belongs to each
cluster:

πj · ϕ(7; µj , σ 2 = 9)
γ(z = j | x = 7) = P3
2
k=1 πk · ϕ(7; µk , σ = 9)

where
(x − µ)2
 
2 1
ϕ(x; µ, σ ) = √ exp −
2π · σ 2 2σ 2
Make sure the responsibilities across all clusters sum to 1. Round each responsibility to 3
decimal places.

3

Document information

Uploaded on
September 30, 2026
Number of pages
17
Written in
2026/2027
Type
Exam (elaborations)
Contains
Questions & answers
$16.58

Wrong document? Swap it for free Within 14 days of purchase and before downloading, you can choose a different document. You can simply spend the amount again.
Written by students who passed
Immediately available after payment
Read online or as PDF

Seller avatar
Reputation scores are based on the amount of documents a seller has sold for a fee and the reviews they have received for those documents. There are three levels: Bronze, Silver and Gold. The better the reputation, the more your can rely on the quality of the sellers work.
wisenurse
5.0
(1)
Sold
44
Followers
2
Items
1156
Last sold
2 days ago



Why students choose Stuvia

Created by fellow students, verified by reviews

Quality you can trust: written by students who passed their tests and reviewed by others who've used these notes.

Didn't get what you expected? Choose another document

No worries! You can instantly pick a different document that better fits what you're looking for.

Pay as you like, start learning right away

No subscription, no commitments. Pay the way you're used to via credit card and download your PDF document instantly.

Student with book image

“Bought, downloaded, and aced it. It really can be that simple.”

Alisha Student

Working on your references?

Create accurate citations in APA, MLA and Harvard with our free citation generator.

Working on your references?

Frequently asked questions