WGU D499 Machine Learning | Project 2 - Identify Customer Segments
WGU D499
MACHINE LEARNING
Project 2: Identify Customer Segments
Arvato Demographics | Unsupervised Learning | PCA | K-Means Clustering
2026/2027 Updated
Complete Project Analysis
Notebook-aligned preprocessing, feature engineering, PCA interpretation, clustering, and customer-segment analysis.
Independent reference resource. The proprietary Arvato datasets are not reproduced or distributed in this file.
Page 1
, WGU D499 Machine Learning | Project 2 - Identify Customer Segments
1. Project Purpose and Verified Scope
This D499 project applies unsupervised learning to German demographic data supplied for the Udacity Arvato project.
The analysis first builds demographic clusters from the general population, then maps the mail-order company's
customer records into those same clusters to identify segments that are overrepresented or underrepresented among
customers.
Current verification: Udacity's current Unsupervised Learning curriculum still lists Identify Customer Segments. The
recovered notebook uses four supporting files: the general-population dataset, customer dataset, feature-summary file, and
data dictionary. The datasets are proprietary and should not be redistributed.
Naming note: The broader Udacity nanodegree may display this as its customer-segmentation project after an image-
classifier project. 2026 WGU student reporting confirms D499 itself contains two projects; this resource therefore labels
customer segmentation as D499 Project 2.
2. Project Files and Core Deliverables
Item Purpose Handling
General population of Germany; about Use inside the authorized project
Udacity_AZDIAS_Subset.csv
891k persons x 85 features. environment.
Mail-order customer demographics; about Use inside the authorized project
Udacity_CUSTOMERS_Subset.csv
192k persons x 85 features. environment.
Feature metadata: information level, type,
AZDIAS_Feature_Summary.csv Used to drive cleaning decisions.
and missing/unknown codes.
Human-readable definitions of
Data_Dictionary.md Use for interpretation.
demographic features.
Primary Jupyter notebook containing the
Complete code, outputs, and written
Identify_Customer_Segments.ipynb project framework and discussion
interpretations.
checkpoints.
Rendered copy of the completed Keep alongside the notebook for
HTML export
notebook. review/sharing where permitted.
Page 2
WGU D499
MACHINE LEARNING
Project 2: Identify Customer Segments
Arvato Demographics | Unsupervised Learning | PCA | K-Means Clustering
2026/2027 Updated
Complete Project Analysis
Notebook-aligned preprocessing, feature engineering, PCA interpretation, clustering, and customer-segment analysis.
Independent reference resource. The proprietary Arvato datasets are not reproduced or distributed in this file.
Page 1
, WGU D499 Machine Learning | Project 2 - Identify Customer Segments
1. Project Purpose and Verified Scope
This D499 project applies unsupervised learning to German demographic data supplied for the Udacity Arvato project.
The analysis first builds demographic clusters from the general population, then maps the mail-order company's
customer records into those same clusters to identify segments that are overrepresented or underrepresented among
customers.
Current verification: Udacity's current Unsupervised Learning curriculum still lists Identify Customer Segments. The
recovered notebook uses four supporting files: the general-population dataset, customer dataset, feature-summary file, and
data dictionary. The datasets are proprietary and should not be redistributed.
Naming note: The broader Udacity nanodegree may display this as its customer-segmentation project after an image-
classifier project. 2026 WGU student reporting confirms D499 itself contains two projects; this resource therefore labels
customer segmentation as D499 Project 2.
2. Project Files and Core Deliverables
Item Purpose Handling
General population of Germany; about Use inside the authorized project
Udacity_AZDIAS_Subset.csv
891k persons x 85 features. environment.
Mail-order customer demographics; about Use inside the authorized project
Udacity_CUSTOMERS_Subset.csv
192k persons x 85 features. environment.
Feature metadata: information level, type,
AZDIAS_Feature_Summary.csv Used to drive cleaning decisions.
and missing/unknown codes.
Human-readable definitions of
Data_Dictionary.md Use for interpretation.
demographic features.
Primary Jupyter notebook containing the
Complete code, outputs, and written
Identify_Customer_Segments.ipynb project framework and discussion
interpretations.
checkpoints.
Rendered copy of the completed Keep alongside the notebook for
HTML export
notebook. review/sharing where permitted.
Page 2