WGU D499 Machine Learning | Project 1 - Finding Donors for CharityML
WGU D499
MACHINE LEARNING
Project 1: Finding Donors for CharityML
Supervised Learning | Model Evaluation | GridSearchCV | Feature Importance
2026/2027 Updated
Complete Project Resource
Built from the recovered D499/Udacity notebook structure and current scikit-learn conventions.
Important: This resource is independently prepared and is not an official WGU or Udacity publication. Model outputs can vary with
software versions, parameter choices, and environment.
2026/2027 Updated | Project-aligned reference resource
, WGU D499 Machine Learning | Project 1 - Finding Donors for CharityML
1. Project Purpose and Verified Scope
Project 1 applies supervised machine learning to cleaned U.S. Census data to predict whether an individual earns more
than $50,000 per year. The business objective is to help the fictitious nonprofit CharityML focus outreach on higher-
probability donors. The recovered D499 notebook requires implementation work plus eight written analytical responses.
Verified project package: The D499 repository identifies finding_donors.ipynb as the main notebook and requires the
supplied census.csv and visuals.py support files. The notebook itself contains the implementation checkpoints and
Q1-Q8 response sequence.
2. Completion Checklist
# Checkpoint Evidence in this resource
Compute record counts and the >$50K
1 Data exploration
percentage.
Log-transform skewed capital features;
2 Feature preprocessing
normalize numerical fields.
One-hot encode categorical variables and
3 Encoding
map income to 0/1.
Use the verified 80/20 split with
4 Train/test split
random_state=0.
Calculate accuracy and F0.5 for an all-
5 Naive benchmark
positive predictor.
Select and justify three supervised
6 Model rationale
algorithms.
Measure train/predict time, accuracy, and
7 Training pipeline
F0.5.
Evaluate each model on 1%, 10%, and
8 Initial evaluation
100% of training data.
Use GridSearchCV and an F0.5 scorer;
9 Model tuning tune at least one important parameter
across at least three values.
Rank the strongest predictors using a
10 Feature importance
model exposing feature_importances_.
Retrain using only the top five features
11 Feature selection
and compare performance.
Answer Q1-Q8 with evidence from
12 Written analysis
executed results.
Save the completed notebook and a
13 Export
rendered HTML copy.
2026/2027 Updated | Project-aligned reference resource
WGU D499
MACHINE LEARNING
Project 1: Finding Donors for CharityML
Supervised Learning | Model Evaluation | GridSearchCV | Feature Importance
2026/2027 Updated
Complete Project Resource
Built from the recovered D499/Udacity notebook structure and current scikit-learn conventions.
Important: This resource is independently prepared and is not an official WGU or Udacity publication. Model outputs can vary with
software versions, parameter choices, and environment.
2026/2027 Updated | Project-aligned reference resource
, WGU D499 Machine Learning | Project 1 - Finding Donors for CharityML
1. Project Purpose and Verified Scope
Project 1 applies supervised machine learning to cleaned U.S. Census data to predict whether an individual earns more
than $50,000 per year. The business objective is to help the fictitious nonprofit CharityML focus outreach on higher-
probability donors. The recovered D499 notebook requires implementation work plus eight written analytical responses.
Verified project package: The D499 repository identifies finding_donors.ipynb as the main notebook and requires the
supplied census.csv and visuals.py support files. The notebook itself contains the implementation checkpoints and
Q1-Q8 response sequence.
2. Completion Checklist
# Checkpoint Evidence in this resource
Compute record counts and the >$50K
1 Data exploration
percentage.
Log-transform skewed capital features;
2 Feature preprocessing
normalize numerical fields.
One-hot encode categorical variables and
3 Encoding
map income to 0/1.
Use the verified 80/20 split with
4 Train/test split
random_state=0.
Calculate accuracy and F0.5 for an all-
5 Naive benchmark
positive predictor.
Select and justify three supervised
6 Model rationale
algorithms.
Measure train/predict time, accuracy, and
7 Training pipeline
F0.5.
Evaluate each model on 1%, 10%, and
8 Initial evaluation
100% of training data.
Use GridSearchCV and an F0.5 scorer;
9 Model tuning tune at least one important parameter
across at least three values.
Rank the strongest predictors using a
10 Feature importance
model exposing feature_importances_.
Retrain using only the top five features
11 Feature selection
and compare performance.
Answer Q1-Q8 with evidence from
12 Written analysis
executed results.
Save the completed notebook and a
13 Export
rendered HTML copy.
2026/2027 Updated | Project-aligned reference resource