WGU D501 Machine Learning DevOps | Project 1 | 2026/2027 Updated
WGU D501
Machine Learning DevOps
Project 1 - Reproducible ML Pipeline for Short-Term Rental Prices in
NYC
MLflow | Hydra | Data Validation | Random Forest | Model Selection | Versioned Retraining
2026/2027 Updated - Current Udacity Workspace Workflow
Current-version note
This resource follows the October 2026 CD16201 starter workflow. The current project runs locally in the Udacity Workspace
with MLflow/SQLite and portable artifacts. W&B, hosted tracking, a public GitHub repository, and remote releases are no longer
required for the recommended submission route.
1
, WGU D501 Machine Learning DevOps | Project 1 | 2026/2027 Updated
1. Project Scope and Current Assessment Workflow
The project builds a reusable regression pipeline for NYC short-term rental prices. The assessment emphasizes
reproducibility, explicit lineage, data validation, consistent model comparison, held-out evaluation, and versioned recovery
when a new data delivery exposes a defect.
Stage Purpose Primary output
Download Load bundled sample1.csv into a new run raw.csv
Apply price/date rules; later add NYC
Basic cleaning clean.csv
geographic filtering
Data check Run deterministic and distribution checks validation/report.xml
Create fixed train/validation source and held-
Data split split/trainval.csv, test.csv, split.json
out test data
Fit preprocessing + Random Forest; log
Train RF model/, metrics.json, feature_importance.png
validation metrics
Compare >=3 distinct configurations on same
Compare/select comparison.csv, selected_model.json
split
Held-out evaluation Evaluate only the explicitly selected model evaluation/metrics.json
Version replay Replay 1.0.0 and 1.0.1 against sample2.csv version manifests and source.bundle
Controlling metric
Select the candidate with the lowest validation MAE. R2 is logged and interpreted, but the held-out test set is not used to
choose hyperparameters.
2
WGU D501
Machine Learning DevOps
Project 1 - Reproducible ML Pipeline for Short-Term Rental Prices in
NYC
MLflow | Hydra | Data Validation | Random Forest | Model Selection | Versioned Retraining
2026/2027 Updated - Current Udacity Workspace Workflow
Current-version note
This resource follows the October 2026 CD16201 starter workflow. The current project runs locally in the Udacity Workspace
with MLflow/SQLite and portable artifacts. W&B, hosted tracking, a public GitHub repository, and remote releases are no longer
required for the recommended submission route.
1
, WGU D501 Machine Learning DevOps | Project 1 | 2026/2027 Updated
1. Project Scope and Current Assessment Workflow
The project builds a reusable regression pipeline for NYC short-term rental prices. The assessment emphasizes
reproducibility, explicit lineage, data validation, consistent model comparison, held-out evaluation, and versioned recovery
when a new data delivery exposes a defect.
Stage Purpose Primary output
Download Load bundled sample1.csv into a new run raw.csv
Apply price/date rules; later add NYC
Basic cleaning clean.csv
geographic filtering
Data check Run deterministic and distribution checks validation/report.xml
Create fixed train/validation source and held-
Data split split/trainval.csv, test.csv, split.json
out test data
Fit preprocessing + Random Forest; log
Train RF model/, metrics.json, feature_importance.png
validation metrics
Compare >=3 distinct configurations on same
Compare/select comparison.csv, selected_model.json
split
Held-out evaluation Evaluate only the explicitly selected model evaluation/metrics.json
Version replay Replay 1.0.0 and 1.0.1 against sample2.csv version manifests and source.bundle
Controlling metric
Select the candidate with the lowest validation MAE. R2 is logged and interpreted, but the held-out test set is not used to
choose hyperparameters.
2