WGU D502 Data Analytics Capstone | BHN1 Task 3: Project Report
WGU D502
Data Analytics Capstone
BHN1 Task 3: Project Report
Predicting Online Purchase Intention from E-Commerce
Session Behavior Using Machine Learning
Primary method: Random Forest classification
Dataset: UCI Online Shoppers Purchasing Intention Dataset
Benchmark: Held-out F1 >= 0.60
2026/2027
D502 | BHN1 Task 3 | Page 1
, WGU D502 Data Analytics Capstone | BHN1 Task 3: Project Report
Project Results Snapshot
This report concludes the approved capstone project by documenting execution, data preparation, supervised
classification, model evaluation, practical significance, recommendations, and the required Panopto
presentation plan. The Random Forest model remained the primary hypothesis-supporting method established
in Task 2.
Measure Result / conclusion
Cleaned analytical data 12,205 sessions after removal of 125 exact duplicate rows;
no source-reported missing values.
Primary model Random Forest classifier with class-balanced training and a
leakage-controlled preprocessing pipeline.
Primary success benchmark Held-out F1 >= 0.60.
Verified Random Forest benchmark run Precision 0.6267; recall 0.7251; F1 0.6723; ROC-AUC
0.9296; accuracy 0.8894.
Hypothesis conclusion Supported: the Random Forest F1 of 0.6723 exceeds the
prespecified 0.60 benchmark by 0.0723.
Strongest predictive signal PageValues ranked first and accounted for approximately
38% of reported Random Forest importance in the
documented reference run.
Business interpretation The model is useful for prioritization and low-cost targeting,
but predictions remain probabilistic and should not be
treated as causal or certain.
Analytical evidence note: The numerical model results reproduced in this resource come from a documented
public implementation that uses the same UCI dataset, duplicate-removal rule, 80/20 stratified split,
random_state=42, preprocessing pattern, and class-balanced modeling approach. They are included so the
resource does not invent metrics. A student submission should rerun the companion analysis code and
replace the benchmark metrics/screenshots with evidence generated in the student environment.
D502 | BHN1 Task 3 | Page 2
WGU D502
Data Analytics Capstone
BHN1 Task 3: Project Report
Predicting Online Purchase Intention from E-Commerce
Session Behavior Using Machine Learning
Primary method: Random Forest classification
Dataset: UCI Online Shoppers Purchasing Intention Dataset
Benchmark: Held-out F1 >= 0.60
2026/2027
D502 | BHN1 Task 3 | Page 1
, WGU D502 Data Analytics Capstone | BHN1 Task 3: Project Report
Project Results Snapshot
This report concludes the approved capstone project by documenting execution, data preparation, supervised
classification, model evaluation, practical significance, recommendations, and the required Panopto
presentation plan. The Random Forest model remained the primary hypothesis-supporting method established
in Task 2.
Measure Result / conclusion
Cleaned analytical data 12,205 sessions after removal of 125 exact duplicate rows;
no source-reported missing values.
Primary model Random Forest classifier with class-balanced training and a
leakage-controlled preprocessing pipeline.
Primary success benchmark Held-out F1 >= 0.60.
Verified Random Forest benchmark run Precision 0.6267; recall 0.7251; F1 0.6723; ROC-AUC
0.9296; accuracy 0.8894.
Hypothesis conclusion Supported: the Random Forest F1 of 0.6723 exceeds the
prespecified 0.60 benchmark by 0.0723.
Strongest predictive signal PageValues ranked first and accounted for approximately
38% of reported Random Forest importance in the
documented reference run.
Business interpretation The model is useful for prioritization and low-cost targeting,
but predictions remain probabilistic and should not be
treated as causal or certain.
Analytical evidence note: The numerical model results reproduced in this resource come from a documented
public implementation that uses the same UCI dataset, duplicate-removal rule, 80/20 stratified split,
random_state=42, preprocessing pattern, and class-balanced modeling approach. They are included so the
resource does not invent metrics. A student submission should rerun the companion analysis code and
replace the benchmark metrics/screenshots with evidence generated in the student environment.
D502 | BHN1 Task 3 | Page 2