Problem 3.1a
For this problem, I tried the leave one out cross validation approach with different splits. The
leave one out cross validation gives the “best k” which I used to find the accuracy of the train
and test data. I have also tried splits for the data for this problem which I have shown in the table.
Methodology is included in my R script.
Split Train accuracy Test accuracy Best k
7 0% train, 30% test 0.8468271 or 84.7% 0.8020305 or 80.2% 12
6 0% train, 40% test 0.8545918 or 85.5% 0.8053435 or 80.5% 5
5 0% train, 50% test 0.8440367 or 84.4% 0.82263 or 82.3% 5
Conclusion: Looking at the test accuracies, the 50% train and 50% test split lead to the best test
accuracy in this case. The 70% train and 30% test split had the least accuracy which could be due
to overfitting.
Problem 3.1b
I have not fully completed this problem but this was my taught process.
Conclusion: Incomplete
Problem 4.1
I go to the gym daily and realize everyone has a different routine. Some people are using the
treadmill and bicycle for cardio, some are lifting weights, and others are going to a yoga class in
the morning. We can apply a clustering model for all the different groups of gym members and
their workout routines. Some of the predictors include:
Time of day – some people prefer to go to the gym during the morning, afternoon, or
evening
Average workout time – some people go to the gym for a short or long-duration
Preferred activity type - some people do cardio, weightlift, or classes for a workout
Frequency of visits per week – occasional gym-goers vs. regular gym-goers
Calories burned per visit – depends on workout intensity and workout