ISYE 6501 Homework 7
Question 10.1a Methodology & Analysis
I made a model named “crime_tree” for the regression tree and these are the results. There are
four variables that the tree used to split the data which are Po1, Pop, LF, and NW. The residuals
are centered around zero, which means the tree captures general trends well. Next, I plotted
crime_tree and used $frame and $where.
Po1 has the strongest split as well as higher police spending doesn’t mean lower crime. The
lowest crime rates are found in areas with low police spending, small populations, and low labor
force participation. The $where shows where each leaf falls into in the 47 observations.
https://www.stuvia.com/user/nursecare
,https://www.stuvia.com/user/nursecare
I pruned the dataset using the prune.tree() function and I used 5 as the best.
I did the summary() and the model is worse when pruned since the residual mean deviance
increased.
Next, I did the cross validation on the pruned model. I also got the list sizes of pruned trees and
training deviance per tree and the cross validated deviance. Lastly, I plotted the
cv_crime$deviance against cv_crime$size to find the smallest tree size where deviance stops
decreasing at.
https://www.stuvia.com/user/nursecare