ISYE 6501 – Homework 10 Solutions & Answers (Verified, A+ Grade) ISYE
https://www.stuvia.com/user/nursecare
6501 - Homework 10 GUARANTEED 100% COMPLETE SOLUTIONS 2026 ,.
The breast cancer data set breast-cancer-wisconsin.data.txt has missing values. Let’s load it in and see what
is wrong
with it
# initilize libraries library(kknn)
# Set seed for reproducibility
set.seed(123)
# To begin, we import the data and save as a dataframe for kknn functon data <- as.data.frame(read.table("~/GT
OMSA/Intro
Analytcs to
Modeling/HW 10/breast-cancer-wisconsin.da
header=FALSE,sep=","))
head(data)
## V1 V2 V3 V4 V5 V6 V7 V8 V9 V10
V11
## 1 1000025 5 1 1 1 2 1 3 1 1 2
## 2 1002945 5 4 4 5 7 10 3 2 1 2
## 3 1015425 3 1 1 1 2 2 3 1 1 2
## 4 1016277 6 8 8 1 3 4 3 7 1 2
## 5 1017023 4 1 1 3 2 1 3 1 1 2
## 6 1017122 8 10 10 8 7 10 9 1 4
7
summary(data)
## V1 V2 V3 V4
## Min. : 61634 Min. : 1.000 Min. : 1.000 Min. : 1.000
## 1st Qu.: 870688 1st Qu.: 2.000 1st Qu.: 1.000 1st Qu.: 1.000
## Median : 1171710 Median : 4.000 Median : 1.000 Median : 1.000
## Mean : 1071704 Mean : 4.418 Mean : 3.134 Mean : 3.207
## 3rd Qu.: 1238298 3rd Qu.: 6.000 3rd Qu.: 5.000 3rd Qu.: 5.000
## Max. :13454352 Max. :10.000 Max. :10.000 Max. :10.000
## V5 V6 V7 V8
## Min. : 1.000 Min. : 1.000 Length:699 Min. : 1.000
## 1st Qu.: 1.000 1st Qu.: 2.000 Class :character 1st Qu.: 2.000
## Median : 1.000 Median : 2.000 Mode :character Median : 3.000
## Mean : 2.807 Mean : 3.216 Mean : 3.438
## 3rd Qu.: 4.000 3rd Qu.: 4.000 3rd Qu.: 5.000
https://www.stuvia.com/user/nursecare
, lOMoARcPSD|63116700
https://www.stuvia.com/user/nursecare
## Max. :10.000 Max. :10.000 Max. :10.000
## V9 V10 V11
## Min. : 1.000 Min. : 1.000 Min. :2.00
## 1st Qu.: 1.000 1st Qu.: 1.000 1st Qu.:2.00
## Median : 1.000 Median : 1.000 Median :2.00
## Mean : 2.867 Mean : 1.589 Mean :2.69
## 3rd Qu.: 4.000 3rd Qu.: 1.000 3rd Qu.:4.00
## Max. :10.000 Max. :10.000 Max. :4.00
V7 looks weird, it’s being listed as a character class. Looking through the actual data set in another tab, I see
some “?”
characters throughout. Per lecture, we don’t want too much of our data to be missing (<5%)
error_list <- which(data$V7 == "?") error_list #list
of
indices that have a ?
## [1] 24 41 140 146 159 165 236 250 276 293 295 298 316 322
412 618
length(error_list)/nrow(data)
## [1] 0.02288984
About ~2% of our data in column V7 is missing. We should also check if any of our data is biased by breaking
out our
response variable.
data_good <- data[-error_list,] data_bad
<-
data[error_list,] table(data$V11)
##
## 2 4
## 458 241
table(data_good$V11)
##
## 2 4
## 444 239
table(data_bad$V11)
##
## 2 4
## 14 2
That doesn’t look good, the corrupted data is skewed towards one of the outcomes 2 (benign) - something to
keep
mindingoing forward. Anyways, we should be good to contnue forward with our different methods of
replacement.
https://www.stuvia.com/user/nursecare