STATISTICS II
SOLUTIONS MANUAL
KOEN HANEGREEFS
VUB
,Table of Contents
Statistics II — Worked Solutions Manual ......................................................................... 4
How to use this manual............................................................................................... 4
Chapter 10 — Worked Solutions (WPO 2: Confidence Intervals for Proportions) ................ 5
Exercise 4 — Women in Latvia ..................................................................................... 5
Exercise 5 — Store Customers .................................................................................... 6
Exercise 34 — Mortgages 2013 .................................................................................... 6
Exercise 35 — Loans ................................................................................................... 7
Exercise 71 — Smartphones ........................................................................................ 8
Exercise 75 — Pilot Study (Emissions) .......................................................................... 9
RStudio Exercise — Holiday Shopping ......................................................................... 9
LEGO Example — Summary (Conceptual Illustration) ................................................. 10
Chapter 11 — Worked Solutions (WPO 3: Confidence Intervals for Means)...................... 12
Exercise 24 — Confidence Intervals and Sample Size ................................................. 12
Exercise 29 — Parking Garage ................................................................................... 13
Exercise 36 — Late Flight Arrivals............................................................................... 14
Exercise 41 — Wire Manufacturing............................................................................. 14
Exercise 47 (R Studio) — Credit Card Spending .......................................................... 15
Chapter 12 — Worked Solutions (WPO 4: Testing Hypotheses About Proportions) ........... 17
Exercise 11 — Pizza (p-value interpretation) (At home)................................................ 17
Exercise 15 — Setting Up Hypotheses (At home) ........................................................ 17
Exercise 33 — Twins (In class) ................................................................................... 18
Exercise 27 — Absentees (In class) ............................................................................ 19
Exercise R Studio — Flower Hill Real Estate (In class) ................................................. 20
Chapter 13 — Worked Solutions (WPO 5: More About Tests & Intervals (paired t-test,
errors, power)) ............................................................................................................. 23
Exercise 27 — Loans (at home) (Important) ................................................................ 23
Exercise 29 — Second Loan (at home) (Important)...................................................... 24
Exercise 33 — Testing Cars (Important) ...................................................................... 24
Exercise 37 — Gender Discrimination (Important) ...................................................... 25
Exercise 47 — Two Coins (Neutral) ............................................................................ 26
R Exercise — Flower Hill (Important) .......................................................................... 27
Chapter 14 — Worked Solutions (WPO 6: Comparing Two Means) .................................. 30
Koen Hanegreefs 1
, Exercise 26 — Sports (at home) ................................................................................. 30
Exercise 32 — Soccer................................................................................................ 31
Exercise 34 — Living Cost .......................................................................................... 31
Exercise 55 — Income (at home)................................................................................ 33
R Studio Exercise — CardioX Drug Trial ...................................................................... 33
Chapter 15 — Worked Solutions (WPO 7: Chi-Square Tests) ........................................... 36
Exercise 25 — Cranberry Juice (Homogeneity) (Important) .......................................... 36
Exercise 32 — Labour Force Immigration (Homogeneity / 2×2) (Important) .................. 38
Exercise 41 — Racial Steering (Independence) (Important) ......................................... 40
Exercise 43 — Racial Steering Revisited: Confidence Interval (Important) .................... 41
R Studio Exercise — New Drug (CardioX vs StandardCare) (Neutral) ............................ 42
Question 1: Do severe side effects differ across hospitals (site)? ................................ 42
Question 2: Do severe side effects differ by drug? ...................................................... 42
LEGO Example — Chi-Square Test of Independence with Expected Count Problem
(Neutral)................................................................................................................... 43
Chapter 16 — Worked Solutions (WPO 8: Linear Regression) .......................................... 45
Exercise: WPO 8 Part 1 — Elderly Care Survey (Q5: Regression) (Important)................. 45
Exercise 30 (At Home) — Automobiles (Important) ..................................................... 46
Exercise 37 — Labor Productivity 2020 (Important) ..................................................... 47
Exercise 39 — Labor Productivity 2020, Part 2 (Important) .......................................... 48
Exercise 56 — Youth Employment 2016 (Important).................................................... 48
R Studio Exercise — Graduate Employment (Important) ............................................. 49
LEGO Example — Full Worked Regression Pipeline (Neutral) ...................................... 50
Chapter 17 — Worked Solutions (WPO 9: Inference for Regression & Residual Analysis) .. 53
Exercise 12 — Durbin-Watson for Mail Orders (Important) .......................................... 53
Exercise 21 — Palm Oil Prices (Scatterplot Interpretation) (Important) ........................ 53
Exercise 32 — Unusual Points in Scatterplots (Important) ........................................... 54
Exercise 40 — Palm Oil Part 2: Price Difference Regression (Important) ....................... 55
Exercise 53 — Lobster Catch Value 1950–2016 (Important) ........................................ 56
R Exercise — Energy Consumption at a Belgian Factory (Important) ............................ 57
Summary Table (Important) ....................................................................................... 61
Chapter 23 — Worked Solutions (Seminar 10: Non-Parametric Methods)........................ 62
Exercise 6 — Choosing the Right Test (Important) ....................................................... 62
Koen Hanegreefs 2
, Exercise 17 — Trophy Sizes (Mann-Whitney) (Important) ............................................. 62
Exercise 22 — Scanner Errors (Wilcoxon Signed-Rank) (Important) ............................. 63
Exercise 26 — Cholesterol Drugs (Kruskal-Wallis) (Important) .................................... 63
Exercise 30 — Carbon Footprint (Spearman’s 𝜌) (Neutral) .......................................... 64
23.9 What Can Go Wrong (Important) ........................................................................ 65
Koen Hanegreefs 3
,Statistics II — Worked Solutions Manual
Course: Statistics for Business and Economics II (1009724BNR) — Vrije Universiteit
Brussel (VUB) Handbook: Business Statistics, 4th global edition — Sharpe, De Veaux &
Velleman
How to use this manual
This is the companion to Statistics II — Course Summary. It contains a step-by-step
worked solution for every exercise covered in the WPO seminars (WPO 1 R-recap
excluded), grouped by chapter.
Every numerical solution follows the same four-step pedagogy:
1. Setup / Hypotheses — state 𝐻! and 𝐻" in words and symbols
2. Conditions — check independence, randomization, 10 % condition, sample-size or
Normality requirements
3. Mechanics — substitute numbers into the formula, show every intermediate value,
give R code
4. Conclusion — translate the statistical decision into the original context
Conceptual exercises (e.g., “what is a Type I error?”) use a simpler restatement-and-
explanation format.
Cross-reference: each chapter heading here matches the chapter number in Statistics II —
Course Summary; pair the two documents while studying.
Koen Hanegreefs 4
,Chapter 10 — Worked Solutions (WPO 2: Confidence Intervals for
Proportions)
Exercise 4 — Women in Latvia
Problem. The proportion of women in Latvia is approximately 54%. A company conducting
a marketing survey telephones 400 people in Latvia at random.
(a) What is the sampling distribution of the observed proportion that are women?
Because 𝑛 = 400 is large and conditions are met (see below), the sampling distribution is
approximately Normal:
0.54 × 0.46
𝑝̂ ∼ 𝑁 +0.54, 0 3
400
(b) What is the standard deviation of that proportion?
𝑝𝑞 0.54 × 0.46 0.2484
𝑆𝐷(𝑝̂ ) = 8 = 0 =0 = √0.000621 = 0.025
𝑛 400 400
(c) Would you be surprised to find 56% women in a sample of 400?
𝑝̂ − 𝑝 0.56 − 0.54 0.02
𝑧= = = = 0.8
𝑆𝐷(𝑝̂ ) 0.025 0.025
𝑝̂ = 0.56 lies only 0.8 standard deviations above the mean. This is well within 1 SD of 𝑝.
Not surprising.
(d) Would you be surprised to find 51% women in a sample of 400?
0.51 − 0.54 −0.03
𝑧= = = −1.2
0.025 0.025
𝑝̂ = 0.51 lies 1.2 SDs below the mean. Still within a common range. Not surprising.
(e) Would you be surprised if fewer than 180 women were in the sample?
Convert the count to a proportion:
180
𝑝̂ = = 0.45
400
0.45 − 0.54 −0.09
𝑧= = = −3.6
0.025 0.025
Koen Hanegreefs 5
,𝑝̂ = 0.45 is more than 3 standard deviations below the mean. By the 68–95–99.7 rule, fewer
than 0.3% of samples would produce a result this extreme. Very surprising.
Exercise 5 — Store Customers
Problem. A store manager samples 100 customers from last month and finds that 15 of
them spent more than £1000. Are all assumptions and conditions for the sampling
distribution of the proportion satisfied?
Observed: 𝑛 = 100, 𝑝̂ = 15/100 = 0.15, 𝑞C = 0.85.
Condition Check Met?
Randomization Assume customers selected randomly and Yes
independently (assumed)
10% condition Assume 100 < 10% of monthly customers Yes
(assumed)
Success/Failure 𝑛𝑝̂ = 100 × 0.15 = 15 ≥ 10 and 𝑛𝑞C = 100 × 0.85 = Yes
85 ≥ 10
All conditions are satisfied. The sampling distribution of 𝑝̂ can be modelled as Normal.
Exercise 34 — Mortgages 2013
Problem. Bloomberg reports Spanish mortgage default rate is 5% in 2013. A large Spanish
bank holds 𝑛 = 9 455 mortgages.
(a) Can you use the Normal model? Check conditions.
Condition Check Result
Randomization Assume the 9 455 mortgages are a random sample of all Assumed
Spanish mortgages
10% condition Assume 9 455 < 10% of all Spanish mortgages Assumed
Success/Failure 𝑛𝑝 = 9455 × 0.05 = 472.75 ≥ 10 and 𝑛𝑞 = 9455 × 0.95 = Met
8982.25 ≥ 10
Standard deviation:
𝑝𝑞 0.05 × 0.95
𝑆𝐷(𝑝̂ ) = 8 = 0 = √0.000005026 = 0.00224
𝑛 9455
The sampling distribution is 𝑁(0.05, 0.00224).
Koen Hanegreefs 6
,(b) Sketch and label using the 68–95–99.7 rule.
Applying the rule to 𝑁(0.05,0.00224):
Range Interval
𝜇 ± 1 𝑆𝐷 [0.04776, 0.05224]
𝜇 ± 2 𝑆𝐷 [0.04552, 0.05448]
𝜇 ± 3 𝑆𝐷 [0.04328, 0.05672]
A bell curve centred at 0.05 would be labelled with these values at ±1, ±2, ±3 standard
deviations.
(c) How many homeowners might the bank expect to default?
Using 95% certainty (±2 𝑆𝐷):
Lower bound: 0.05 − 2 × 0.00224 = 0.04552 ⟹ 0.04552 × 9455 ≈ 430
Upper bound: 0.05 + 2 × 0.00224 = 0.05448 ⟹ 0.05448 × 9455 ≈ 515
The bank expects between 430 and 515 of the 9 455 mortgages to default (with
approximately 95% confidence).
Exercise 35 — Loans
Problem. A bank believes 7% of loan recipients will not make timely payments. The bank
has recently approved 𝑛 = 200 loans.
(a) Mean and standard deviation of the proportion who may not pay on time.
𝜇#$ = 𝑝 = 0.07
𝑝𝑞 0.07 × 0.93 0.0651
𝑆𝐷(𝑝̂ ) = 8 =0 =0 = √0.0003255 = 0.018
𝑛 200 200
(b) Assumptions and conditions.
Condition Check Met?
Randomization Assume 200 loans selected randomly from the bank’s Assumed
portfolio
10% condition Assume 200 < 10% of all bank loans Assumed
Success/Failure 𝑛𝑝 = 200 × 0.07 = 14 ≥ 10 and 𝑛𝑞 = 200 × 0.93 = 186 ≥ Met
10
The sampling distribution is 𝑁(0.07, 0.018).
Koen Hanegreefs 7
,(c) Probability that over 10% will not make timely payments.
𝑝̂ − 𝑝 0.10 − 0.07 0.03
𝑧= = = = 1.663
𝑆𝐷(𝑝̂ ) 0.018 0.018
𝑃(𝑝̂ > 0.10) = 𝑃(𝑧 > 1.663) = 1 − 𝛷(1.663) = 0.048
The probability that more than 10% of clients will not pay on time is approximately
4.8%.
Exercise 71 — Smartphones
Problem. A survey of 500 teens finds 57% prefer iOS.
(a) Create a 95% confidence interval for this percentage.
Given: 𝑝̂ = 0.57, 𝑞C = 0.43, 𝑛 = 500.
Conditions check (for CI — use 𝑝̂ ):
Condition Check Met?
Randomization Assume random sample < 10% of all teens Assumed
Success/Failure 𝑛𝑝̂ = 500 × 0.57 = 285 ≥ 10 and 𝑛𝑞C = 500 × 0.43 = 215 ≥ Met
10
𝑝̂ 𝑞C 0.57 × 0.43 0.2451
𝑆𝐸(𝑝̂ ) = 0 =0 =0 = √0.0004902 = 0.0221
𝑛 500 500
95% 𝐶𝐼 = 𝑝̂ ± 𝑧 ∗ ⋅ 𝑆𝐸(𝑝̂ ) = 0.57 ± 1.96 × 0.0221 = 0.57 ± 0.0433
95% 𝐶𝐼 = [52.67%, 61.33%]
Interpretation: We are 95% confident that the true proportion of teens who prefer iOS is
between 52.67% and 61.33%.
(b) How many teens must be surveyed if the margin of error is cut to half?
Current ME:
𝑀𝐸 = 𝑧 ∗ ⋅ 𝑆𝐸(𝑝̂ ) = 1.96 × 0.0221 = 0.0433
To halve the ME: 𝑀𝐸new = 𝑀𝐸/2.
Since 𝑀𝐸 = 𝑧 ∗ Y𝑝̂ 𝑞C/𝑛, halving ME requires √𝑛 to double, meaning 𝑛 must be multiplied by
4:
𝑛new = 4 × 𝑛 = 4 × 500 = 2000
Koen Hanegreefs 8
, Verification: 𝑀𝐸new = 1.96Y0.57 × 0.43/2000 = 1.96 × 0.01564 ≈ 0.0217 ≈ 𝑀𝐸/2 ✓
They must survey 2000 teens.
General rule: Halving the margin of error requires quadrupling the sample size.
Exercise 75 — Pilot Study (Emissions)
Problem. An environmental agency wants to estimate the proportion of cars failing
emissions standards with ME = 3% and 90% confidence. A pilot study of 𝑛! = 60 cars finds
9 with faulty emissions systems. How many cars should be sampled for the full
investigation?
Given: 𝑀𝐸 = 0.03, confidence = 90% so 𝑧 ∗ = 1.645, pilot estimate 𝑝̂ = 9/60 = 0.15, 𝑞C =
0.85.
(𝑧 ∗ )& ⋅ 𝑝̂ 𝑞C (1.645)& × 0.15 × 0.85
𝑛= =
𝑀𝐸 & (0.03)&
2.706 × 0.1275 0.345
𝑛= = = 383.35
0.0009 0.0009
Round up: 𝑛 = 384.
The agency should sample at least 384 cars for the full investigation.
RStudio Exercise — Holiday Shopping
Problem. A credit card company wants to determine whether the true proportion of “High
Spenders” (spending > $1000) in the North Region in December is at least 50%, as the
Regional Manager claims. A sample of 600 customers is available
(holiday_spending.csv). Use a 95% confidence interval.
Step 1 — Data cleaning and subsetting:
# Import data
raw_data <- read.csv("holiday_spending.csv")
# A. Remove missing values
clean_data <- na.omit(raw_data)
# B. Convert dates and extract month
clean_data$month <- months(as.Date(clean_data$date))
# C. Filter for North Region, December only
north_data <- subset(clean_data, region == "North" & date == "2023-12-15")
Koen Hanegreefs 9