ISYE 6414 MIDTERM EXAM 2 – PART 2
DATA ANALYSIS | COMPLETE PRACTICE
QUESTIONS, SOLUTIONS & R CODE |
2026/2027 UPDATED | QUESTIONS 1–
100
Study-use note: The questions below are newly written practice material designed to
develop skills relevant to ISYE 6414-style statistical modeling and R-based data
analysis. They are not official Georgia Institute of Technology exam questions or
answer keys.
Introduction
This practice bank emphasizes applied statistical modeling and data-analysis skills,
with particular attention to interpreting model output, selecting appropriate models,
diagnosing assumptions, and implementing analyses in R. Questions are intentionally
scenario-based and require more than memorizing commands. You will encounter
regression interpretation, categorical variables, interactions, transformations, model
diagnostics, multicollinearity, model comparison, prediction intervals, confidence
intervals, ANOVA, residual analysis, influential observations, and practical R
implementation. Several questions require reasoning from hypothetical R output
rather than calculating every quantity from scratch. The goal is to practice recognizing
which statistical procedure is appropriate, interpreting coefficients correctly,
identifying violations of assumptions, and choosing defensible remedies. Where R
code is relevant, the questions focus on understanding what commands such as lm(),
summary(), anova(), predict(), and diagnostic functions actually accomplish.
Work through each question before checking the answer and rationale.
Core Domains Tested
1.
Linear Regression — Model specification, estimation, fitted values, and
coefficient interpretation.
2.
3.
Multiple Regression — Partial effects and adjustment for other predictors.
4.
5.
1
,Categorical Predictors — Dummy-variable construction and reference-group
interpretation.
6.
7.
Interactions — Understanding how one predictor changes another predictor's
effect.
8.
9.
Transformations — Logarithmic and polynomial transformations.
10.
11.
Inference — t-tests, F-tests, confidence intervals, and p-values.
12.
13.
ANOVA & Model Comparison — Nested-model testing and variance
decomposition.
14.
15.
Diagnostics — Residuals, leverage, influence, heteroscedasticity, and
nonlinearity.
16.
17.
Multicollinearity — Correlated predictors and consequences for inference.
18.
19.
R Data Analysis — Correct use and interpretation of common R modeling
commands.
20.
2
, Q1: A researcher fits lm(Y ~ X1 + X2, data = d). The estimated
coefficient for X1 is 4.2. Which interpretation is most
appropriate?
A) Increasing X1 by one unit increases Y by exactly 4.2 units for every observation.
B) Holding X2 constant, a one-unit increase in X1 is associated with an estimated
4.2-unit increase in the mean response Y.
C) X1 causes Y to increase by 4.2 units.
D) The correlation between X1 and Y is 4.2.
Rationale: B is correct because a multiple-regression coefficient represents the
estimated change in the conditional mean response associated with a one-unit change
in that predictor while the other included predictors are held constant. A is too
deterministic because regression describes an average relationship rather than
guaranteeing individual outcomes. C incorrectly interprets association as causation.
D is incorrect because a regression coefficient is not generally a correlation and may
have units.
Q2: In R, which command most directly provides coefficient
estimates, standard errors, t-statistics, and p-values for a
fitted linear model fit?
A) coef(fit)
B) anova(fit)
C) summary(fit)
D) resid(fit)
Rationale: C is correct because summary(fit) provides the main inferential output
for an lm object, including coefficient estimates, standard errors, t-statistics, and
corresponding p-values. A primarily extracts coefficient estimates. B provides
ANOVA information rather than the standard coefficient table. D extracts residuals.
Q3: A 95% confidence interval for a regression coefficient is
(1.3, 5.7). Which conclusion is justified?
A) There is a 95% probability that the fixed coefficient lies between 1.3 and 5.7.
B) 95% of observations have coefficients between 1.3 and 5.7.
C) The coefficient must equal 3.5.
D) The interval provides a 95% confidence procedure whose resulting interval
excludes zero, providing evidence against a zero coefficient at the 5% level.
Rationale: D is correct. Because zero is excluded, a two-sided 5% test of the null
hypothesis that the coefficient equals zero would reject the null under the usual
assumptions. A is a common but technically incorrect frequentist interpretation. B is
unrelated to coefficient intervals. C incorrectly treats the midpoint as the known
coefficient.
3
DATA ANALYSIS | COMPLETE PRACTICE
QUESTIONS, SOLUTIONS & R CODE |
2026/2027 UPDATED | QUESTIONS 1–
100
Study-use note: The questions below are newly written practice material designed to
develop skills relevant to ISYE 6414-style statistical modeling and R-based data
analysis. They are not official Georgia Institute of Technology exam questions or
answer keys.
Introduction
This practice bank emphasizes applied statistical modeling and data-analysis skills,
with particular attention to interpreting model output, selecting appropriate models,
diagnosing assumptions, and implementing analyses in R. Questions are intentionally
scenario-based and require more than memorizing commands. You will encounter
regression interpretation, categorical variables, interactions, transformations, model
diagnostics, multicollinearity, model comparison, prediction intervals, confidence
intervals, ANOVA, residual analysis, influential observations, and practical R
implementation. Several questions require reasoning from hypothetical R output
rather than calculating every quantity from scratch. The goal is to practice recognizing
which statistical procedure is appropriate, interpreting coefficients correctly,
identifying violations of assumptions, and choosing defensible remedies. Where R
code is relevant, the questions focus on understanding what commands such as lm(),
summary(), anova(), predict(), and diagnostic functions actually accomplish.
Work through each question before checking the answer and rationale.
Core Domains Tested
1.
Linear Regression — Model specification, estimation, fitted values, and
coefficient interpretation.
2.
3.
Multiple Regression — Partial effects and adjustment for other predictors.
4.
5.
1
,Categorical Predictors — Dummy-variable construction and reference-group
interpretation.
6.
7.
Interactions — Understanding how one predictor changes another predictor's
effect.
8.
9.
Transformations — Logarithmic and polynomial transformations.
10.
11.
Inference — t-tests, F-tests, confidence intervals, and p-values.
12.
13.
ANOVA & Model Comparison — Nested-model testing and variance
decomposition.
14.
15.
Diagnostics — Residuals, leverage, influence, heteroscedasticity, and
nonlinearity.
16.
17.
Multicollinearity — Correlated predictors and consequences for inference.
18.
19.
R Data Analysis — Correct use and interpretation of common R modeling
commands.
20.
2
, Q1: A researcher fits lm(Y ~ X1 + X2, data = d). The estimated
coefficient for X1 is 4.2. Which interpretation is most
appropriate?
A) Increasing X1 by one unit increases Y by exactly 4.2 units for every observation.
B) Holding X2 constant, a one-unit increase in X1 is associated with an estimated
4.2-unit increase in the mean response Y.
C) X1 causes Y to increase by 4.2 units.
D) The correlation between X1 and Y is 4.2.
Rationale: B is correct because a multiple-regression coefficient represents the
estimated change in the conditional mean response associated with a one-unit change
in that predictor while the other included predictors are held constant. A is too
deterministic because regression describes an average relationship rather than
guaranteeing individual outcomes. C incorrectly interprets association as causation.
D is incorrect because a regression coefficient is not generally a correlation and may
have units.
Q2: In R, which command most directly provides coefficient
estimates, standard errors, t-statistics, and p-values for a
fitted linear model fit?
A) coef(fit)
B) anova(fit)
C) summary(fit)
D) resid(fit)
Rationale: C is correct because summary(fit) provides the main inferential output
for an lm object, including coefficient estimates, standard errors, t-statistics, and
corresponding p-values. A primarily extracts coefficient estimates. B provides
ANOVA information rather than the standard coefficient table. D extracts residuals.
Q3: A 95% confidence interval for a regression coefficient is
(1.3, 5.7). Which conclusion is justified?
A) There is a 95% probability that the fixed coefficient lies between 1.3 and 5.7.
B) 95% of observations have coefficients between 1.3 and 5.7.
C) The coefficient must equal 3.5.
D) The interval provides a 95% confidence procedure whose resulting interval
excludes zero, providing evidence against a zero coefficient at the 5% level.
Rationale: D is correct. Because zero is excluded, a two-sided 5% test of the null
hypothesis that the coefficient equals zero would reject the null under the usual
assumptions. A is a common but technically incorrect frequentist interpretation. B is
unrelated to coefficient intervals. C incorrectly treats the midpoint as the known
coefficient.
3