COGS 108 ASSIGNMENT 4: DATA
ANALYSIS | 2026 UPDATED WITH
COMPLETE SOLUTIONS.
149 Questions with Answers and Detailed Rationales
100 PERCENT GUARANTEED PASS
INSTANT DOWNLOAD ANSWERS INCLUDED
IMPORTANCE OF THIS DOCUMENT
This comprehensive examination preparation guide has been meticulously developed to help you succeed in the
COGS 108 ASSIGNMENT 4: DATA ANALYSIS | 2026 UPDATED WITH COMPLETE SOLUTIONS.. It contains
149 carefully selected questions that reflect the most current exam content and testing strategies. Each question
is accompanied by a correct answer and a detailed rationale that explains the underlying pathophysiology,
pharmacology, or clinical reasoning.
Self-Assessment – Test your knowledge and Exam Preparation – Familiarize yourself with the
identify areas requiring further question format and content
study areas
Concept Reinforcement – Deepen your Confidence Building – Develop test-taking
understanding through strategies and reduce
evidence-based exam anxiety
rationales
Time Management – Practice answering
questions under simulated
exam conditions
Review Summary 149 Questions
Foundations - Application - COGS 108 Assignment 4 DATA Analysis 2026 Updated WITH Complete
Solutions Cognitive Science / DATA Analysis Undergraduate YEAR 3
All answers with rationales
,Table of Contents
Content Area Questions Key Topics
DATA Manipulation AND 1-25 Researcher, Appropriate, Regression, Missing, Hours
Cleaning WITH Pandas
Exploratory DATA Analysis 26-50 Regression, Researcher, Dataset, Interpretation, Appropriate
AND Visualization
Probability AND Sampling 51-75 Appropriate, Regression, Interpretation, Values, Missing
Distributions
Statistical Inference 76-100 Hypothesis, Values, Results, Appropriate, Study
Confidence Intervals AND
Hypothesis Testing
Simulation AND Resampling 101-125 Appropriate, Researcher, Dataset, Study, Hours
Methods Bootstrapping
Permutation Tests
Linear Regression AND 126-149 Dataset, Researcher, Appropriate, Distribution, Regression
Model Evaluation
TOTAL 149 All questions include answers and detailed rationales
,Section A - DATA Manipulation AND Cleaning WITH
Pandas
Q1.
A cognitive science researcher collects reaction time data that is heavily right-skewed.
Which transformation is most appropriate to reduce skewness before conducting a t-test?
A. Square root transformation B. Logarithmic transformation
C. Reciprocal transformation D. Box-Cox transformation with lambda = 2
Correct: B - Logarithmic transformation
Rationale:Logarithmic transformation is commonly used to reduce right skewness in
positively skewed data, such as reaction times. Square root is milder and may not sufficiently
reduce heavy skew, reciprocal is too extreme and reverses order, and Box-Cox with
lambda=2 increases skew.
Why the other answers are wrong:
A. Square root transformation is less effective for heavily skewed data and may not achieve
normality.
C. Reciprocal transformation is too drastic and can invert the scale, complicating interpretation.
D. Box-Cox with lambda=2 is for left-skewed data; it would worsen right skew.
Reference: Field, A. (2024). Discovering Statistics Using R, 2nd Ed., Ch. 5
Q2.
In a multiple regression model predicting memory performance from age, education, and
sleep quality, the VIF for age is 6.5. What does this indicate?
A. Age is not a significant predictor. B. There is multicollinearity involving age.
C. The model has heteroscedasticity. D. Age should be removed from the model.
Correct: B - There is multicollinearity involving age.
Rationale:A VIF above 5 (or sometimes 10) indicates problematic multicollinearity, meaning
age is highly correlated with other predictors. It does not directly indicate significance,
heteroscedasticity, or that age must be removed without further analysis.
Why the other answers are wrong:
A. VIF measures collinearity, not significance; significance is assessed via p-values.
C. Heteroscedasticity is about unequal variance of residuals, not VIF.
D. Removal is not automatic; one might combine variables or use regularization.
Reference: James, G. et al. (2023). An Introduction to Statistical Learning, 2nd Ed., Ch. 3
Page 3
, Section A - DATA Manipulation AND Cleaning WITH Pandas
Q3.
A researcher wants to visualize the distribution of a categorical variable with many levels.
Which plot is most effective?
A. Histogram B. Boxplot
C. Bar chart D. Scatterplot
Correct: C - Bar chart
Rationale:Bar charts are ideal for displaying frequencies or proportions of categorical
variables with multiple levels. Histograms and boxplots are for continuous data, and
scatterplots show relationships between two continuous variables.
Why the other answers are wrong:
A. Histograms are for continuous data, not categorical.
B. Boxplots summarize continuous distributions, not categories.
D. Scatterplots require two continuous variables.
Reference: Healy, K. (2024). Data Visualization: A Practical Introduction, 2nd Ed., Ch. 2
Q4.
In a study with 20 participants, a researcher conducts 10 independent t-tests without
correction. What is the approximate family-wise error rate?
A. 0.05 B. 0.10
C. 0.40 D. 0.60
Correct: C - 0.40
Rationale:The family-wise error rate is approximately 1 - (1 - 0.05)^10 "H 0.40. This illustrates
the multiple comparisons problem, where uncorrected tests inflate Type I error.
Why the other answers are wrong:
A. 0.05 is the per-comparison error rate, not family-wise.
B. 0.10 would be for two tests, not ten.
D. 0.60 overestimates; the correct calculation yields about 0.40.
Reference: Shadish, W.R. et al. (2022). Experimental and Quasi-Experimental Designs, Ch. 5
Q5.
Which of the following best describes the purpose of cross-validation in predictive
modeling?
A. To estimate the model's performance on B. To increase the model's fit to the training
unseen data. data.
Page 4
ANALYSIS | 2026 UPDATED WITH
COMPLETE SOLUTIONS.
149 Questions with Answers and Detailed Rationales
100 PERCENT GUARANTEED PASS
INSTANT DOWNLOAD ANSWERS INCLUDED
IMPORTANCE OF THIS DOCUMENT
This comprehensive examination preparation guide has been meticulously developed to help you succeed in the
COGS 108 ASSIGNMENT 4: DATA ANALYSIS | 2026 UPDATED WITH COMPLETE SOLUTIONS.. It contains
149 carefully selected questions that reflect the most current exam content and testing strategies. Each question
is accompanied by a correct answer and a detailed rationale that explains the underlying pathophysiology,
pharmacology, or clinical reasoning.
Self-Assessment – Test your knowledge and Exam Preparation – Familiarize yourself with the
identify areas requiring further question format and content
study areas
Concept Reinforcement – Deepen your Confidence Building – Develop test-taking
understanding through strategies and reduce
evidence-based exam anxiety
rationales
Time Management – Practice answering
questions under simulated
exam conditions
Review Summary 149 Questions
Foundations - Application - COGS 108 Assignment 4 DATA Analysis 2026 Updated WITH Complete
Solutions Cognitive Science / DATA Analysis Undergraduate YEAR 3
All answers with rationales
,Table of Contents
Content Area Questions Key Topics
DATA Manipulation AND 1-25 Researcher, Appropriate, Regression, Missing, Hours
Cleaning WITH Pandas
Exploratory DATA Analysis 26-50 Regression, Researcher, Dataset, Interpretation, Appropriate
AND Visualization
Probability AND Sampling 51-75 Appropriate, Regression, Interpretation, Values, Missing
Distributions
Statistical Inference 76-100 Hypothesis, Values, Results, Appropriate, Study
Confidence Intervals AND
Hypothesis Testing
Simulation AND Resampling 101-125 Appropriate, Researcher, Dataset, Study, Hours
Methods Bootstrapping
Permutation Tests
Linear Regression AND 126-149 Dataset, Researcher, Appropriate, Distribution, Regression
Model Evaluation
TOTAL 149 All questions include answers and detailed rationales
,Section A - DATA Manipulation AND Cleaning WITH
Pandas
Q1.
A cognitive science researcher collects reaction time data that is heavily right-skewed.
Which transformation is most appropriate to reduce skewness before conducting a t-test?
A. Square root transformation B. Logarithmic transformation
C. Reciprocal transformation D. Box-Cox transformation with lambda = 2
Correct: B - Logarithmic transformation
Rationale:Logarithmic transformation is commonly used to reduce right skewness in
positively skewed data, such as reaction times. Square root is milder and may not sufficiently
reduce heavy skew, reciprocal is too extreme and reverses order, and Box-Cox with
lambda=2 increases skew.
Why the other answers are wrong:
A. Square root transformation is less effective for heavily skewed data and may not achieve
normality.
C. Reciprocal transformation is too drastic and can invert the scale, complicating interpretation.
D. Box-Cox with lambda=2 is for left-skewed data; it would worsen right skew.
Reference: Field, A. (2024). Discovering Statistics Using R, 2nd Ed., Ch. 5
Q2.
In a multiple regression model predicting memory performance from age, education, and
sleep quality, the VIF for age is 6.5. What does this indicate?
A. Age is not a significant predictor. B. There is multicollinearity involving age.
C. The model has heteroscedasticity. D. Age should be removed from the model.
Correct: B - There is multicollinearity involving age.
Rationale:A VIF above 5 (or sometimes 10) indicates problematic multicollinearity, meaning
age is highly correlated with other predictors. It does not directly indicate significance,
heteroscedasticity, or that age must be removed without further analysis.
Why the other answers are wrong:
A. VIF measures collinearity, not significance; significance is assessed via p-values.
C. Heteroscedasticity is about unequal variance of residuals, not VIF.
D. Removal is not automatic; one might combine variables or use regularization.
Reference: James, G. et al. (2023). An Introduction to Statistical Learning, 2nd Ed., Ch. 3
Page 3
, Section A - DATA Manipulation AND Cleaning WITH Pandas
Q3.
A researcher wants to visualize the distribution of a categorical variable with many levels.
Which plot is most effective?
A. Histogram B. Boxplot
C. Bar chart D. Scatterplot
Correct: C - Bar chart
Rationale:Bar charts are ideal for displaying frequencies or proportions of categorical
variables with multiple levels. Histograms and boxplots are for continuous data, and
scatterplots show relationships between two continuous variables.
Why the other answers are wrong:
A. Histograms are for continuous data, not categorical.
B. Boxplots summarize continuous distributions, not categories.
D. Scatterplots require two continuous variables.
Reference: Healy, K. (2024). Data Visualization: A Practical Introduction, 2nd Ed., Ch. 2
Q4.
In a study with 20 participants, a researcher conducts 10 independent t-tests without
correction. What is the approximate family-wise error rate?
A. 0.05 B. 0.10
C. 0.40 D. 0.60
Correct: C - 0.40
Rationale:The family-wise error rate is approximately 1 - (1 - 0.05)^10 "H 0.40. This illustrates
the multiple comparisons problem, where uncorrected tests inflate Type I error.
Why the other answers are wrong:
A. 0.05 is the per-comparison error rate, not family-wise.
B. 0.10 would be for two tests, not ten.
D. 0.60 overestimates; the correct calculation yields about 0.40.
Reference: Shadish, W.R. et al. (2022). Experimental and Quasi-Experimental Designs, Ch. 5
Q5.
Which of the following best describes the purpose of cross-validation in predictive
modeling?
A. To estimate the model's performance on B. To increase the model's fit to the training
unseen data. data.
Page 4