COGS 108 ASSIGNMENT 3: DATA
ANALYSIS | 2026 UPDATE WITH
COMPLETE SOLUTIONS.
150 Questions with Answers and Detailed Rationales
100 PERCENT GUARANTEED PASS
INSTANT DOWNLOAD ANSWERS INCLUDED
IMPORTANCE OF THIS DOCUMENT
This comprehensive examination preparation guide has been meticulously developed to help you succeed in the
COGS 108 ASSIGNMENT 3: DATA ANALYSIS | 2026 UPDATE WITH COMPLETE SOLUTIONS.. It contains 150
carefully selected questions that reflect the most current exam content and testing strategies. Each question is
accompanied by a correct answer and a detailed rationale that explains the underlying pathophysiology,
pharmacology, or clinical reasoning.
Self-Assessment – Test your knowledge and Exam Preparation – Familiarize yourself with the
identify areas requiring further question format and content
study areas
Concept Reinforcement – Deepen your Confidence Building – Develop test-taking
understanding through strategies and reduce
evidence-based exam anxiety
rationales
Time Management – Practice answering
questions under simulated
exam conditions
Review Summary 150 Questions
Foundations - Application - COGS 108 Assignment 3 DATA Analysis 2026 Update WITH Complete
Solutions DATA Analysis FOR Cognitive Science COGS 108 Undergraduate YEAR 3 DATA Science /
Cognitive Science
All answers with rationales
,Table of Contents
Content Area Questions Key Topics
DATA Manipulation WITH 1-25 Dataset, Study, Researcher, Distribution, Hours
Pandas
DATA Visualization AND 26-50 Appropriate, Interval, Confidence, Linear, Regression
Exploratory DATA Analysis
Probability AND Simulation 51-75 Regression, Hours, Appropriate, Analysis, Column
Sampling AND Study Design 76-100 Appropriate, Regression, Study, Student, Hours
Hypothesis Testing AND 101-125 Column, Reports, Analysis, Values, Difference
Statistical Inference
Confidence Intervals AND 126-150 Appropriate, Hours, Dataset, Wants, Regression
Effect SIZE
TOTAL 150 All questions include answers and detailed rationales
,Section A - DATA Manipulation WITH Pandas
Q1.
A cognitive-science dataset records reaction time (ms) for a flanker task, but 4% of trials
show values of 0 ms. What is the most defensible first step before analysis?
A. Impute the 0 ms values with the column B. Investigate whether 0 ms represents
mean so no data are lost. equipment failure or a coding artifact, then
treat as missing.
C. Winsorize the 0 ms values to the 5th D. Exclude all participants who have any 0
percentile of the distribution. ms trial to preserve data integrity.
Correct: B - Investigate whether 0 ms represents equipment failure or a coding artifact,
then treat as missing.
Rationale:Physiologically impossible reaction times signal a data-quality problem, so the
analyst must first diagnose the source (device error, coding artifact) before deciding how to
handle it. Mean imputation (A) fabricates plausible values from impossible ones, winsorizing
(C) preserves an artifact as if it were real, and dropping entire participants (D) discards valid
data and biases the sample.
Why the other answers are wrong:
A. Imputing impossible values with the mean invents data and distorts the distribution.
C. Winsorizing retains an artifact as if it were a legitimate extreme observation.
D. Dropping all participants with any artifact over-excludes and biases the sample.
Reference: COGS 108 Assignment 3, Data Cleaning module; Wickham & Grolemund, R for Data
Science (2nd Ed.), Ch. 5
Q2.
A researcher wants to test whether mean accuracy differs between two independent
groups, but the accuracy distribution is heavily left-skewed with many ceiling scores.
Which test is most appropriate?
A. Independent-samples t-test, because it is B. Paired-samples t-test, because it controls
robust to skew with n > 30. for individual differences.
C. Permutation test on the difference in D. Chi-square test of independence on the
means, because it makes no normality raw accuracy scores.
assumption.
Correct: C - Permutation test on the difference in means, because it makes no normality
assumption.
Page 3
, Section A - DATA Manipulation WITH Pandas
Rationale: A permutation test compares the observed mean difference to a null distribution
generated by reshuffling group labels, requiring no distributional assumptions-ideal for
skewed, bounded data. The t-test (A) still assumes approximately normal sampling
distributions, the paired test (B) is for within-subject designs, and chi-square (D) is for
categorical counts, not continuous accuracy.
Why the other answers are wrong:
A. The t-test's robustness does not fully rescue heavily skewed, bounded accuracy data.
B. A paired test applies to repeated measures, not two independent groups.
D. Chi-square requires categorical counts, not continuous accuracy scores.
Reference: COGS 108 Assignment 3, Hypothesis Testing module; Ernst, Permutation Methods (2004)
Q3.
A scatterplot of study hours vs. exam score shows a strong positive relationship, but one
participant reported 60 hours of study in a 24-hour period. How should this point be
handled?
A. Remove it immediately, since outliers B. Keep it, because removing data points is
always distort regression estimates. never scientifically justified.
C. Verify the value against source records D. Recode it to the next-highest observed
and analyze results with and without it. value to reduce its leverage.
Correct: C - Verify the value against source records and analyze results with and without
it.
Rationale:An implausible value requires verification first; then a sensitivity analysis (with vs.
without) shows whether conclusions depend on that point. Automatic removal (A) may discard
valid extremes, keeping blindly (B) risks leverage-driven distortion, and recoding (D) silently
fabricates a value.
Why the other answers are wrong:
A. Not all outliers are errors; automatic removal can bias results.
B. Outliers that are data-entry errors should not be kept uncritically.
D. Recoding to the next-highest value fabricates data and hides the issue.
Reference: COGS 108 Assignment 3, Exploratory Data Analysis module; Tukey, Exploratory Data
Analysis (1977)
Q4.
In a bootstrap analysis of median response time, the 95% confidence interval is [412, 489]
ms. Which interpretation is correct?
A. 95% of individual response times fall B. There is a 95% probability the true
between 412 and 489 ms. population median lies in this interval.
Page 4
ANALYSIS | 2026 UPDATE WITH
COMPLETE SOLUTIONS.
150 Questions with Answers and Detailed Rationales
100 PERCENT GUARANTEED PASS
INSTANT DOWNLOAD ANSWERS INCLUDED
IMPORTANCE OF THIS DOCUMENT
This comprehensive examination preparation guide has been meticulously developed to help you succeed in the
COGS 108 ASSIGNMENT 3: DATA ANALYSIS | 2026 UPDATE WITH COMPLETE SOLUTIONS.. It contains 150
carefully selected questions that reflect the most current exam content and testing strategies. Each question is
accompanied by a correct answer and a detailed rationale that explains the underlying pathophysiology,
pharmacology, or clinical reasoning.
Self-Assessment – Test your knowledge and Exam Preparation – Familiarize yourself with the
identify areas requiring further question format and content
study areas
Concept Reinforcement – Deepen your Confidence Building – Develop test-taking
understanding through strategies and reduce
evidence-based exam anxiety
rationales
Time Management – Practice answering
questions under simulated
exam conditions
Review Summary 150 Questions
Foundations - Application - COGS 108 Assignment 3 DATA Analysis 2026 Update WITH Complete
Solutions DATA Analysis FOR Cognitive Science COGS 108 Undergraduate YEAR 3 DATA Science /
Cognitive Science
All answers with rationales
,Table of Contents
Content Area Questions Key Topics
DATA Manipulation WITH 1-25 Dataset, Study, Researcher, Distribution, Hours
Pandas
DATA Visualization AND 26-50 Appropriate, Interval, Confidence, Linear, Regression
Exploratory DATA Analysis
Probability AND Simulation 51-75 Regression, Hours, Appropriate, Analysis, Column
Sampling AND Study Design 76-100 Appropriate, Regression, Study, Student, Hours
Hypothesis Testing AND 101-125 Column, Reports, Analysis, Values, Difference
Statistical Inference
Confidence Intervals AND 126-150 Appropriate, Hours, Dataset, Wants, Regression
Effect SIZE
TOTAL 150 All questions include answers and detailed rationales
,Section A - DATA Manipulation WITH Pandas
Q1.
A cognitive-science dataset records reaction time (ms) for a flanker task, but 4% of trials
show values of 0 ms. What is the most defensible first step before analysis?
A. Impute the 0 ms values with the column B. Investigate whether 0 ms represents
mean so no data are lost. equipment failure or a coding artifact, then
treat as missing.
C. Winsorize the 0 ms values to the 5th D. Exclude all participants who have any 0
percentile of the distribution. ms trial to preserve data integrity.
Correct: B - Investigate whether 0 ms represents equipment failure or a coding artifact,
then treat as missing.
Rationale:Physiologically impossible reaction times signal a data-quality problem, so the
analyst must first diagnose the source (device error, coding artifact) before deciding how to
handle it. Mean imputation (A) fabricates plausible values from impossible ones, winsorizing
(C) preserves an artifact as if it were real, and dropping entire participants (D) discards valid
data and biases the sample.
Why the other answers are wrong:
A. Imputing impossible values with the mean invents data and distorts the distribution.
C. Winsorizing retains an artifact as if it were a legitimate extreme observation.
D. Dropping all participants with any artifact over-excludes and biases the sample.
Reference: COGS 108 Assignment 3, Data Cleaning module; Wickham & Grolemund, R for Data
Science (2nd Ed.), Ch. 5
Q2.
A researcher wants to test whether mean accuracy differs between two independent
groups, but the accuracy distribution is heavily left-skewed with many ceiling scores.
Which test is most appropriate?
A. Independent-samples t-test, because it is B. Paired-samples t-test, because it controls
robust to skew with n > 30. for individual differences.
C. Permutation test on the difference in D. Chi-square test of independence on the
means, because it makes no normality raw accuracy scores.
assumption.
Correct: C - Permutation test on the difference in means, because it makes no normality
assumption.
Page 3
, Section A - DATA Manipulation WITH Pandas
Rationale: A permutation test compares the observed mean difference to a null distribution
generated by reshuffling group labels, requiring no distributional assumptions-ideal for
skewed, bounded data. The t-test (A) still assumes approximately normal sampling
distributions, the paired test (B) is for within-subject designs, and chi-square (D) is for
categorical counts, not continuous accuracy.
Why the other answers are wrong:
A. The t-test's robustness does not fully rescue heavily skewed, bounded accuracy data.
B. A paired test applies to repeated measures, not two independent groups.
D. Chi-square requires categorical counts, not continuous accuracy scores.
Reference: COGS 108 Assignment 3, Hypothesis Testing module; Ernst, Permutation Methods (2004)
Q3.
A scatterplot of study hours vs. exam score shows a strong positive relationship, but one
participant reported 60 hours of study in a 24-hour period. How should this point be
handled?
A. Remove it immediately, since outliers B. Keep it, because removing data points is
always distort regression estimates. never scientifically justified.
C. Verify the value against source records D. Recode it to the next-highest observed
and analyze results with and without it. value to reduce its leverage.
Correct: C - Verify the value against source records and analyze results with and without
it.
Rationale:An implausible value requires verification first; then a sensitivity analysis (with vs.
without) shows whether conclusions depend on that point. Automatic removal (A) may discard
valid extremes, keeping blindly (B) risks leverage-driven distortion, and recoding (D) silently
fabricates a value.
Why the other answers are wrong:
A. Not all outliers are errors; automatic removal can bias results.
B. Outliers that are data-entry errors should not be kept uncritically.
D. Recoding to the next-highest value fabricates data and hides the issue.
Reference: COGS 108 Assignment 3, Exploratory Data Analysis module; Tukey, Exploratory Data
Analysis (1977)
Q4.
In a bootstrap analysis of median response time, the 95% confidence interval is [412, 489]
ms. Which interpretation is correct?
A. 95% of individual response times fall B. There is a 95% probability the true
between 412 and 489 ms. population median lies in this interval.
Page 4