Week 1
● What is Statistics?
○ The science (the math) and art of identifying meaningful patterns in data.
- Analyzing basic data
○ Formally: Mathematical methods for collecting, organizing, and analyzing data.
○ "Statistic" can also mean a summary of sample data.
● Importance of Statistics
○ Used in science, policy, and daily decision-making (ex. Calculating gpa (mean))
○ Helps in understanding the world and making fact-based decisions.
○ Supports "data-driven decision-making."
● Three Key Aspects of Statistics
○ Study Design (Data Collection)
○ Descriptive Statistics – Summarizing data within a sample.
- Answers the “what”
○ Inferential Statistics – Making generalizations from a sample to a population.
● Descriptives vs. Inference
○ Descriptive Statistics: Summarizing and describing a dataset
○ Inferential Statistics: Making generalizations from a sample to a population.
- Sample statistics are likely close to true population parameters
(representative)
- Inference helps quantify the potential sampling error.
○ Sample Summaries: Statistics (from data)
- Numbers (ex. Average age)
○ Population Summaries: Population Parameters (true values in the
population).
- People/sample (ex. Western students)
● Sampling
○ The process of selecting cases for a study to understand a population.
○ Population: The entire group being studied.
- Ex. western students as a whole
○ Sample: A subset of the population used for analysis.
- Ex. 50 western students (you want to represent a group that represents
all of western not only 1st year eng students)
○ The goal is to obtain a representative sample.
- We achieve this by having a random sample
● Random vs. Non-Random Sampling
○ Random Sampling: Each case has a known probability of being selected.
■ Simple Random Sample (SRS): Each case has an equal probability.
■ Other methods: Clustered Sampling, Stratified Sampling.
○ Non-Random Sampling: Convenience, snowball, expert selection, etc.
, ■ May be biased (systematically different from the population).
● Sampling Error
○ Every sample varies naturally.
- When I select 100 people it will be different than your 100 selected people
○ Sampling error: normal expected, inevitable variation across samples (error
does not mean mistake)
- Ex. If Jim, Nik, and Dan have 5 apples as their samples each, every
sample will differ, some will have bigger apples, some will be more red.
(that’s the error)
● Two Main Types of Variables
1. Categorical (Qualitative or Discrete)
■ Nominal: Labels without a meaningful order (e.g., gender, province, field
of study).
- No math operations are possible
■ Ordinal: Categories with a ranked order, but without meaningful
numerical differences
- (e.g., education levels, survey ratings).
2. Continuous (Quantitative)
■ Interval/Ratio: Numerical values where math operations are meaningful
- (e.g., age in years, temperature, income).
- NOTE* Ordinal variables with 4+ categories can often be analyzed
as continuous.
● Structure of Datasets
○ A dataset is an organized collection of data.
○ Data points = individual pieces of data.
- Usually numerical.
○ Variables (columns) represent characteristics.
○ Cases (rows) represent individual units (people, companies, etc.).
● What is Statistics?
○ The science (the math) and art of identifying meaningful patterns in data.
- Analyzing basic data
○ Formally: Mathematical methods for collecting, organizing, and analyzing data.
○ "Statistic" can also mean a summary of sample data.
● Importance of Statistics
○ Used in science, policy, and daily decision-making (ex. Calculating gpa (mean))
○ Helps in understanding the world and making fact-based decisions.
○ Supports "data-driven decision-making."
● Three Key Aspects of Statistics
○ Study Design (Data Collection)
○ Descriptive Statistics – Summarizing data within a sample.
- Answers the “what”
○ Inferential Statistics – Making generalizations from a sample to a population.
● Descriptives vs. Inference
○ Descriptive Statistics: Summarizing and describing a dataset
○ Inferential Statistics: Making generalizations from a sample to a population.
- Sample statistics are likely close to true population parameters
(representative)
- Inference helps quantify the potential sampling error.
○ Sample Summaries: Statistics (from data)
- Numbers (ex. Average age)
○ Population Summaries: Population Parameters (true values in the
population).
- People/sample (ex. Western students)
● Sampling
○ The process of selecting cases for a study to understand a population.
○ Population: The entire group being studied.
- Ex. western students as a whole
○ Sample: A subset of the population used for analysis.
- Ex. 50 western students (you want to represent a group that represents
all of western not only 1st year eng students)
○ The goal is to obtain a representative sample.
- We achieve this by having a random sample
● Random vs. Non-Random Sampling
○ Random Sampling: Each case has a known probability of being selected.
■ Simple Random Sample (SRS): Each case has an equal probability.
■ Other methods: Clustered Sampling, Stratified Sampling.
○ Non-Random Sampling: Convenience, snowball, expert selection, etc.
, ■ May be biased (systematically different from the population).
● Sampling Error
○ Every sample varies naturally.
- When I select 100 people it will be different than your 100 selected people
○ Sampling error: normal expected, inevitable variation across samples (error
does not mean mistake)
- Ex. If Jim, Nik, and Dan have 5 apples as their samples each, every
sample will differ, some will have bigger apples, some will be more red.
(that’s the error)
● Two Main Types of Variables
1. Categorical (Qualitative or Discrete)
■ Nominal: Labels without a meaningful order (e.g., gender, province, field
of study).
- No math operations are possible
■ Ordinal: Categories with a ranked order, but without meaningful
numerical differences
- (e.g., education levels, survey ratings).
2. Continuous (Quantitative)
■ Interval/Ratio: Numerical values where math operations are meaningful
- (e.g., age in years, temperature, income).
- NOTE* Ordinal variables with 4+ categories can often be analyzed
as continuous.
● Structure of Datasets
○ A dataset is an organized collection of data.
○ Data points = individual pieces of data.
- Usually numerical.
○ Variables (columns) represent characteristics.
○ Cases (rows) represent individual units (people, companies, etc.).