Chapter 2
Data Collection
🎯 LEARNING OBJECTIVES
By the end of this chapter, you should understand how data is collected and how
different study designs affect the results. You will learn to distinguish between
populations and samples, identify sampling techniques, recognize observational
studies and experiments, identify different types of variables, and recognize common
sources of bias.
Now that we understand what statistics does, we need to talk about something incredibly
important: Where does the data come from?
You can have the fanciest graph in the world and the most impressive-looking spreadsheet… but if
the data was collected poorly, your conclusion can still be wrong.
POPULATION VS. SAMPLE
Let's imagine we want to know how college students feel about having classes on Friday. Do we
need to ask every single college student in the country? Obviously not. Instead, I could choose a
smaller group of students and ask them.
The population (denoted by N) is the entire group we are interested in studying. If I'm studying
college students at a particular school, the population might be all students at that school. If I'm
studying American college students, the population might be all college students in the United
States.
The sample (denoted by n) is the smaller group we actually collect data from. So if there are
20,000 students in our population but we survey 500 of them, those 500 students are our
sample.
WHY DO WE USE SAMPLES?
Sometimes, we are able to survey an entire population. But usually it's expensive, time-consuming,
or impossible. A sample allows us to gather information from a smaller group and use that
information to learn about a larger population.
pg. 1
, Elementary Statistics
SAMPLING TECHNIQUES
Different sampling methods can produce different results.
● Simple random sample (most ideal): every member of the population has an equal
chance of being selected.
Ex: Putting 100 names into a hat and choosing a random name.
● Systematic sample (ideal): select members using a regular interval, such as every 10th
person.
Ex: A retail company wants to survey every 5th customer.
● Stratified sample (ideal): divide the population into strata (groups) and randomly
sample from each group.
Ex: Dividing students into groups—freshmen, sophomores, juniors, and seniors—and
then taking a simple random sample of each group.
● Cluster sample (ideal): divide the population into groups, randomly select groups, and
survey everyone in selected groups.
Ex: A school has 30 classes; we randomly choose 5 classes, and we survey every student
in those 5 classes.
● Convenience sample (less ideal): choose people who are easiest to reach. This sampling
technique may introduce bias and not represent the population accurately.
Ex: You’re standing in a mall and surveying the first couple of people you see.
● Voluntary response sample (less ideal): people choose whether to participate. People
with negative opinions are more likely to volunteer.
Ex: Online polls and surveys
💡 KEY IDEA
It’s easy to confuse stratified with cluster sampling. Stratified sampling means to
randomly take some people from each group. Cluster sampling means to choose some
groups and take everyone within those groups.
OBSERVATIONAL STUDIES VS. EXPERIMENTS
Suppose researchers want to know whether students who drink coffee tend to study more
hours. They could simply observe students and record whether they drink coffee or how much
they study. They're just observing what naturally happens—that's an observational study.
Now imagine researchers randomly assign students to drink either coffee or a caffeine-free
drink and then compare their study performance. Now researchers are actively assigning
something. That's an experiment.
VARIABLES
pg. 2