1. Investigator A takes a random sample of 100 men age 18-24 in a community. Investigator B
takes a random sample of 1,000 such men.
Introduction
Biostatistics forms the cornerstone of evidence-based public health practice, providing the essential
tools for collecting, analysing, and interpreting data that inform health policy and clinical
decision-making (Kirkwood & Sterne, 2003). Among the most fundamental concepts in statistical
inference are sampling distributions and confidence intervals, which enable researchers to draw
conclusions about populations based on information gathered from samples (Altman, 1991).
Understanding how sample size influences these statistical measures is critical for designing robust
studies and correctly interpreting research findings.
The relationship between sample size and statistical precision has profound implications for public
health research, where resource constraints often necessitate careful consideration of how many
participants to include in a study (Rothman, Greenland & Lash, 2008). A sample that is too small
may fail to detect important health effects, while an unnecessarily large sample wastes resources and
may raise ethical concerns about exposing more participants to research procedures than necessary
(Armitage, Berry & Matthews, 2001).
This assignment examines a comparative scenario involving two investigators sampling from the
same population of men aged 18-24 in a community. Investigator A takes a random sample of 100
men, while Investigator B takes a random sample of 1,000 men. Through systematic analysis of how
sample size affects various statistical measures, this essay elucidates the fundamental principles of
sampling distributions, standard errors, confidence intervals, and the concept of the sample mean.
a. Which investigator will tend to get a bigger standard deviation (SD) for the heights of the men
in his sample? Or, can it not be determined?
Conceptual Understanding of Standard Deviation
The standard deviation is a measure of the dispersion or spread of individual values within a dataset
(Field, 2018). For the heights of men aged 18-24 in a community, the standard deviation reflects the
natural variability in height among individuals in that population (Kuzawa et al., 2010). This
variability is determined by biological factors including genetics, nutrition, and environmental
influences (Rothman, 2012).
Sample Size and Standard Deviation
It is a common misconception that larger samples produce smaller standard deviations (Bland, 2015).
In reality, the standard deviation of a sample is an estimate of the population standard deviation and
is not systematically affected by sample size (Armitage, Berry & Matthews, 2001). The formula for
the sample standard deviation is:
s (xi - x) (n - 1)
This formula divides by (n - 1), which is the degrees of freedom, to provide an unbiased estimate of
the population standard deviation (Daniel & Cross, 2018). As sample size increases, the estimate
becomes more precise, but the expected value of the standard deviation remains approximately equal
to the population standard deviation regardless of sample size (Altman, 1991).