Introduction to Data Analysis
Population vs Sample- The population is everyone/possible case you are considering while the
sample is a collection of individuals or items under consideration in a study. The sample is a
smaller part of the population. The sample is meant to be representative enough of the
population so you can use it to make inferences about a population
Descriptive vs Inferential- Descriptive just describes the data but Inferential is regarding coming
to a conclusion
PICOT (Useful when designing or evaluating a study)
Data needs to be quality. It cannot be anecdotal or based on belief. Several data points are
needed before an inference can be made
Statistics- Evidence-based practice (data-based) of collecting/gathering, analyzing and
evaluating/interpreting numerical data
The target population is all the people you are interested in and the study design has criteria so
that remove a bunch of people. All the people who are eligible are the study population
according to the study criteria. From the eligible people, not everyone can be enrolled because
they did not consent to be a part of the study. This then becomes the study sample. The study
sample conclusions cannot be generalizable to the study population that met all the study
design criteria, not the population at large. This is after the inclusion and exclusion criteria are
met (characteristics and features of the participants that make they included or excluded in a
study)
Consort Flow Diagram: This is a diagram that shows how you went from the general population
to the final study sample. It shows all the exclusion and inclusion criteria for the study.
Tables provide descriptive characteristics calculated from sample data. Data points usually have
numerical counts and percentages/proportions
Randomization of participants often works to make the different treatment groups comparable
(age, race, and gender to each other so that no group has an advantage) over the other
Parameter vs Statistic: Parameter that describes a population (fixed, usually unknown because
you need the whole population to get the number). Statistics are numbers that describe sample
populations (they vary from one sample to the next). There is something called sampling
variability. The statistic can be used to estimate the parameter
EOR- Exposure, Outcome, Relationship. This is a principle used in epidemiology
RCT- Randomized clinical trial
When comparing exposed group to non exposed group, you have to make sure that the
exposed group get the treatment at baseline e.g, a heart transplant recipient, gets their heart
transplant immediately they get the diagnosis
When writing the methods section, you need to be able to determine and describe the design of
the study. Various Study Design:
- Descriptive vs Analytic. In analytic studies, there is a hypothesis that needs to be tested
- Observational vs Interventional
- Nonrandomized vs Randomized: Randomization makes the study groups more
comparable to one another
- Retrospective vs Prospective (looking at the past, looking at the future). Prospective is
preferred to retrospective
, - Cross-sectional vs Longitudinal: Cross-section (data from various groups gathered at the
same point in time). Longitudinal (same group over a period of time)
- Controlled(has control group: placebo or standard of care)vs Uncontrolled
- Blinded vs Open Label: Blinding can be single (pt does not know the study group they
are in) or double-blinded (neither researcher nor participant knows)
- Case-Control vs Cohort: Cohort (start with a group of people, measuring their
exposures, and measure their outcomes over time, measuring relative risk). In case-
control (comparing those that have disease to those who do not, then you measure their
exposures, retrospective in nature, calculating odds ratio)
Study the design of data in Table 1 of the Introductory slides
This class will focus on unadjusted analysis on cross-sectional studies
ANS: The children are the target population to look at. All mothers should be HIV positive and
we can look at the childrens’ outcomes, whether they get HIV or not through vertical
transmission
PICOT- Population, Intervention, Comparison, Outcome, and Time (For randomized studies,
baseline is typically the date of randomization. For nonrandomization, baseline date has to be
clearly specified)
. This is a method of creating research questions, especially clinical questions in healthcare
All-cause mortality- Death for any cause or reason within the study sample
Competing risk- events that prevent the primary outcome in a study from taking place. Death is
the most common competing risk because it prevents many healthcare outcome we wish to
study such as hospital readmission, disease occurence etc.
“Any conclusions drawn from subgroup hypotheses not explicitly stated in the protocol should
be given much less credibility than those from hypotheses stated a priori. Retrospective
subgroup analysis should serve primarily to generate new hypotheses for subsequent testing”.
Note that the sample size for a subgroup from a given study is usually small and so hypotheses
for just that small group are not as reliable. Results are not conclusive and should merely be
used to generate hypotheses for further analyses
,A case-control study is the standard disease often used to evaluate the risk factors associated
with disease manifestation.
Cohort studies are preferred to case-control studies because there is removed bias but case-
control studies are faster and cheaper
Randomized controlled trials are the standard. Prospectively assigning an individual or a group
of individuals to either a study/intervention group or a control group. The groups of patients
should be comparable in every respect except treatment.
Randomization usually accounts for both the recognized and unrecognized risk factors to
ensure an even distribution of those risk factors in each study group
Types: Fixed allocation (same probability of being assigned to the treatment group through out
study) and adaptive randomization (probability of being assigned the treatment changes as the
study progresses)
There are three types of fixed allocation randomization:simple, blocked, and stratified.
The allocation ratio may be equal or unequal.
Simplest trial is a single arm trial: People with disease at baseline are given a treatment. The
participants after the trial are compared to how they were before the trial. Participants at
baseline function as the control. They don’t determine treatment the true efficacy (though it can
show preliminary info about efficacy) but can be used to determine safety information about
treatment
Participants/patients from previous research efforts can be used as controls on a current study.
Those are historical controls
In a randomized trial, there is a principle called Intention to Treat (ITT): The idea is that you stay
in the group you were assigned to regardless of whether you adhered to the treatment or did not
, receive the assigned treatment of your group. This is the prefered statistics method to avoid
introducing bias in post-randomization exclusions.
Randomization guarantees that statistical tests will have valid false positive error rates. WHAT?? Don’t understand
Figure: Study design organised according to level of evidence
Parallel group: Treatment and control group where participants are randomly assigned
Case control studies are retrospective
Studies that do not use randomized assignment are generally referred to simply as ‘clinical
trials,’ with no mention of randomization
Observational studies can reveal an association, whereas RCTs can help establish causation
Data Visualization
Types of variables: Numerical variables (continuous and discrete [countable]).
Categorical/qualitative (high/low, male/female, cancer stage 1-4) variables on the other hand
need larger sample sizes
Continuous variables require lower sample sizes because changes in outcome are more easily
detectible.
Categorical variables can be nominal (no meaning in the order (name only) e.g race, sex (F=1,
M=0) and ordinal (there is an order to these e.g mild-> moderate -> severe)
The scale and range used to draw the graphs and axes have to be considered when presenting
data
A variable is a characteristic of an individual
Individuals are the people described by a data set. They do not have to be people. They can
can also be animals, hospitals, etc.
T-tests do not work on longitudinal data
The type of variable you have affects the type of analytical tests that are done
Time to event is a common variable.g time to death
Continuous is measured while discrete are counted
Risk is a probability. It is a pure number. It does not have a unit. Rate is not a pure number as it
has units (e.g m/s for speed).
For a living person, the time to death is unknown so you can saynthe time to death is
‘censored’. The time to death is the time to the last date of observation. Time to death is a
continuous in nature but is not classed with other continuous variables because the event/death
does not always take place within the study period
Population vs Sample- The population is everyone/possible case you are considering while the
sample is a collection of individuals or items under consideration in a study. The sample is a
smaller part of the population. The sample is meant to be representative enough of the
population so you can use it to make inferences about a population
Descriptive vs Inferential- Descriptive just describes the data but Inferential is regarding coming
to a conclusion
PICOT (Useful when designing or evaluating a study)
Data needs to be quality. It cannot be anecdotal or based on belief. Several data points are
needed before an inference can be made
Statistics- Evidence-based practice (data-based) of collecting/gathering, analyzing and
evaluating/interpreting numerical data
The target population is all the people you are interested in and the study design has criteria so
that remove a bunch of people. All the people who are eligible are the study population
according to the study criteria. From the eligible people, not everyone can be enrolled because
they did not consent to be a part of the study. This then becomes the study sample. The study
sample conclusions cannot be generalizable to the study population that met all the study
design criteria, not the population at large. This is after the inclusion and exclusion criteria are
met (characteristics and features of the participants that make they included or excluded in a
study)
Consort Flow Diagram: This is a diagram that shows how you went from the general population
to the final study sample. It shows all the exclusion and inclusion criteria for the study.
Tables provide descriptive characteristics calculated from sample data. Data points usually have
numerical counts and percentages/proportions
Randomization of participants often works to make the different treatment groups comparable
(age, race, and gender to each other so that no group has an advantage) over the other
Parameter vs Statistic: Parameter that describes a population (fixed, usually unknown because
you need the whole population to get the number). Statistics are numbers that describe sample
populations (they vary from one sample to the next). There is something called sampling
variability. The statistic can be used to estimate the parameter
EOR- Exposure, Outcome, Relationship. This is a principle used in epidemiology
RCT- Randomized clinical trial
When comparing exposed group to non exposed group, you have to make sure that the
exposed group get the treatment at baseline e.g, a heart transplant recipient, gets their heart
transplant immediately they get the diagnosis
When writing the methods section, you need to be able to determine and describe the design of
the study. Various Study Design:
- Descriptive vs Analytic. In analytic studies, there is a hypothesis that needs to be tested
- Observational vs Interventional
- Nonrandomized vs Randomized: Randomization makes the study groups more
comparable to one another
- Retrospective vs Prospective (looking at the past, looking at the future). Prospective is
preferred to retrospective
, - Cross-sectional vs Longitudinal: Cross-section (data from various groups gathered at the
same point in time). Longitudinal (same group over a period of time)
- Controlled(has control group: placebo or standard of care)vs Uncontrolled
- Blinded vs Open Label: Blinding can be single (pt does not know the study group they
are in) or double-blinded (neither researcher nor participant knows)
- Case-Control vs Cohort: Cohort (start with a group of people, measuring their
exposures, and measure their outcomes over time, measuring relative risk). In case-
control (comparing those that have disease to those who do not, then you measure their
exposures, retrospective in nature, calculating odds ratio)
Study the design of data in Table 1 of the Introductory slides
This class will focus on unadjusted analysis on cross-sectional studies
ANS: The children are the target population to look at. All mothers should be HIV positive and
we can look at the childrens’ outcomes, whether they get HIV or not through vertical
transmission
PICOT- Population, Intervention, Comparison, Outcome, and Time (For randomized studies,
baseline is typically the date of randomization. For nonrandomization, baseline date has to be
clearly specified)
. This is a method of creating research questions, especially clinical questions in healthcare
All-cause mortality- Death for any cause or reason within the study sample
Competing risk- events that prevent the primary outcome in a study from taking place. Death is
the most common competing risk because it prevents many healthcare outcome we wish to
study such as hospital readmission, disease occurence etc.
“Any conclusions drawn from subgroup hypotheses not explicitly stated in the protocol should
be given much less credibility than those from hypotheses stated a priori. Retrospective
subgroup analysis should serve primarily to generate new hypotheses for subsequent testing”.
Note that the sample size for a subgroup from a given study is usually small and so hypotheses
for just that small group are not as reliable. Results are not conclusive and should merely be
used to generate hypotheses for further analyses
,A case-control study is the standard disease often used to evaluate the risk factors associated
with disease manifestation.
Cohort studies are preferred to case-control studies because there is removed bias but case-
control studies are faster and cheaper
Randomized controlled trials are the standard. Prospectively assigning an individual or a group
of individuals to either a study/intervention group or a control group. The groups of patients
should be comparable in every respect except treatment.
Randomization usually accounts for both the recognized and unrecognized risk factors to
ensure an even distribution of those risk factors in each study group
Types: Fixed allocation (same probability of being assigned to the treatment group through out
study) and adaptive randomization (probability of being assigned the treatment changes as the
study progresses)
There are three types of fixed allocation randomization:simple, blocked, and stratified.
The allocation ratio may be equal or unequal.
Simplest trial is a single arm trial: People with disease at baseline are given a treatment. The
participants after the trial are compared to how they were before the trial. Participants at
baseline function as the control. They don’t determine treatment the true efficacy (though it can
show preliminary info about efficacy) but can be used to determine safety information about
treatment
Participants/patients from previous research efforts can be used as controls on a current study.
Those are historical controls
In a randomized trial, there is a principle called Intention to Treat (ITT): The idea is that you stay
in the group you were assigned to regardless of whether you adhered to the treatment or did not
, receive the assigned treatment of your group. This is the prefered statistics method to avoid
introducing bias in post-randomization exclusions.
Randomization guarantees that statistical tests will have valid false positive error rates. WHAT?? Don’t understand
Figure: Study design organised according to level of evidence
Parallel group: Treatment and control group where participants are randomly assigned
Case control studies are retrospective
Studies that do not use randomized assignment are generally referred to simply as ‘clinical
trials,’ with no mention of randomization
Observational studies can reveal an association, whereas RCTs can help establish causation
Data Visualization
Types of variables: Numerical variables (continuous and discrete [countable]).
Categorical/qualitative (high/low, male/female, cancer stage 1-4) variables on the other hand
need larger sample sizes
Continuous variables require lower sample sizes because changes in outcome are more easily
detectible.
Categorical variables can be nominal (no meaning in the order (name only) e.g race, sex (F=1,
M=0) and ordinal (there is an order to these e.g mild-> moderate -> severe)
The scale and range used to draw the graphs and axes have to be considered when presenting
data
A variable is a characteristic of an individual
Individuals are the people described by a data set. They do not have to be people. They can
can also be animals, hospitals, etc.
T-tests do not work on longitudinal data
The type of variable you have affects the type of analytical tests that are done
Time to event is a common variable.g time to death
Continuous is measured while discrete are counted
Risk is a probability. It is a pure number. It does not have a unit. Rate is not a pure number as it
has units (e.g m/s for speed).
For a living person, the time to death is unknown so you can saynthe time to death is
‘censored’. The time to death is the time to the last date of observation. Time to death is a
continuous in nature but is not classed with other continuous variables because the event/death
does not always take place within the study period