Ginett e Lafi t – 2023-2024
Inhoudstafel
1
, LES 1 – The data analysis workflow: an illustration
Boxplot: shows distribution of the dependent variable
Median (not average!!)
1st and 2th quartile
Max score, min score
Outliers (uitschieters) = a value that
is more than 1.5 times the
interquartile range (IQR) below the
first quartile or above the third
quartile
Symmetric or skewed
The larger the 2 tails, the bigger the standard deviation (= bad) best = when both tails are
very short
Terminology
H0 = null hypothesis
H1 = alternative hypothesis
α = significant level
p = probability of exceedance (overschrijdingskans)
Yij = score of person i in group j of the dependent variable, with j equal to 1 or 2
nj = number of observations in group j = sample size
Ȳj = sample mean (steekproefgemiddelde) in group j (for total group)
µ = mean (for total population)
σ² = variance
σ = standard deviation
df = degrees of freedom
Working model
1) Formulate model and hypothesis
2) Select the significant level α
3) Test statistics: choice and calculation
2
,4) Derive sample distribution and determine the p-value
df = degrees of freedom (nodig voor tabel)
The larger the p-value, the less evidence to reject the null hypothesis
The smaller the p-value, the more evidence to reject the null hypothesis
P-value: table D
5) Decision
Is it a one-sided or two-sided test?
Compare p-value with α: when p < α => reject H0
6) Determine effect size
Confidence Interval (CI) (betrouwbaarheidsinterval – BI) for difference between two
averages
CI = 100*(1- α) %
CI is always by two-sided tests
CI of 99% means that if we conduct this experiment over and over again, the score
will be in the interval 99% of the times
Ȳ2 - Ȳ1 = center of CI = (lower limit of CI + upper limit of CI)/2
t*(n1+n2-2) x SE(Ȳ2 - Ȳ1) = m = upper limit of CI - center of CI
7) Interpretation
Formulate conclusion
o Answer research questions
o Use substantive terminology
Summarize results using plots
o Not necessary when only 2 groups compared
o Useful with multiple groups
State findings’ limitations of the study
o Randomization: causal inference possible
o Random samples: assumption questionable; strictly no inference to population
possible
3
, LES 2 – Analysis of the variance with one factor (ANOVA)
When there are 1 or 2 groups => t-test is also possible
When there are more than 2 groups => use ANOVA
Terminology
ANOVA = ANalysis Of VAriance
IV = independent variable (most of the time: amount of groups)
DV = dependent variable (most of the time: number of items)
STOS = modality-specific speech and language development disorder
Yij = score of person i in group j of the DV
nj = number of observations in group j = sample size
N = total number of observations
a = number of groups
Ȳj = sample mean in group j
Ȳ = µ = overall sample mean = population mean
S’ = standard deviation
df = degrees of freedom
CI / d / R² / ω² = effectgrootte
MSeffect = tussengroepsvariabiliteit
MSerror|full = binnengroepsvariabiliteit
Working model
1. Preparation
N?
a?
ni?
2. Exploratory data analysis
Ȳj for all groups?
S’i for all groups?
Ȳ?
S’?
Formulate presumption on research question
View assumptions and conditions
2 models:
3 assumptions
Normality
Independence of the observations of the groups
Equal population standard deviations within groups (=
homoscedasticity)
1) Violations of assumptions
Robustness against violations of assumptions
A statistical method is robust against a certain violation, if the
statistical inferences are valid, even if the assumption is violated
A valid inference means that the stated uncertainty (e.g., p-value) is
the actual level of uncertainty
4