VERSCHIL MET VORIGE LES
- Daar hadden we 2 groepen
- Nu hebben we meer dan 2 groepen : μ1=μ2=μ3?
ONE-WAY ANOVA = EENRICHTINGS-ANOVA
- = Er is 1 factor met meer dan 2 niveaus (want bij 2 niveaus kan je ook t-test gebruiken)
- = Eigenlijk verzameling v onafhankelijk groepen t-test
GOAL OF THE LESSON
- Learn how to use one-way analysis of variance (ANOVA) to test if there are differences in the population’s means
between two or more groups
o !!! Anova gaat over verschil in gemiddelde tss 2/meer groepen
HC 2: ONE WAY ANOVA PART 1
DATA ANALYSIS WORKFLOW
STEP 1: STATE THE RESEARCH QUESTION. EXAMPLE
Goal is to test whether 2/more groups have different population means
- We don’t care of the 4 are different we want to know if at least 1 is different
EXAMPLE: response to eye color
- Investigate differences in attitudes towards a brand expressed by 4 groups
o ‘Blue’ = Model with blue eyes
o ‘Brown’ = Model with brown eyes
o ‘Green’ = Model with green eyes
o ‘Down’ = Model’s eye color cannot be seen
- Each group saw the same advertisement, except for the manipulated aspect:
o The model’s eye color
- Dependent variabel = attitude toward brand
EXPLORATIEVE DATA
- Voordat je inferentiële statistiek uitvoert (hypothesetoetsen, p-waarden berekenen, enz.)
o Moet je altijd eerst naar je data kijken = SP-gegevens
- Hulpmiddel
o Visualisatie kan helpen om:
Codeerfouten te ontdekken
Extreme waarden te vinden
Interessante patronen te zien
- Hoe?
o Gegevensmatrix ; historgrammen, boxplot, etc.
- Intercolural trauma test
o = Patroon in de exploratieve (SP) data zijn zo duidelijk dat er geen inferentiële statistiek nodig is
POPULATIEGEMIDDELDEN zijn parameters.
- We kennen ze niet → we schatten ze met een steekproef.
- Geschatte parameter wordt aangegeven met een hoedje: ^μ
,STEP 2: STATISTICAL HYPOTHESIS
HYPOTHESES IN ONE-WAY ANOVA:
- Use to investigate if there are differences in the population mean between 2/ more groups
- It does not specify which means are different or the direction of the difference.
- Null hypothesis H0:
o ALL GROUP MEANS ARE EQUAL (with “a” being the number of groups) = BALANCED model
o H0 : µ1 = µ2 = ... = µa
- Alternative hypothesis (note H1 = HA)
o H1 : At least 1 population mean is different from the others !! let op: w in woorden geschreven dus
o Significance level ( = Type I error rate): α In our example: α= 0.05
STATISTICAL MODEL = GENERATIEVE MODELLEN (= specifi eren exact hoe scores op criteria-var. w gegenereerd )
- Y ij = De score van persoon i in conditie j Voorbeeld
i = persoon binnen groep - Y 11 = score van persoon
j = groep/conditie vd DV 1 in groep 1
- Y 23 = score van persoon
- nj = aantal observaties in groep j
- a = aantal groepen
- Ý j = Het groepsgemiddelde van groep j
- Y = Globaal gemiddelde
a
- N= ∑ n j = n1+ n2+ n3+ … n a = Totaal aantal observaties
j=1
- S ' 2Y j
= SP variantie in groep j
2
- S 'Y = Totale SP variantie
IN OUR EXAMPLE = PARTICIPANT DATA SET
DV = DEPENDENT VARIABLE
- Y ij = score of participant i in group j of DV attitudes towards the brand, with j=1 ,2 , 3 , or 4 , and i=1 , … , n j
- a=4 :
o There are 4 groups:
‘Blue’ = Model with blue eyes,
‘Brown’ = Model with brown eyes,
‘Green’ = Model with green eyes,
‘Down’ = Model’s eye color cannot be seen
- n1=67 is the number of participants in group 1 (‘Blue’),
n2 =37 is the number of participants in group 2 (‘Brown’),
n3 =77 is the number of participants in group 3 (‘Green’), and
n 4=41 is the number of participants in group 4 (‘Down’)
- N=222 is the total number of participants
,- Ý 1=3 . 194 is the sample mean in group 1 (‘Blue’),
Ý 2=3 . 724 is the sample mean in group 2 (‘Brown’),
Ý 3=3 . 860 is the sample mean in group 3 (‘Green’), and
Ý 4 =3 .107 is the sample mean in group 4 (‘Down’)
- Ý =3. 497 is the overall sample mean
, VERSCHIL TUSSEN EEN GEBALANCEERD EN ONGEBALANCEERD DESIGN .
- Gebalanceerd als elke groep evenveel deelnemers heeft.
o Bijvoorbeeld: placebo: 60 - ADM: 60 - CT: 60
- Ongebalanceerd als de groepsgroottes verschillen.
o Bijvoorbeeld: placebo: 60 - CT: 60 - ADM: 120
- Waarom is dat belangrijk?
o Omdat sommige formules eenvoudiger zijn in het gebalanceerde geval.
o Maar ANOVA WERKT OOK bij ongebalanceerde ontwerpen.
WAT DOEN WE NU EIGENLIJK?
- Anova vergelijkt eigenlijk 2 modellen
O Gereduceerde model stelt: er is 1 gemeenschappelijk gemiddelde voor ALLE groepen H0
O Volledig model stelt: elke groep heeft eigen gemiddelde H1
STATISTICAL MODEL: ASSUMING THE NULL HYPOTHESIS IS TRUE
Reduced model = verwijzen naar statistisch model v H0 = WAAR
Statistisch model
- Y ij = μ + ε ij met ε ij N (0; σ 2) voor j = 1, …, a EN i = 1, … n j
Griekse letter (ε ¿ = error/voorspellingsfout / residu
- Volgt normaal verdeling
- = Individueel verschil tss waarde Y ij & μ
(want μ (tot) = gem. v groep)
- = Willekeurige afwijking / random variabele
o Geven aan of iemand boven of onder het gemiddelde ligt
+ als boven gemiddelde
-als onder gemiddelde
o Absolute waarden van epsilon geeft weer hoe ver iemand van het gemiddelde zit
Groot: ver van gemiddelde
Klein: dicht bij gemiddelde
Dit statistic model = linear model
We weten dat
- Als error normaal verdeeld is
dan is score v Y ij ook normaal verdeel
Schematisch plot v data als H0 waar is
- Puntjes = data verzameling
- Onder H0
o Mean/ gemiddelde
= Exact gelijk
= Gelijk aan μ v die