Introduction
Statistics
Main ideas:
What are statistical methods?
● collection, presentation, description, analyses and interpretation of data
Why is collecting data important? ->
● The research questions lays the foundation for identifying variables, relationships to test and the
statistical tools needed to produce meaningful results. Therefore, if the research question is vaguely
or wrongly worded this can heavily impact the rest of the research process
What makes a good research question? ->
● Theory driven, specific (using a specific population or variable), testable (can be answered with
data) and meaningful (contributes to psychological knowledge)
What are descriptive statistics? ->
● Techniques for summarising, classifying and describing numerical data
Populations and Samples ->
● Populations are a complete set of events of interest whilst samples are a subset of the population.
We actually thoroughly observe the sample and make estimates about the population based on
them.
Sample variability ->
● There’s always going to be some variation within and between samples due to chance
Sampling bias ->
● This happens when the sample we collect systematically favours certain results so the sample is
less representative than ideal e.g collecting only basketball players’ heights in order to work out the
average height of ALL students
Inferential statistics ->
● These are procedures for making generalisations about a population based on samples. They tell
us how to make generalisations (how likely is it that our sample is representative) and how good
they are (have the findings occurred by chance or are they meaningful?)
Probability ->
● How confident can we be that something didn’t occur by chance? Probability describes the
likelihood of observing certain sample data
Null hypotheses statistical testing
● Is a hypothesis about a population, is the sample data strong enough to support a claim or has it
occurred by chance?
● Research hypotheses (H1 or HA) -> something you want to test whether that be a
difference/inequality in relationships or across groups
● Directional hypotheses require one tailed tests whilst non directional hypotheses require two tailed
tests
, ● Null hypotheses (H0) -> No difference or effect or relationship, so nothing is happening. It can be a
starting point before any information is collected to counter and a benchmark to compare the
outcome of our study if it is due to factors other than chance. So, start with the assumption that it’s
true and then find evidence against it. You can either reject or fail to reject it.
Extra Notes:
● Data predicts which statistical test to use -> for example comparison data uses t tests and ANOVA,
relationship questions use correlation and regression analysis and frequency questions require Chi
Squared tests
● Research questions need to be specific and testable e.g. ‘Sleep helps people remember things’ ->
Do students who get at least 8 hours of sleep a night before a word test recall more words than
students who get 4 hours of sleep?’
● In regards to sampling, random sampling gives each population member an equal chance of
selection. The exact population and sample depends on research goals
● Probability is measured with 0 to 1 where 1 is certain and 0 means the event is certain not to occur
● NEVER ACCEPT a null hypothesis. The data we collect could reasonably come from a world where
the hypothesis is true and just because the null appears to be true doesn't mean it is. Hypotheses
testing also cannot prove whether two groups are the same and would be the same as accepting
the null hypotheses. This would need different methods of testing including the Bayesian tests
(forming a set of defined hypotheses and testing them) or confidence intervals (parameters that
could plausibly produce the data are all close to H0)
Summary:
● All the way from the beginning, the research question, data and stats depend on this in order for the
whole scientific method to work. The research question is predictive of the type of statistical tools to
use. It’s also ultimately important for working out whether results are truly valid i.e. show a
meaningful pattern or are by chance.
Howell, D. C. (2017). Fundamental Statistics for the Behavioral Sciences (9th ed.). Belmont, CA: Duxbury
Press. Chapter 8.
What are sample statistics, population statistics and conditional probability?
Samples statistics refers to the mean and standard deviation obtained from a sample, population statistics
refers to the mean and standard deviations obtained from a population and conditional probability is the
probability of an event occurring given another event has occurred.
What are sampling distributions and sampling error ?
Sampling distribution refers to the distribution of the sample statistics from multiple samples in a
population, sampling error refers to the way that these statistics vary from one another based on random
variability rather than by mistake.
What is the standard error of the mean?
Standard error is just the standard deviation of the sampling distribution
How does all this relate to hypothesis testing?
The main point of sampling distributions is to test a hypothesis. An important aspect of testing a hypothesis
is to know the probability of obtaining said predicted score. We want to look at a research hypothesis (H1),
we set up a null hypothesis (H0), obtain a random sample under the assumption that H0 is true, we
calculate the probability of the mean to be at least as large as the actual sample mean, based on that
decided to either fail or not fail H0