Arzoo Chaudhry Genomics
Association Analysis
Genetic Association
Basically when you have a particular variant present in people with a
disease than people who don’t have the disease.
Case-Control Study
Cases are people with the disease
you’re looking into
The definition of the disease must
be applied in a consistent way
Controls must be matched as
similar as possible for the non-
disease traits – like age, sex,
location.
Case-Control Genetic Study
Large numbers of well-defined cases like
thousands
Equal numbers of matched controls
Reliable genotyping technology (SNP
microarray)
Standard statistical analysis (PLINK)
Positive associations should be replicated
Genetic Markers
A genetic marker is a gene or DNA sequence with a known location on
a chromosome that can be used to identify individuals or species.
• Individuals in a population are genetically far more diverse than
individuals in a single family.
• To capture this genetic diversity, we need to use >500 000 genetic
markers
, Arzoo Chaudhry Genomics
Single Nucleotide Polymorphism (SNP)
• Common in the genome ~1/300 nucleotides
• 12 million common SNPs identified in human genome
• Generated by mismatch repair during mitosis
In this example we have a short sequence as
an example that we are going to replicate
• During DNA replication the two strands
will separate and will be used as
templates to synthesise complementary
strands.
• If that goes well then we should end up with two identical copies.
• However, when synthesising this strand, instead of incorporating an
A, a G has been incorporated.
• The mismatch repair mechanism will identify this mistake and
correct it so that the bases are a standard Watson-Crick base pair
• However, in this instance it hasn’t corrected the G, it’s replaced the
T with a C.
• And what we end up with is at this position there’s either a T or a C
• If this change occurs in the gametes and isn’t deleterious then it will
get passed on to the next generation
• As time goes on it can spread through the population.
SNPs can be found in….
• Gene (coding region)
– No amino acid change (synonymous)
– Amino acid change (non-synonymous)
– New stop codon (nonsense)
• Gene (non-coding region)
– Promoter – mRNA and protein level changed
– Terminator - mRNA and protein level changed
– Splice site – Altered mRNA, altered protein
• Intergenic region (98% of genome)
dbSNP
The database of SNPs and multiple small-scale variations that include
insertions/deletions, microsatellites and non-polymorphic variants.
The rs number is a unique identifier given to each SNP
Minor Allele Frequency (MAF)
Basically when you get a SNP, e.g. it could be C to G, well there is going to
be probabilities of each one occurring, the smaller probability is the minor
allele freq. the minor and major allele freq add up to 1.
In a population of 1000, those probabilities are multiplied by 1000 and
TIMES BY 2(each person has two allele’s)
Association Analysis
Genetic Association
Basically when you have a particular variant present in people with a
disease than people who don’t have the disease.
Case-Control Study
Cases are people with the disease
you’re looking into
The definition of the disease must
be applied in a consistent way
Controls must be matched as
similar as possible for the non-
disease traits – like age, sex,
location.
Case-Control Genetic Study
Large numbers of well-defined cases like
thousands
Equal numbers of matched controls
Reliable genotyping technology (SNP
microarray)
Standard statistical analysis (PLINK)
Positive associations should be replicated
Genetic Markers
A genetic marker is a gene or DNA sequence with a known location on
a chromosome that can be used to identify individuals or species.
• Individuals in a population are genetically far more diverse than
individuals in a single family.
• To capture this genetic diversity, we need to use >500 000 genetic
markers
, Arzoo Chaudhry Genomics
Single Nucleotide Polymorphism (SNP)
• Common in the genome ~1/300 nucleotides
• 12 million common SNPs identified in human genome
• Generated by mismatch repair during mitosis
In this example we have a short sequence as
an example that we are going to replicate
• During DNA replication the two strands
will separate and will be used as
templates to synthesise complementary
strands.
• If that goes well then we should end up with two identical copies.
• However, when synthesising this strand, instead of incorporating an
A, a G has been incorporated.
• The mismatch repair mechanism will identify this mistake and
correct it so that the bases are a standard Watson-Crick base pair
• However, in this instance it hasn’t corrected the G, it’s replaced the
T with a C.
• And what we end up with is at this position there’s either a T or a C
• If this change occurs in the gametes and isn’t deleterious then it will
get passed on to the next generation
• As time goes on it can spread through the population.
SNPs can be found in….
• Gene (coding region)
– No amino acid change (synonymous)
– Amino acid change (non-synonymous)
– New stop codon (nonsense)
• Gene (non-coding region)
– Promoter – mRNA and protein level changed
– Terminator - mRNA and protein level changed
– Splice site – Altered mRNA, altered protein
• Intergenic region (98% of genome)
dbSNP
The database of SNPs and multiple small-scale variations that include
insertions/deletions, microsatellites and non-polymorphic variants.
The rs number is a unique identifier given to each SNP
Minor Allele Frequency (MAF)
Basically when you get a SNP, e.g. it could be C to G, well there is going to
be probabilities of each one occurring, the smaller probability is the minor
allele freq. the minor and major allele freq add up to 1.
In a population of 1000, those probabilities are multiplied by 1000 and
TIMES BY 2(each person has two allele’s)