Week 1:
Big Data use = cut down costs in healthcare + prioritize patient
outcome
It is a paradigm shift of going from professional opinions → evidence
based, large scale data approaches
What is data?
o Set of values of variables
o Collected to be examined to help decision making
What is data science?
o Interdisciplinary field that uses scientific methods, and
algorithms to extract knowledge from structured and
unstructured data
What is big data?
o Very large sets of data, produced from all sorts of research.
We need special tools to analyze it
The 3 V’s of Big Data:
1. Volume (scale of data, very large)
2. Variety (different types of data need to be integrated)
3. Velocity (speed of new data streaming)
4. The 4th V => Veracity (uncertainty of data; we may have
missing data/background noise)
5. The 5th V => Value (impact of the data; go from data →
information)
Big data/data science field is emerging since data collection scale is
increasing, and getting cheaper
Week 2:
There are two ways to report data:
1. Scientific paper
2. Data report (newer way)
A good data report has:
o Clear type of data report
o Audience described
o Plan for message, tools and reasoning
o Objective
o Contains conclusions (and tells what actions can be taken)
Genetics + GWAS
Heritability = proportion of trait variance attributable to genetic
variance.
o Meaning the extent to which observed individual differences
can be traced back to genetic differences
o Humans share 99.9% of their genes
Twin studies helped to untangle the effects of genetics and
environment on human traits since:
, o Monozygotic twins share 100% of the 0.1% of variation in
genes among people
o Dizygotic twins share 50% of these
o Both share 100% of the environment
So, if a trait is 100% heritable, similarities in MZ = 2 x DZ twins
If trait is 100% environmental, similarities between MZ = DZ
The meta-analysis from twin studies concluded that:
o All traits are heritable to some extent
o Influence of shared environment (c2) is relatively small
o Majority of traits are consistent w the model where genetic
variance is additive
Every single cell contains the same DNA but not every gene is
expressed in every cell; the specific genes expressed determines
the cell type
What are SNPs? The 0.1% of genetic variation among people; can be
inside or outside genes (in the non-functional part; probably with
regulatory function).
o The result of SNPs can be harmful/harmless, latent (depending
on other factors) or silent
o SNPs are caused by mutations, recombination (scale of
crossing over) or segregation (scale of chromosome
combination)
Monogenic disorders = single gene responsible
Polygenic disorders = genetics + environment responsible
To find the genes responsible for such diseases you can conduct:
o Candidate Gene Studies:
1-10 preselected genes based on prior knowledge are selected
and then you test them for associations with traits by
comparing allele frequencies.
Cons: false positives can be found; better to do an
explorative study (GWAS)
o Genome Wide Association Studies:
Genotype a large set of individuals on 1 million SNPs (not all 3
x 109 nucleotides, since SNPs are closely located). For each
SNP compare the allele frequency across cases and controls
and conduct a statistical test for a difference in frequency. The
outcome is a Manhattan Plot (plot of chromosome number v.s.
the -log10(p value) on the axis). Past a certain threshold of -
log10(p value), there is evidence for association.
Pros: may identify multiple possible loci
Cons: increased likelihood of false positives, population
stratification may occur (need to be corrected for), large
samples are needed to detect the smaller effects
e.g. the proof for schizophrenia being heritable was found with
GWAS studies, and now novel genes are being discovered, as
data increases
4 major issues with GWAS:
Big Data use = cut down costs in healthcare + prioritize patient
outcome
It is a paradigm shift of going from professional opinions → evidence
based, large scale data approaches
What is data?
o Set of values of variables
o Collected to be examined to help decision making
What is data science?
o Interdisciplinary field that uses scientific methods, and
algorithms to extract knowledge from structured and
unstructured data
What is big data?
o Very large sets of data, produced from all sorts of research.
We need special tools to analyze it
The 3 V’s of Big Data:
1. Volume (scale of data, very large)
2. Variety (different types of data need to be integrated)
3. Velocity (speed of new data streaming)
4. The 4th V => Veracity (uncertainty of data; we may have
missing data/background noise)
5. The 5th V => Value (impact of the data; go from data →
information)
Big data/data science field is emerging since data collection scale is
increasing, and getting cheaper
Week 2:
There are two ways to report data:
1. Scientific paper
2. Data report (newer way)
A good data report has:
o Clear type of data report
o Audience described
o Plan for message, tools and reasoning
o Objective
o Contains conclusions (and tells what actions can be taken)
Genetics + GWAS
Heritability = proportion of trait variance attributable to genetic
variance.
o Meaning the extent to which observed individual differences
can be traced back to genetic differences
o Humans share 99.9% of their genes
Twin studies helped to untangle the effects of genetics and
environment on human traits since:
, o Monozygotic twins share 100% of the 0.1% of variation in
genes among people
o Dizygotic twins share 50% of these
o Both share 100% of the environment
So, if a trait is 100% heritable, similarities in MZ = 2 x DZ twins
If trait is 100% environmental, similarities between MZ = DZ
The meta-analysis from twin studies concluded that:
o All traits are heritable to some extent
o Influence of shared environment (c2) is relatively small
o Majority of traits are consistent w the model where genetic
variance is additive
Every single cell contains the same DNA but not every gene is
expressed in every cell; the specific genes expressed determines
the cell type
What are SNPs? The 0.1% of genetic variation among people; can be
inside or outside genes (in the non-functional part; probably with
regulatory function).
o The result of SNPs can be harmful/harmless, latent (depending
on other factors) or silent
o SNPs are caused by mutations, recombination (scale of
crossing over) or segregation (scale of chromosome
combination)
Monogenic disorders = single gene responsible
Polygenic disorders = genetics + environment responsible
To find the genes responsible for such diseases you can conduct:
o Candidate Gene Studies:
1-10 preselected genes based on prior knowledge are selected
and then you test them for associations with traits by
comparing allele frequencies.
Cons: false positives can be found; better to do an
explorative study (GWAS)
o Genome Wide Association Studies:
Genotype a large set of individuals on 1 million SNPs (not all 3
x 109 nucleotides, since SNPs are closely located). For each
SNP compare the allele frequency across cases and controls
and conduct a statistical test for a difference in frequency. The
outcome is a Manhattan Plot (plot of chromosome number v.s.
the -log10(p value) on the axis). Past a certain threshold of -
log10(p value), there is evidence for association.
Pros: may identify multiple possible loci
Cons: increased likelihood of false positives, population
stratification may occur (need to be corrected for), large
samples are needed to detect the smaller effects
e.g. the proof for schizophrenia being heritable was found with
GWAS studies, and now novel genes are being discovered, as
data increases
4 major issues with GWAS: