Introduction to bioinformatics
Learning objectives:
1. Understand the scale of biological information and the need for bioinformaticians
2. Make information and instructions computable
3. Describe important biological databases
4. Chose bioinformatics software
5. Use
- FASTA format
- Genes names
- Protein families databases(pfam)
- Gene ontologies
- NCBI databases
- NCBI BLAST (web portal)
The scale of biological information
The amount of biological data we have is huge and needs to be stored and processed for use.
- i.e. the human body comprises of 50 trillion cells
- 33 billion of 50 trillion are brain cells
- In each cell there is 23 chromosomes
- The human genome is the blue print/ software to design the cells and determine function
- There are 4 bases in the genome and a total of 3.2 billion bases in the 50 trillion cells
Bioinformatics:
- The branch of science concerned with information and information flow in biological systems.
- Especially the use of computational methods in genetics and genomics.
- Using information systems to understand deep biological data sets
- Take biological data and place into a form that the computer will understand, operate on biological data.
- The science of collecting and analysing complex biological data such as genetic codes.
Bill gates said
- ‘DNA is like a computer programme (because it tells the cell what to do) but far more advanced than any software
ever created.’
Computer systems think/deal differently to the human brain
- Computer systems cannot understand ambiguity
- Therefore, need to be accurate and precise
- Computers think in a binary manner
Human brain thinks differently
- There are 33 billion cells in the human brain that have multiple connects
- The multiple connections allows the human brain to not think in a binary manner
- Whereas a computer has, a large number of switches but only occur on or off in the computer chip and, therefore,
think in a binary manner.
The more switches, the faster the computer
Moore’s law is a prediction that the transistors / switches on a computer doubles every 2 years
o Therefore, computers are getting faster
Another way of thinking about it is the cost of a computer switch is halving every 2 years (white line on the
blue graph)
Blue line = Cost of generating DNA sequence data
o Which follows the Moore’s law
o But then there is a large drop in price
o This presents challenges in bioinformatics
o Due to the increase in biological data compared to computer chip it brings challenges, such as
placing the data in a suitable place to be easily accessible
, Storing and sharing information
Blue = the cumulative number of base pairs in the DNA stored in this database
Brief overview of ‘omics’
Omics = following the flow of information in a living cell from DNA – RNA-Proteins –metabolites
- DNA = genomics
- RNA = Transcriptonomics
- Protein = proteomics
- Metabolites =metabolomics
Learning objectives:
1. Understand the scale of biological information and the need for bioinformaticians
2. Make information and instructions computable
3. Describe important biological databases
4. Chose bioinformatics software
5. Use
- FASTA format
- Genes names
- Protein families databases(pfam)
- Gene ontologies
- NCBI databases
- NCBI BLAST (web portal)
The scale of biological information
The amount of biological data we have is huge and needs to be stored and processed for use.
- i.e. the human body comprises of 50 trillion cells
- 33 billion of 50 trillion are brain cells
- In each cell there is 23 chromosomes
- The human genome is the blue print/ software to design the cells and determine function
- There are 4 bases in the genome and a total of 3.2 billion bases in the 50 trillion cells
Bioinformatics:
- The branch of science concerned with information and information flow in biological systems.
- Especially the use of computational methods in genetics and genomics.
- Using information systems to understand deep biological data sets
- Take biological data and place into a form that the computer will understand, operate on biological data.
- The science of collecting and analysing complex biological data such as genetic codes.
Bill gates said
- ‘DNA is like a computer programme (because it tells the cell what to do) but far more advanced than any software
ever created.’
Computer systems think/deal differently to the human brain
- Computer systems cannot understand ambiguity
- Therefore, need to be accurate and precise
- Computers think in a binary manner
Human brain thinks differently
- There are 33 billion cells in the human brain that have multiple connects
- The multiple connections allows the human brain to not think in a binary manner
- Whereas a computer has, a large number of switches but only occur on or off in the computer chip and, therefore,
think in a binary manner.
The more switches, the faster the computer
Moore’s law is a prediction that the transistors / switches on a computer doubles every 2 years
o Therefore, computers are getting faster
Another way of thinking about it is the cost of a computer switch is halving every 2 years (white line on the
blue graph)
Blue line = Cost of generating DNA sequence data
o Which follows the Moore’s law
o But then there is a large drop in price
o This presents challenges in bioinformatics
o Due to the increase in biological data compared to computer chip it brings challenges, such as
placing the data in a suitable place to be easily accessible
, Storing and sharing information
Blue = the cumulative number of base pairs in the DNA stored in this database
Brief overview of ‘omics’
Omics = following the flow of information in a living cell from DNA – RNA-Proteins –metabolites
- DNA = genomics
- RNA = Transcriptonomics
- Protein = proteomics
- Metabolites =metabolomics