Introduction to bioinformatics
Learning objectives:
1. Understand the scale of biological information and the need for bioinformaticians
2. Make information and instructions computable
3. Describe important biological databases
4. Chose bioinformatics software
5. Use
- FASTA format
- Genes names
- Protein families databases(pfam)
- Gene ontologies
- NCBI databases
- NCBI BLAST (web portal)
The scale of biological information
The amount of biological data we have is huge and needs to be stored and processed for use.
- i.e. the human body comprises of 50 trillion cells
- 33 billion of 50 trillion are brain cells
- In each cell there is 23 chromosomes
- The human genome is the blue print/ software to design the cells and determine function
- There are 4 bases in the genome and a total of 3.2 billion bases in the 50 trillion cells
Bioinformatics:
- The branch of science concerned with information and information flow in biological systems.
- Especially the use of computational methods in genetics and genomics.
- Using information systems to understand deep biological data sets
- Take biological data and place into a form that the computer will understand, operate on biological data.
- The science of collecting and analysing complex biological data such as genetic codes.
Bill gates said
- ‘DNA is like a computer programme (because it tells the cell what to do) but far more advanced than any software
ever created.’
Computer systems think/deal differently to the human brain
- Computer systems cannot understand ambiguity
- Therefore, need to be accurate and precise
- Computers think in a binary manner
Human brain thinks differently
- There are 33 billion cells in the human brain that have multiple connects
- The multiple connections allows the human brain to not think in a binary manner
- Whereas a computer has, a large number of switches but only occur on or off in the computer chip and, therefore,
think in a binary manner.
The more switches, the faster the computer
Moore’s law is a prediction that the transistors / switches on a computer doubles every 2 years
o Therefore, computers are getting faster
Another way of thinking about it is the cost of a computer switch is halving every 2 years (white line on the
blue graph)
Blue line = Cost of generating DNA sequence data
o Which follows the Moore’s law
o But then there is a large drop in price
o This presents challenges in bioinformatics
o Due to the increase in biological data compared to computer chip it brings challenges, such as
placing the data in a suitable place to be easily accessible.
, Storing and sharing information
Blue = the cumulative (increase) number of base pairs of the DNA stored in this database
To find what you want in this big database, need to use tools of bioinformatics
So this is the end of the section that answers why we do bioinformatics
Learning objectives:
1. Understand the scale of biological information and the need for bioinformaticians
2. Make information and instructions computable
3. Describe important biological databases
4. Chose bioinformatics software
5. Use
- FASTA format
- Genes names
- Protein families databases(pfam)
- Gene ontologies
- NCBI databases
- NCBI BLAST (web portal)
The scale of biological information
The amount of biological data we have is huge and needs to be stored and processed for use.
- i.e. the human body comprises of 50 trillion cells
- 33 billion of 50 trillion are brain cells
- In each cell there is 23 chromosomes
- The human genome is the blue print/ software to design the cells and determine function
- There are 4 bases in the genome and a total of 3.2 billion bases in the 50 trillion cells
Bioinformatics:
- The branch of science concerned with information and information flow in biological systems.
- Especially the use of computational methods in genetics and genomics.
- Using information systems to understand deep biological data sets
- Take biological data and place into a form that the computer will understand, operate on biological data.
- The science of collecting and analysing complex biological data such as genetic codes.
Bill gates said
- ‘DNA is like a computer programme (because it tells the cell what to do) but far more advanced than any software
ever created.’
Computer systems think/deal differently to the human brain
- Computer systems cannot understand ambiguity
- Therefore, need to be accurate and precise
- Computers think in a binary manner
Human brain thinks differently
- There are 33 billion cells in the human brain that have multiple connects
- The multiple connections allows the human brain to not think in a binary manner
- Whereas a computer has, a large number of switches but only occur on or off in the computer chip and, therefore,
think in a binary manner.
The more switches, the faster the computer
Moore’s law is a prediction that the transistors / switches on a computer doubles every 2 years
o Therefore, computers are getting faster
Another way of thinking about it is the cost of a computer switch is halving every 2 years (white line on the
blue graph)
Blue line = Cost of generating DNA sequence data
o Which follows the Moore’s law
o But then there is a large drop in price
o This presents challenges in bioinformatics
o Due to the increase in biological data compared to computer chip it brings challenges, such as
placing the data in a suitable place to be easily accessible.
, Storing and sharing information
Blue = the cumulative (increase) number of base pairs of the DNA stored in this database
To find what you want in this big database, need to use tools of bioinformatics
So this is the end of the section that answers why we do bioinformatics