evolution at the molecular level.
How common is neutral evolution?
In the last lecture we saw that a viable contender with selection to explain evolution, especially of
small changes such as point mutations, is neutral evolution. We are then left with a problem,
namely which of these two forces is responsible for the sequence evolution that we see. Let us
suppose we have sequences from a gene from many species. There has been some divergence
between the sequences. What we want to know is whether this was due mainly to selection or
drift. Selection if directional favouring change could have caused the differences but so could drift
- which was it? Does selection usually affect mutations that alter proteins, does it affect silent
changes in exons that do not change the protein, does it care about the sequence between
genes?
How to tell if a protein is neutrally evolving?
Understanding how common neutral (or nearly neutral) evolution might be has been one of the
major tasks of molecular evolutionists. There are at least two ways to use the amount of
sequence divergence to answer these questions.
A)Ka/Ks diagnoses whether mutations that change a protein are always neutral
B)Dispersion index tests whether those mutations that are fixed are likely to have
spread by drift
, A) Ka/Ks
We start with the alignment of the same gene in 2 species. Differences are where evolution has
happened.
Imagine we have the sequence of the same gene in for example, mouse and rat. The sequences
will not be identical as there will have been some evolution between the species. This evolution
will be seen as mismatches in the alignment (see fig 1).
Two sorts of changes can happen to genes:
- First the changes could change the DNA but not the amino acids. These are known as
synonymous substitutions.
- Second the changes could alter both the DNA and the amino acids. These are non-
synonymous substitutions.
We count the number of each (Ls and Ln). We then make 2 corrections:
1. We correct Ln and Ls to allow for the possibility that there might have been multiple changes
that we cannot see.
2. We normalise the number of inferred changes to the number of non-synonymous sites and
number of synonymous sites. If mutations were random in the gene, about 75% would change
the protein owing to the structure of the code.
If the evolution of this sequence was perfectly neutral then a new mutation that altered the
protein would be as likely to go to fixation as one that had no effect. It is this property that we are
going to use. It doesn’t follow that will be as many nonsynonymous changes as synonymous
ones. This is because many more changes (those are the first two sites in a codon) will change
the amino acid than will not. So we need to calculate the number of synonymous changes per