Sign in

Libre University uses your GitHub account. Signing in is only needed to sit a final test, so the score is kept on your profile.

How fast selection works

Once evolution is defined as a change in allele frequency, the question of whether selection is strong enough to matter stops being a matter of opinion and becomes a calculation.

The previous lesson established that a population left alone does nothing: allele frequencies sit where they are, generation after generation, and the assumptions behind that result are a list of the ways a population can be made to move. This lesson breaks the first of them, the assumption that every genotype survives and reproduces equally well, and works out how fast the frequency changes when it does.

Fitness is a ratio, not a virtue

The word fitness carries two centuries of unhelpful baggage, so it is worth stating flatly what it means in the algebra. The absolute fitness of a genotype is the average number of offspring an individual of that genotype leaves. Only ratios matter for the frequency of an allele, so the convention is to divide through by the largest, giving a relative fitness w with a maximum of one, and to write the shortfall of any other genotype as a selection coefficient s=1-w.

Three things follow at once. Fitness is a property of a genotype in a particular environment, not of an organism, and certainly not of a species: the same allele can have w=1 in one valley and w=0.6 in the next. It is an average over individuals, so a genotype with high fitness contains individuals who leave nothing. And it counts descendants, not strength, health or longevity, which is why a peacock's tail can raise fitness while lowering nearly everything else about its bearer's condition.

It also has nothing to do with dominance. The previous lesson's answer to Yule was that a dominant allele does not spread because it is dominant, and the same holds here: dominance changes how selection sees a genotype, not how strongly it acts.

The recursion, and what it says about speed

Take one locus with alleles A and a at frequencies p and q. Assign relative fitnesses wAA, wAa and waa. Random mating produces zygotes in Hardy-Weinberg proportions, selection then multiplies each class by its fitness, and renormalising by the mean fitness

w=p2wAA+2pqwAa+q2waa

gives the frequencies among the survivors. Counting allele copies among those survivors gives the frequency in the next generation,

p=p2wAA+pqwAaw

That single line is the whole of one-locus selection theory. Everything below is a special case of it.

The simplest case is the clearest. Suppose the organism is haploid, or equivalently that the heterozygote sits exactly halfway between the homozygotes, and let the favoured type have fitness 1+s against the other's 1. Then the ratio p/q is multiplied by exactly 1+s every generation, so ln(p/q) increases by ln(1+s) per generation and the frequency traces a logistic curve. That is a useful thing to know, because the odds ratio is a straight line in time even though the frequency is not.

Example. An allele with a 1 per cent advantage starts at a frequency of 1 per cent. How many generations until it reaches 99 per cent, and how does that change if the advantage is 10 per cent?

The odds go from 0.01/0.99 to 0.99/0.01, a factor of (99)2=9801. The number of generations is ln(9801)/ln(1.01)=9.190/0.00995=924. With s=0.1 the denominator is ln(1.1)=0.0953 and the answer is 97 generations. A 1 per cent advantage is far too small to detect in any field study anyone could run, and it sweeps an allele through a population in under a thousand generations, which for an annual plant is under a thousand years and geologically instantaneous. This is the single most important number in the subject: selection so weak that it is invisible is still overwhelmingly fast on the timescale the previous lessons established.

Now you. The same allele, again at 1 per cent with a 1 per cent advantage. How long does it take to get from 1 per cent to 2 per cent, and from 50 per cent to 51 per cent? Why are the answers so different?

Answer

From 1 to 2 per cent the odds go from 0.010101 to 0.020408, a factor of 2.02, so the time is ln(2.02)/ln(1.01)=0.703/0.00995=71 generations. From 50 to 51 per cent the odds go from 1 to 1.0408, a factor of 1.0408, and the time is 0.0400/0.00995=4 generations. Selection is roughly eighteen times faster in the middle than at the bottom, because what selection acts on is the difference between the two types weighted by how often they meet the environment, and pq is largest at one half. The two tails are where alleles spend nearly all of their time, which is why a sweep looks like nothing at all for a long while and then happens suddenly.

Melanism, costed

The peppered moth, Biston betularia, is normally pale and speckled, and rests on lichen-covered bark where it is very hard to see. A black form named carbonaria was first recorded near Manchester in 1848. By the end of the century it made up something like 98 per cent of the moths caught in the industrial districts, while remaining rare in rural Dorset and Cornwall. The black form is produced by a dominant allele, so a heterozygote is black.

J. B. S. Haldane used this case in 1924 as the first calculation of a selection coefficient from field data, and it is worth repeating because his conclusion is checkable.

Example. Take the melanic phenotype at 0.1 per cent in 1848 and 99 per cent in 1898, with one generation a year. The melanic allele is dominant. What relative fitness does the pale form need?

Because melanism is dominant, a melanic phenotype frequency of 0.001 means q2=0.999, so q=0.9995 and the allele frequency p is 0.0005. Give the melanic genotypes fitness 1 and the pale homozygote 1-s, so that w=1-sq2 and p=p/w. Iterating that recursion for 50 generations and solving for the s that lands the melanic phenotype on 0.99 gives s=0.333. The pale form is two-thirds as fit as the black, or equivalently the black form is 1/0.667=1.5 times as fit. Haldane's published figure was 50 per cent, and the point he drew from it is the one that mattered in 1924: a selective advantage large enough to do this is small enough that no naturalist watching a wood would ever notice it.

Now you. The Clean Air Act of 1956 reversed the conditions. Around Manchester the melanic form fell from roughly 90 per cent in 1960 to roughly 10 per cent by 1995, again about one generation a year. What is the selection coefficient now acting against melanics, and why is it smaller than the one that put them there?

Answer

Now the melanic genotypes carry the cost. With w=(1-s)(1-q2)+q2 and p=p(1-s)/w, starting from p=0.684 (the value giving 90 per cent melanic phenotypes) and requiring 10 per cent melanic phenotypes after 35 generations, the answer is s=0.15. Published estimates from mark-release-recapture and from the frequency series itself sit between 0.1 and 0.2, so the arithmetic and the fieldwork agree.

It is smaller for a structural reason worth keeping. Selection against a dominant allele is efficient, because every copy of it is exposed in a black moth, so a moderate coefficient moves the frequency quickly. Selection against a recessive, which is what the pale allele suffered on the way up, is inefficient once the recessive is rare, because most copies are hidden in heterozygotes. Getting the melanic allele from 0.0005 to near fixation therefore required a much larger coefficient than removing it does. The asymmetry is not biology, it is arithmetic.

The case has been attacked, and the attacks are worth knowing. Bernard Kettlewell's 1950s experiments, which released marked moths and recovered them, used higher densities than occur naturally and sometimes placed moths on tree trunks in daylight, whereas the moths in fact rest mostly on the undersides of branches. Michael Majerus ran a seven-year experiment answering those criticisms, published after his death in 2012: 4,864 moths released on natural resting positions, with predation scored directly, giving a survival advantage to the pale form of about 9 per cent per day in an unpolluted wood. The mechanism is bird predation, the coefficient is real, and the older experiments were sloppy rather than wrong.

Dominance decides what selection can see

The asymmetry in the moth case is general, and it has a consequence that matters far beyond moths.

Example. A recessive allele is lethal in homozygotes, so waa=0 and everything else is 1. Starting from q=0.01, how many generations of complete lethality are needed to halve the frequency?

For a lethal recessive the recursion simplifies to q=q/(1+q), which rearranges to 1/q=1/q+1: the reciprocal of the frequency rises by exactly one per generation. Halving q from 0.01 to 0.005 means taking 1/q from 100 to 200, which takes 100 generations. Killing every homozygote, in every generation, for a century of human generations, removes half the allele. The reason is in the previous lesson's arithmetic: at q=0.01 the ratio of heterozygous carriers to affected individuals is 2p/q=198, so more than 99 per cent of the copies are in people selection cannot touch.

Now you. What does this imply about proposals to eliminate a recessive disease allele by preventing affected people from reproducing?

Answer

It implies they cannot work, and the numbers were available to the people who proposed them. The eugenic sterilisation laws passed in the United States from 1907 and in several European countries after were justified partly by the claim that heritable defects could be bred out of a population. For a rare recessive, sterilising every affected individual is exactly the lethal-recessive model above, and it takes 100 generations, roughly 2,500 years, to halve an allele already at 1 per cent, while mutation replaces some of what is removed. R. A. Fisher and Lancelot Hogben both made versions of this argument in the 1930s. The policies were unjust on grounds that have nothing to do with algebra, and they were also, on their own stated terms, arithmetically futile.

When selection keeps variation instead of spending it

Every case so far ends with one allele at fixation and the variation gone. Selection is normally a consumer of variation, and that is a problem for a theory that needs variation to keep working. There is one important configuration where it is not.

If the heterozygote is fitter than both homozygotes, neither allele can be eliminated, because whichever becomes rare finds itself mostly in heterozygotes and therefore mostly in the fittest class. Writing wAA=1-s, wAa=1 and waa=1-t, the frequency settles at

q*=ss+t

The sickle-cell polymorphism is the textbook case because all three fitnesses can be estimated. The β-globin allele HbS differs from the normal allele by one base, changing glutamic acid to valine at the sixth position of the beta chain. Homozygotes have sickle-cell anaemia. Heterozygotes are largely healthy and, as A. C. Allison established in 1954 by comparing malarial lowlands with highlands in East Africa, are strongly protected against falciparum malaria; case-control work in Kenya published in 2005 puts the protection against severe malaria at roughly tenfold.

Using the conventional estimates wAA=0.89, wAS=1 and wSS=0.20, the coefficients are s=0.11 against the normal homozygote and t=0.80 against the sickle homozygote, so

q*=0.110.11+0.80=0.121

At that frequency 1.5 per cent of births are SS and about 21 per cent of the population carries the trait, which is close to what is measured across the malarial belt of West and Central Africa. Note the price. The mean fitness at equilibrium is w=0.903, so the population pays a permanent 9.7 per cent reduction in mean fitness to hold the polymorphism. Selection is not an optimiser: it is a bookkeeper that stops where the ledger balances, and here it balances at a point that kills a percentage of every generation's children.

What the algebra assumes, and what it cannot do

The recursion assumes one locus, constant fitnesses, an infinite population and random mating. Real fitnesses vary between years, sites and densities, most characters involve many loci, and populations are finite. Later lessons take those apart. Even so, the one-locus model gets the industrial melanism coefficient right to within the spread of the field estimates, which is the sort of agreement a simple model earns.

The deeper limitation is structural rather than numerical, and it is where this lesson hands over. The recursion moves frequencies. It has no term that makes a new allele. Feed it a population with one allele at a locus and it will return that population unchanged forever, however strong the selection, because p=1 gives p=1. Selection is a filter, and a filter is empty until something is poured into it. Where the alleles come from in the first place is the subject of the next lesson.