Sign in

Libre University uses your GitHub account. Signing in is only needed to sit a final test, so the score is kept on your profile.

Populations and frequencies

A question asked from the floor of a scientific meeting in 1908 exposed the fact that nobody yet knew what Mendel's rules imply about a whole population, as opposed to a single family.

The previous lesson ended with two camps convinced they had incompatible sciences: discrete factors giving sharp ratios on one side, smooth continuous variation and the resemblance of relatives on the other. Both were right. Reconciling them requires changing the object of study from the pedigree to the population, and once that change is made the theory acquires something it had lacked since 1859: a number to measure.

Yule's question

At a meeting of the Royal Society of Medicine, Reginald Punnett presented Mendelism, and G. Udny Yule asked an awkward question. Brachydactyly, a condition producing short fingers, is inherited as a dominant. If dominants are three times as common as recessives in an F2, should not brachydactyly gradually rise until three-quarters of the population has short fingers? It plainly has not. Does that not tell against the whole scheme?

Punnett could not answer on the spot, and took the question to his cricketing acquaintance G. H. Hardy, a Cambridge pure mathematician who regarded the problem as trivial and the fuss as evidence that biologists could not do algebra. Hardy's reply appeared as a short letter in Science on 10 July 1908 under the title "Mendelian proportions in a mixed population". Wilhelm Weinberg, a physician in Stuttgart, had published the same result six months earlier in German, and the result carries both names.

Yule's error is a good one to have made, because it is the natural reading. The 3:1 ratio is a statement about the offspring of a particular cross, two heterozygotes, and says nothing about how common that cross is. To answer the question you have to ask what fraction of the population carries each allele, which is a quantity nobody had thought to define.

The result

Consider one locus with two alleles, A and a. Let p be the fraction of all the gene copies in the population that are A, and q=1-p the fraction that are a. Assume mating is at random with respect to this locus, that the population is large enough for sampling error to be negligible, and that nothing kills, mutates or immigrates differentially.

Under random mating, a zygote is formed by drawing two gametes independently from the pool. The chance of drawing A twice is p2, the chance of a twice is q2, and the chance of one of each is 2pq, the factor of two counting the two orders. So the genotype frequencies are

AA:Aa:aa=p2:2pq:q2

The second half of the result is the important half. Work out the allele frequency in this new generation. Every AA contributes two A copies and every Aa contributes one, so

p=2p2+2pq2(p2+2pq+q2)=p2+pq1=p(p+q)=p

The frequency is unchanged. It will be unchanged next generation and the one after. Nothing happens.

That is the answer to Yule. Dominance describes how an allele is expressed when paired with another, not how often it occurs, and expression has no bearing on transmission. A dominant allele at 1 per cent stays at 1 per cent forever unless something acts on it, and a recessive at 99 per cent stays there.

Why a null model is worth having

It is easy to dismiss the Hardy-Weinberg result as an accounting identity, and in one sense it is. Its value is that it is the first statement in biology of what happens when nothing happens, and every science needs one of those. Newton's first law is not interesting because objects often move at constant velocity; it is interesting because it identifies which observations require a force.

The list of assumptions is therefore not a weakness but the actual content. Genotype frequencies depart from p2:2pq:q2, or allele frequencies move between generations, only if at least one of the following fails: random mating, a population large enough to ignore sampling, no selection on the locus, no mutation at the locus, no migration in or out, and equal frequencies in the two sexes at an autosomal locus. Each failure is a named evolutionary force, and the remaining lessons of the first half of this course take them one at a time.

The result also supplies a definition. Evolution is a change in allele frequency in a population across generations. That is deliberately narrow and it is what makes the subject quantitative. It says nothing about progress, complexity or species, and it locates evolution in a population rather than in an individual, which is why the statement that an individual organism does not evolve is not a pedantic quibble but a consequence of the definition.

Example. Cystic fibrosis is recessive and affects about one in 2,500 births in northern European populations. What fraction of that population carries one copy?

Affected individuals are the aa class, so q2=1/2500=0.0004 and q=0.02. Then p=0.98, and the carrier frequency is

2pq=2×0.98×0.02=0.0392

about 3.9 per cent, or one person in 26. Note the ratio of carriers to affected individuals, 2p/q=2×0.98/0.02=98. For every child with the disease there are 98 unaffected carriers, so almost all copies of the allele are in people who will never show it. That fact governs how quickly selection can remove a rare recessive, and the next lesson makes it quantitative.

Now you. Phenylketonuria affects about one in 10,000 births in the same populations. What is the carrier frequency, and how many carriers are there per affected person?

Answer

q2=10-4, so q=0.01 and p=0.99. Carriers are 2pq=2×0.99×0.01=0.0198, about 2 per cent, or one person in 51. The ratio of carriers to affected is 2p/q=1.98/0.01=198. Making the disease ten times rarer does not make the allele ten times rarer, only about three times, because the affected frequency goes as the square of the allele frequency. This square is the single most useful thing about the Hardy-Weinberg result in practice: it converts a countable clinical incidence into an allele frequency you could not otherwise observe.

Testing a population against the null

Because it predicts genotype frequencies from allele frequencies, the result is testable on any sample where genotypes can be scored directly.

Example. A survey types 1,000 people at a locus with two codominant alleles, so all three genotypes can be distinguished, and finds 298 MM, 489 MN and 213 NN. Does the sample fit Hardy-Weinberg proportions?

First get the allele frequency by counting copies. There are 2×298+489=1085 copies of M out of 2000, so p=0.5425 and q=0.4575. The expected counts are then p2N=294.3, 2pqN=496.4 and q2N=209.3, and

χ2=3.72294.3+7.42496.4+3.72209.3=0.047+0.110+0.065=0.22

On one degree of freedom, since the allele frequency was estimated from the same data, that is a very good fit. Human blood group loci generally do fit this well, which is worth knowing: it means that for most loci, most of the time, none of the forces is acting strongly enough to see in a sample of a thousand.

Now you. A different sample of 1,000 gives 357, 485 and 158. Compute the fit, and say what a large excess of homozygotes would have indicated had one appeared.

Answer

Copies of the first allele number 2×357+485=1199, so p=0.5995, q=0.4005, and the expected counts are 359.4, 480.2 and 160.4. Then χ2=0.016+0.048+0.036=0.10, again a good fit. An excess of homozygotes in both directions, with a deficit of heterozygotes, has three standard causes and they are hard to tell apart from one sample: inbreeding, which pairs relatives and therefore pairs identical alleles; the Wahlund effect, where the sample unknowingly pools two populations with different allele frequencies; and a null allele that fails to amplify in the assay, so that heterozygotes are misread as homozygotes. The third is a laboratory artefact and is the commonest explanation in practice, which is a useful corrective to reading every departure as biology.

Where continuous variation comes from

The Hardy-Weinberg result settles Yule's question, but the deeper quarrel of the previous lesson was about continuous characters. Human height has no classes: it is a smooth distribution with no gaps, and no amount of pea breeding produces anything like it. How can discrete factors give that?

The answer, suggested by Yule himself in 1902 and worked out completely by R. A. Fisher in 1918, is that they give it as soon as more than one locus affects the same character. Suppose a character is influenced by n loci, at each of which one allele adds a unit to the measurement and the other adds nothing, and that the alleles are at frequency one half. An individual carries 2n alleles, of which some number are the adding kind, so the character takes one of 2n+1 values with binomial frequencies.

Example. Take one locus, then two, then six. How many phenotypic classes are there, in what proportions, and how common is the most extreme individual?

With one locus there are three classes in the proportions 1:2:1, and the extreme occurs at frequency 1/4. With two loci there are five classes, 1:4:6:4:1, and the extreme is 1/16. With six loci there are thirteen classes, 1:12:66:220:495:792:924:792:495:220:66:12:1, and the extreme is 1/4096. Thirteen classes spread over a range of twelve units, with a standard deviation of 12×0.25=1.73 units, is already a bell-shaped distribution in which neighbouring classes differ by less than the measurement error of most instruments. Add any environmental contribution and the classes smear into each other completely.

Now you. Two tall parents of the same height sometimes produce a child taller than either. Blending inheritance cannot allow this. Why does the multiple-factor model allow it easily?

Answer

Because the parents' identical measurements can rest on different combinations of alleles. If height is set by six loci and each parent carries seven adding alleles out of twelve, they look the same, but one may carry them at loci 1 to 4 and the other at loci 3 to 6. Their child draws one allele from each parent at each locus and can easily assemble eight or nine adding alleles, exceeding both parents. This is transgressive segregation, it is routine in plant and animal breeding, and it is a direct prediction of particulate inheritance that blending forbids. It is also the reason selection can push a population beyond the range of any individual it started with, which is precisely what Fleeming Jenkin said could not happen.

What is now available and what is missing

The two camps were arguing about nothing. Continuous characters are Mendelian characters counted several at a time, the resemblance between relatives that the biometricians measured is exactly what the multiple-factor model predicts, and Fisher's 1918 paper derived the correlations between parents, siblings and cousins from Mendelian assumptions and matched them to Pearson's own data. The synthesis that name-checks Fisher, J. B. S. Haldane and Sewall Wright is built on this reconciliation, and the field it created is population genetics.

What the reconciliation delivers to Darwin's argument is the missing premise. Heredity is particulate, so variation is not destroyed by transmission; it is conserved exactly, under the stated assumptions, forever. Selection therefore has a permanent supply to work on rather than a fading one, and the swamping objection is dead.

What it does not deliver is any movement at all. Hardy-Weinberg is a statement that populations sit still. To make anything happen, one of its assumptions has to be broken deliberately, and the first and most important of the breakages is the subject of the next lesson.