A question asked from the floor of a scientific meeting in 1908 exposed the fact that nobody yet knew what Mendel's rules imply about a whole population, as opposed to a single family.
The previous lesson ended with two camps convinced they had incompatible sciences: discrete factors giving sharp ratios on one side, smooth continuous variation and the resemblance of relatives on the other. Both were right. Reconciling them requires changing the object of study from the pedigree to the population, and once that change is made the theory acquires something it had lacked since 1859: a number to measure.
Yule's question
At a meeting of the Royal Society of Medicine, Reginald Punnett presented Mendelism, and G. Udny Yule asked an awkward question. Brachydactyly, a condition producing short fingers, is inherited as a dominant. If dominants are three times as common as recessives in an F2, should not brachydactyly gradually rise until three-quarters of the population has short fingers? It plainly has not. Does that not tell against the whole scheme?
Punnett could not answer on the spot, and took the question to his cricketing acquaintance G. H. Hardy, a Cambridge pure mathematician who regarded the problem as trivial and the fuss as evidence that biologists could not do algebra. Hardy's reply appeared as a short letter in Science on 10 July 1908 under the title "Mendelian proportions in a mixed population". Wilhelm Weinberg, a physician in Stuttgart, had published the same result six months earlier in German, and the result carries both names.
Yule's error is a good one to have made, because it is the natural reading. The 3:1 ratio is a statement about the offspring of a particular cross, two heterozygotes, and says nothing about how common that cross is. To answer the question you have to ask what fraction of the population carries each allele, which is a quantity nobody had thought to define.
The result
Consider one locus with two alleles, and . Let be the fraction of all the gene copies in the population that are , and the fraction that are . Assume mating is at random with respect to this locus, that the population is large enough for sampling error to be negligible, and that nothing kills, mutates or immigrates differentially.
Under random mating, a zygote is formed by drawing two gametes independently from the pool. The chance of drawing twice is , the chance of twice is , and the chance of one of each is , the factor of two counting the two orders. So the genotype frequencies are
The second half of the result is the important half. Work out the allele frequency in this new generation. Every contributes two copies and every contributes one, so
The frequency is unchanged. It will be unchanged next generation and the one after. Nothing happens.
That is the answer to Yule. Dominance describes how an allele is expressed when paired with another, not how often it occurs, and expression has no bearing on transmission. A dominant allele at 1 per cent stays at 1 per cent forever unless something acts on it, and a recessive at 99 per cent stays there.
Why a null model is worth having
It is easy to dismiss the Hardy-Weinberg result as an accounting identity, and in one sense it is. Its value is that it is the first statement in biology of what happens when nothing happens, and every science needs one of those. Newton's first law is not interesting because objects often move at constant velocity; it is interesting because it identifies which observations require a force.
The list of assumptions is therefore not a weakness but the actual content. Genotype frequencies depart from , or allele frequencies move between generations, only if at least one of the following fails: random mating, a population large enough to ignore sampling, no selection on the locus, no mutation at the locus, no migration in or out, and equal frequencies in the two sexes at an autosomal locus. Each failure is a named evolutionary force, and the remaining lessons of the first half of this course take them one at a time.
The result also supplies a definition. Evolution is a change in allele frequency in a population across generations. That is deliberately narrow and it is what makes the subject quantitative. It says nothing about progress, complexity or species, and it locates evolution in a population rather than in an individual, which is why the statement that an individual organism does not evolve is not a pedantic quibble but a consequence of the definition.
Example. Cystic fibrosis is recessive and affects about one in 2,500 births in northern European populations. What fraction of that population carries one copy?
Affected individuals are the class, so and . Then , and the carrier frequency is
about 3.9 per cent, or one person in 26. Note the ratio of carriers to affected individuals, . For every child with the disease there are 98 unaffected carriers, so almost all copies of the allele are in people who will never show it. That fact governs how quickly selection can remove a rare recessive, and the next lesson makes it quantitative.
Now you. Phenylketonuria affects about one in 10,000 births in the same populations. What is the carrier frequency, and how many carriers are there per affected person?
Answer
, so and . Carriers are , about 2 per cent, or one person in 51. The ratio of carriers to affected is . Making the disease ten times rarer does not make the allele ten times rarer, only about three times, because the affected frequency goes as the square of the allele frequency. This square is the single most useful thing about the Hardy-Weinberg result in practice: it converts a countable clinical incidence into an allele frequency you could not otherwise observe.
Testing a population against the null
Because it predicts genotype frequencies from allele frequencies, the result is testable on any sample where genotypes can be scored directly.
Example. A survey types 1,000 people at a locus with two codominant alleles, so all three genotypes can be distinguished, and finds 298 , 489 and 213 . Does the sample fit Hardy-Weinberg proportions?
First get the allele frequency by counting copies. There are copies of out of , so and . The expected counts are then , and , and
On one degree of freedom, since the allele frequency was estimated from the same data, that is a very good fit. Human blood group loci generally do fit this well, which is worth knowing: it means that for most loci, most of the time, none of the forces is acting strongly enough to see in a sample of a thousand.
Now you. A different sample of 1,000 gives 357, 485 and 158. Compute the fit, and say what a large excess of homozygotes would have indicated had one appeared.
Answer
Copies of the first allele number , so , , and the expected counts are 359.4, 480.2 and 160.4. Then , again a good fit. An excess of homozygotes in both directions, with a deficit of heterozygotes, has three standard causes and they are hard to tell apart from one sample: inbreeding, which pairs relatives and therefore pairs identical alleles; the Wahlund effect, where the sample unknowingly pools two populations with different allele frequencies; and a null allele that fails to amplify in the assay, so that heterozygotes are misread as homozygotes. The third is a laboratory artefact and is the commonest explanation in practice, which is a useful corrective to reading every departure as biology.
Where continuous variation comes from
The Hardy-Weinberg result settles Yule's question, but the deeper quarrel of the previous lesson was about continuous characters. Human height has no classes: it is a smooth distribution with no gaps, and no amount of pea breeding produces anything like it. How can discrete factors give that?
The answer, suggested by Yule himself in 1902 and worked out completely by R. A. Fisher in 1918, is that they give it as soon as more than one locus affects the same character. Suppose a character is influenced by loci, at each of which one allele adds a unit to the measurement and the other adds nothing, and that the alleles are at frequency one half. An individual carries alleles, of which some number are the adding kind, so the character takes one of values with binomial frequencies.
Example. Take one locus, then two, then six. How many phenotypic classes are there, in what proportions, and how common is the most extreme individual?
With one locus there are three classes in the proportions 1:2:1, and the extreme occurs at frequency . With two loci there are five classes, 1:4:6:4:1, and the extreme is . With six loci there are thirteen classes, 1:12:66:220:495:792:924:792:495:220:66:12:1, and the extreme is . Thirteen classes spread over a range of twelve units, with a standard deviation of units, is already a bell-shaped distribution in which neighbouring classes differ by less than the measurement error of most instruments. Add any environmental contribution and the classes smear into each other completely.
Now you. Two tall parents of the same height sometimes produce a child taller than either. Blending inheritance cannot allow this. Why does the multiple-factor model allow it easily?
Answer
Because the parents' identical measurements can rest on different combinations of alleles. If height is set by six loci and each parent carries seven adding alleles out of twelve, they look the same, but one may carry them at loci 1 to 4 and the other at loci 3 to 6. Their child draws one allele from each parent at each locus and can easily assemble eight or nine adding alleles, exceeding both parents. This is transgressive segregation, it is routine in plant and animal breeding, and it is a direct prediction of particulate inheritance that blending forbids. It is also the reason selection can push a population beyond the range of any individual it started with, which is precisely what Fleeming Jenkin said could not happen.
What is now available and what is missing
The two camps were arguing about nothing. Continuous characters are Mendelian characters counted several at a time, the resemblance between relatives that the biometricians measured is exactly what the multiple-factor model predicts, and Fisher's 1918 paper derived the correlations between parents, siblings and cousins from Mendelian assumptions and matched them to Pearson's own data. The synthesis that name-checks Fisher, J. B. S. Haldane and Sewall Wright is built on this reconciliation, and the field it created is population genetics.
What the reconciliation delivers to Darwin's argument is the missing premise. Heredity is particulate, so variation is not destroyed by transmission; it is conserved exactly, under the stated assumptions, forever. Selection therefore has a permanent supply to work on rather than a fading one, and the swamping objection is dead.
What it does not deliver is any movement at all. Hardy-Weinberg is a statement that populations sit still. To make anything happen, one of its assumptions has to be broken deliberately, and the first and most important of the breakages is the subject of the next lesson.