A theory is only trustworthy once someone has said clearly what it does not explain, and this is the lesson that does that.
The previous thirteen built a mechanism and tested it: variation arises at a measured rate, heredity conserves it, selection changes its frequency at a rate the algebra predicts, drift decides the fate of most of it, splitting produces a tree, and the tree is confirmed by anatomy, by the fossil sequence and by shared genomic errors. That is a strong position. It is not an unlimited one, and the parts of it that are overstated in popular accounts are worth separating from the parts that hold.
What the theory does not claim
Four claims get attached to it that are not in it.
It says nothing about the origin of life. Natural selection requires entities that reproduce with heritable variation, so it can only start once such entities exist. How they arose is a separate and currently unsolved problem, and its difficulty is not evidence against anything in this course.
It contains no notion of progress. The definition from the fifth lesson is change in allele frequency, and nothing in the algebra prefers complexity: parasites lose organs routinely, cave fish lose eyes, and the most abundant lineages on earth are prokaryotes that have not changed body plan in three billion years. The tree has no main trunk and no summit.
It is not a claim that organisms are optimal. Selection is a hill-climbing process with no foresight, working on the variation that happens to be present, in a population of finite size where a beneficial mutation is lost 98 times in 100.
And it carries no moral content. That a behaviour is common because it raised reproductive success is a statement about causes and licenses nothing about how anyone should act. The inference from one to the other, made by social Darwinists and by the eugenic programmes the sixth lesson showed to be arithmetically futile as well as unjust, is a straightforward logical error rather than a controversial interpretation.
Constraint: what selection cannot reach
The clearest limits are visible in the products. Selection modifies what is there, so an inherited arrangement that has become inconvenient is usually elaborated rather than replaced.
The vertebrate retina is installed backwards. Photoreceptors face away from the light, so incoming photons pass through the nerve fibre layer and the blood supply first, and the fibres must gather and leave through a hole in the retina, producing a blind spot in every vertebrate eye. Cephalopods, which evolved a camera eye independently, have theirs the right way round and no blind spot. Both work; only one is what an engineer starting fresh would draw. The left recurrent laryngeal nerve is the same story: it runs from the brain, down the neck, around the aortic arch, and back up to a larynx sitting a few centimetres from where it started, which in a fish is a direct route past a gill arch and in a giraffe is a detour of about four metres.
Example. Stephen Jay Gould and Richard Lewontin argued in 1979 that biologists too readily assume every character has an adaptive explanation, and borrowed a term from architecture: the spandrels of San Marco, the tapering triangles between arches, are richly decorated but exist because you cannot put a dome on arches without them. State what an adaptive hypothesis has to do to be more than a story.
It has to make a prediction that could fail, and there are three standard ways to supply one. First, comparative: if a character is an adaptation to a circumstance, it should appear independently in unrelated lineages facing that circumstance and be absent in relatives that do not, which is a test across a tree rather than an argument about one species. Second, manipulative: alter the character and measure the fitness consequence directly, as the widowbird experiment below does. Third, genetic: show that the character's variation is heritable and that its bearers currently differ in reproductive success, which is the breeder's equation of the previous lesson used as a test rather than a prediction.
The alternatives an adaptive claim must beat are specific: the character may be a by-product of another that is under selection, a consequence of a developmental or physical constraint, a neutral variant fixed by drift, or an inherited feature that is no longer doing anything.
Now you. Human chins are unique among primates: no other ape has one. Sketch an adaptive hypothesis, then say what would have to be shown for it to beat the alternatives.
Answer
An adaptive hypothesis is easy to produce, which is the trouble. The chin might resist mechanical stress from chewing, or from speech, or it might be a sexually selected signal. Each is plausible and each is testable, and the first has largely failed its test: measurements of stress in the mandible during chewing do not show the chin bearing the loads it would need to bear, and human jaws are less stressed than those of apes with no chins, not more.
The leading alternative is not an adaptation at all. The human face has become smaller and has retracted under the braincase over the last few hundred thousand years, and the chin may simply be the part of the mandibular symphysis left projecting once the tooth-bearing part above it shrank, which is a by-product in exactly Gould and Lewontin's sense. To beat that, an adaptive hypothesis would have to show that chin variation is heritable, that it predicts reproductive success now or predicted it in the past, and that the by-product model fails to produce the observed shape from measured changes in facial dimensions. Nobody has shown this, and the honest current answer is that we do not know why humans have chins.
Characters that lower survival
Two classes of character look like direct counterexamples to the whole scheme, and both resolve into it in ways that sharpen the theory rather than patching it.
The first is ornament. A peacock's train is heavy, conspicuous and useless for anything except display, and Darwin, who was clear that natural selection could not produce it, proposed a second process in 1871: sexual selection, in which a character spreads because it raises mating success even at a cost to survival. The mechanism follows directly once fitness is understood as descendants rather than as survival, which is the definition the sixth lesson insisted on.
The prediction is manipulable, and Malte Andersson tested it in 1982 on long-tailed widowbirds in Kenya, where males have tails around half a metre long. He divided 36 territorial males into four groups: tails cut short, tails lengthened with the cut feathers glued on, an untouched control, and a control cut and reglued at the original length to isolate the effect of the handling. Males with lengthened tails acquired roughly four times as many new nests as males with shortened tails, and the two control groups were intermediate and indistinguishable from each other. The character is costly, it is preferred, and the preference is what maintains it. Why it is preferred remains a live argument between Fisher's runaway, in which preference and trait become genetically correlated and drive each other, Zahavi's handicap principle, in which the cost is the signal, and the view that a trait exploits a pre-existing bias in the sensory system.
Characters that lower reproduction to zero
The sharper problem is the one Darwin called insuperable in the third lesson of this course: sterile castes. A worker ant leaves no offspring, so a process defined by differential reproduction cannot favour anything about her.
W. D. Hamilton supplied the arithmetic in 1964. An allele is propagated by any copy of itself, wherever it sits, so the relevant quantity is not the bearer's offspring but the total effect on copies of the allele. Writing for the cost to the actor and for the benefit to the recipient, both in offspring equivalents, and for the probability that the recipient carries the same allele by descent, an altruistic act spreads when
J. B. S. Haldane is supposed to have put it as being willing to lay down his life for two brothers or eight cousins: with and , both give against .
Example. A Belding's ground squirrel that gives an alarm call attracts the predator's attention. Suppose calling costs the caller 0.06 of its expected lifetime offspring and raises each squirrel within earshot by 0.05. How many full siblings must be nearby for calling to be favoured, and how many nieces?
For full siblings , so the condition is , giving : three siblings suffice. For nieces and nephews , so and five are needed. For first cousins at , ten.
Paul Sherman's fieldwork on this species found exactly the predicted asymmetry: females, which remain near where they were born and are therefore surrounded by relatives, give alarm calls far more often than males, which disperse and are surrounded by strangers. The same individual calls more in years when it has kin nearby. This is a quantitative prediction about who should be altruistic to whom, and it is not obtainable from any theory of selection acting only on individuals.
Now you. Ants, bees and wasps are haplodiploid: males develop from unfertilised eggs and are haploid. Full sisters therefore share rather than 0.5. This was for decades the standard explanation of why eusociality evolved so often in this group. Why is it no longer accepted as sufficient?
Answer
Two problems, one arithmetical and one empirical. The arithmetical one is that a worker's relatedness to her brothers is only 0.25, so her average relatedness to siblings is 0.5, exactly as in a diploid species, unless the colony biases its investment towards females. Robert Trivers and Hope Hare showed in 1976 that many ant colonies do bias it, in roughly the predicted 3:1 ratio, which rescues the argument but makes it conditional rather than automatic. The empirical problem is worse: queens in many eusocial species mate with multiple males, which drops relatedness among workers towards 0.25 to 0.3, and eusociality has arisen in termites, which are diploid, and in naked mole rats, which are mammals.
The current view, argued most forcefully by Jacobus Boomsma, is that the key precondition is lifetime monogamy rather than haplodiploidy: if a queen mates once, her offspring are full siblings, so to a sibling equals to one's own offspring and the barrier to helping instead of breeding disappears. This is a hypothesis inside the theory being tested and largely rejected using the theory's own arithmetic, which is what a working research programme looks like from inside.
The tautology charge
The most persistent philosophical objection is that the theory is empty: fitness is defined as reproductive success, so "the fittest survive" says only that those who reproduce, reproduce. Karl Popper endorsed a version of this in 1974, calling Darwinism a metaphysical research programme rather than a testable theory.
Example. Is the charge sound?
No, and the reason is that the objection targets a slogan rather than the theory. If fitness could only be measured after the fact by counting offspring, the criticism would land. It is not measured that way: fitness is predicted in advance from the relation between a structure and a circumstance, and the prediction is then checked against the count.
The lessons of this course are a list of instances. Beak depth was measured in 1976, the mechanical argument said deeper beaks crack Tribulus seeds, and the drought of 1977 killed 85 per cent of the population in the predicted direction. Melanic moths were predicted to survive better on soot-blackened bark, and the coefficient derived from the frequency series matched the one measured by releasing marked moths. Sickle-cell heterozygotes were predicted to resist malaria before anyone had counted their offspring, and the equilibrium frequency computed from independently estimated fitnesses matches the observed one. Each of these could have come out the other way, which is precisely what a tautology cannot do. Popper himself withdrew the claim, writing in 1978 that he had changed his mind about the testability and logical status of the theory of natural selection.
Now you. A weaker version survives that reply: that some particular adaptive explanations are untestable even though the theory is not. Is that version sound, and what follows?
Answer
It is sound, and it is Gould and Lewontin's point in a different vocabulary. A theory can be perfectly testable while a specific application of it is not: "the chin is an adaptation for resisting speech-related stress" is a claim about one structure in one species with no comparative sample, no manipulation available and no fitness measurement, and stating it does not make it science.
What follows is a division of labour rather than a verdict. The general theory is tested by its quantitative predictions about rates, frequencies and patterns, of the kind this course has worked through since the sixth lesson, and those tests are unambiguous. Individual adaptive hypotheses have to earn their status one at a time, by the comparative, manipulative and genetic routes named earlier, and many in the literature have not done so. Keeping the two kinds of claim apart is the whole of the correction Gould and Lewontin asked for, and it is now standard practice even among people who dislike the paper.
Where the argument is currently open
Several disagreements are genuine, and it would be misleading to present the field as finished.
The unit of selection. Selection acts on genes, individuals and groups at once, and these can conflict. Transposable elements make up about 45 per cent of the human genome and largely serve their own replication; the t-haplotype in house mice is transmitted to over 90 per cent of a heterozygous male's offspring instead of 50 per cent, while lowering the fitness of the mice carrying it. Whether group-level selection is ever strong enough to matter for ordinary characters is disputed, and the argument became public in 2010 when Martin Nowak, Corina Tarnita and E. O. Wilson published a critique of inclusive fitness theory that drew a reply signed by 137 biologists.
How much of molecular evolution is selected. Kimura's neutral theory has never been fully settled against its selectionist alternatives, and genome-scale data has sharpened rather than resolved the question. Michael Lynch has argued that much of eukaryotic genome architecture arose because eukaryotic populations are too small for selection to remove mildly deleterious insertions, which would make complexity a consequence of weak selection rather than strong.
Whether the synthesis needs extending. Since around 2015 a group of biologists have argued for an "extended evolutionary synthesis" incorporating developmental bias, plasticity that precedes genetic change, niche construction and non-genetic inheritance. Their opponents reply that all of these are already accommodated and that the proposal is a change of emphasis rather than of theory. Be wary of accounts from either side that describe it as settled.
The Cambrian. Most animal phyla appear in the fossil record within roughly 24 million years after 538.8 million years ago. Molecular clocks put the underlying divergences earlier, in the Ediacaran, and the mechanisms proposed for the event, including rising oxygen, the origin of predation and the evolution of developmental gene regulation, are not mutually exclusive and none is established. Note that 24 million years is a long time on the scale this course has used: about sixty-five times the 364,000 generations that even a pessimistic model of eye evolution demands.
What a finisher can do
The two facts of the first lesson now have an account. Adaptation, the quantitative fit of a structure to a job that Paley stated better than anyone, comes from cumulative selection on heritable variation, at rates measured in moths, finches, maize and bacteria that match the algebra. The nested hierarchy, which he could not use, comes from repeated splitting, and is confirmed by anatomy, by the order of the fossil record, and by shared genomic errors that no other account explains without abandoning falsifiability.
What matters more for reading any further work in the field is the ability to say what each body of evidence establishes and what it does not: that the fossil record cannot identify ancestors, that a tree topology carries no dates, that a real-time experiment cannot reach common descent, that an adaptive story is not an adaptive hypothesis until it can fail, and that a theory whose limits are stated this precisely is more trustworthy than one without them.