Sign in

Libre University uses your GitHub account. Signing in is only needed to sit a final test, so the score is kept on your profile.

Evidence in the genome

A genome is a document that has been copied continuously for four billion years, and like every long-copied document its errors are more informative than its text.

The previous lesson closed on the limits of the fossil record: it is 99.99 per cent absent, it cannot identify ancestors, and it cannot resolve tempo below its sampling interval. The record this lesson uses has none of those problems. Every living organism carries a complete copy, nothing has been eroded away, and the entries can be read to the base. What makes it decisive is a specific feature of it, which is that a great deal of what it contains is broken, and the breakages are shared in exactly the pattern the tree predicts.

The code itself

The first fact is the arrangement that translates DNA into protein. Sixty-four triplets of bases specify twenty amino acids and a stop signal, and the assignment is essentially the same in bacteria, archaea, plants, fungi and animals. About thirty variant codes are known, in mitochondria, in ciliates and in a few bacterial groups, and every one is a small modification of the standard set rather than an independent system.

The number of conceivable assignments is 2164=4.2×1084. One assignment is used by everything, which is exactly what descent from a single population that had already fixed a code predicts, and it is the reason a human gene can be inserted into a bacterium and produce a working human protein, which is how insulin has been manufactured since 1978.

The honest qualification is that the code is not arbitrary. Similar amino acids tend to have similar codons, so that a single-base error often substitutes a chemically similar residue and does less damage. Simulations by Stephen Freeland and Laurence Hurst in 1998 found the natural code better at this than around 999,999 of a million randomly generated alternatives. So the code's structure is partly explained by selection on error tolerance, and its universality is the part that carries the ancestry argument, not its details.

Shared errors

The argument that does the real work is older than molecular biology and comes from textual scholarship. If two manuscripts of the same text contain the same unique misspelling in the same word, they are copies of a common exemplar, and this holds whatever the text says. Correct readings prove nothing, because both scribes could have copied correctly from different sources; a shared error has no explanation except shared descent.

Genomes are full of errors, and several kinds of them cannot plausibly arise twice in the same place.

Example. Endogenous retroviruses are the remains of viral infections of germ cells. A retrovirus inserts its genome into a host chromosome at a position that is essentially arbitrary, and if the cell is one that makes gametes, the insertion is inherited. About 8 per cent of the human genome consists of such sequences, roughly 200,000 identifiable elements, nearly all of them mutated past the point of producing a virus. Almost every one of them sits at the same position in the chimpanzee genome. What is the chance of that arising independently?

Take the target as the 3.1×109 bases of the genome. Two independent insertions landing at the same base have probability 3.2×10-10. Insertion is not uniform, since retroviruses favour open chromatin and some sequence contexts, so be generous and suppose only a million sites are ever usable: the probability is then 10-6 per insertion, and the chance that 200,000 of them coincide is 10-6 raised to the 200,000th power.

The number is not worth writing out. What makes the argument work is not its size but its structure: the insertions are shared in a nested pattern, so that some are found in all primates, some in apes but not monkeys, some in humans and chimpanzees only, and the pattern reproduces the tree built from anatomy and from working genes. A million independent accidents agreeing on one diagram out of 1020 is the same test as the previous lessons, run on characters whose position carries no function at all.

Now you. A critic replies that retroviruses might have preferred insertion sites, so the same sites could be hit repeatedly in different lineages. What observation answers this without appealing to probability?

Answer

Look at the sequence of the shared insertion, not just its position. An inserted element begins to accumulate mutations the moment it lands, and since it is usually non-functional those mutations are neutral, so it decays at the mutation rate of the eighth lesson.

Two elements independently inserted at a favoured site would be two independent copies of the ancestral viral sequence and would differ from each other by however much the virus itself had changed between the two infections, with no particular pattern. Two elements inherited from a common ancestor should differ by exactly the neutral divergence between the two species, roughly 1.2 per cent between human and chimpanzee, and should share every mutation that arose before the split and none that arose after. That is what is found. The same test applies to where each element sits relative to its neighbours: independent insertions would land in different genomic contexts, and inherited ones sit in identical flanking sequence, including the short target-site duplication the insertion mechanism creates.

There is also a decisive special case. Some ERV insertions are present in one individual human and absent in another, which shows the process is ongoing and that its products are inherited as ordinary alleles.

The broken gene for vitamin C

Ascorbic acid is required by every vertebrate, and almost all of them make it in the liver or kidney from glucose, in four enzymatic steps. The last step is performed by L-gulonolactone oxidase, encoded by the gene GULO.

Humans cannot do it. Neither can any other haplorhine primate, meaning monkeys, apes and tarsiers, nor guinea pigs, nor several bat lineages, nor most passerine birds. We must eat vitamin C, and if we do not we get scurvy, which killed more sailors than combat did for three centuries. A goat, which makes its own, produces something like 13 g a day under stress, against a human dietary requirement around 90 mg.

The gene is still there. Nobuyo Nishikimi and colleagues reported in 1994 that the human genome carries GULO on chromosome 8 as a pseudogene, recognisable by sequence similarity to the working versions in other mammals, with several exons missing and the remainder riddled with mutations that would prevent a functional protein even if the missing exons were restored.

Example. Why is a broken GULO stronger evidence for common descent than a working one would be?

Because a working gene has a functional explanation and a broken one does not. Any account of why humans and rats both possess a functional GULO can appeal to the fact that both need vitamin C, so the shared feature is explained by shared requirements rather than shared ancestry.

A pseudogene has no such escape. The human genome contains the machinery for making vitamin C, in the right place, in a form that cannot work, in an animal that would benefit from it working. On the descent account this is exactly what should be found: an ancestor with a working gene, a lineage that happened to eat enough fruit for the loss to cost nothing, a disabling mutation that drifted to fixation, and subsequent neutral decay. On any account in which each species was arranged independently for its needs, the sequence has to be explained as something deliberately included and deliberately disabled.

Now you. Guinea pigs also lack vitamin C synthesis. What does the descent account predict about how the guinea pig's GULO is broken compared with the human one, and what would falsify it?

Answer

It predicts different lesions. Primates and rodents separated long before either lost the function, so the two losses are independent events and the mutations that caused them should have nothing in common: different exons missing, different frameshifts, different stop codons. That is what is observed. The primate pseudogene is missing a particular set of exons, and the guinea pig pseudogene is disabled by different changes in different places.

More sharply, it predicts a nested pattern within the primates. The disabling mutations in humans, chimpanzees, orangutans and macaques should be the same mutations, because they were inherited from one common ancestor in which the gene broke once, and the differences between these sequences should be the ordinary neutral divergence accumulated since. That is also what is observed.

Falsification would be straightforward. If the human and guinea pig pseudogenes carried the same disabling mutations at the same positions, or if humans and chimpanzees carried different ones, the shared-error argument would collapse, because the errors would no longer track the tree. This is worth stating because it shows the evidence is not a story fitted after the fact: the pattern of breakages has a specific predicted shape and could have had any other.

The genome is full of the same pattern. Humans carry roughly 400 working olfactory receptor genes and roughly 470 broken ones, a loss shared with other primates in a nested arrangement; toothless baleen whales carry disabled enamel genes with frameshifts shared across the baleen whales and absent in toothed ones; and placental mammals including humans retain decayed remnants of the egg yolk protein genes their egg-laying ancestors used.

A chromosome count that does not match

There is one more case worth working because it was a genuine prediction with a plain physical answer.

Humans have 23 pairs of chromosomes. Chimpanzees, gorillas and orangutans all have 24. Since all four descend from a common ancestor, one of two things happened: the great apes independently gained a chromosome by splitting one, or the human lineage lost one by fusing two.

Example. Take the fusion hypothesis. What must be true of a human chromosome if it is a fusion of two ancestral ones, and where exactly should the evidence lie?

Chromosomes end in telomeres, tandem repeats of the sequence TTAGGG in vertebrates, and each has one centromere, the region where the spindle attaches. If two chromosomes joined end to end, then the resulting chromosome must contain, somewhere in its middle, the remains of two telomeres facing each other, since the joined ends were previously chromosome tips. It must also contain the remains of two centromeres, one functional and one that has been silenced, because a chromosome with two active centromeres is pulled in both directions and torn apart at cell division. And the banding pattern of the fused chromosome must match the two ape chromosomes laid end to end.

Now you. All three were checked. What was found, and how much does it establish?

Answer

Human chromosome 2 is the fusion. Its banding pattern corresponds to chimpanzee chromosomes 2A and 2B placed end to end, which was noticed in the 1980s. In 1991 a team led by Jonathan IJdo reported the sequence at band 2q13: a stretch of degenerate TTAGGG repeats arranged head to head, exactly the inverted arrangement two joined chromosome tips would produce, and degenerate in the way sequence no longer maintained by telomerase would become. And at 2q21 there is a region of the alpha satellite DNA characteristic of centromeres, corresponding in position to the centromere of chimpanzee 2B, inactive. The whole structure was confirmed in full sequence when the chimpanzee genome was published in 2005.

What it establishes is narrow and strong. It establishes that the human lineage underwent a specific chromosomal event, and it removes the chromosome-count difference from the list of objections, since the count now has a mechanism with physical remains. It does not by itself establish common descent, which rests on the accumulated pattern rather than on any single case. Its value is that it was a prediction with three independent components, each of which could have come out otherwise, and the fusion site is precisely the sort of thing that is not needed by, and is a positive nuisance to, any organism that has it.

What the alternative would have to say

It is worth stating plainly what the shared-error evidence demands of a design account, because this is where the argument of the first lesson closes.

Paley's inference was from the co-adaptation of parts to an end, and it is a good inference about eyes. It has nothing to say about a disabled gene for a vitamin its bearer must otherwise eat, about 200,000 fragments of dead virus at matching addresses in two species, or about a chromosome carrying the scar of a join it did not need. To keep the design account, each of these has to be attributed to a designer who inserted broken machinery, copied one lineage's specific accidents into another's genome, and arranged the whole collection so that it reproduces, across hundreds of thousands of independent features, the same branching diagram that anatomy and the fossil sequence give.

That is not impossible. It is unfalsifiable, which is a different and worse property, and it was the ground on which the argument was actually decided.

What this evidence does not settle

Genomic evidence establishes relationship and history extremely well and mechanism hardly at all. That two species share an ancestor is one claim; that the differences between them accumulated by mutation, drift and selection at the rates the earlier lessons measured is another, and the sequences alone do not prove it. A pseudogene shows that a gene broke; it does not show that natural selection built the gene in the first place.

Which is why the next lesson leaves the record entirely and looks at the process running, in populations where the starting frequencies were measured, the selective agent is known, and the change happened while somebody was watching.