A structure that suggests a copying mechanism is a hypothesis, and by 1957 there were three competing versions of how DNA might be duplicated.
The previous lesson ended with Watson and Crick's remark that base pairing suggests a way to copy the molecule. Take that seriously and three things could happen when a double helix is duplicated. Semiconservative: the strands separate, each templates a new partner, and every daughter molecule has one old strand and one new. Conservative: the original double helix somehow directs the synthesis of an entirely new one and stays intact itself. Dispersive: the molecule is copied in pieces that are interleaved, so both daughter molecules are patchworks of old and new along their length.
All three are consistent with base pairing. Only an experiment can choose.
The most beautiful experiment in biology
Matthew Meselson and Franklin Stahl, at Caltech in 1958, found a way to weigh DNA precisely enough to tell the three apart.
The tool is caesium chloride density gradient centrifugation. Spin a concentrated caesium chloride solution at very high speed for long enough and the heavy caesium ions redistribute until the solution has a smooth density gradient down the tube. DNA in that tube migrates to the depth where the solution density matches its own and forms a sharp band there, which can be photographed by ultraviolet absorption. The resolution is remarkable: it separates molecules differing in density by well under one per cent.
The label is nitrogen. DNA bases are rich in nitrogen, and the heavy stable isotope nitrogen-15 is not radioactive, so bacteria can simply be grown on it. Meselson and Stahl grew E. coli for many generations on ammonium chloride made with nitrogen-15 until essentially all its DNA was heavy, then abruptly transferred the culture to ordinary nitrogen-14 medium and took samples at intervals.
The predictions differ sharply. After exactly one round of replication, conservative copying gives two populations, one fully heavy and one fully light, so two bands. Semiconservative gives every molecule one heavy strand and one light one, so a single band exactly halfway between. Dispersive also gives a single band halfway between, so one generation cannot distinguish the last two.
The observation after one generation was a single intermediate band, at the density expected for a hybrid. Conservative replication was dead.
The second generation separates the survivors. Semiconservative copying predicts that each hybrid molecule gives one hybrid and one fully light molecule, so half hybrid and half light, in two distinct bands. Dispersive copying predicts a single band that has moved to three-quarters light, since every molecule is still a uniform patchwork. What appeared was two bands in equal amounts, one at hybrid density and one at light. Dispersive was dead too.
Meselson and Stahl added a further check that is often left out and which closes the argument properly. They heated the hybrid DNA to separate the strands, and the two single strands banded at two different densities, one fully heavy and one fully light. Each strand was uniform along its length, which no dispersive scheme allows.
Semiconservative replication was established with a design that answers the question directly, needs no assumptions about mechanism, and can be understood from the photograph of the tubes alone.
Why the machinery cannot be as tidy as the picture
The clean image is two strands unzipping while polymerases run along behind them. The actual chemistry forbids it, for a reason worth deriving rather than memorising.
DNA polymerase adds a nucleotide by having the free 3' hydroxyl of the growing chain attack the innermost phosphate of an incoming nucleoside triphosphate, displacing pyrophosphate. The energy for the new bond comes from the incoming nucleotide, not from the chain. Synthesis therefore runs only in the 5' to 3' direction, because that is the end that carries a free hydroxyl to attack with.
Could a polymerase have evolved to run the other way, carrying the triphosphate at the growing end of the chain and attacking incoming nucleotides with it? Chemically, perhaps. But then any removal of a wrongly incorporated base would take the triphosphate away with it and leave a chain that cannot be extended, and the whole proofreading system described below becomes impossible. One-directional synthesis is the price of being able to correct mistakes, and every DNA polymerase in every organism pays it.
The consequence is that the two strands of a replication fork cannot be copied in the same way, because they are antiparallel. As the fork opens, one template runs 3' to 5' into the fork, and its new strand can be made continuously in the direction the fork is moving. This is the leading strand. The other template runs the wrong way, and its new strand has to be made in short pieces pointing backwards, each started afresh as more template is exposed. This is the lagging strand, and the pieces are Okazaki fragments, found by Reiji and Tsuneko Okazaki in 1968 by labelling replicating DNA for a few seconds and finding the label first in short pieces and only later in long ones. They run about 1000 to 2000 nucleotides in bacteria and 100 to 200 in eukaryotes, and DNA ligase seals them together afterwards.
There is a second awkwardness. No DNA polymerase can start a chain; all of them can only extend an existing 3' end. Every fragment therefore needs a primer, and the primer is laid down by primase, which makes a short piece of RNA. Using RNA looks like a complication and is a solution: primase is much less accurate than DNA polymerase, and the cell can afford that only because the primer is later recognised as RNA, excised, and replaced with DNA made properly. Marking the low-quality start with a chemically different material is what makes its removal possible.
Example. E. coli replicates its 4,641,652 base pair genome from one origin, in both directions, with each fork running at about 1000 nucleotides per second. Work out the time required, and then reconcile it with the fact that the organism can divide every 20 minutes.
Two forks share the genome, so each covers about 2.32 million bases at 1000 per second, which is 2320 s, or 38.7 minutes. Call it 40. That is nearly twice the fastest doubling time, which appears impossible: the cell divides before it has finished copying its own DNA. The resolution is that replication rounds overlap. A fast-growing cell initiates a new round at the origin before the previous round has reached the terminus, so at any moment the chromosome carries several forks, and the region near the origin is present in four or eight copies while the terminus is present in one. Division then happens every 20 minutes even though each individual round takes 40, in the same way a factory with a two-hour assembly time can still ship a unit every hour by having two on the line. This has a testable consequence and it is observed: genes near the origin are at higher copy number in fast-growing cells than genes near the terminus, and are correspondingly more highly expressed. Bacteria exploit this, placing ribosomal RNA genes close to the origin, which is exactly where a fast-growing cell wants extra copies.
Now you. A human cell must copy 6.4 billion base pairs, and its polymerases run at roughly 50 nucleotides per second, five times slower than the bacterial enzyme. S phase lasts about eight hours. What does this force to be true, and roughly how many of whatever it is are needed?
Answer
It forces many origins rather than one. A single fork at 50 nucleotides per second covers million bases in eight hours, so copying 6.4 billion needs at least forks, which is about 2200 origins working simultaneously. Measured estimates for a human cell run to 30,000 to 50,000 origins, an order of magnitude more than the minimum, and the excess matters: not all origins fire in every cell cycle, and the dormant ones are a reserve that can be activated if a fork stalls. Two further requirements follow. The origins must be coordinated, since firing the same stretch twice would duplicate a region and firing none would leave a gap, and eukaryotes solve this with a licensing system that marks each origin once per cycle and cannot re-mark it until the cell has passed through mitosis. And the forks must eventually meet, so the chromosome ends up as one continuous molecule, which requires thousands of ligation events per chromosome per division.
Getting to one in a billion
Fidelity is the number that makes heredity possible, and it is achieved by three filters in series rather than by any single accurate step.
The first is base selection by the polymerase. The enzyme's active site is shaped to fit a correct Watson-Crick pair, and it closes around the incoming nucleotide only when the geometry is right; a mismatched pair has the wrong width and the wrong hydrogen bond pattern, so the catalytic residues do not line up. This alone gives about one error in . Notice what the mechanism implies: the discrimination is geometric rather than energetic, which is why it works for all four correct pairs despite A-T having two hydrogen bonds and G-C having three.
The second is proofreading. Most DNA polymerases carry a second active site, a 3' to 5' exonuclease, positioned about 3 nm from the polymerising site. A correctly paired 3' end sits in the polymerase site; a mismatched end is frayed and unstable, and the single strand flops across into the exonuclease site, where the last nucleotide is chopped off before synthesis resumes. This is the mechanism that one-directional synthesis exists to permit, and it improves fidelity by a further factor of about 100.
The third is mismatch repair, which operates after the fork has passed. A separate protein complex scans the new duplex for the distortion a mismatched pair makes, excises a stretch of the new strand containing it, and resynthesises. The hard part is knowing which strand is new, since both look chemically identical: bacteria use transient undermethylation of the new strand, and eukaryotes appear to use the nicks between Okazaki fragments and the association with the replication machinery itself. This filter is worth another factor of 100 to 1000.
Multiply them: , one error per billion bases copied. For E. coli that is mutations per genome per replication, or one mutation somewhere in the chromosome roughly every 200 divisions.
For humans the measured germline rate, from sequencing parents and children directly, is about per base per generation, which across 6.4 billion base pairs gives about 77 new mutations in every child. The observed figure is around 70. That rate is higher than the per-replication rate because a generation involves many cell divisions and decades of chemical damage in between, and it rises with the father's age because sperm are produced by continuing division while eggs are not.
The fidelity is not maximal, and that is not an accident of engineering. A population with a zero mutation rate cannot adapt to anything, and the rates observed across organisms cluster near a value that balances the cost of deleterious mutations against the need for variation. Losing mismatch repair is a real and well-characterised human condition, Lynch syndrome, which raises the mutation rate roughly a hundredfold and causes early colorectal and other cancers. The connection between copying fidelity and cancer is the subject of the last lesson of this course.
Damage, which never stops
Copying accurately would be sufficient if DNA were chemically inert between copies, and it is not. A human cell suffers on the order of ten thousand depurinations a day, where a purine base simply falls off its sugar, a few hundred cytosine deaminations, and a continuous background of oxidative lesions from the by-products of the respiration described earlier in this course. An hour of bright sunlight generates tens of thousands of ultraviolet-induced pyrimidine dimers in an exposed skin cell.
Every one of those would be a mutation without repair, so a cell runs several repair systems continuously. Base excision repair cuts out a single damaged base and replaces it. Nucleotide excision repair removes a stretch of a dozen or more nucleotides around a bulky distortion such as a pyrimidine dimer. Mismatch repair, described above, corrects replication errors. Double-strand breaks, the most dangerous lesion because no intact template remains on either strand, are handled either by homologous recombination, which copies the sequence from the sister chromatid and is accurate but only available after replication, or by non-homologous end joining, which simply ligates the ends and often loses a few bases.
The common architecture is worth stating: every one of these works because the molecule is double-stranded, so the information lost from one strand is still present on the other. That is the redundancy the previous lesson identified as the point of the structure, now doing its second job.
Example. People with xeroderma pigmentosum lack functional nucleotide excision repair, and develop skin cancers, often hundreds of them, at a rate roughly a thousand times the normal rate, beginning in early childhood and confined almost entirely to sun-exposed skin. Why is this a much stronger argument that ultraviolet light causes skin cancer than any epidemiological correlation could be?
Because it identifies the mechanism and it predicts the pattern. An epidemiological association between sun exposure and skin cancer is consistent with many explanations, including confounding by outdoor occupation, skin type or something else that tracks sunlight. Xeroderma pigmentosum specifies exactly which repair pathway is missing, that pathway is exactly the one that removes exactly the lesion ultraviolet light makes, and the consequence is exactly the tumours predicted, in exactly the places predicted, at an enormously magnified rate. The chain from cause to lesion to failed repair to mutation to tumour is complete and each link is independently established. It is also quantitative in a useful direction: if removing one repair pathway multiplies the risk a thousandfold, the pathway is normally correcting nearly all of the damage, which tells you that the ordinary rate of ultraviolet lesions is very high and that our tolerance of sunlight is bought entirely by repair rather than by resistance.
Now you. Tumours carrying inherited mutations in BRCA1 or BRCA2, which are needed for homologous recombination, are treated with drugs that inhibit PARP, an enzyme involved in repairing single-strand breaks. Why does inhibiting a different repair pathway kill these tumour cells while largely sparing the patient's normal cells?
Answer
Because the patient's normal cells still carry one working copy of BRCA, and the tumour has lost both. Blocking PARP leaves single-strand breaks unrepaired; when a replication fork reaches one it collapses into a double-strand break, which a cell with intact homologous recombination repairs accurately and a cell without it cannot. Losing either pathway alone is survivable and losing both is not, which is what synthetic lethality means, and here the first loss was supplied by the tumour's own genetics and the second by the drug. Two things follow. The therapeutic window comes from a difference between tumour and host that already exists rather than from any selectivity of the molecule, which is why the drug must be prescribed on the basis of a genetic test rather than a tissue of origin. And the resistance mechanism is predictable and is observed: tumours that restore homologous recombination, sometimes by a second mutation that repairs the reading frame of the broken BRCA gene, become resistant, which is the same selection argument that appears wherever a therapy kills most of a variable population.
Unwinding, and what it does to the rest of the molecule
A helix cannot be opened without consequence. Helicase separates the strands at the fork, using ATP, and every turn it opens must go somewhere, because the DNA beyond the fork is not free to spin: it is long, entangled, and in bacteria a closed circle.
The result is that torsional strain accumulates ahead of the fork as positive supercoiling. At 1000 nucleotides per second and about 10.5 base pairs per turn, the fork is generating roughly 95 turns per second, which if the molecule did spin freely would be a rotation at 5700 rpm. Topoisomerases relieve this by cutting one or both strands, letting the molecule rotate or pass through the break, and resealing. Bacterial DNA gyrase does more, actively introducing negative supercoils using ATP, which keeps the chromosome underwound and makes it easier to open.
This is a good drug target precisely because the bacterial and human enzymes differ. Fluoroquinolone antibiotics such as ciprofloxacin inhibit bacterial gyrase, and several anticancer drugs including etoposide and doxorubicin inhibit human topoisomerase II. Both classes work by trapping the enzyme after it has cut the DNA and before it has resealed, so the drug converts a housekeeping enzyme into a machine that makes double-strand breaks, which is a more interesting mechanism than simple inhibition and explains why the cells most affected are the ones replicating fastest.
The ends
A linear chromosome has a problem a circular one does not. The lagging strand is made in fragments, each needing an RNA primer, and the primer at the very end of the chromosome cannot be replaced with DNA once removed, because replacement requires a 3' end upstream to extend from and there is none. Every round of replication therefore shortens the chromosome.
The solution is a repeated sequence at each end, the telomere, which in humans is thousands of copies of TTAGGG, together with proteins that cap it so that the cell does not read a chromosome end as a broken chromosome. Losing 50 to 200 base pairs of a 10 kilobase telomere per division allows on the order of a hundred divisions before functional sequence is reached, which is the right order of magnitude for the roughly fifty divisions Leonard Hayflick measured for cultured human fibroblasts in 1961.
Elizabeth Blackburn and Carol Greider found the enzyme that rebuilds telomeres in 1985, working on a ciliate with an unusually large number of chromosome ends. Telomerase carries its own short RNA template and uses it to extend the chromosome end, which makes it a reverse transcriptase: an enzyme that writes DNA from RNA. It is active in germ cells and stem cells and largely off in ordinary somatic cells, and it is reactivated in around ninety per cent of human cancers, which is what allows a tumour cell lineage to keep dividing past the point where a normal cell would stop.
Example. Some organisms, including bacteria and many viruses, have circular genomes and no telomeres at all. Does this mean the end replication problem is an evolutionary accident that could have been avoided?
Not quite, and the trade is worth spelling out. A circle has no ends, so it needs no telomeres, no telomerase and no cap to distinguish an end from a break. What it cannot easily do is recombine and segregate at the scale a eukaryotic nucleus requires: crossing over between two circles produces a single larger circle rather than two exchanged chromosomes, and a circular chromosome with an odd number of crossovers becomes a catenated dimer that must be resolved before division. Linear chromosomes make meiosis, recombination and the whole apparatus of sexual reproduction tractable, and they allow a genome to be divided into many separately segregating pieces. The end replication problem is the cost of that, and the interesting observation is that the cost turned out to be useful: because telomeres shorten with division, they function as a division counter, and cells that have divided too many times can be retired rather than allowed to accumulate mutations indefinitely. A constraint became a safeguard, which is a common shape in evolution and a poor argument for design.
Now you. A biotechnology company proposes switching telomerase on in adult human tissues to prevent ageing. Argue the case against, using only what is in this lesson.
Answer
The central objection is that telomere shortening is one of the mechanisms limiting the proliferation of a cell that has accumulated damage, and roughly ninety per cent of human cancers reactivate telomerase precisely because they need to escape it. Switching it on everywhere removes a barrier that tumours normally have to break through, in exactly the cells most likely to be on the way to becoming tumours, since those are the ones that have divided most and therefore have the shortest telomeres and the most accumulated mutations. The mutation arithmetic earlier in this lesson sharpens the point: at per base, a lineage that divides many more times accumulates proportionally more mutations, so extending division capacity and raising mutation load are the same intervention. A second objection is that telomere length is not the only thing limiting cell lifespan, since senescent cells also accumulate oxidative and protein damage, so the intervention might buy divisions without buying function. The honest conclusion is not that the idea is worthless but that it is a trade between two failure modes, and that anyone proposing it has to say why the cancer risk is acceptable rather than treat telomerase as a repair.
The sequence is now copied, checked, and passed on with about one error per billion bases. Nothing so far has read it. A gene sitting in a chromosome does nothing at all, and the machinery of the first six lessons is all protein, made of a different chemical alphabet in a different compartment. The next lesson is the first half of the connection: a working copy of the message, made of a material that is deliberately built not to last.