Sign in

Libre University uses your GitHub account. Signing in is only needed to sit a final test, so the score is kept on your profile.

What the gene is made of

Every mechanism in the previous six lessons was a specific protein, inherited as a specification from the cell that divided to make this one, and until 1944 almost everyone assumed the specification was itself protein.

That assumption was reasonable, which is worth saying before it is demolished. Chromosomes are made of protein and DNA together. Proteins are built from twenty different amino acids and can be arranged in astronomically many ways, while DNA has only four bases and looked monotonous. Phoebus Levene, who identified the components of nucleic acid, had proposed that DNA was a repeating tetranucleotide, a dull structural polymer with the four bases in fixed rotation. A molecule with no variety cannot carry a message, and for thirty years the field believed DNA had no variety.

A mouse experiment nobody understood

Frederick Griffith was a public health bacteriologist studying pneumococcus during the pandemics of the 1920s, and he was not looking for genes. The bacterium comes in two forms: a smooth strain with a polysaccharide capsule, which is lethal to mice, and a rough strain without one, which is harmless because white blood cells can engulf it.

His 1928 result was a four-way comparison, and its power comes from the controls rather than from the striking case. Live rough bacteria: the mouse lives. Live smooth: the mouse dies. Heat-killed smooth: the mouse lives, so heat has destroyed the organism. Heat-killed smooth mixed with live rough: the mouse dies, and living smooth bacteria can be recovered from its blood.

Something had passed from the dead cells into the living ones and permanently changed them, and it bred true in the descendants. Griffith called it the transforming principle and did not speculate much about what it was. Notice how little the experiment assumes: no chemistry at all, only the observation that a heritable trait can be transferred between cells by a substance from a corpse.

Purifying the principle

Oswald Avery, Colin MacLeod and Maclyn McCarty spent more than a decade at the Rockefeller Institute reproducing transformation in a test tube and then asking which chemical fraction did it. Their 1944 paper is a model of subtractive reasoning.

They prepared extracts of heat-killed smooth cells and purified the active material. Its chemistry looked like DNA: the ultraviolet absorption, the elemental ratio of nitrogen to phosphorus, and behaviour on precipitation with alcohol all matched. Then the removals. Proteases destroyed the protein and transformation still worked. Ribonuclease destroyed the RNA and transformation still worked. Lipid extraction changed nothing. Deoxyribonuclease abolished it completely.

The logical structure is clean: everything else can be destroyed without effect, and only the one thing cannot. The reception was not. The dominant objection was that a trace of protein could have survived the treatment and been the real agent, since a very small amount of a very active substance would be undetectable, and Alfred Mirsky pressed this in print for years. Avery himself was cautious in a way the paper's later reputation obscures, and he was never awarded a Nobel Prize.

The objection was not unreasonable. It is genuinely hard to prove that a purified preparation contains none of a contaminant, and the counter-argument in the end was not chemical purity but the accumulation of independent lines of evidence.

The blender

Alfred Hershey and Martha Chase gave one of those lines in 1952, using a virus that infects bacteria. Bacteriophage T2 consists of nothing but protein and DNA, and its life cycle offered a natural separation: the phage attaches to the outside of the bacterium and something enters, after which the cell produces hundreds of new phage.

The trick is that protein and DNA can be labelled separately by elements each contains and the other does not. Protein contains sulfur, in cysteine and methionine, and no phosphorus. DNA contains phosphorus in its backbone and no sulfur. So they grew one batch of phage with radioactive sulfur-35 and another with radioactive phosphorus-32, let each infect bacteria, then sheared the attached phage coats off the cell surfaces in a kitchen blender and spun the mixture so that the heavy bacteria pelleted and the light phage coats stayed in the supernatant.

Most of the sulfur label ended up in the supernatant, with the discarded coats. Most of the phosphorus label stayed with the cells, and around thirty per cent of it appeared in the next generation of phage. Protein stays outside, DNA goes in.

Example. In the Hershey and Chase experiment roughly twenty per cent of the sulfur label stayed with the bacteria and a substantial fraction of the phosphorus was washed off. Given that the folklore version reports a clean separation, how much weight should the experiment carry?

Less than it is usually given, and it is worth being precise about why. The result is a difference in distribution, not a clean partition: some protein enters or sticks, some DNA fails to. Hershey's own paper is correspondingly hedged, and it concludes that the protein has no function in growth rather than that DNA is the genetic material. Taken alone the experiment is suggestive rather than decisive, and a determined critic could have argued that the twenty per cent of protein entering the cell was the important twenty per cent. What made it persuasive was not its internal cleanliness but its independence: it used a different organism, a different technique and a different logic from Avery's, and it pointed the same way. That is the honest structure of the case for DNA. No single experiment closed it. Three did, together with the fact that the structure found the following year immediately explained something no protein model could. It is also a useful correction to how experiments get remembered, because the blender is memorable and the decade of careful enzymology at the Rockefeller is not.

Now you. Suppose Avery, MacLeod and McCarty's preparation really had contained a trace of protein that was the true transforming agent. Their deoxyribonuclease result would then need another explanation. Can you construct one, and what would you do to test between the two?

Answer

A defender of protein could argue that the hypothetical active protein is bound to DNA and requires that DNA as a carrier or scaffold, so destroying the DNA would inactivate it without the DNA itself being the message. The argument is not absurd, since DNA-binding proteins are real. Two tests distinguish the possibilities. The first is dose response and specificity of the enzyme: deoxyribonuclease abolishes activity while proteases at high concentration do not, and if the carrier account were right one would expect at least partial loss with proteases too, since a protein stripped from its scaffold should also fail. The second and stronger test is to make the material rather than purify it. Transformation with chemically synthesised DNA of known sequence, or with DNA that has passed through a step no protein could survive, removes the contamination argument entirely. That is essentially what the field eventually did, and it is the general escape from any purification argument: stop subtracting and start constructing, because a contaminant that survives a synthesis you performed yourself is not a contaminant anyone can appeal to.

Chargaff's ratios

Erwin Chargaff, at Columbia, applied paper chromatography to hydrolysed DNA from many organisms and published in 1950 two findings, one of which is famous and one of which mattered more at the time.

The famous one is that within any species the amount of adenine equals the amount of thymine and the amount of guanine equals the amount of cytosine. In human thymus DNA the four bases come out at roughly 30.9, 29.4, 19.9 and 19.8 per cent, so A/T=1.05 and G/C=1.005. Within experimental error, one to one.

The finding that mattered more in 1950 is that the ratio of the AT pairs to the GC pairs varies enormously between species. Human DNA is about 40 per cent G plus C, Plasmodium falciparum around 20 per cent, some Streptomyces over 70. That range killed the tetranucleotide hypothesis outright. DNA is not a monotonous repeat; its composition differs from organism to organism, which is the minimum requirement for carrying information.

Example. A bacterium's DNA is 62 per cent G plus C. Work out the percentage of each of the four bases, and say what you can and cannot infer from a high figure.

Chargaff's rules give G=C and A=T, so the 62 per cent divides equally into 31 per cent guanine and 31 per cent cytosine, and the remaining 38 per cent gives 19 per cent adenine and 19 per cent thymine. What you can infer is a physical property: a GC pair makes three hydrogen bonds and an AT pair two, and GC-rich DNA also stacks more favourably, so a GC-rich genome has a higher melting temperature, which is the basis of every calculation anyone does when designing a primer. What you cannot safely infer is that the organism lives somewhere hot. The correlation between whole-genome GC content and growth temperature across prokaryotes is weak, and the strong correlation is with the GC content of the paired stems of ribosomal and transfer RNA, which are structural molecules that must stay folded. Genomic GC content correlates better with other things entirely, including which repair and mutational biases the lineage has. This is a good instance of a real physical mechanism supporting an inference at one level and not at another.

Now you. Two viral genomes are analysed. The first gives A 32, T 32, G 18, C 18 per cent. The second gives A 25, T 33, G 24, C 18. What is the most important difference, and what would you conclude?

Answer

The first obeys Chargaff's rules, with A=T and G=C, and the second does not: adenine and thymine differ by eight percentage points and guanine and cytosine by six. Since the rules follow from every base being physically paired with its complement along a double-stranded molecule, a genome that violates them is not double-stranded. The conclusion is that the second virus has a single-stranded genome, and this is exactly how such genomes were first recognised, in the phage φX174, whose measured composition is close to the second set of numbers. Two further points are worth taking. This is a case where a rule is more useful for its exceptions than for its instances, since the rule holding tells you only what you already assumed while the rule failing identifies something new. And the inference is quantitative rather than qualitative, so it requires knowing the measurement error: base composition determined by chromatography carried an uncertainty of a per cent or so, which is why a discrepancy of six to eight points is convincing and one of one point would not be.

Chargaff had the essential clue and did not see what it meant. The pairing rules are a chemical fact in search of a structural explanation, and supplying that explanation is what the following three years were about.

The structure, and who supplied the data

Rosalind Franklin and Maurice Wilkins at King's College London were taking X-ray diffraction photographs of DNA fibres. Franklin, with her student Raymond Gosling, obtained in May 1952 the photograph known as Photo 51, of the hydrated B form, and it is one of the most information-dense images in science.

Three things can be read off it almost directly. The bold X-shaped cross of reflections is the diffraction signature of a helix, and the angle of its arms gives the pitch relative to the diameter. The spacing of the layer lines gives a repeat of 3.4 nm along the fibre axis. And a strong reflection at 0.34 nm gives the spacing between successive stacked bases, so a turn of the helix contains 3.4/0.34=10 base pairs. The missing fourth layer line indicates two strands offset from each other rather than one, which is where the major and minor grooves come from. Franklin's own analysis had also established that the phosphate backbone lies on the outside, which ruled out the three-chain models with the bases outward that Linus Pauling published and that James Watson and Francis Crick had earlier attempted.

Watson and Crick, at Cambridge, built models rather than taking data, and in early 1953 they saw Photo 51 and an unpublished Medical Research Council progress report containing Franklin's numerical results. This was done without her knowledge or agreement. Their paper in Nature on 25 April 1953 cites her only as having "unpublished" work that "stimulated" them, and she died of ovarian cancer in 1958 at 37, four years before the Nobel Prize went to Watson, Crick and Wilkins. The structure is correct and the model-building was a genuine achievement; the credit was not distributed honestly, and any account that omits this is telling the story wrong.

Why the pairing has to be what it is

The model's central constraint is uniform width, and this is where Chargaff's ratios stop being a coincidence.

The bases come in two sizes. Adenine and guanine are purines, double-ring systems. Thymine and cytosine are pyrimidines, single rings and roughly half the width. If the two backbones are to run at a constant separation, every rung of the ladder must be the same length, and the only way to achieve that with these four components is to pair one purine with one pyrimidine. Two purines would bulge, two pyrimidines would pinch, and either would kink the helix.

That narrows it to four possible pairings, and hydrogen bonding chooses between them. In their correct tautomeric forms, adenine presents a donor and an acceptor in positions that match thymine exactly, making two hydrogen bonds; guanine and cytosine match at three positions. Adenine against cytosine puts two acceptors opposite each other and does not bond. So A pairs with T and G with C, and this is precisely Chargaff's result: the ratios are one to one because the bases are physically paired one to one along the molecule.

The geometry has a further consequence that is easy to pass over and is the reason DNA can carry information at all. Because every rung has the same width, any sequence of pairs makes an equally good helix. The molecule imposes no constraint on the order of its own letters. Compare a protein, whose sequence determines whether it folds at all, and where most random sequences are useless. DNA is a structurally indifferent medium, which is exactly what a message wants to be written on.

Count the capacity. Four possibilities per position is two bits, so the E. coli genome of 4,641,652 base pairs holds 9.3×106 bits, about 1.16 megabytes. A haploid human genome of 3.2 billion base pairs holds 800 megabytes. These figures are upper bounds, since real genomes are highly redundant and much of a human genome is repeated sequence, but they establish the order of magnitude: everything needed to specify a human being would fit on a compact disc.

Example. A student says the two DNA strands carry the same information twice over, so the molecule is fifty per cent redundant and evolution should have removed the duplication. What is wrong with this?

Nothing about the fact, and everything about the conclusion. The two strands are indeed redundant in the information-theoretic sense: given one, the other is fully determined by the pairing rules, so the molecule stores 2 bits per base pair rather than 4. But redundancy is what the structure is for. It gives a copying mechanism, since each strand is a template for the other, which is the subject of the next lesson. It gives a repair mechanism, since damage to one strand can be corrected by reading the intact partner, and a cell suffers thousands of DNA lesions a day that would be uncorrectable in a single-stranded molecule. And it gives chemical stability, because the reactive base edges are turned inward and stacked, protected from water and from oxidation. Single-stranded genomes do exist, in many viruses, and they show the expected properties: much higher mutation rates and much smaller genomes. RNA viruses run mutation rates around 10-4 per base per replication against 10-9 for a cell, and no RNA virus has a genome much beyond thirty thousand bases. The redundancy is not waste. It is what makes a large genome possible.

Now you. Photo 51 shows a strong reflection at a spacing of 0.34 nm and a repeat of 3.4 nm, giving ten base pairs per turn. Suppose an X-ray photograph of a different nucleic acid gave 0.28 nm and 3.1 nm instead. What would you conclude, and what would you want to check next?

Answer

You would conclude that the bases are stacked more closely, at 0.28 nm, and that the helix contains 3.1/0.28=11 base pairs per turn, so it is a different helical form: more tightly wound and shorter per base. This is close to the real A form of DNA, which is what B-DNA becomes when it is dehydrated, and it is also the form double-stranded RNA takes, since the extra hydroxyl on ribose prevents the sugar from adopting the conformation B-DNA requires. The check to run next is humidity, which is exactly what the King's group did: Franklin's key experimental contribution was recognising that DNA fibres exist in two distinct forms and that the sharp, interpretable pictures came from the hydrated B form, so controlling water content was the difference between an uninterpretable smear and Photo 51. The wider point is that a diffraction pattern gives dimensions rather than a structure, and turning dimensions into a structure needs chemistry: bond lengths, ring geometry, what can hydrogen bond to what. Watson and Crick's contribution was that second step, and it needed the first.

The sentence at the end of the paper

Watson and Crick closed their 1953 paper with one of the most quoted sentences in science: "It has not escaped our notice that the specific pairing we have postulated immediately suggests a possible copying mechanism for the genetic material."

The claim is that the structure does not merely accommodate heredity, it explains it. Separate the two strands, and each carries the complete information needed to rebuild its partner, because every base specifies what must sit opposite it. A gene is a sequence, a copy is made by templating, and the same base pairing that holds the molecule together is the mechanism by which it is duplicated.

That is a hypothesis, and a strong one, because it makes a prediction that could fail. If each new double helix is built from one old strand and one new one, then after a single round of copying in labelled medium every molecule should be a hybrid, and after two rounds half should be hybrid and half entirely new, in a specific and measurable ratio. Two other schemes were live at the time and predicted different outcomes. The next lesson is the experiment that decided between them, which is often called the most beautiful in biology, and then the machinery that turns out to carry out the copying, which is stranger than the elegant picture suggests.