Sign in

Libre University uses your GitHub account. Signing in is only needed to sit a final test, so the score is kept on your profile.

Reading it out

A gene sitting in a chromosome does nothing, and the first thing a cell does with one is make a copy of it out of a different and deliberately unstable material.

The reason there has to be an intermediate at all can be argued before any experiment, and the first lesson of this course did so. A bacterium holds one copy of each gene and about three million protein molecules, so the average gene has been turned into hundreds or thousands of products. Reading a gene cannot consume it, and one template cannot serve as the direct assembly site for a thousand simultaneous copies. Something has to amplify, and whatever amplifies has to be removable when the product is no longer wanted.

In a eukaryote there is a second argument, purely geographical. DNA is inside the nucleus, ribosomes are outside it, and a chromosome does not leave. Whatever carries the message must be able to cross the nuclear envelope, and DNA does not.

Finding the messenger

Elliot Volkin and Lazarus Astrachan noticed in 1956 that bacteria infected with a bacteriophage make a burst of RNA whose base composition resembles the phage DNA rather than the host's, and which turns over rapidly. It was the right observation and nobody knew what to do with it, because the reigning model held that each ribosome was a dedicated machine specialised for one protein.

Sydney Brenner, François Jacob and Matthew Meselson tested that model directly in 1961, using the density labelling trick from the previous lesson. They grew E. coli on heavy isotopes so that all its ribosomes were dense, then switched to light medium and infected with phage. If ribosomes were specialised, the phage would have to build new, light ribosomes to make its own proteins. If instead ribosomes are generic readers, the phage would make only a new message and feed it to the old ribosomes.

The new, rapidly labelled RNA was found associated with the pre-existing heavy ribosomes, and phage protein was made on them. Ribosomes are not specialised. They are interchangeable machines that read whatever message is loaded into them, and the specificity lives in the message. That is the messenger RNA hypothesis, and the experiment establishes both halves of it at once.

Why the copy is RNA, and why RNA is a poor archive

RNA differs from DNA in two ways, and both are consequences rather than accidents.

The sugar carries a hydroxyl at the 2' position. That hydroxyl sits next to the phosphodiester backbone and can attack it, so RNA hydrolyses spontaneously far faster than DNA, particularly at alkaline pH. A molecule with a built-in self-destruct is a bad archive and an excellent temporary message.

RNA uses uracil where DNA uses thymine, and thymine is simply uracil with a methyl group. Making that methyl group costs the cell energy at every one of the billions of thymines in a genome, which looks like waste until you ask what cytosine does when left alone. Cytosine deaminates spontaneously to uracil, and in a human cell this happens on the order of a hundred times a day. If DNA contained uracil normally, a repair system would have no way to tell an original U from a C that had decayed, and every deamination would become a permanent C to T mutation. Because DNA uses thymine, any uracil found in DNA is damage by definition, and a dedicated enzyme, uracil-DNA glycosylase, removes it. The cell pays for a methyl group on every thymine in order to make an entire class of chemical damage detectable. RNA, which is discarded within minutes, does not need the protection and does not pay for it.

So the division of labour is chemical. DNA is methylated, double-stranded and 2'-deoxy, which is to say built for permanence and correctability. RNA is single-stranded, unmethylated at that position, and 2'-hydroxylated, which is to say built to be made quickly and destroyed quickly.

Where to start, and how the polymerase knows

RNA polymerase copies one strand of the DNA, running 5' to 3' along the growing RNA exactly as DNA polymerase does, and using the same base pairing except that adenine on the template calls for uracil. It needs no primer, which is a striking difference from DNA polymerase and follows from the difference in what a mistake costs: a wrong base in a message that lives five minutes is thrown away with the message, while a wrong base in a genome is inherited forever. RNA polymerase accordingly runs at an error rate around 10-4 to 10-5, thousands of times worse than replication, and it does not proofread anything like as hard.

The harder problem is where to start. A bacterial genome of 4.6 million base pairs contains a few thousand genes, and the polymerase must find their beginnings and no other position. It does so by recognising a promoter: in E. coli, two short sequences upstream of the start, centred about 10 and 35 base pairs before it, with consensus sequences TATAAT and TTGACA. The recognition is done not by the polymerase itself but by a detachable subunit, the sigma factor, which binds the promoter, positions the enzyme, and falls off once transcription is under way.

That architecture is worth pausing on because it is a general design. A cell holding several different sigma factors can redirect its entire transcriptional programme by changing which one is loaded, since each recognises a different promoter consensus. E. coli switches to a heat shock sigma factor within seconds of a temperature rise, and Bacillus subtilis runs sporulation through a timed cascade of them. One interchangeable part on a generic machine controls which subset of the genome is read.

Eukaryotes are more elaborate. There are three polymerases, of which RNA polymerase II makes all messenger RNA, and it cannot recognise a promoter at all on its own. A set of general transcription factors assembles at the promoter first and recruits it, and the whole assembly is further controlled by regulatory proteins binding at enhancers, which may be tens of thousands of base pairs away and are brought close by looping of the DNA. The next lesson is about what that elaboration buys.

Rates are moderate. Bacterial RNA polymerase runs at 50 to 90 nucleotides per second, so a typical 1000-base gene is transcribed in about 20 seconds. RNA polymerase II runs slower, around 2000 nucleotides per minute.

Example. In bacteria, ribosomes attach to a messenger RNA and begin making protein while RNA polymerase is still transcribing the far end of it, so transcription and translation are physically coupled. In eukaryotes the nuclear envelope makes this impossible. What does the separation make available, and what does it cost?

What it makes available is processing. A message that must be finished, modified and exported before it can be read is a message that can be edited in between, and every eukaryotic modification described in the rest of this lesson, capping, splicing, polyadenylation and quality control, depends on there being a stage at which the transcript exists but is not yet being translated. Coupled bacteria cannot splice, because a ribosome would already have read the intron before it could be removed. The cost is speed and simplicity. A bacterium responds to a change in its environment by transcribing a gene and having protein appear seconds later, while a eukaryote takes minutes. It is a real trade, and it maps onto how the two kinds of organism live: a bacterium competes on how fast it can respond and divide, and a multicellular eukaryote is playing a slower game where regulatory sophistication is worth more than a few minutes of delay. The general point is that a compartment boundary is not only a barrier but an opportunity, because it creates a stage in a process where something can be inserted.

Now you. Messenger RNA in E. coli has a median half-life of about five minutes, while the median human mRNA lasts around ten hours. Take a cell that abruptly stops transcribing a gene. Work out what fraction of the message remains after thirty minutes in each case, and say what the difference is for.

Answer

Thirty minutes is six half-lives for the bacterium, leaving 0.56=0.016, under two per cent. For the human message thirty minutes is one twentieth of a half-life, leaving 0.50.05=0.97, essentially all of it. The difference is a difference in what the two cells use transcriptional control for. A bacterium switches genes off by ceasing transcription and letting the existing message decay, and this only works if decay is fast, so a short half-life is what makes the switch sharp. Its environment can change in seconds and it must be able to stop making a protein almost as fast as it started. A human cell in a tissue is not making that kind of decision: its expression programme is largely stable over hours to days, so a durable message is cheaper, since it is amplified more per transcription event. The general principle is that the response time of any regulated system is set by the lifetime of its components, not by how fast it can be switched, and a cell that needs to respond quickly must be willing to throw things away quickly. Half-lives are not uniform within either organism, and the exceptions prove the rule: the human messages with the shortest half-lives, minutes rather than hours, encode cytokines, cell cycle regulators and transcription factors, which are exactly the proteins whose levels must be able to fall fast.

Genes are not continuous

In 1977, working independently, Richard Roberts and Phillip Sharp did an experiment on adenovirus that nobody expected to produce a surprise. They hybridised a mature messenger RNA to the DNA of the gene that made it and looked at the result in an electron microscope.

If a gene were a continuous stretch of DNA matching its message, the hybrid would be a simple double-stranded line. What they saw instead was a hybrid interrupted by large loops of DNA hanging out unpaired. Stretches of the gene were simply absent from the message. Roberts and Sharp shared a Nobel Prize in 1993 for what turned out to be a general feature of eukaryotic genes.

The absent stretches are introns and the retained ones exons. The primary transcript contains both, and the spliceosome, a large complex of small nuclear RNAs and proteins, cuts out each intron and joins the flanking exons. It finds the boundaries partly by short consensus sequences, almost always GU at the start of an intron and AG at the end, and partly by an internal branch point, and the chemistry runs through a lariat intermediate in which the intron's 5' end is joined to the branch point before the exons are ligated.

The scale is startling. A typical human gene occupies about 27,000 base pairs of DNA and yields a mature coding message of a little over 1000 nucleotides, so the great majority of what is transcribed is discarded within minutes. The dystrophin gene is 2.4 million base pairs with 79 exons, and at RNA polymerase II's elongation rate transcribing it once takes over twelve hours. Around 1.5 per cent of the human genome codes for protein.

That is not obviously good engineering, and for a while introns were widely described as junk. Three things they buy are now clear.

The first is alternative splicing. If exons can be joined in more than one combination, one gene can specify several proteins. Around ninety-five per cent of human multi-exon genes are alternatively spliced, and the extreme case is the Dscam gene of the fruit fly, which offers 12, 48, 33 and 2 mutually exclusive alternatives at four positions, giving 12×48×33×2=38{,}016 possible messages from a single gene, more than the fly has genes. This is the main reason the human genome's roughly 20,000 protein-coding genes, a number that shocked people when it was published in 2001, is compatible with a far larger number of distinct proteins.

The second is regulation, since splicing itself can be controlled, so which protein a gene makes can depend on the cell type or the signal received.

The third is evolutionary. Exons often correspond to structural or functional modules of a protein, and recombination within introns can shuffle those modules between genes without disrupting either, which is a much more promising way to invent a new protein than accumulating point mutations.

None of this shows that introns arose because they were useful. The honest position is that they are ancient, that they impose a real cost in transcription and in splicing errors, and that lineages under pressure to be small and fast, including yeast and most bacteria, have lost nearly all of them.

The rest of the processing

Three further modifications happen to a eukaryotic message and each has a job.

A modified guanine cap is added to the 5' end within seconds of transcription starting. It protects that end from exonucleases and is the mark the ribosome recognises when it loads.

A poly(A) tail of a couple of hundred adenines is added to the 3' end after the transcript is cut at a signal sequence. It also protects against degradation, and its gradual shortening in the cytoplasm is one of the clocks that sets a message's lifetime.

Export through the nuclear pore is selective, and a transcript that has not been properly capped, spliced and polyadenylated is retained and degraded. The pore is a quality control gate, which matters because an incompletely spliced message would be translated into a wrong protein.

The RNA that is never translated

Messenger RNA is a minority product. By mass, most of the RNA in any cell is ribosomal RNA, which is transcribed by its own polymerase, is never translated, and turns out in a later lesson to be the catalyst of protein synthesis rather than its scaffold. Transfer RNA is likewise a final product. The small nuclear RNAs of the spliceosome are another, and they carry out the cutting described above.

Beyond those, eukaryotes make regulatory RNA. Micro RNAs, of which humans have several hundred well-supported examples, are short transcripts that base pair with target messages and suppress them, either by blocking translation or by triggering degradation. Long non-coding RNAs do a variety of jobs, one of which, shutting down an entire X chromosome, appears in a later lesson. And some RNAs are catalysts outright: Thomas Cech found in 1982 that an intron in Tetrahymena splices itself out with no protein present, and Sidney Altman showed that the RNA component of ribonuclease P is the catalytic part. They shared a Nobel Prize in 1989 for establishing that RNA can be an enzyme.

The phrase "one gene, one protein" was a useful approximation in 1941 and is not a description of a eukaryotic cell.

Example. Andrew Fire and Craig Mello injected RNA into the nematode C. elegans to try to suppress a gene. Injecting the antisense strand alone gave weak silencing, injecting the sense strand alone gave weak silencing, and injecting both together as a duplex silenced the gene powerfully at very low doses. What does the dose tell you?

That the mechanism is catalytic rather than stoichiometric. If a suppressing RNA worked simply by pairing with its target and blocking it, one molecule could disable at most one message, and the effect would scale with the amount injected. Silencing at a few molecules per cell means each injected duplex must be responsible for destroying many messages, which requires an enzymatic machine that is guided by the RNA and reused. That is what was found: the duplex is cut into short fragments, one strand of which is loaded into a protein complex that then cleaves every message matching it, repeatedly. Fire and Mello received a Nobel Prize in 2006. The wider point is a general way of reading an experiment, and it recurs throughout molecular biology: a potency far higher than the number of molecules present is the signature of catalysis, and it distinguishes a guide from a blocker without knowing anything about the proteins involved.

Now you. A messenger RNA vaccine has to deliver an intact message into the cytoplasm of a human cell. Using this lesson and the lesson on membranes, name the two problems that have to be solved and how each is addressed.

Answer

The first is stability and immune detection. RNA is intrinsically short-lived, as this lesson argued it is built to be, and cells additionally carry sensors that detect foreign RNA and shut down translation. Both are addressed chemically: the messages are capped and polyadenylated as a natural transcript would be, and uridine is replaced throughout by a modified nucleoside, which greatly reduces recognition by those sensors and raises the protein yield. Katalin Kariko and Drew Weissman published that finding in 2005 and received a Nobel Prize in 2023. The second problem is delivery, and it is a membrane problem: an RNA molecule is large and carries a phosphate charge on every residue, so it is exactly the class of molecule the second lesson of this course showed cannot cross a lipid bilayer. The solution is to package it in a lipid nanoparticle containing an ionisable lipid, which is neutral outside the cell and becomes positively charged in the acidic interior of an endosome, where it destabilises the endosomal membrane and releases the cargo. Both halves of the design are direct applications of results in this course, which is a reasonable answer to anyone who asks what the chemistry of a bilayer is good for knowing.

Example. Beta thalassaemia is often caused by mutations that create a new GU sequence inside an intron of the beta-globin gene, or destroy an existing one at a real boundary. Explain how a single base change in a stretch of DNA that is thrown away can abolish a protein.

Because what is thrown away is decided by sequence, and changing the sequence changes the decision. Create a plausible splice site inside an intron and the spliceosome may use it, so part of the intron is retained in the message; destroy a real one and the spliceosome skips to the next available site, so an exon is lost or intron sequence is read through. In either case the reading frame downstream is usually shifted, since exon lengths are not multiples of three, and a frameshift produces a stop codon within a few dozen codons. The result is no functional beta-globin at all from that allele, from a mutation that touches no codon. The lesson generalises beyond this disease: the information in a gene is not confined to the parts that code, and a substantial fraction of the disease-causing mutations found by sequencing patients lie in splice sites, promoters and regulatory regions rather than in coding sequence. Any analysis that looks only at codons will miss them, which is a practical reason the intron discovery matters clinically and not only conceptually.

Now you. A drug called nusinersen treats spinal muscular atrophy. The disease is caused by loss of the SMN1 gene, but patients retain a nearly identical gene, SMN2, which differs by a single base that causes exon 7 to be skipped most of the time, giving a non-functional protein. Given only that, what kind of molecule would you design?

Answer

Something that changes the splicing decision rather than the gene, and the natural candidate is a short synthetic nucleic acid complementary to a specific sequence on the SMN2 transcript. Nusinersen is exactly that: an antisense oligonucleotide that base pairs with an element in the intron downstream of exon 7 where a repressor protein normally binds, blocking that binding and causing the spliceosome to include exon 7. The result is full-length functional SMN protein made from a gene the patient already has. Two features of the approach follow from the reasoning. It does not need to deliver a gene, only to occupy a site, which is a much smaller molecule and a much less risky intervention than gene therapy. And it must be delivered where it is needed and repeated, because an oligonucleotide is eventually degraded and does not replicate: nusinersen is injected into the spinal fluid every four months. The general principle is worth carrying away, because it is a direct consequence of this lesson. Splicing is a decision made on a molecule that exists for minutes, and a decision is something a drug can lean on.

The message now exists: capped, spliced, exported, and ready to be read. What it is not is protein. A message is a sequence of four kinds of base and a protein is a sequence of twenty kinds of amino acid, and nothing in the chemistry of a base has any affinity for an amino acid. There must be a dictionary, and it must be embodied in physical objects rather than merely written down. The next lesson is how the dictionary was worked out, why it has the structure it has, and what reads it.