Sign in

Libre University uses your GitHub account. Signing in is only needed to sit a final test, so the score is kept on your profile.

Choosing what to make

A cell that transcribed and translated every gene it carries would spend its entire energy budget making proteins it has no use for.

The previous lesson finished the path from gene to protein and priced it: four high-energy phosphates per peptide bond, and the majority of a growing cell's income spent on translation. The question this lesson answers is which genes get read, when, and on what evidence, and the answer is one of the few pieces of molecular biology that was worked out almost completely before any of the molecules involved had been seen.

The cost of making something you do not need

The argument for regulation is usually stated qualitatively and it can be measured. E. coli carries about 4,300 genes and expresses perhaps a third of them at any time. Fully induced, the enzyme beta-galactosidase can reach several per cent of the cell's total protein, and the three genes of the lactose operon together are a real fraction of a cell's manufacturing capacity.

Antony Dean, Daniel Dykhuizen and Daniel Hartl measured the cost directly in the 1980s by growing bacteria in a chemostat with and without unnecessary expression of the lactose genes. The fitness cost of expressing them when there was no lactose to use came out at a few per cent of growth rate. That sounds negligible until it is compounded: a strain growing a few per cent slower is displaced by its competitor within a few hundred generations, which for a bacterium is a matter of days. Regulation is not a refinement. It is the difference between a lineage that persists and one that does not.

The lactose switch

Jacques Monod had observed in 1941 that E. coli given both glucose and lactose grows in two phases: a first burst on glucose, a lag of an hour or so, and then a second burst on lactose. He called it diauxie, and it says two things at once. The cell prefers glucose. And it can only use lactose after some delay, during which it is evidently making something.

What it makes are the products of three adjacent genes, transcribed as one message: lacZ, beta-galactosidase, which cleaves lactose; lacY, a permease that carries lactose into the cell; and lacA, a transacetylase whose role is peripheral. Monod and François Jacob, at the Pasteur Institute, worked out how their expression is controlled, and published the operon model in 1961.

The model has three elements. A promoter, where RNA polymerase binds. An operator, a short sequence overlapping the promoter. And a separate gene, lacI, encoding a repressor protein that binds the operator and blocks transcription. Lactose, or more precisely its isomer allolactose, binds the repressor and changes its shape so that it releases the operator. So the default state is off, and the substrate switches it on by disabling the thing that was holding it off.

Two features of this are worth stating plainly because they recur everywhere. Control is negative: the natural state of the promoter is active and a protein is spent to suppress it. And the signal is the substrate itself, so the switch is a direct report of what is available rather than an inference.

The numbers make the mechanism vivid. A cell contains roughly ten repressor tetramers, against a genome of 4.6 million base pairs, and they must find one 21 base pair operator among them. Induction changes expression of the operon by around a thousandfold.

Cis and trans

The genetics that established the model is the part worth learning, because it is a way of reasoning rather than a fact.

Jacob and Monod could make partial diploids, bacteria carrying a second copy of the lactose region on an F' plasmid, and then ask how two different alleles behave in the same cell. The results split the elements into two classes.

A cell with a broken lacI gene expresses the operon constitutively, all the time. Add a good copy of lacI on a plasmid and regulation is restored, including at the chromosomal operon. Whatever lacI makes is diffusible: it is made in one place and acts anywhere in the cell. Such an element is trans-acting, and it must therefore be a product, a protein or an RNA.

A cell with a mutated operator also expresses its operon constitutively, but adding a good copy of the region on a plasmid does not fix it. The plasmid's operon is regulated normally and the chromosomal one is still stuck on. Whatever the operator is, it acts only on the DNA it is physically part of. Such an element is cis-acting, and it must therefore be a site rather than a product.

That single distinction, made without knowing what either element was, tells you that lacI encodes a diffusible molecule and that the operator is a binding site on the DNA. It remains the standard first test applied to any new regulatory element.

Arthur Pardee, Jacob and Monod added a third result in 1959, in the experiment named PaJaMo after them. They mated a donor carrying working lacZ and lacI into a recipient lacking both. Beta-galactosidase appeared immediately at full rate and then, an hour or so later, shut down. The interpretation is that the recipient initially has the structural gene but no repressor, so transcription runs freely, and expression falls only once enough repressor has been synthesised from the newly arrived lacI. The experiment shows repression is a positive act by a product rather than the absence of an activator, and it also showed that expression stops within minutes of the repressor arriving, which requires the message to be short-lived, a result that fed directly into the messenger RNA experiments of two years later.

Example. A mutant lacI allele is found whose repressor protein binds the operator normally but can no longer bind allolactose. Predict the phenotype, and predict what happens in a partial diploid carrying this allele together with a normal lacI.

The protein binds the operator and can never be released, so the operon is permanently off and the cell cannot use lactose at all. This is the lacI-s superrepressor phenotype. In a partial diploid it is dominant, which is the interesting part: a cell carrying both a normal and a superrepressor allele still cannot induce, because the mutant protein is present, binds the operator, and does not care what the normal protein is doing. Compare this with an ordinary loss-of-function lacI allele, which is recessive, since a normal repressor made from the other copy can act on both operators. The general principle is that loss of function in a trans-acting negative regulator is recessive, while loss of the ability to be switched off is dominant, and a geneticist seeing a dominant regulatory mutation should immediately suspect a regulator that has become insensitive to its signal rather than one that has stopped working. That reasoning transfers directly to cancer genetics, where the distinction between a tumour suppressor that must be lost twice and an oncogene that acts when one copy is altered is exactly this distinction.

Now you. A cell contains about ten repressor tetramers, so their concentration is roughly 1.7×10-8 M. Diffusion-limited binding for a protein and a small target, once the requirement for correct orientation is taken into account, is at best around 107 to 108 M⁻¹s⁻¹. Riggs, Bourgeois and Cohn measured the repressor's association rate in 1970 and got about 1010 M⁻¹s⁻¹. What must be wrong with the assumption?

Answer

The assumption that the repressor searches by three-dimensional diffusion, colliding with the operator directly out of solution. A measured rate a hundred to a thousand times above the three-dimensional limit cannot be explained by making the protein faster or the target bigger, so the search itself must be different in kind. The resolution, developed by Otto Berg, Robert Winter and Peter von Hippel around 1981, is facilitated diffusion: the repressor binds DNA non-specifically with modest affinity and then slides along it, so a single collision anywhere on the chromosome scans a stretch of hundreds of base pairs rather than testing one. The search alternates between one-dimensional sliding, which is thorough but slow at covering distance, and three-dimensional hopping, which covers distance but samples sparsely, and the combination beats either alone. This has since been watched directly with single-molecule fluorescence. The general point is worth keeping, because it recurs whenever a measured rate exceeds a diffusion limit: the limit is not wrong, the assumed geometry of the search is. Reducing the dimensionality of a search is one of the few ways to beat diffusion, and cells use it repeatedly.

An AND gate

Negative control by the repressor is only half the switch. If lactose alone were sufficient, a cell with both sugars available would make lactose enzymes it does not need, and Monod's diauxic curve shows it does not.

The other half is positive. When glucose is scarce, the cell's cyclic AMP concentration rises. Cyclic AMP binds a protein called CAP, the catabolite activator protein, and the complex binds just upstream of the lactose promoter and helps recruit RNA polymerase. The lactose promoter is intrinsically weak, and without CAP even a fully derepressed operon transcribes at only a small fraction of its maximum.

So the operon computes a logical function of two inputs: transcribe if lactose is present AND glucose is absent. Neither input alone is sufficient. This is the first well-characterised biological logic gate, and its architecture, one negative and one positive input converging on one promoter, is the ancestor of every combinatorial regulation scheme in the rest of this lesson.

The design also answers a question the negative-only version leaves open. Why bother with a weak promoter and an activator, rather than a strong promoter and a repressor alone? Because a repressor is never perfect: it dissociates from time to time, and a strong promoter fires during those moments. Making the promoter weak lowers the leak, and adding an activator restores the maximum when it is genuinely wanted. Combining a poor promoter with a recruitable activator gives a much larger ratio between the off and on states than either mechanism can give alone.

Example. Timothy Gardner, Charles Cantor and James Collins built a synthetic genetic toggle switch in E. coli in 2000, using only components of the kind described above. What is the minimal design, and what makes it hold its state?

Two repressors, each transcribed from a promoter that the other one represses. If the first repressor is being made, it shuts off the second's gene, so the second is absent and the first's own promoter is unrepressed, which keeps the first being made. The mirror image is equally stable. The circuit therefore has two stable states and remembers which one it is in, and it can be flipped from outside by transiently inducing whichever repressor is currently off, after which it stays flipped with no further input. Two features are worth extracting. Memory here is a property of the wiring rather than of any molecule: nothing in the circuit is permanently altered, and the state is held by an ongoing pattern of activity. And the design is exactly the lac operon's negative control used twice, which is the point of the experiment, since building a device from characterised parts and having it behave as predicted is a much stronger test of understanding than any observation of a natural circuit.

Now you. The same toggle fails to be bistable if each repressor binds its target promoter as a simple monomer with no cooperativity, and works when the repressors act as dimers or tetramers binding cooperatively. Why should cooperativity matter?

Answer

Because bistability requires the feedback to be steeper than the process it opposes. With a simple non-cooperative repressor, output falls off gradually as repressor rises, and the two mutually repressing branches settle at a single intermediate compromise: one stable state, half on, which is the useless middle. Cooperative binding makes the response sigmoidal, so a small change in repressor concentration produces a large change in transcription, and once the loop's gain exceeds one the intermediate state becomes unstable and the system falls to one extreme or the other. This is why the natural repressors described in this lesson are oligomers, the lac repressor being a tetramer that binds two operators at once, and it connects to the observation about the lac operon in single cells: a graded input can only be converted into an all-or-none decision by a nonlinearity somewhere. The general rule is that positive feedback alone gives amplification, and positive feedback plus nonlinearity gives a switch.

The same problem in a nucleus

Eukaryotic regulation does everything bacterial regulation does and adds several layers, all of which exist because the genome is much larger and because the cell must hold a decision for a lifetime rather than for a generation.

The first layer is packaging. Eukaryotic DNA is wound around histone octamers, 147 base pairs in about 1.65 turns per nucleosome, repeating roughly every 200 base pairs, which puts something like 32 million nucleosomes on a diploid human genome. DNA wrapped on a nucleosome is largely inaccessible, so the default state of a eukaryotic gene is off in a way a bacterial gene never is, and a great deal of regulation consists of moving or modifying nucleosomes rather than of competing with polymerase for a site.

The second layer is chemical marking. Histone tails carry acetyl, methyl and other groups added and removed by dedicated enzymes; acetylation of lysines neutralises their positive charge, loosens the grip on DNA, and is generally associated with active genes. DNA itself is methylated at cytosines in CG dinucleotides, and dense methylation of a promoter region is generally associated with silence. Both kinds of mark can be copied to daughter cells after replication, which gives them something bacterial regulation lacks: memory that survives division.

The third layer is distance and combination. Eukaryotic regulatory sites, enhancers, may sit tens or hundreds of kilobases from the gene they control, in either direction, and act by looping the intervening DNA. Humans have roughly 1,600 transcription factors, and a typical gene integrates inputs from many of them, so the same factor contributes to different outcomes depending on which others are present. This is why a modest number of regulators can specify a large number of distinct cell states.

One genome, many cells

A human body contains a few hundred recognisably different cell types with, in almost every case, identical DNA. That claim needed proving, because the obvious alternative, that differentiation works by discarding genes no longer needed, was seriously held.

John Gurdon tested it in 1962 by transplanting a nucleus from an intestinal cell of a feeding tadpole into an enucleated frog egg. Some of those eggs developed into normal, fertile adult frogs. A nucleus from a fully differentiated cell therefore still contained everything needed to build an entire animal, and differentiation had not removed anything. Ian Wilmut's team extended this to mammals with Dolly the sheep in 1996. Shinya Yamanaka closed the circle in 2006 by showing that four transcription factors introduced into an adult mouse fibroblast could reprogramme it into a pluripotent stem cell, so the reversal does not even require an egg. Gurdon and Yamanaka shared a Nobel Prize in 2012.

There are genuine exceptions and they are informative. Mature red blood cells eject their nuclei entirely. Lymphocytes physically cut and rejoin their antibody and receptor genes, permanently and irreversibly, which is how a limited genome specifies an enormous repertoire of receptors. Both are real changes to DNA in the service of a cell's function, and their rarity is what makes the general rule interesting.

The most visible demonstration of stable regulatory states is on the back of a cat. Female mammals carry two X chromosomes and males one, and the imbalance is corrected by inactivating one X in each cell of a female, early in development, at random. The inactivation is carried out by a long non-coding RNA, Xist, transcribed from the chromosome it silences and coating it. Once made, the choice is inherited by every descendant of that cell. In a cat heterozygous for a coat colour gene on the X, each patch of fur is a clone descended from one early cell, and the patchwork of orange and black is a map of a regulatory decision taken in an embryo and remembered through hundreds of divisions. Calico cats are almost always female for exactly this reason.

Example. Aaron Novick and Milton Weiner reported in 1957 that at intermediate inducer concentrations a population of bacteria does not consist of cells each half induced. It consists of fully induced cells and fully uninduced cells in some proportion. What kind of mechanism produces that, and what is it good for?

Positive feedback produces it. One of the genes the operon switches on is lacY, the permease that brings lactose into the cell, so a cell that is slightly induced imports more inducer, which induces it further. Above a threshold the loop runs to saturation and below it the loop collapses, giving two stable states and nothing stable in between, which is bistability. Whether an individual cell goes up or down at an intermediate external concentration depends on chance fluctuations in the small number of molecules involved, which is why the population splits. Two things make this valuable. It converts a graded and noisy input into a clean decision, so the cell is not left half-committed with the costs of expression and few of the benefits. And it gives memory, since a cell that has switched on stays on even if the inducer falls somewhat, so a fluctuating environment does not cause repeated expensive switching. The same architecture, positive feedback producing two stable states, is how developmental decisions are made permanent in animals, and the calico cat's X inactivation is an instance of it.

Now you. Identical twins have identical genomes and are nonetheless distinguishable, and the differences between them grow with age. Using only what is in this lesson, suggest where the differences could come from.

Answer

Three sources, all consistent with identical DNA sequence. The first is somatic mutation: the fidelity arithmetic from an earlier lesson gives roughly one error per billion bases copied, so after decades of cell division two bodies that started identical carry different mutations, and this accumulates. The second is regulatory state. DNA methylation and histone marks are copied through division but not perfectly, and they respond to diet, smoking, infection and stress, so the epigenetic profiles of twins diverge measurably with age, which has been shown directly by comparing young and old twin pairs. The third is developmental chance. X inactivation in females is random per cell, and many developmental decisions have the same bistable character described above, so which cell takes which path is decided by molecular noise rather than by genotype, and it happens independently in two embryos. The general conclusion is the one this lesson has been building towards. A genome is not a blueprint from which a body is read off; it is a set of instructions whose execution depends on state, and identical instructions run twice do not produce identical results. Twins are the cleanest available demonstration that regulation is not a detail on top of genetics.

Regulation decides which proteins a cell makes and in what quantity, and it does not decide where they end up. A digestive enzyme has to reach the outside of the cell, a respiratory complex has to reach the inner mitochondrial membrane, a histone has to reach the nucleus, and a ribosome makes all of them in the same cytosol. The next lesson follows a newly made protein to its destination, and finds along the way that two of the compartments it can be sent to carry their own genomes, their own ribosomes, and an origin quite unlike the rest of the cell.